Inverted Index for Similar Document Detection: A Case Study at Can Tho University Journal of Science
摘要
The problem of text copying has been a matter of great concern in research and creativity. This study proposes a novel document similarity computation method based on cosine similarity and inverted index. In the research, we deploy the method in two approaches: the traditional approach, which combines cosine similarity with term frequency-inverse document frequency (TF-IDF), and our approach, which combines cosine similarity with inverted index, for calculating the similarity between documents and comparing the effectiveness of the two approaches. The method applies to a dataset collected from Vietnamese articles published in the Can Tho University Journal of Science (CTUJS) from 2015 to 2020. The experimental results show that using Cosine similarity combining the inverted index achieves benefit accuracy and execution time compared to the traditional approach in document similarity measurement tasks.