<p>The problem of text copying has been a matter of great concern in research and creativity. This study proposes a novel document similarity computation method based on cosine similarity and inverted index. In the research, we deploy the method in two approaches: the traditional approach, which combines cosine similarity with term frequency-inverse document frequency (TF-IDF), and our approach, which combines cosine similarity with inverted index, for calculating the similarity between documents and comparing the effectiveness of the two approaches. The method applies to a dataset collected from Vietnamese articles published in the Can Tho University Journal of Science (CTUJS) from 2015 to 2020. The experimental results show that using Cosine similarity combining the inverted index achieves benefit accuracy and execution time compared to the traditional approach in document similarity measurement tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inverted Index for Similar Document Detection: A Case Study at Can Tho University Journal of Science

  • Hai Thanh Nguyen,
  • Ky Hoa Duong,
  • Linh Thuy Thi Pham,
  • Phuong Ha Dang Bui,
  • Nguyen Thai-Nghe,
  • Tran Thanh Dien

摘要

The problem of text copying has been a matter of great concern in research and creativity. This study proposes a novel document similarity computation method based on cosine similarity and inverted index. In the research, we deploy the method in two approaches: the traditional approach, which combines cosine similarity with term frequency-inverse document frequency (TF-IDF), and our approach, which combines cosine similarity with inverted index, for calculating the similarity between documents and comparing the effectiveness of the two approaches. The method applies to a dataset collected from Vietnamese articles published in the Can Tho University Journal of Science (CTUJS) from 2015 to 2020. The experimental results show that using Cosine similarity combining the inverted index achieves benefit accuracy and execution time compared to the traditional approach in document similarity measurement tasks.