<p>Plagiarism is a major problem in education, especially in higher education environments. To address this problem, a comprehensive detection method is proposed, utilizing cutting-edge models like Bidirectional Encoder Representations from Transformers (BERT) and cosine similarity for detecting direct copy and semantic changes. The methodology involves evaluating the BERT model and cosine similarity first, then performing pairwise comparisons to determine how similar a student’s work is to a reference text. If the semantic changes are higher, the content is classified as paraphrased; otherwise, it is considered direct copy in cases where the similarity threshold is greater than 30%. When the similarity threshold is less than or equal to 30%, the method moves to online plagiarism detection. This includes web scraping, generating search queries, and employing Term Frequency-Inverse Document Frequency (TF-IDF) vectorization with cosine similarity to assess the similarities between students’ content and online content. As a result, the study detects online plagiarism and distinguishes between plagiarizing and authentic content. The purpose is to provide an effective method to maintain academic integrity and enhance originality in student work. In similarity testing on MIT plagiarism detection datasets, this study demonstrated an average recall of 80%, precision of 68%, accuracy of 71%, and f1-score of 74%. With a threshold of 0.2, the average precision for online plagiarism detection reached 65%, accompanied by a 56% recall, an accuracy of 60%, and an f1-score of 60%. The results of the study demonstrate the efficiency of the strategy in enhancing plagiarism detection accuracy while reducing false positives.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A comprehensive strategy for identifying plagiarism in academic submissions

  • Deblina Mazumder Setu,
  • Tania Islam,
  • Md. Erfan,
  • Samrat Kumar Dey,
  • Md. Rashid Al Asif,
  • Md. Samsuddoha

摘要

Plagiarism is a major problem in education, especially in higher education environments. To address this problem, a comprehensive detection method is proposed, utilizing cutting-edge models like Bidirectional Encoder Representations from Transformers (BERT) and cosine similarity for detecting direct copy and semantic changes. The methodology involves evaluating the BERT model and cosine similarity first, then performing pairwise comparisons to determine how similar a student’s work is to a reference text. If the semantic changes are higher, the content is classified as paraphrased; otherwise, it is considered direct copy in cases where the similarity threshold is greater than 30%. When the similarity threshold is less than or equal to 30%, the method moves to online plagiarism detection. This includes web scraping, generating search queries, and employing Term Frequency-Inverse Document Frequency (TF-IDF) vectorization with cosine similarity to assess the similarities between students’ content and online content. As a result, the study detects online plagiarism and distinguishes between plagiarizing and authentic content. The purpose is to provide an effective method to maintain academic integrity and enhance originality in student work. In similarity testing on MIT plagiarism detection datasets, this study demonstrated an average recall of 80%, precision of 68%, accuracy of 71%, and f1-score of 74%. With a threshold of 0.2, the average precision for online plagiarism detection reached 65%, accompanied by a 56% recall, an accuracy of 60%, and an f1-score of 60%. The results of the study demonstrate the efficiency of the strategy in enhancing plagiarism detection accuracy while reducing false positives.