Text similarity is important in natural language processing (NLP), including information retrieval, text categorization, and plagiarism detection. This study investigates the field of NLP, especially addressing the growing problem of plagiarism enabled by social networking (SN) platforms. Plagiarism, a global issue, damages academic and professional integrity, with social media serving as a fertile ground for uncredited material duplication. The investigation thoroughly assesses a variety of text similarity metrics, including Jaccard, Cosine, Euclidean, and Hamming distances, as well as novel techniques such as Minhash, and N-gram analysis. The approaches are used to find and measure similarities between documents, which adds to the arsenal of plagiarism detection algorithms. This study identifies essential variables in constructing effective plagiarism detectors, such as proper text representation, similarity metric selection, parameter tuning, and ensemble approaches. The study stresses the significance of strong algorithms in fighting plagiarism, given the huge terrain of digital information. The findings demonstrate limits in existing metrics, with each technique highlighting strengths and flaws. The Jaccard and Cosine similarity emphasizes exact overlaps but ignores semantic subtleties, whereas Euclidean distance and Hamming distance struggle with variable-length texts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Directing Natural Language Processing Text Similarity Challenges in Social Media with AI Techniques

  • N. Madhan,
  • S. Dheva Rajan,
  • Madhuri Jain

摘要

Text similarity is important in natural language processing (NLP), including information retrieval, text categorization, and plagiarism detection. This study investigates the field of NLP, especially addressing the growing problem of plagiarism enabled by social networking (SN) platforms. Plagiarism, a global issue, damages academic and professional integrity, with social media serving as a fertile ground for uncredited material duplication. The investigation thoroughly assesses a variety of text similarity metrics, including Jaccard, Cosine, Euclidean, and Hamming distances, as well as novel techniques such as Minhash, and N-gram analysis. The approaches are used to find and measure similarities between documents, which adds to the arsenal of plagiarism detection algorithms. This study identifies essential variables in constructing effective plagiarism detectors, such as proper text representation, similarity metric selection, parameter tuning, and ensemble approaches. The study stresses the significance of strong algorithms in fighting plagiarism, given the huge terrain of digital information. The findings demonstrate limits in existing metrics, with each technique highlighting strengths and flaws. The Jaccard and Cosine similarity emphasizes exact overlaps but ignores semantic subtleties, whereas Euclidean distance and Hamming distance struggle with variable-length texts.