The primary objective of this paper is to provide a thorough examination and comparative analysis of various methodologies for computing sentence similarity within the domain of natural language processing (NLP). By exploring a wide range of approaches—string-based, syntax-based, graph-based, transformer-based, and prompt engineering-based methods—the paper aims to evaluate the effectiveness and limitations of each technique. Specific methods discussed include Word2Vec, WordNet, GUSUM, BERT, SGPT, and AnglE. Our findings highlight the broader implications of semantic similarity over string-based techniques, with a notable shift towards Language Model Models (LLMs) alongside transformer dominance. Future research suggestions include developing algorithms for unsupervised sentence embeddings, handling datasets exceeding sequence length limits, and exploring context regularization methods. Analyzing challenges in sentence-based similarity via graph-based approaches and experimenting with prompt engineering techniques are also recommended, providing a roadmap for advancing sentence similarity computation in NLP.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Methodologies for Computing Sentence Similarity in Natural Language Processing

  • Sagar Mondal,
  • Abirami Gurushanker,
  • Mirudhula Loganath,
  • Rishima Chowdhury,
  • Sankari Karthik,
  • Lekshmi Kalinathan,
  • Janaki Meena Murugan,
  • Marimuthu Marimuthu,
  • Saravanan Palani

摘要

The primary objective of this paper is to provide a thorough examination and comparative analysis of various methodologies for computing sentence similarity within the domain of natural language processing (NLP). By exploring a wide range of approaches—string-based, syntax-based, graph-based, transformer-based, and prompt engineering-based methods—the paper aims to evaluate the effectiveness and limitations of each technique. Specific methods discussed include Word2Vec, WordNet, GUSUM, BERT, SGPT, and AnglE. Our findings highlight the broader implications of semantic similarity over string-based techniques, with a notable shift towards Language Model Models (LLMs) alongside transformer dominance. Future research suggestions include developing algorithms for unsupervised sentence embeddings, handling datasets exceeding sequence length limits, and exploring context regularization methods. Analyzing challenges in sentence-based similarity via graph-based approaches and experimenting with prompt engineering techniques are also recommended, providing a roadmap for advancing sentence similarity computation in NLP.