<p>Sentence similarity detection provides significant advantages across different applications, such as customer support applications, e-commerce customer service, educational platforms, community opportunities and question-answering systems. This study presents a comparative analysis of various machine learning and deep learning models for sentence similarity detection, including cosine similarity, adaptive boosting (AdaBoost), Extreme Gradient Boosting (XGBoost), Convolutional Neural Network with Long Short-Term Memory, and Bidirectional Encoder Representations from Transformers with Long Short-Term Memory. This research also evaluates the impact of various vectorization techniques, such as Term Frequency-Inverse Document Frequency, OpenAI embeddings and Topic Modeling, on the performance of these models. The proposed research validates the effectiveness of these approaches in enhancing the accuracy of similarity detection. The findings offer valuable insights into an optimal combination of vectorization methods and models for improved sentence similarity detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative analysis of sentence similarity detection using machine and deep learning with vectorization techniques

  • Gayatri Girish Asalkar,
  • Bechoo Lal,
  • Nilesh B. Korade

摘要

Sentence similarity detection provides significant advantages across different applications, such as customer support applications, e-commerce customer service, educational platforms, community opportunities and question-answering systems. This study presents a comparative analysis of various machine learning and deep learning models for sentence similarity detection, including cosine similarity, adaptive boosting (AdaBoost), Extreme Gradient Boosting (XGBoost), Convolutional Neural Network with Long Short-Term Memory, and Bidirectional Encoder Representations from Transformers with Long Short-Term Memory. This research also evaluates the impact of various vectorization techniques, such as Term Frequency-Inverse Document Frequency, OpenAI embeddings and Topic Modeling, on the performance of these models. The proposed research validates the effectiveness of these approaches in enhancing the accuracy of similarity detection. The findings offer valuable insights into an optimal combination of vectorization methods and models for improved sentence similarity detection.