Semantic search is an evolution in both accuracy and flexibility of the search engine. Unlike traditional keyword-based searches, it delves deeper by comprehending the semantics behind words and user queries, resulting in more accurate and relevant search outcomes. In this research, we conduct experiments on how Mobifone News data differs from general model training data and how these differences affect the model feature extraction output. Furthermore, we conduct investigations of the method to finetune a feature extraction BERT on processed data and the use of Multiple Negative Ranking Loss when the data for the Semantic Textual Similarity training task is not in the ideal format of Premise-Hypothesis-Label. Furthermore, we address the model performance on different pooling methods. Finally, we evaluate the model inference performance on a complete pipeline to ensure compelling business requirements. The results show that when data is not ideal, a semantic search model based on Transformers still achieves great retrieval rates on a human-evaluated keywords dataset. We also succeeded in creating a highly accurate model while compelling to the required speed on a low-end system with no GPU.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Semantic Search Through Domain Adaptation: A Case Study on the Mobifone News Dataset

  • Thach Duc Long,
  • Nguyen Minh Hieu,
  • Nguyen Trieu Ngoc Huyen,
  • Phan Duy Hung

摘要

Semantic search is an evolution in both accuracy and flexibility of the search engine. Unlike traditional keyword-based searches, it delves deeper by comprehending the semantics behind words and user queries, resulting in more accurate and relevant search outcomes. In this research, we conduct experiments on how Mobifone News data differs from general model training data and how these differences affect the model feature extraction output. Furthermore, we conduct investigations of the method to finetune a feature extraction BERT on processed data and the use of Multiple Negative Ranking Loss when the data for the Semantic Textual Similarity training task is not in the ideal format of Premise-Hypothesis-Label. Furthermore, we address the model performance on different pooling methods. Finally, we evaluate the model inference performance on a complete pipeline to ensure compelling business requirements. The results show that when data is not ideal, a semantic search model based on Transformers still achieves great retrieval rates on a human-evaluated keywords dataset. We also succeeded in creating a highly accurate model while compelling to the required speed on a low-end system with no GPU.