Abstract <p>The paper explores methods for aggregating the results of three embedding models to estimate semantic similarity of texts: cosine similarities averaging, vector concatenation, and selection of one of the three cosine similarities based on principal component analysis and singular value decomposition methods. Statistical analysis showed that averaging cosine proximity as an aggregation method shows the best results. The obtained results can be used to improve the robustness of models for semantic proximity evaluation of texts when using embedding ensembles.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis and Evaluation of the Consistency of Methods for Integrating Semantic Representations of Texts

  • A. V. Ilina,
  • P. V. Zrelov

摘要

Abstract

The paper explores methods for aggregating the results of three embedding models to estimate semantic similarity of texts: cosine similarities averaging, vector concatenation, and selection of one of the three cosine similarities based on principal component analysis and singular value decomposition methods. Statistical analysis showed that averaging cosine proximity as an aggregation method shows the best results. The obtained results can be used to improve the robustness of models for semantic proximity evaluation of texts when using embedding ensembles.