Comparative Analysis and Evaluation of the Consistency of Methods for Integrating Semantic Representations of Texts
摘要
Abstract
The paper explores methods for aggregating the results of three embedding models to estimate semantic similarity of texts: cosine similarities averaging, vector concatenation, and selection of one of the three cosine similarities based on principal component analysis and singular value decomposition methods. Statistical analysis showed that averaging cosine proximity as an aggregation method shows the best results. The obtained results can be used to improve the robustness of models for semantic proximity evaluation of texts when using embedding ensembles.