This work introduces a novel framework for training Arabic nested embedding models through Matryoshka Embedding Learning, utilizing multilingual, Arabic-specific, and English-based models to demonstrate the power of nested embeddings in various Arabic NLP downstream tasks. Our contributions include the translation of key sentence similarity datasets into Arabic, enabling a robust evaluation framework to compare these models across multiple dimensions. We trained several nested embedding models on the Arabic Natural Language Inference triplet dataset, utilizing the power of anchor-positive-negative sampling to enhance semantic differentiation, and assessed their performance using various evaluation metrics including Pearson and Spearman correlations for cosine similarity, Manhattan distance, Euclidean distance, and dot product similarity. The results show that Matryoshka embedding models significantly outperform both base models and large language models like OpenAI’s text embedding models, with improvements of up to 20–25% across various similarity metrics, as evidenced by their performance on MTEB benchmarks such as STS17, STS22, and STS22-v2. These findings underscore the effectiveness of language-specific training and highlight the potential of Matryoshka models in advancing semantic textual similarity tasks for Arabic NLP.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Semantic Similarity Understanding in Arabic NLP with Nested Embedding Learning

  • Omer Nacar,
  • Anis Koubaa

摘要

This work introduces a novel framework for training Arabic nested embedding models through Matryoshka Embedding Learning, utilizing multilingual, Arabic-specific, and English-based models to demonstrate the power of nested embeddings in various Arabic NLP downstream tasks. Our contributions include the translation of key sentence similarity datasets into Arabic, enabling a robust evaluation framework to compare these models across multiple dimensions. We trained several nested embedding models on the Arabic Natural Language Inference triplet dataset, utilizing the power of anchor-positive-negative sampling to enhance semantic differentiation, and assessed their performance using various evaluation metrics including Pearson and Spearman correlations for cosine similarity, Manhattan distance, Euclidean distance, and dot product similarity. The results show that Matryoshka embedding models significantly outperform both base models and large language models like OpenAI’s text embedding models, with improvements of up to 20–25% across various similarity metrics, as evidenced by their performance on MTEB benchmarks such as STS17, STS22, and STS22-v2. These findings underscore the effectiveness of language-specific training and highlight the potential of Matryoshka models in advancing semantic textual similarity tasks for Arabic NLP.