<p>In the era of large language models, capturing fine-grained semantics remains critical, as these models often overlook subtle semantic nuances. Sememes, the smallest units of meaning, are essential for enriching semantic representations. However, existing sememe prediction methods rely solely on intrinsic word features or dictionary definitions, neglecting the potential of subword information to bridge the gap between them. This limitation results in poor performance in predicting out-of-vocabulary (OOV) and low-frequency words. To address this, we propose the Sememe Prediction through Semantic Synthesis&#xa0;(SPSY) framework, which integrates subword-level information with dictionary definitions. This approach enhances sensitivity to subtle semantic variations, significantly improving prediction accuracy. Evaluations on the HowNet and WordNet datasets show that our framework outperforms existing models, achieving a 2.91% gain in mean average precision for the Chinese dataset and a 5.54% gain for the English dataset. It also achieves state-of-the-art performance, surpassing previous models by at least 2.88% across all word frequencies and by 4.11% for OOV words. Furthermore, the framework demonstrates its versatility through successful applications in industrial knowledge graph verification and entity recognition.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SPSY: a semantic synthesis framework for lexical sememe prediction and its applications

  • Tao Wen,
  • Jianpeng Hu,
  • Jin Zhao,
  • Xiaolong Gong,
  • Shuqun Yang

摘要

In the era of large language models, capturing fine-grained semantics remains critical, as these models often overlook subtle semantic nuances. Sememes, the smallest units of meaning, are essential for enriching semantic representations. However, existing sememe prediction methods rely solely on intrinsic word features or dictionary definitions, neglecting the potential of subword information to bridge the gap between them. This limitation results in poor performance in predicting out-of-vocabulary (OOV) and low-frequency words. To address this, we propose the Sememe Prediction through Semantic Synthesis (SPSY) framework, which integrates subword-level information with dictionary definitions. This approach enhances sensitivity to subtle semantic variations, significantly improving prediction accuracy. Evaluations on the HowNet and WordNet datasets show that our framework outperforms existing models, achieving a 2.91% gain in mean average precision for the Chinese dataset and a 5.54% gain for the English dataset. It also achieves state-of-the-art performance, surpassing previous models by at least 2.88% across all word frequencies and by 4.11% for OOV words. Furthermore, the framework demonstrates its versatility through successful applications in industrial knowledge graph verification and entity recognition.