Most topic detection research focuses on cross-language information processing in languages with abundant resources, such as English-Chinese, English-German, and others. There are also some studies targeting low-resource languages such as Vietnamese-Chinese and Tibetan-Chinese cross-language topic detection research. However, there is very little research on Mongolian-Chinese cross-language studies. The primary reasons are the lack of corpora in the Mongolian language and the limitations of conventional topic detection methods in text representation, clustering, and topic representation within the Mongolian language domain. Therefore, this paper proposes a Mongolian-Chinese cross-lingual topic detection method based on knowledge distillation and contrastive learning. Firstly, the pre-trained language model is fine-tuned through knowledge distillation tasks to enable the model to integrate the semantic representation of Mongolian-Chinese cross-lingual texts. Then, the effect of cross-lingual topic detection is improved by combining contrastive learning training between news of different topics. Experimental results show that the proposed method combining knowledge distillation and contrastive learning outperforms other baseline models, achieving at least 2, 0.8, and 3% points higher in F1 score, topic diversity, and topic consistency, respectively. This approach effectively enhances the accuracy of Mongolian-Chinese cross-lingual topic detection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mongolian-Chinese Cross-Lingual Topic Detection Based on Knowledge Distillation and Contrastive Learning Methods

  • Yanli Wang,
  • Yatu Ji,
  • Baolei Sun,
  • Nier Wu,
  • Qing-Dao-Er-Ji Ren,
  • Bailun Wang

摘要

Most topic detection research focuses on cross-language information processing in languages with abundant resources, such as English-Chinese, English-German, and others. There are also some studies targeting low-resource languages such as Vietnamese-Chinese and Tibetan-Chinese cross-language topic detection research. However, there is very little research on Mongolian-Chinese cross-language studies. The primary reasons are the lack of corpora in the Mongolian language and the limitations of conventional topic detection methods in text representation, clustering, and topic representation within the Mongolian language domain. Therefore, this paper proposes a Mongolian-Chinese cross-lingual topic detection method based on knowledge distillation and contrastive learning. Firstly, the pre-trained language model is fine-tuned through knowledge distillation tasks to enable the model to integrate the semantic representation of Mongolian-Chinese cross-lingual texts. Then, the effect of cross-lingual topic detection is improved by combining contrastive learning training between news of different topics. Experimental results show that the proposed method combining knowledge distillation and contrastive learning outperforms other baseline models, achieving at least 2, 0.8, and 3% points higher in F1 score, topic diversity, and topic consistency, respectively. This approach effectively enhances the accuracy of Mongolian-Chinese cross-lingual topic detection.