The cross-lingual topic-essay generation task (CTEG) aims to generate sentence-level text in a target language based on input topic words from a source language. Recently, research on the generation of essays from topic words has primarily focused on monolingual settings. Extending this to cross-language scenarios requires overcoming challenges related to cross-language alignment and mitigating topic drift in the generated text. To address these challenges, we propose a novel cross-lingual topic-essay generation approach based on knowledge and topic consistency constraints. This approach extracts semantic alignment knowledge from a source language essay to a target language essay using a translation teacher model, which builds cross-lingual semantic alignment and guides the generation in the student model. Additionally, a cosine similarity-based topic consistency loss enhances the generated essays’ topic consistency relative to the input topic words. To validate the effectiveness of the proposed model, we constructed a dataset of 160,000 Chinese topic-Vietnamese essay pairs and a dataset of 350,000 Chinese topic-English essay pairs. Experimental results show that our model outperforms various baseline models in terms of various evaluation metrics on both the Chinese-Vietnamese and Chinese-English datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Cross-Lingual Topic-Essay Generation with Knowledge and Topic Consistency Constraints

  • Huailing Gu,
  • Yuxin Huang,
  • Zhengtao Yu,
  • Cunli Mao

摘要

The cross-lingual topic-essay generation task (CTEG) aims to generate sentence-level text in a target language based on input topic words from a source language. Recently, research on the generation of essays from topic words has primarily focused on monolingual settings. Extending this to cross-language scenarios requires overcoming challenges related to cross-language alignment and mitigating topic drift in the generated text. To address these challenges, we propose a novel cross-lingual topic-essay generation approach based on knowledge and topic consistency constraints. This approach extracts semantic alignment knowledge from a source language essay to a target language essay using a translation teacher model, which builds cross-lingual semantic alignment and guides the generation in the student model. Additionally, a cosine similarity-based topic consistency loss enhances the generated essays’ topic consistency relative to the input topic words. To validate the effectiveness of the proposed model, we constructed a dataset of 160,000 Chinese topic-Vietnamese essay pairs and a dataset of 350,000 Chinese topic-English essay pairs. Experimental results show that our model outperforms various baseline models in terms of various evaluation metrics on both the Chinese-Vietnamese and Chinese-English datasets.