<p>Natural language generation has achieved remarkable performance on various tasks, including dialogue generation, summarization, translation, etc. Nevertheless, the repetition problem exists in nearly all the generated tasks mentioned above. Several methods based on the token-level probabilities have been proposed to solve the repetition problem. However, for the task of dialogue generation, as the continuation, i.e., the next utterance has strong coherence with the dialogue history, directly applying a repetition penalty may violate the coherence. Additionally, the response generated by the generative model has a high chance of repeating the dialogue history. To address these problems, we propose a novel repetition penalty approach with intraresponse n-gram repetition penalty (IRNRP) and interconversation repetition penalty (ICRP) for Chinese dialogue generation. The experiments on large-scale Chinese dialogue datasets have shown the effectiveness of our proposed approach.(Our code will be released at <a href="https://git.openi.org.cn/PCL-Platform.Intelligence/PanGu-Dialog">https://git.openi.org.cn/PCL-Platform.Intelligence/PanGu-Dialog</a>)</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Alleviating Chinese repetitive generation via intra and intersentence penalty

  • Jiacheng Yang,
  • Fangqing Jiang,
  • Hui Wang,
  • Hai-Tao Zheng,
  • Hong-Gee Kim

摘要

Natural language generation has achieved remarkable performance on various tasks, including dialogue generation, summarization, translation, etc. Nevertheless, the repetition problem exists in nearly all the generated tasks mentioned above. Several methods based on the token-level probabilities have been proposed to solve the repetition problem. However, for the task of dialogue generation, as the continuation, i.e., the next utterance has strong coherence with the dialogue history, directly applying a repetition penalty may violate the coherence. Additionally, the response generated by the generative model has a high chance of repeating the dialogue history. To address these problems, we propose a novel repetition penalty approach with intraresponse n-gram repetition penalty (IRNRP) and interconversation repetition penalty (ICRP) for Chinese dialogue generation. The experiments on large-scale Chinese dialogue datasets have shown the effectiveness of our proposed approach.(Our code will be released at https://git.openi.org.cn/PCL-Platform.Intelligence/PanGu-Dialog)