Hallucinations significantly impact the practical use of abstractive summarization models, a problem previously attributed to the model's inability to generate accurate words. This study identifies an overlooked factor: the model's reckless syntactic structure planning, which accounts for most hallucinated content. Our findings suggest that the model inadequately considers the document’s available information when selecting syntactic structures. Instead, it erroneously and directly replicates the co-occurrence relationship between document fragments and summary syntactic structures in the training data, which leads to the inclusion of hallucinated content. To address this issue, we propose a new method, named TSSP, that guides the model to plan more targeted syntactic structures. The experimental results demonstrate a 90.4% reduction in hallucinated content and a 17.6% improvement in factual consistency (as evaluated by FactCC) compared to the strongest baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TSSP: Reducing Hallucinations in Abstractive Summarization Through Targeted Syntactic Structure Planning

  • Dongsheng Chen,
  • Dingxin Hu,
  • Lei Li

摘要

Hallucinations significantly impact the practical use of abstractive summarization models, a problem previously attributed to the model's inability to generate accurate words. This study identifies an overlooked factor: the model's reckless syntactic structure planning, which accounts for most hallucinated content. Our findings suggest that the model inadequately considers the document’s available information when selecting syntactic structures. Instead, it erroneously and directly replicates the co-occurrence relationship between document fragments and summary syntactic structures in the training data, which leads to the inclusion of hallucinated content. To address this issue, we propose a new method, named TSSP, that guides the model to plan more targeted syntactic structures. The experimental results demonstrate a 90.4% reduction in hallucinated content and a 17.6% improvement in factual consistency (as evaluated by FactCC) compared to the strongest baselines.