<p>Multimodal abstractive summarization integrates information from diverse modalities, such as text and images, to generate concise, coherent summaries. Despite its advancements, extrinsic hallucination, a phenomenon where generated summaries include content not present in the source, remains a critical challenge, leading to inaccuracies and reduced reliability. Existing techniques largely focus on intrinsic hallucinations and often require substantial architectural changes, leaving extrinsic hallucination inadequately addressed. This paper proposes a novel post-processing approach to mitigate extrinsic hallucination by leveraging external knowledge vocabularies as a corrective mechanism. The framework identifies hallucinated tokens in generated summaries using a cosine similarity metric and replaces them with factually consistent tokens sourced from external knowledge, ensuring improved coherence and faithfulness. By operating at the token level, the approach preserves linguistic, semantic, and syntactic structures, effectively reducing unfaithful content. The proposed method incorporates domain knowledge through a Word2Vec-based vocabulary, offering a scalable solution to enhance factual consistency without modifying model architectures. The contributions of this study include a detailed methodology for identifying and correcting hallucinated tokens, a robust post-processing pipeline for refining summaries, and a demonstration of the method’s effectiveness in reducing extrinsic hallucinations. Experimental results on MSMO dataset highlight the approach’s potential to improve the accuracy, coherence, and reliability of multimodal abstractive summarization systems, addressing a significant gap in the field and shows the state-of-the-art results with RMS-Prop Optimizer.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reducing extrinsic hallucination in multimodal abstractive summaries with post-processing technique

  • Shaik Rafi,
  • Lenin Laitonjam,
  • Ranjita Das

摘要

Multimodal abstractive summarization integrates information from diverse modalities, such as text and images, to generate concise, coherent summaries. Despite its advancements, extrinsic hallucination, a phenomenon where generated summaries include content not present in the source, remains a critical challenge, leading to inaccuracies and reduced reliability. Existing techniques largely focus on intrinsic hallucinations and often require substantial architectural changes, leaving extrinsic hallucination inadequately addressed. This paper proposes a novel post-processing approach to mitigate extrinsic hallucination by leveraging external knowledge vocabularies as a corrective mechanism. The framework identifies hallucinated tokens in generated summaries using a cosine similarity metric and replaces them with factually consistent tokens sourced from external knowledge, ensuring improved coherence and faithfulness. By operating at the token level, the approach preserves linguistic, semantic, and syntactic structures, effectively reducing unfaithful content. The proposed method incorporates domain knowledge through a Word2Vec-based vocabulary, offering a scalable solution to enhance factual consistency without modifying model architectures. The contributions of this study include a detailed methodology for identifying and correcting hallucinated tokens, a robust post-processing pipeline for refining summaries, and a demonstration of the method’s effectiveness in reducing extrinsic hallucinations. Experimental results on MSMO dataset highlight the approach’s potential to improve the accuracy, coherence, and reliability of multimodal abstractive summarization systems, addressing a significant gap in the field and shows the state-of-the-art results with RMS-Prop Optimizer.