Existing multimodal summarization methods primarily focus on multimodal fusion to efficiently utilize the visual information for summarization. However, they fail to exploit the deep interaction between textual and visual modality. Moreover, optimizing the model by maximum likelihood estimation (MLE) leads to exposure bias, causing the model to generate the next word that is based on the previously generated erroneous words during inference. To address these challenges, we propose a novel modality-aware fusion module (MAF) with a summarization ranking (SumR) training objective. Specifically, the MAF module exploits the interaction in the multimodal input through multiple fusion layers, and SumR aims to align the probability order predicted by the model with actual quality metrics, therefore it is able to reduce the exposure bias problem during inference. Extensive experiments on a large-scale dataset demonstrate that our method outperforms existing models, achieving superior results in both automatic and human evaluation metrics. The generated multimodal summaries provide richer context and enhance user’s comprehension by combining the key textual information with the relevant visual content.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Summarization with Modality-Aware Fusion and Summarization Ranking

  • Xuming Ye,
  • Chaomurilige,
  • Zheng Liu,
  • Haoyu Luo,
  • Jun Dong,
  • Yingzhe Luo

摘要

Existing multimodal summarization methods primarily focus on multimodal fusion to efficiently utilize the visual information for summarization. However, they fail to exploit the deep interaction between textual and visual modality. Moreover, optimizing the model by maximum likelihood estimation (MLE) leads to exposure bias, causing the model to generate the next word that is based on the previously generated erroneous words during inference. To address these challenges, we propose a novel modality-aware fusion module (MAF) with a summarization ranking (SumR) training objective. Specifically, the MAF module exploits the interaction in the multimodal input through multiple fusion layers, and SumR aims to align the probability order predicted by the model with actual quality metrics, therefore it is able to reduce the exposure bias problem during inference. Extensive experiments on a large-scale dataset demonstrate that our method outperforms existing models, achieving superior results in both automatic and human evaluation metrics. The generated multimodal summaries provide richer context and enhance user’s comprehension by combining the key textual information with the relevant visual content.