Multimodal Summarization with Modality-Aware Fusion and Summarization Ranking
摘要
Existing multimodal summarization methods primarily focus on multimodal fusion to efficiently utilize the visual information for summarization. However, they fail to exploit the deep interaction between textual and visual modality. Moreover, optimizing the model by maximum likelihood estimation (MLE) leads to exposure bias, causing the model to generate the next word that is based on the previously generated erroneous words during inference. To address these challenges, we propose a novel modality-aware fusion module (MAF) with a summarization ranking (SumR) training objective. Specifically, the MAF module exploits the interaction in the multimodal input through multiple fusion layers, and SumR aims to align the probability order predicted by the model with actual quality metrics, therefore it is able to reduce the exposure bias problem during inference. Extensive experiments on a large-scale dataset demonstrate that our method outperforms existing models, achieving superior results in both automatic and human evaluation metrics. The generated multimodal summaries provide richer context and enhance user’s comprehension by combining the key textual information with the relevant visual content.