<p>Image captioning is a multimodal task that involves both computer vision and natural language processing. In recent years, to address the issue of insufficient visual information, multiple features are often used. However, existing image description methods based on the encoder–decoder framework fail to achieve effective interaction and fusion of different modality features. In this paper, we propose a <i>Bidirectional Multimodal Fusion Network</i> (<b>BMFNet</b>) for image captioning, which provides deep interaction and fusion of multiple features throughout the encoding and decoding process. BMFNet first enhances grid features and region features at the channel level, then incorporates an effective bidirectional channel-level multimodal feature fusion module (CMFF) in the encoder. This module effectively compensates for differences between features and integrates their respective advantages. In the decoder, we introduce a cross-attention mechanism that adopts a dual-path information flow architecture for image caption generation. This design enables multi-stage interaction and progressive fusion between visual and semantic features, which significantly enhances the model’s cross-modal reasoning capability. Qualitative and quantitative experiments conducted on the MSCOCO dataset show that, compared to traditional methods using dual features, our model improves the CIDEr score by 2.8% on the Karpathy test split, achieving a score of 136.6 CIDEr.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BMFNet: Bidirectional Multimodal Fusion Network for image captioning

  • Lixia Xue,
  • ZiQian Jin,
  • Ronggui Wang,
  • Juan Yang

摘要

Image captioning is a multimodal task that involves both computer vision and natural language processing. In recent years, to address the issue of insufficient visual information, multiple features are often used. However, existing image description methods based on the encoder–decoder framework fail to achieve effective interaction and fusion of different modality features. In this paper, we propose a Bidirectional Multimodal Fusion Network (BMFNet) for image captioning, which provides deep interaction and fusion of multiple features throughout the encoding and decoding process. BMFNet first enhances grid features and region features at the channel level, then incorporates an effective bidirectional channel-level multimodal feature fusion module (CMFF) in the encoder. This module effectively compensates for differences between features and integrates their respective advantages. In the decoder, we introduce a cross-attention mechanism that adopts a dual-path information flow architecture for image caption generation. This design enables multi-stage interaction and progressive fusion between visual and semantic features, which significantly enhances the model’s cross-modal reasoning capability. Qualitative and quantitative experiments conducted on the MSCOCO dataset show that, compared to traditional methods using dual features, our model improves the CIDEr score by 2.8% on the Karpathy test split, achieving a score of 136.6 CIDEr.