Remote Sensing Visual Question Answering (RSVQA) aims to generate accurate answers to questions about remote sensing images. The inherent complexity of RS data, featuring rich local details and global contextual information, poses challenges in effectively modeling the deep interactions between these features. Additionally, significant cross-modal differences between remote sensing imagery and textual descriptions make it difficult for existing methods to capture deep semantic correlations across modalities. To address these issues, we propose the Multimodal Co-Attention Fusion Network (MCAF-Net), which incorporates a novel Dual-Stage Collaborative Attention Mechanism (DSCA) to enhance both intra-modal and cross-modal interactions. In the first stage, DSCA dynamically integrates local features extracted by a Convolutional Neural Network (CNN) with global features captured by a Vision Transformer (ViT), ensuring complementary feature synergy. In the second stage, DSCA performs deep fusion of the aggregated visual features with textual features extracted by BERT, capturing complex cross-modal dependencies. Extensive experiments on three public RSVQA datasets demonstrate that MCAF-Net outperforms state-of-the-art methods across all benchmarks, validating its effectiveness and robustness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Collaborative Attention Fusion Network for Remote Sensing Visual Question Answering

  • Ke Hu,
  • Wenzhen Zhang,
  • Shichao Zhang

摘要

Remote Sensing Visual Question Answering (RSVQA) aims to generate accurate answers to questions about remote sensing images. The inherent complexity of RS data, featuring rich local details and global contextual information, poses challenges in effectively modeling the deep interactions between these features. Additionally, significant cross-modal differences between remote sensing imagery and textual descriptions make it difficult for existing methods to capture deep semantic correlations across modalities. To address these issues, we propose the Multimodal Co-Attention Fusion Network (MCAF-Net), which incorporates a novel Dual-Stage Collaborative Attention Mechanism (DSCA) to enhance both intra-modal and cross-modal interactions. In the first stage, DSCA dynamically integrates local features extracted by a Convolutional Neural Network (CNN) with global features captured by a Vision Transformer (ViT), ensuring complementary feature synergy. In the second stage, DSCA performs deep fusion of the aggregated visual features with textual features extracted by BERT, capturing complex cross-modal dependencies. Extensive experiments on three public RSVQA datasets demonstrate that MCAF-Net outperforms state-of-the-art methods across all benchmarks, validating its effectiveness and robustness.