Remote Sensing Visual Question Answering (RS VQA) seeks to generate precise and context-aware responses to questions about remote sensing (RS) imagery. Recent studies in remote sensing have harnessed the complementarity of multi-scale information, achieving notable advancements in tasks. However, in RS VQA, existing approaches often fail to fully exploit the complementary information embedded in multi-level visual features during feature extraction. To this end, we propose the Multi-Level Dynamic Gated Interaction Fusion Network (MDGIF). Specifically, our approach utilizes outputs from multiple stages of the Swin Transformer to extract visual features, leveraging the strengths of features at varying hierarchical levels. By integrating a dynamic gating mechanism with cross-level interaction and hierarchical aggregation, MDGIF effectively captures complementary information across feature levels while dynamically adjusting feature weights based on the question context. Extensive experiments demonstrate that MDGIF achieves state-of-the-art performance on the RSVQA-LR, RSVQA-HR, and RSIVQA datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-level Dynamic Gated Interaction Fusion Network for Remote Sensing Visual Question Answering

  • Ke Hu,
  • Wenzhen Zhang,
  • Shichao Zhang

摘要

Remote Sensing Visual Question Answering (RS VQA) seeks to generate precise and context-aware responses to questions about remote sensing (RS) imagery. Recent studies in remote sensing have harnessed the complementarity of multi-scale information, achieving notable advancements in tasks. However, in RS VQA, existing approaches often fail to fully exploit the complementary information embedded in multi-level visual features during feature extraction. To this end, we propose the Multi-Level Dynamic Gated Interaction Fusion Network (MDGIF). Specifically, our approach utilizes outputs from multiple stages of the Swin Transformer to extract visual features, leveraging the strengths of features at varying hierarchical levels. By integrating a dynamic gating mechanism with cross-level interaction and hierarchical aggregation, MDGIF effectively captures complementary information across feature levels while dynamically adjusting feature weights based on the question context. Extensive experiments demonstrate that MDGIF achieves state-of-the-art performance on the RSVQA-LR, RSVQA-HR, and RSIVQA datasets.