Multi-level Dynamic Gated Interaction Fusion Network for Remote Sensing Visual Question Answering
摘要
Remote Sensing Visual Question Answering (RS VQA) seeks to generate precise and context-aware responses to questions about remote sensing (RS) imagery. Recent studies in remote sensing have harnessed the complementarity of multi-scale information, achieving notable advancements in tasks. However, in RS VQA, existing approaches often fail to fully exploit the complementary information embedded in multi-level visual features during feature extraction. To this end, we propose the Multi-Level Dynamic Gated Interaction Fusion Network (MDGIF). Specifically, our approach utilizes outputs from multiple stages of the Swin Transformer to extract visual features, leveraging the strengths of features at varying hierarchical levels. By integrating a dynamic gating mechanism with cross-level interaction and hierarchical aggregation, MDGIF effectively captures complementary information across feature levels while dynamically adjusting feature weights based on the question context. Extensive experiments demonstrate that MDGIF achieves state-of-the-art performance on the RSVQA-LR, RSVQA-HR, and RSIVQA datasets.