<p>Multimodal Graph Neural Networks (GNNs) have emerged as a powerful paradigm for integrating heterogeneous Earth Observation (EO) data-such as optical and SAR imagery, in-situ sensor readings, geospatial vector layers, and socio-economic records-into unified relational representations. By jointly modeling complementary spatial and temporal modalities, multimodal GNNs enable more accurate, robust, and interpretable solutions for environmental monitoring and sustainable resource management. This article presents, to the best of the authors’ knowledge, the first comprehensive survey dedicated to multimodal GNN methods for EO applications, covering literature up to 2025. A detailed taxonomy of fusion strategies is proposed, including early, late, hybrid, cross-modal attention, and foundation-model-based approaches, and representative architectures are analyzed across agriculture, water resources, forestry, and urban infrastructure. Publicly available datasets, benchmark tasks, and evaluation metrics are systematically compared, with critical issues such as label scarcity, domain shift, scalability, and real-time deployment examined in depth. Special emphasis is placed on explainability, uncertainty quantification, and energy-efficient “green AI” practices. Finally, open research challenges are identified and a forward-looking roadmap is outlined toward standardized benchmarks, reproducible experimental protocols, and operational multimodal GNN systems capable of delivering actionable, policy-relevant EO insights at scale.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal graph neural networks for earth observation and sustainable resource management: a comprehensive review and research roadmap

  • Simran Kaur,
  • Harshit Sharma

摘要

Multimodal Graph Neural Networks (GNNs) have emerged as a powerful paradigm for integrating heterogeneous Earth Observation (EO) data-such as optical and SAR imagery, in-situ sensor readings, geospatial vector layers, and socio-economic records-into unified relational representations. By jointly modeling complementary spatial and temporal modalities, multimodal GNNs enable more accurate, robust, and interpretable solutions for environmental monitoring and sustainable resource management. This article presents, to the best of the authors’ knowledge, the first comprehensive survey dedicated to multimodal GNN methods for EO applications, covering literature up to 2025. A detailed taxonomy of fusion strategies is proposed, including early, late, hybrid, cross-modal attention, and foundation-model-based approaches, and representative architectures are analyzed across agriculture, water resources, forestry, and urban infrastructure. Publicly available datasets, benchmark tasks, and evaluation metrics are systematically compared, with critical issues such as label scarcity, domain shift, scalability, and real-time deployment examined in depth. Special emphasis is placed on explainability, uncertainty quantification, and energy-efficient “green AI” practices. Finally, open research challenges are identified and a forward-looking roadmap is outlined toward standardized benchmarks, reproducible experimental protocols, and operational multimodal GNN systems capable of delivering actionable, policy-relevant EO insights at scale.