Crossmodal Hierarchical Heterogeneous Graph Guided Mamba Network for remote sensing semantic segmentation
摘要
Driven by the rapid development of Earth observation sensor technologies, the Digital Surface Model (DSM), with its unique elevation features, has become a key auxiliary modality for enhancing the semantic segmentation performance of optical remote sensing images. Most existing cross-modal fusion studies tend to aggregate all modalities indiscriminately, overlooking the modal heterogeneity between the geometric and topological features represented by DSM and the spectral response features of high-resolution optical images in the representation space. Additionally, deep semantic interactions are constrained by the secondary computational complexity of Transformer architectures, making it difficult to establish fine-grained long-range dependencies in high-resolution scenarios. In this work, we propose a Hierarchical Heterogeneous Graph-guided Mamba Network (HGGMNet). HGGMNet utilizes a novel Cross-modal Heterogeneous Semantic Alignment Module (HSAM), which builds a learnable heterogeneous feature similarity metric through heterogeneous graph convolutions, alleviating feature heterogeneity across modalities. For deep semantic fusion, we innovatively construct the Cross-modal Joint Mamba Module (CMJMamba). Using the proposed cross-modal joint scanning mechanism, it effectively captures higher-level fused features from both modalities. Extensive experiments on two large-scale high-resolution remote sensing datasets demonstrate that the proposed HGGMNet outperforms other mainstream methods.