<p>Driven by the rapid development of Earth observation sensor technologies, the Digital Surface Model (DSM), with its unique elevation features, has become a key auxiliary modality for enhancing the semantic segmentation performance of optical remote sensing images. Most existing cross-modal fusion studies tend to aggregate all modalities indiscriminately, overlooking the modal heterogeneity between the geometric and topological features represented by DSM and the spectral response features of high-resolution optical images in the representation space. Additionally, deep semantic interactions are constrained by the secondary computational complexity of Transformer architectures, making it difficult to establish fine-grained long-range dependencies in high-resolution scenarios. In this work, we propose a Hierarchical Heterogeneous Graph-guided Mamba Network (HGGMNet). HGGMNet utilizes a novel Cross-modal Heterogeneous Semantic Alignment Module (HSAM), which builds a learnable heterogeneous feature similarity metric through heterogeneous graph convolutions, alleviating feature heterogeneity across modalities. For deep semantic fusion, we innovatively construct the Cross-modal Joint Mamba Module (CMJMamba). Using the proposed cross-modal joint scanning mechanism, it effectively captures higher-level fused features from both modalities. Extensive experiments on two large-scale high-resolution remote sensing datasets demonstrate that the proposed HGGMNet outperforms other mainstream methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Crossmodal Hierarchical Heterogeneous Graph Guided Mamba Network for remote sensing semantic segmentation

  • Rong Gao,
  • Qingyang Feng,
  • Jinshan Cao,
  • Xiaoxiao Feng,
  • Jing Wang,
  • Lingyu Yan

摘要

Driven by the rapid development of Earth observation sensor technologies, the Digital Surface Model (DSM), with its unique elevation features, has become a key auxiliary modality for enhancing the semantic segmentation performance of optical remote sensing images. Most existing cross-modal fusion studies tend to aggregate all modalities indiscriminately, overlooking the modal heterogeneity between the geometric and topological features represented by DSM and the spectral response features of high-resolution optical images in the representation space. Additionally, deep semantic interactions are constrained by the secondary computational complexity of Transformer architectures, making it difficult to establish fine-grained long-range dependencies in high-resolution scenarios. In this work, we propose a Hierarchical Heterogeneous Graph-guided Mamba Network (HGGMNet). HGGMNet utilizes a novel Cross-modal Heterogeneous Semantic Alignment Module (HSAM), which builds a learnable heterogeneous feature similarity metric through heterogeneous graph convolutions, alleviating feature heterogeneity across modalities. For deep semantic fusion, we innovatively construct the Cross-modal Joint Mamba Module (CMJMamba). Using the proposed cross-modal joint scanning mechanism, it effectively captures higher-level fused features from both modalities. Extensive experiments on two large-scale high-resolution remote sensing datasets demonstrate that the proposed HGGMNet outperforms other mainstream methods.