MDFNet: Multimodal Remote Sensing Image Segmentation Method Based on Multilevel Discrepancy Feature Fusion
摘要
In recent years, the semantic segmentation of remote sensing data through multimodal fusion has garnered significant research interest. Nevertheless, most of these methods extract public information from source images for fusion operations, resulting in inadequate ability to consider discrepancy information. In this work, we propose a multilevel discrepancy fusion network, termed MDFNet, which efficiently integrates Convolutional Neural Network (CNN) and Vision Transformer (ViT) into a unified framework for multimodal semantic segmentation. Firstly, the interactive feature calibration module of the channel and space provides an overall calibration, which solves different noise and uncertainties in different modalities to achieve better multi-modal feature extraction and interaction. Secondly, by redesigning the cross-attention mechanism, a novel ViT fusion network, the cross-modal feature fusion module is constructed to simultaneously extract discrepancy and common information from the two modal images for fusion. Finally, a deep fusion module is designed by alternately integrating Self-Attention (SA) and Cross-Modal Feature Fusion (CMFF) layers to capture cross-modal features with enhanced inter-class discrimination and reduced intra-class variation.