DB-MFNet: A Dual-Branch Cross-Modal Fusion Network for High-Resolution Remote Sensing Semantic Segmentation
摘要
Combining land cover semantic segmentation with deep learning neural networks and applying it is a very important research direction. In the past few years, The multimodal fusion model has made great progress and attracted extensive attention. Most models achieve this by integrating visual Transformers with Convolutional Neural Networks (CNNs), This results in an increased overall computational load and potential loss of contextual data, particularly during feature fusion. The self-attention mechanism’s inherent deficiency in local receptive field constraints hinders precise localization of high-frequency spatial patterns. To further improve the effect of network in remote sensing segmentation, this paper proposes a dual-branch multi-scale fusion network (DB-MFNet). In feature-level fusion, a novel multi-scale attention module (MSAF) is used to construct a multi-scale feature fusion module (MFF), which extracts modality-specific features at different scales and combines them with channel-space mixed attention for dynamic weighted fusion. Furthermore, this paper introduces a spatial-channel mutual fusion aggregation module (SCMFA) and, through testing, develops an encoder (SCMF-Former) capable of spatial-channel mutual fusion of high-level features. Finally, a dynamic upsampling module (Dysample) is introduced into the decoder, forming a dynamic upsampling cascade decoder that adaptively focuses on important areas. Experimental results of DB-MFNet on the Potsdam and Vaihingen datasets indicate that this approach demonstrates good performance in remote sensing semantic segmentation.