A heterogeneous image fusion network model based on lightweight Swin+ self-learning fusion strategy
摘要
Image fusion aims to integrate complementary information from multi-source images to generate a single image with enhanced interpretability. To address challenges such as feature redundancy, high computational cost, and low efficiency, this paper proposes a heterogeneous image fusion network named Swin+, incorporating a multi-branch feature extraction backbone, a class activation map (CAM)-based self-learning fusion strategy, and a lightweight decoder. The model combines window-based spatial self-attention with cross-channel attention to capture both spatial dependencies and channel relationships, enabling more informative feature representation. A multi-scale integration scheme ensures structural consistency and detail preservation across layers. Furthermore, an improved Monarch Butterfly Optimization (MBO) algorithm is used to reduce redundant parameters and search for optimal lightweight configurations. Compared with conventional Swin backbones, our method reduces FLOPs and parameter size significantly while preserving accuracy. Extensive experiments on multiple datasets demonstrate that Swin + outperforms several state-of-the-art fusion methods in both objective metrics (MI, SCD, MS-SSIM) and subjective evaluations. The visual results also include zoom-in regions to highlight texture fidelity and contrast enhancement. The proposed model provides a practical balance between fusion quality and computational efficiency, making it highly applicable for real-time or resource-constrained deployment.
Graphical abstract