MARCFusion: adaptive residual cross-domain fusion network for medical image fusion
摘要
The purpose of multimodal medical image fusion (MMIF) is to obtain a comprehensive fused image by merging complementary information from medical images, which facilitates clinical diagnosis and medical research. However, deep learning-based fusion methods always neglect to fully utilize the frequency domain features, which may lead to insufficient retention of texture details and high-frequency information in the fusion results. To address the mentioned issue, in this paper, we propose an adaptive residual cross-domain fusion network for medical image fusion named MARCFusion. In the proposed method, the context feature extraction module (CFEM) consisting of context feature selection modules (CFSM) and convolutional attention residual modules (CARM) is devised to capture multimodal deep features from different scales. In addition, a residual cross-domain fusion module (RCFM) consisting of a residual spatial-domain fusion block (RSFB) and a residual frequency-domain fusion block (RFFB) is elaborated to further communicate and integrate multiscale deep features, where RSFB is constructed to extract spatial domain context features and RFFB is designed to capture frequency domain complementary features. Finally, a decoder with a nest connection architecture is constructed to reconstruct the fused image. Furthermore, a joint loss function consisting of detail loss and saliency loss is defined to train MARCFusion in an unsupervised manner, which can preserve richer information in the fused image. Extensive experimental results on mainstream datasets show that MARCFusion surpasses some state-of-the-art methodologies in terms of visual observation and objective evaluation.