<p>Current image fusion methods employing CNNs or Transformers face inherent limitations. CNN-based approaches lack the ability to effectively represent distant feature dependencies and global semantics. While Transformer-based methods, despite their ability to model long-range interactions, suffer from insufficient computational efficiency. To address these constraints, we propose MCLSC-Fusion, an innovative network for fusing infrared and visible images, which comprises two separate encoders, a single decoder and skip connections. In the encoders, we propose CNN-based Long-Short residual Feature extraction Blocks (LSFBs) that use short- and long-distance residual connections to refine features and improve representation. The decoder employs Lite Transformers, which reduce computational cost compared to standard Transformers, to produce a fused image that maintains information in both foreground and background regions. To connect the encoders and the decoder, we propose the multi-scale Skip-Connection Feature Fusion Blocks (SFFBs), which integrate the Enhanced Cross-modality Feature Extraction Transformer (ECFET). The ECFET is specifically designed to enhance cross-modal feature representations and eliminate redundant information through an improved Lite Transformer architecture. Moreover, we creatively incorporate Contrast Limited Adaptive Histogram Equalization (CLAHE) into the SSIM loss, which enhances contrast and salient object information. Extensive experimental results demonstrate that MCLSC-Fusion consistently outperforms existing fusion approaches, generalizes well across diverse modalities, and contributes positively to downstream object detection task.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCLSC-Fusion: a multi-scale cross-modality long-short connection fusion network for infrared and visible images

  • Hongyuan Lu,
  • Huiqian Du,
  • Min Xie

摘要

Current image fusion methods employing CNNs or Transformers face inherent limitations. CNN-based approaches lack the ability to effectively represent distant feature dependencies and global semantics. While Transformer-based methods, despite their ability to model long-range interactions, suffer from insufficient computational efficiency. To address these constraints, we propose MCLSC-Fusion, an innovative network for fusing infrared and visible images, which comprises two separate encoders, a single decoder and skip connections. In the encoders, we propose CNN-based Long-Short residual Feature extraction Blocks (LSFBs) that use short- and long-distance residual connections to refine features and improve representation. The decoder employs Lite Transformers, which reduce computational cost compared to standard Transformers, to produce a fused image that maintains information in both foreground and background regions. To connect the encoders and the decoder, we propose the multi-scale Skip-Connection Feature Fusion Blocks (SFFBs), which integrate the Enhanced Cross-modality Feature Extraction Transformer (ECFET). The ECFET is specifically designed to enhance cross-modal feature representations and eliminate redundant information through an improved Lite Transformer architecture. Moreover, we creatively incorporate Contrast Limited Adaptive Histogram Equalization (CLAHE) into the SSIM loss, which enhances contrast and salient object information. Extensive experimental results demonstrate that MCLSC-Fusion consistently outperforms existing fusion approaches, generalizes well across diverse modalities, and contributes positively to downstream object detection task.