<p>To address the challenges in infrared-visible light and polarization-visible light image fusion tasks-such as insufficient feature extraction, difficulty in simultaneously capturing global dependencies and local spatial information, and the lack of multi-task adaptability in existing models optimized for single tasks-this paper proposes an end-to-end solution named LGMFuse. The core innovation of LGMFuse lies in its Local and Global Mamba (LGM) module, which significantly enhances multi-directional perception through an eight-directional scanning mechanism. The LGM module comprises two parallel branches: global four-directional scanning and local multi-scale four-directional scanning, designed to extract spatial local features and capture global dependencies, respectively. In the encoding stage, LGMFuse employs a three-stage feature extraction architecture, progressively extracting multi-scale multimodal features via the Local and Global Mamba Encode Block (LGME), while acquiring higher-level semantic information as the network depth increases. In the fusion stage, the Local and Global Mamba Fusion Block (LGMF) focuses on the deep fusion of feature maps from different modalities at the same scale, preserving complementary characteristics while minimizing redundancy and noise to ensure the precision of the fused features. Experimental results demonstrate that LGMFuse achieves state-of-the-art performance in multimodal fusion accuracy and object detection across three public datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LGMFuse: A Multi-modal Image Fusion Method Based on Local and Global Mamba

  • Xuedong He,
  • Jin Duan,
  • Meiling Gao,
  • Yong Zhu,
  • Xiaoyu Jin,
  • Jialin Wang

摘要

To address the challenges in infrared-visible light and polarization-visible light image fusion tasks-such as insufficient feature extraction, difficulty in simultaneously capturing global dependencies and local spatial information, and the lack of multi-task adaptability in existing models optimized for single tasks-this paper proposes an end-to-end solution named LGMFuse. The core innovation of LGMFuse lies in its Local and Global Mamba (LGM) module, which significantly enhances multi-directional perception through an eight-directional scanning mechanism. The LGM module comprises two parallel branches: global four-directional scanning and local multi-scale four-directional scanning, designed to extract spatial local features and capture global dependencies, respectively. In the encoding stage, LGMFuse employs a three-stage feature extraction architecture, progressively extracting multi-scale multimodal features via the Local and Global Mamba Encode Block (LGME), while acquiring higher-level semantic information as the network depth increases. In the fusion stage, the Local and Global Mamba Fusion Block (LGMF) focuses on the deep fusion of feature maps from different modalities at the same scale, preserving complementary characteristics while minimizing redundancy and noise to ensure the precision of the fused features. Experimental results demonstrate that LGMFuse achieves state-of-the-art performance in multimodal fusion accuracy and object detection across three public datasets.