<p>Multimodal image fusion (MMIF) generates high-quality images with rich details for downstream tasks by integrating information contained in different modalities. Models based on local Convolutional Neural Networks (CNN) struggle to capture global information, while Transformer-based methods enable global modeling but suffer from high computational complexity. Mamba processes long-range dependencies with linear complexity, effectively addressing these limitations. In this paper, we propose STMFusion, a novel Mamba-based semantic-guided fusion network. This framework adopts a Cross-Modulated Fusion Block (CMFB), which leverages semantic information to guide the fusion process and integrates cross-modal features through a shared state modulation strategy. It not only achieves adaptive dynamic feature alignment but also significantly suppresses redundant features. Furthermore, we develop a new module named a Texture-Enhanced Visual State Space (TE-VSS). It integrates the Dynamic Gradient Operator (DGO) to effectively solve the problem of high-frequency detail loss existing in state-space models. Experiments demonstrate that STMFusion achieves competitive or superior performance in multimodal image fusion and downstream tasks, showing wide applicability and superiority.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

STMFusion: a semantic-guided and texture-enhanced Mamba network for multimodal image fusion

  • Huan Gao,
  • Weihua Su,
  • Dongyuan Zang,
  • Lehao Yang,
  • Haoyu Guan,
  • Zhe Li,
  • Zunlun Peng

摘要

Multimodal image fusion (MMIF) generates high-quality images with rich details for downstream tasks by integrating information contained in different modalities. Models based on local Convolutional Neural Networks (CNN) struggle to capture global information, while Transformer-based methods enable global modeling but suffer from high computational complexity. Mamba processes long-range dependencies with linear complexity, effectively addressing these limitations. In this paper, we propose STMFusion, a novel Mamba-based semantic-guided fusion network. This framework adopts a Cross-Modulated Fusion Block (CMFB), which leverages semantic information to guide the fusion process and integrates cross-modal features through a shared state modulation strategy. It not only achieves adaptive dynamic feature alignment but also significantly suppresses redundant features. Furthermore, we develop a new module named a Texture-Enhanced Visual State Space (TE-VSS). It integrates the Dynamic Gradient Operator (DGO) to effectively solve the problem of high-frequency detail loss existing in state-space models. Experiments demonstrate that STMFusion achieves competitive or superior performance in multimodal image fusion and downstream tasks, showing wide applicability and superiority.