<p>To address the issues of detail blurring and artifact interference in low-resolution infrared images during multimodal fusion tasks, an end-to-end super-resolution fusion collaborative model (DRA-Net) is proposed. This model jointly optimizes super-resolution reconstruction and image fusion tasks by integrating a multi-scale feature extraction block dominated by residual dense connections and a multi-scale attention aggregation module, thereby achieving hierarchical extraction and dynamic weighted fusion of cross-modal features. The method enhances the efficiency of shallow detail propagation to deeper layers through the cascaded structure of residual dense blocks while extracting rich semantic information at multiple scales. Additionally, the attention mechanism adaptively integrates infrared thermal radiation features and visible light texture information from both spatial and channel dimensions at multiple scales. Experiments conducted on the M3FD dataset demonstrate that this method outperforms mainstream cascaded models in terms of visual fidelity, edge detail preservation, and structural similarity, achieving an average improvement of 16.38% compared to the second-best model. Furthermore, the inference speed of this method is significantly faster than conventional solutions, with a speedup of 210.40% compared to the second-fastest model. This provides an effective solution for high-quality fusion of low-resolution multimodal images.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A super-resolution fusion model for infrared-visible light images based on multi-scale features and attention networks

  • Chengkai Zhu,
  • Bo Peng,
  • Gaoqiang Wang,
  • Gang Xiao

摘要

To address the issues of detail blurring and artifact interference in low-resolution infrared images during multimodal fusion tasks, an end-to-end super-resolution fusion collaborative model (DRA-Net) is proposed. This model jointly optimizes super-resolution reconstruction and image fusion tasks by integrating a multi-scale feature extraction block dominated by residual dense connections and a multi-scale attention aggregation module, thereby achieving hierarchical extraction and dynamic weighted fusion of cross-modal features. The method enhances the efficiency of shallow detail propagation to deeper layers through the cascaded structure of residual dense blocks while extracting rich semantic information at multiple scales. Additionally, the attention mechanism adaptively integrates infrared thermal radiation features and visible light texture information from both spatial and channel dimensions at multiple scales. Experiments conducted on the M3FD dataset demonstrate that this method outperforms mainstream cascaded models in terms of visual fidelity, edge detail preservation, and structural similarity, achieving an average improvement of 16.38% compared to the second-best model. Furthermore, the inference speed of this method is significantly faster than conventional solutions, with a speedup of 210.40% compared to the second-fastest model. This provides an effective solution for high-quality fusion of low-resolution multimodal images.