Semantic-Guided Multi-attention Model for Infrared and Visible Image Fusion: A Deep Learning Approach
摘要
The goal of visible and infrared image fusion is to create a single, higher-quality picture by combining complimentary information from both kinds of images. However, many existing image fusion techniques overlook high-level semantic information within the images and fail to extract global features from the source images. To address these issues, this paper proposes a network based on semantic-guided multi-attention, termed SGMAFusion. This network leverages a transformer module (Tran) and the VGG-19 network during the encoder stage to extract global and multi-scale features from the source images and employs a multi-level contextual attention mechanism (MLCM) in the image fusion stage. To enrich the semantic information of the model, the semantic segmentation network is coupled with the image fusion network during training. According to experimental data, SGMAFusion performs better than other advanced algorithms in terms of richer semantic information and textural features.