Fa-yolo: multi-scale feature fusion for spectral image object detection in complex scenes
摘要
Spectral imagery object detection involves identifying and localizing specific targets in multispectral (MSI) or hyperspectral (HSI) imagery. Leveraging more spectral bands than visible light, spectral images provide superior recognition capabilities in complex environments such as low visibility or adverse weather, benefiting applications in remote sensing, medical imaging, agriculture, and security. Visible light imaging delivers high-resolution spatial details, including fine textures and object shapes, while spectral imaging captures thermal radiation and material properties, exhibiting strong robustness under low-light conditions. Fusing these two modalities can significantly improve detection accuracy. However, existing methods often suffer from spectral redundancy, weak inter-modal interactions, and limited multi-scale local correlations, leading to performance degradation. To address these issues, we propose FA-YOLO, which adopts a Feature Aggregation and Attention Network (FAANet) as its backbone. FAANet integrates a Feature Interaction Transformer (FIT) module that explicitly reduces spectral redundancy by reconstructing and refining modality-specific features, thereby enhancing cross-modal information exchange under high dimensionality and modality discrepancies. In addition, we introduce the Multi-Scale Hybrid Attention Module (MSHAM) to robustly handle objects of varying scales. MSHAM combines spatial and channel attention across multiple receptive fields, strengthening local correlations and multi-scale feature representation, effectively capturing cross-modal interactions and establishing long-range dependencies. This integration improves both feature extraction and detection accuracy. Experimental results show that in low-visibility scenarios such as nighttime and foggy conditions, FA-YOLO achieves a 7.7% mAP50 gain over a single-modal baseline and a 3.1% mAP50 gain over the best existing bimodal method on the FLIR dataset. It achieves an optimal balance between model size and performance across multiple public datasets, providing a robust bimodal detection solution.