Multispectral pedestrian detection model based on cross-attention and self-adapting differential fusion
摘要
In recent years, multispectral pedestrian detection has been widely applied in the field of autonomous driving. Multispectral images provide complementary visual information, which effectively improves the robustness and reliability of pedestrian detection systems. However, efficiently fusing information from different modalities to reduce the miss rate remains a key challenge. To address this, we propose a novel multispectral pedestrian detection model. First, a cross-attention feature extraction (CAFE) module is introduced to enhance target feature representations by capturing the common information from both spectra. Then, a self-adaptive differential fusion (SDF) module is designed to significantly improve the model’s ability to parse and integrate information from different sources, thereby enhancing the overall robustness and accuracy of the system. Experiments on the KAIST dataset demonstrate that the proposed method achieves logarithmic average miss rates (