<p>In recent years, multispectral pedestrian detection has been widely applied in the field of autonomous driving. Multispectral images provide complementary visual information, which effectively improves the robustness and reliability of pedestrian detection systems. However, efficiently fusing information from different modalities to reduce the miss rate remains a key challenge. To address this, we propose a novel multispectral pedestrian detection model. First, a cross-attention feature extraction (CAFE) module is introduced to enhance target feature representations by capturing the common information from both spectra. Then, a self-adaptive differential fusion (SDF) module is designed to significantly improve the model’s ability to parse and integrate information from different sources, thereby enhancing the overall robustness and accuracy of the system. Experiments on the KAIST dataset demonstrate that the proposed method achieves logarithmic average miss rates (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7587_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="44" /> </InlineMediaObject> <EquationSource Format="TEX">\(\hbox {MR}^{-2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mtext>MR</mtext> <mrow> <mo>-</mo> <mn>2</mn> </mrow> </msup> </math></EquationSource> </InlineEquation>) of 5.29, 5.06, and 5.73 on the three main subsets, respectively. Compared to other state-of-the-art models, our approach reduces the miss rate by 0.64, 1.18, and 1.24, respectively, confirming the effectiveness of the proposed model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multispectral pedestrian detection model based on cross-attention and self-adapting differential fusion

  • Jinjin Wang,
  • Meng Li

摘要

In recent years, multispectral pedestrian detection has been widely applied in the field of autonomous driving. Multispectral images provide complementary visual information, which effectively improves the robustness and reliability of pedestrian detection systems. However, efficiently fusing information from different modalities to reduce the miss rate remains a key challenge. To address this, we propose a novel multispectral pedestrian detection model. First, a cross-attention feature extraction (CAFE) module is introduced to enhance target feature representations by capturing the common information from both spectra. Then, a self-adaptive differential fusion (SDF) module is designed to significantly improve the model’s ability to parse and integrate information from different sources, thereby enhancing the overall robustness and accuracy of the system. Experiments on the KAIST dataset demonstrate that the proposed method achieves logarithmic average miss rates ( \(\hbox {MR}^{-2}\) MR - 2 ) of 5.29, 5.06, and 5.73 on the three main subsets, respectively. Compared to other state-of-the-art models, our approach reduces the miss rate by 0.64, 1.18, and 1.24, respectively, confirming the effectiveness of the proposed model.