<p>Multimodal imaging technology improves the performance of pedestrian detection systems by complementing visible-light and infrared modalities, but problems such as modal imbalance and misalignment still need to be solved. Aiming at these issues, we propose an innovative modality balancing network for multimodal pedestrian detection to better integrate and collaborate the information of visible-light and infrared modalities. This network employs a cross-modal compensation fusion module to enhance image details through upsampling and downsampling operations and realize feature cross-utilization by sharing feature maps from different channels. It also utilizes a multimodal feature alignment mechanism to select complementary features according to lighting conditions, adaptively aligning the features of different modalities. Experimental results demonstrate that on the challenging KAIST multispectral pedestrian dataset, this network performs well in terms of detection accuracy and efficiency and is able to deal with the issue of lighting changes effectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Modality balancing network for pedestrian detection based on cross-modal compensation fusion and multimodal feature alignment

  • Zuhe Li,
  • Ruochong Fu,
  • Xiang Guo,
  • Minghui Zhu,
  • Yifan Gao,
  • Zhiyang Zhao,
  • Penghao Ouyang,
  • Zhijie Xu,
  • Yushan Pan

摘要

Multimodal imaging technology improves the performance of pedestrian detection systems by complementing visible-light and infrared modalities, but problems such as modal imbalance and misalignment still need to be solved. Aiming at these issues, we propose an innovative modality balancing network for multimodal pedestrian detection to better integrate and collaborate the information of visible-light and infrared modalities. This network employs a cross-modal compensation fusion module to enhance image details through upsampling and downsampling operations and realize feature cross-utilization by sharing feature maps from different channels. It also utilizes a multimodal feature alignment mechanism to select complementary features according to lighting conditions, adaptively aligning the features of different modalities. Experimental results demonstrate that on the challenging KAIST multispectral pedestrian dataset, this network performs well in terms of detection accuracy and efficiency and is able to deal with the issue of lighting changes effectively.