In the underground space scenario, the surrounding environment is complex and volatile. Relying solely on a single sensor is prone to interference from environmental changes (such as lighting and occlusion), resulting in the inability of single-modal information to meet the detection requirements in complex environments. To address these issues, we have proposed the Multi-Modal Information Cross-Attention Fusion (MMIF) module and constructed a novel multi-modal target detection model. The proposed MMIF mechanism enhances the characteristic multimodal information and fully integrates the unique multimodal information to resist the interference of environmental changes and improve the detection accuracy. Simultaneously, the existing attention mechanism can not effectively combine channel and spatial attentionthe limits the improvement of target detection performance. Therefore, this paper presents an efficient channel space adaptive weighted convolution method, which adaptively allocates the spatial attention weights of each channel. By employing the multi-scale depthwise separable convolution module, spatial relationships are effectively extracted to cope with complex environmental changes. Extensive experimental studies on the LLVIP dataset, which demonstrate that the method we proposed is effective and achieves advanced detection performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Two-Stream Multi-modal Target Detection Based on Yolov8 Model Under Complex Lighting Conditions

  • Long Qian,
  • Jizhou Lai,
  • Wenxuan Zeng,
  • Ruien Mao,
  • Shen Xu,
  • Shun Dong

摘要

In the underground space scenario, the surrounding environment is complex and volatile. Relying solely on a single sensor is prone to interference from environmental changes (such as lighting and occlusion), resulting in the inability of single-modal information to meet the detection requirements in complex environments. To address these issues, we have proposed the Multi-Modal Information Cross-Attention Fusion (MMIF) module and constructed a novel multi-modal target detection model. The proposed MMIF mechanism enhances the characteristic multimodal information and fully integrates the unique multimodal information to resist the interference of environmental changes and improve the detection accuracy. Simultaneously, the existing attention mechanism can not effectively combine channel and spatial attentionthe limits the improvement of target detection performance. Therefore, this paper presents an efficient channel space adaptive weighted convolution method, which adaptively allocates the spatial attention weights of each channel. By employing the multi-scale depthwise separable convolution module, spatial relationships are effectively extracted to cope with complex environmental changes. Extensive experimental studies on the LLVIP dataset, which demonstrate that the method we proposed is effective and achieves advanced detection performance.