Crowded scenes present significant challenges for pedestrian detection due to multi-scale pedestrian instances, complex backgrounds, and varying occlusion patterns, ranging from minimal to severe occlusion. To overcome these difficulties, this paper introduces EMDC-YOLO, a residual multi-scale attention and cross-scale fusion-based method for pedestrian detection in crowded scenes. Firstly, an iRB_EMA model integrates channel and contextual information to minimize background noise and improve pixel-level focus on occluded pedestrians, thereby improving the model’s occlusion handling capability. Secondly, the DCPAN network is proposed, optimizing multi-scale feature fusion and accurately detecting severely occluded pedestrians. Finally, the ASFF Head network is proposed to generate features with the same resolution and channel number, thereby enhancing the recognition performance of pedestrian targets at different scales. Compared to the baseline, EMDC-YOLO achieves a precision that is 2.17% higher and recall improved by 8.4%. mAP@50 shows a 7.2% enhancement, while mAP@50:95 increases by 7.27%, demonstrating its improved performance and superiority.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EMDC-YOLO: A Residual Multi-scale Attention and Cross-Scale Fusion-Based Method for Pedestrian Detection in Crowded Scenes

  • Xinxin Zhou,
  • Chunying Xie,
  • Yucai Li,
  • Chunzhen Li

摘要

Crowded scenes present significant challenges for pedestrian detection due to multi-scale pedestrian instances, complex backgrounds, and varying occlusion patterns, ranging from minimal to severe occlusion. To overcome these difficulties, this paper introduces EMDC-YOLO, a residual multi-scale attention and cross-scale fusion-based method for pedestrian detection in crowded scenes. Firstly, an iRB_EMA model integrates channel and contextual information to minimize background noise and improve pixel-level focus on occluded pedestrians, thereby improving the model’s occlusion handling capability. Secondly, the DCPAN network is proposed, optimizing multi-scale feature fusion and accurately detecting severely occluded pedestrians. Finally, the ASFF Head network is proposed to generate features with the same resolution and channel number, thereby enhancing the recognition performance of pedestrian targets at different scales. Compared to the baseline, EMDC-YOLO achieves a precision that is 2.17% higher and recall improved by 8.4%. mAP@50 shows a 7.2% enhancement, while mAP@50:95 increases by 7.27%, demonstrating its improved performance and superiority.