<p>To address the challenge of accurately detecting whether people are wearing masks in video sequences, particularly in cases of occlusion or missed detection, we propose DTPM-YOLO (Detection and Tracking of Pedestrian Mask-wearing based on You Only Look Once). DTPM-YOLO introduces a cross-stage partial Bottleneck with 2 convolutions feature (CB2CF) into YOLOv5s to reduce the complexity of feature learning, and incorporates a convolutional block attention module (CBAM) to enhance detection capability, especially for small objects. Additionally, the inner-frame relationship module (IFRM) employs the Hungarian algorithm to establish associations among objects within an image, accurately identifying individuals who are wearing masks and those who are not. To address inconsistencies between the historical prediction direction of tracked targets and the direction of newly detected velocities, we integrate a directional difference factor into the association cost of Deep Simple Online and Real-time Tracking (DeepSORT). Our results demonstrate that DTPM-YOLO achieves a detection speed of 65 FPS, with an mAP@0.5 of 72.39%, which is 8.44% higher than the original YOLOv5. Compared to DeepSORT, our approach improved Multiple Object Tracking Accuracy (MOTA) by 18.0% and Multiple Object Tracking Precision (MOTP) by 3.40% on the MOT16 dataset, effectively enabling real-time mask-wearing detection and tracking.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pedestrian mask-wearing detection based on YOLOv5 and DeepSORT

  • Shuai Wang,
  • Abdul Samad Shibghatullah,
  • Kay Hooi Keoy,
  • Javid Iqbal

摘要

To address the challenge of accurately detecting whether people are wearing masks in video sequences, particularly in cases of occlusion or missed detection, we propose DTPM-YOLO (Detection and Tracking of Pedestrian Mask-wearing based on You Only Look Once). DTPM-YOLO introduces a cross-stage partial Bottleneck with 2 convolutions feature (CB2CF) into YOLOv5s to reduce the complexity of feature learning, and incorporates a convolutional block attention module (CBAM) to enhance detection capability, especially for small objects. Additionally, the inner-frame relationship module (IFRM) employs the Hungarian algorithm to establish associations among objects within an image, accurately identifying individuals who are wearing masks and those who are not. To address inconsistencies between the historical prediction direction of tracked targets and the direction of newly detected velocities, we integrate a directional difference factor into the association cost of Deep Simple Online and Real-time Tracking (DeepSORT). Our results demonstrate that DTPM-YOLO achieves a detection speed of 65 FPS, with an mAP@0.5 of 72.39%, which is 8.44% higher than the original YOLOv5. Compared to DeepSORT, our approach improved Multiple Object Tracking Accuracy (MOTA) by 18.0% and Multiple Object Tracking Precision (MOTP) by 3.40% on the MOT16 dataset, effectively enabling real-time mask-wearing detection and tracking.