<p>To address high miss rates and low detection accuracy caused by object scale variations and occlusion in video pedestrian detection, this study proposes a General Tracking-Information-Aided Detection Framework (GTIADF). The framework integrates object detection and tracking, leveraging inter-frame temporal information to improve robustness against scale changes and occlusion. GTIADF consists of two main components: (1) an enhanced multi-object tracking algorithm, Deep SORT+, which incorporates advanced modules, camera motion compensation, and DEMA for improved tracking robustness; and (2) an aided detection module that reduces missed detections via validation, interpolation, and re-scoring. The framework is adaptable and can be seamlessly integrated with other detectors to enhance their performance. Based on GTIADF, we propose the D3F-YOLO algorithm, which enhances the feature extraction module by using deformable convolution for better detection at varying scales. Additionally, a focusing diffusion pyramid network is introduced to improve multiscale feature representation, and the loss function is optimized to boost accuracy. Experiments on the Caltech dataset show that the proposed method achieves a mean mean precision (mAP) of 65. 7% and a miss rate of 32.4%, confirming its effectiveness in challenging scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing video pedestrian detection with tracking-information-aided framework and multi-scale feature optimization

  • Haifeng Sang,
  • Pengkai Suo

摘要

To address high miss rates and low detection accuracy caused by object scale variations and occlusion in video pedestrian detection, this study proposes a General Tracking-Information-Aided Detection Framework (GTIADF). The framework integrates object detection and tracking, leveraging inter-frame temporal information to improve robustness against scale changes and occlusion. GTIADF consists of two main components: (1) an enhanced multi-object tracking algorithm, Deep SORT+, which incorporates advanced modules, camera motion compensation, and DEMA for improved tracking robustness; and (2) an aided detection module that reduces missed detections via validation, interpolation, and re-scoring. The framework is adaptable and can be seamlessly integrated with other detectors to enhance their performance. Based on GTIADF, we propose the D3F-YOLO algorithm, which enhances the feature extraction module by using deformable convolution for better detection at varying scales. Additionally, a focusing diffusion pyramid network is introduced to improve multiscale feature representation, and the loss function is optimized to boost accuracy. Experiments on the Caltech dataset show that the proposed method achieves a mean mean precision (mAP) of 65. 7% and a miss rate of 32.4%, confirming its effectiveness in challenging scenarios.