<p>In traffic scenarios, image data contain a large number of small objects with limited effective information per object, leading to significant performance degradation of conventional object detectors trained on mainstream datasets. To address this issue, this paper proposes the multi-stage feature enhancer (MSFE), whose core framework employs adaptive guided slicing, upscaling, and detector cascading. By using the output of the previous-stage detector to locate regions of interest, subimages are magnified and fed into the posterior-stage detector, with final results merged to improve small object detection performance. Experiments conducted on 18 mainstream detectors (including the YOLO series and DETR series) across multiple scenarios (VisDrone, UAVDT, GCTS, and MP) demonstrate that MSFE can significantly enhance the small object detection performance of conventional detectors—among these, YOLO11-l achieves an AP improvement of at least 26.5%. Notably, MSFE leverages the pre-trained weights of existing detectors and requires no training itself. Meanwhile, it outperforms the similarly unsupervised SAHI in both performance and speed, and also surpasses ESOD, a supervised specialized small object detection algorithm. This paper also evaluates a large number of cascading strategies for conventional detectors to enable out-of-the-box use. As such, MSFE provides an efficient and general solution for small object detection, alleviating the pressure of traffic-scene small object detection on high-performance computing.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving small object detection in traffic scenarios via multi-stage feature enhancement

  • Gengwei Liao,
  • Cheng-Jie Jin

摘要

In traffic scenarios, image data contain a large number of small objects with limited effective information per object, leading to significant performance degradation of conventional object detectors trained on mainstream datasets. To address this issue, this paper proposes the multi-stage feature enhancer (MSFE), whose core framework employs adaptive guided slicing, upscaling, and detector cascading. By using the output of the previous-stage detector to locate regions of interest, subimages are magnified and fed into the posterior-stage detector, with final results merged to improve small object detection performance. Experiments conducted on 18 mainstream detectors (including the YOLO series and DETR series) across multiple scenarios (VisDrone, UAVDT, GCTS, and MP) demonstrate that MSFE can significantly enhance the small object detection performance of conventional detectors—among these, YOLO11-l achieves an AP improvement of at least 26.5%. Notably, MSFE leverages the pre-trained weights of existing detectors and requires no training itself. Meanwhile, it outperforms the similarly unsupervised SAHI in both performance and speed, and also surpasses ESOD, a supervised specialized small object detection algorithm. This paper also evaluates a large number of cascading strategies for conventional detectors to enable out-of-the-box use. As such, MSFE provides an efficient and general solution for small object detection, alleviating the pressure of traffic-scene small object detection on high-performance computing.