<p>Accurate vehicle detection is crucial in the field of intelligent transportation systems (ITS). However, the complexity of real-world road environments often leads to problems such as misdetection and omission, especially due to overlapping targets, occlusions, and small objects. To address these challenges, a vehicle detection model based on an improved YOLOv5 model is proposed in this paper. The detection network in the model uses the Coordinate Attention (CA) mechanism to improve the recognition of vehicle targets in dense object scenes. The Focal-EIOU Loss is also utilized to improve the localization accuracy of the anchor box and accelerate the convergence of the loss function. In addition, the integration of Omni-dimensional Dynamic Convolution (ODConv) enhances feature extraction by using a multidimensional attention mechanism that simultaneously computes four types of attention across all four dimensions of kernel space. Evaluations on a custom-built vehicle dataset show that the proposed model increases the Mean Average Precision (mAP) by 1.5% compared to the original model, while maintaining the detection speed with only a slight reduction of 8 frame·s<sup>−1</sup>. This improvement achieves higher detection accuracy without significantly compromising speed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced detection of small and occluded road vehicle targets using improved YOLOv5

  • Kaibin Zhu,
  • Hongming Lyu,
  • Yanbin Qin

摘要

Accurate vehicle detection is crucial in the field of intelligent transportation systems (ITS). However, the complexity of real-world road environments often leads to problems such as misdetection and omission, especially due to overlapping targets, occlusions, and small objects. To address these challenges, a vehicle detection model based on an improved YOLOv5 model is proposed in this paper. The detection network in the model uses the Coordinate Attention (CA) mechanism to improve the recognition of vehicle targets in dense object scenes. The Focal-EIOU Loss is also utilized to improve the localization accuracy of the anchor box and accelerate the convergence of the loss function. In addition, the integration of Omni-dimensional Dynamic Convolution (ODConv) enhances feature extraction by using a multidimensional attention mechanism that simultaneously computes four types of attention across all four dimensions of kernel space. Evaluations on a custom-built vehicle dataset show that the proposed model increases the Mean Average Precision (mAP) by 1.5% compared to the original model, while maintaining the detection speed with only a slight reduction of 8 frame·s−1. This improvement achieves higher detection accuracy without significantly compromising speed.