Objective: This study focuses on addressing the issue of road damage detection, including the detection of transverse cracks, longitudinal cracks, alligator cracks, potholes, and uneven manhole covers. The goal is to propose an efficient and accurate detection method for instance-level detection tasks in the field of video or image processing, overcoming the limitations of existing technologies, such as the inapplicability of causality to image tasks and over-reliance on multi-stage encoders or decoders. Methods: Using Adaptive Blending technology on five datasets obtained from vehicle-mounted systems, the network architecture was redesigned to include innovative components such as the coordinate attention module and the globally positioned local refinement head. These enhancements improve feature extraction and model inference capabilities, ultimately achieving automatic detection of small targets from a mobile vehicle perspective. Results: On the five target datasets, the algorithm’s average processing time per image is approximately 9.9 ms. The mean Average Precision (mAP) at 0.5 is 91.14%, and the mAP at 0.5:0.95 is 55.46% on the test set. Conclusion: Experimental results show that the algorithm can accurately identify the five types of targets in vehicle-mounted road scenarios while meeting real-time requirements, making it deployable for detecting these five targets on motor vehicles.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Road Damage Target Detection Using Attention Mechanism and Convolutional Dense Scale Feature Fusion

  • Ge Wang,
  • Xiangwei Zhang,
  • Ruiming Wang,
  • Li Wang,
  • Qingrong Li

摘要

Objective: This study focuses on addressing the issue of road damage detection, including the detection of transverse cracks, longitudinal cracks, alligator cracks, potholes, and uneven manhole covers. The goal is to propose an efficient and accurate detection method for instance-level detection tasks in the field of video or image processing, overcoming the limitations of existing technologies, such as the inapplicability of causality to image tasks and over-reliance on multi-stage encoders or decoders. Methods: Using Adaptive Blending technology on five datasets obtained from vehicle-mounted systems, the network architecture was redesigned to include innovative components such as the coordinate attention module and the globally positioned local refinement head. These enhancements improve feature extraction and model inference capabilities, ultimately achieving automatic detection of small targets from a mobile vehicle perspective. Results: On the five target datasets, the algorithm’s average processing time per image is approximately 9.9 ms. The mean Average Precision (mAP) at 0.5 is 91.14%, and the mAP at 0.5:0.95 is 55.46% on the test set. Conclusion: Experimental results show that the algorithm can accurately identify the five types of targets in vehicle-mounted road scenarios while meeting real-time requirements, making it deployable for detecting these five targets on motor vehicles.