<p>In addressing the challenges inherent in traffic sign recognition, such as susceptibility to interference, difficulty in detecting small targets, and the trade-off between real-time performance and accuracy, we propose an Efficient Traffic Sign YOLO (ETS-YOLO) model to fulfill the demands of real-time sign recognition. Built on the baseline YOLOv5 model, we adopt Partial Convolution (PConv) to redesign the Cross Stage Partial bottleneck including 3 convolutional layers (C3) module of the network, resulting in the proposed C3Efficient module, which reduces redundant calculations and memory accesses, making the model more efficient and lightweight. Accordingly, we utilize the Dynamic Upsampler (DySample) up-sampling to enhance the up-sampling effect during the Feature Pyramid Network (FPN) stage and design a new Normalized Wasserstein Distance (NWD) loss to redefine the anchor positional loss function mechanism, leading to better detection accuracy for tiny objects. Also, applying the Soft-NMS algorithm facilitates anchor filtering optimization, notably improving the detection accuracy of adjacent occluded targets. As a result of these improvements, the ETS-YOLO model achieves a mean Average Precision (mAP_0.5) of 0.805 when trained and tested on 45 types of traffic signs within the TT100K dataset, demonstrating a noteworthy improvement of 0.052 compared to the baseline model. In terms of model complexity, the ETS-YOLO model has 6.0 million parameters and a computational load of 14.3 GFLOPs, which represent reductions of 16.0% in model size and 12.3% in computational load, respectively, compared to the baseline model. Meanwhile, the model achieves an inference latency of 26.3 ms, showing a decrease of 10% in inference time over the baseline model. Ultimately, our model achieves a favorable balance between lightweight and accuracy compared to other state-of-the-art models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ETS-YOLO: An Efficient YOLO-based Model for Real-Time Traffic Sign Recognition

  • Mengran Yang,
  • Shuangshuang Han

摘要

In addressing the challenges inherent in traffic sign recognition, such as susceptibility to interference, difficulty in detecting small targets, and the trade-off between real-time performance and accuracy, we propose an Efficient Traffic Sign YOLO (ETS-YOLO) model to fulfill the demands of real-time sign recognition. Built on the baseline YOLOv5 model, we adopt Partial Convolution (PConv) to redesign the Cross Stage Partial bottleneck including 3 convolutional layers (C3) module of the network, resulting in the proposed C3Efficient module, which reduces redundant calculations and memory accesses, making the model more efficient and lightweight. Accordingly, we utilize the Dynamic Upsampler (DySample) up-sampling to enhance the up-sampling effect during the Feature Pyramid Network (FPN) stage and design a new Normalized Wasserstein Distance (NWD) loss to redefine the anchor positional loss function mechanism, leading to better detection accuracy for tiny objects. Also, applying the Soft-NMS algorithm facilitates anchor filtering optimization, notably improving the detection accuracy of adjacent occluded targets. As a result of these improvements, the ETS-YOLO model achieves a mean Average Precision (mAP_0.5) of 0.805 when trained and tested on 45 types of traffic signs within the TT100K dataset, demonstrating a noteworthy improvement of 0.052 compared to the baseline model. In terms of model complexity, the ETS-YOLO model has 6.0 million parameters and a computational load of 14.3 GFLOPs, which represent reductions of 16.0% in model size and 12.3% in computational load, respectively, compared to the baseline model. Meanwhile, the model achieves an inference latency of 26.3 ms, showing a decrease of 10% in inference time over the baseline model. Ultimately, our model achieves a favorable balance between lightweight and accuracy compared to other state-of-the-art models.