<p>The object detection network can achieve real-time performance on high-performance computers with ease, but the large number of parameters and limited computational resources of mobile devices pose significant challenges, leading to suboptimal detection performance. As application scenarios expand to embedded devices, such as autonomous vehicles and unmanned aerial vehicles, higher requirements on the real-time performance and resource consumption of object detection algorithms have been proposed. The traditional YOLOv4 model has huge parameters and computations, resulting in low detection efficiency in complex environments. To address it, we propose an improved YOLOv4-lite lightweight network based on depthwise over-parameterized convolutional layer (DO-Conv). Firstly, we replace CSPDarknet53 backbone network in YOLOv4 with MobileNetV3. The parameter quantity of YOLOv4-lite is only 62.4<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11554_2025_1645_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> of YOLOv4. Secondly, we use DO-Conv to replace the traditional convolution network in YOLOv4 to promote feature extraction effectiveness without extending layers. Meanwhile, we use ReLU6 to replace original Leaky ReLU to improve detection efficiency and obtain good numerical resolution. Our method achieved 70.38<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11554_2025_1645_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> mean average precision (mAP) on Pascal VOC07+12 dataset and 27.63<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11554_2025_1645_Article_IEq3.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> average precision (AP) on MS COCO 2017 dataset. Its model size is only 34MB. The running speed on Titan X reaches 41.82 frames per second (fps), which is 1.7 times that of YOLOv4. The experimental results demonstrate that the proposed method achieves a well-balanced trade-off between speed and accuracy, thereby meeting the real-time requirements for object detection in practical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A lightweight real-time object detection method for complex scenes based on YOLOv4

  • Peng Ding,
  • Tong Li,
  • Huaming Qian,
  • Lin Ma,
  • Zhongfei Chen

摘要

The object detection network can achieve real-time performance on high-performance computers with ease, but the large number of parameters and limited computational resources of mobile devices pose significant challenges, leading to suboptimal detection performance. As application scenarios expand to embedded devices, such as autonomous vehicles and unmanned aerial vehicles, higher requirements on the real-time performance and resource consumption of object detection algorithms have been proposed. The traditional YOLOv4 model has huge parameters and computations, resulting in low detection efficiency in complex environments. To address it, we propose an improved YOLOv4-lite lightweight network based on depthwise over-parameterized convolutional layer (DO-Conv). Firstly, we replace CSPDarknet53 backbone network in YOLOv4 with MobileNetV3. The parameter quantity of YOLOv4-lite is only 62.4 \(\%\) % of YOLOv4. Secondly, we use DO-Conv to replace the traditional convolution network in YOLOv4 to promote feature extraction effectiveness without extending layers. Meanwhile, we use ReLU6 to replace original Leaky ReLU to improve detection efficiency and obtain good numerical resolution. Our method achieved 70.38 \(\%\) % mean average precision (mAP) on Pascal VOC07+12 dataset and 27.63 \(\%\) % average precision (AP) on MS COCO 2017 dataset. Its model size is only 34MB. The running speed on Titan X reaches 41.82 frames per second (fps), which is 1.7 times that of YOLOv4. The experimental results demonstrate that the proposed method achieves a well-balanced trade-off between speed and accuracy, thereby meeting the real-time requirements for object detection in practical applications.