<p>Vehicle detection in closed-circuit television (CCTV) footage has long been a significant area of research due to the unique challenges posed by the nature of camera placement. CCTV cameras are typically mounted at oblique angles, leading to frequent occlusion in dense traffic scenes. Additionally, the varying distances of vehicles from the cameras, combined with the different angles at which the cameras are positioned, introduce the problem of scale variation. To address these challenges and enable accurate and efficient vehicle detection, we propose Global-Guided Mixing (GGMix), a novel module that combines multi-head self-attention (MHSA) and depthwise convolution in a sequential structure designed to enhance feature representation. It integrates multi-attention mechanisms to better capture vehicle-related features while suppressing irrelevant background information. We introduce GGMix blocks to the backbone of YOLOv8 in order to propose VDC-YOLO, a real-time vehicle detection model. Extensive experiments on a proprietary dataset, MLITcctv, and two public datasets, BitVehicle and i2, demonstrate that our model achieves superior performance, in terms of mAP while balancing accuracy and efficiency. Specifically, VDC-YOLO achieved an mAP of 46.9<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="40747_2025_2073_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> at 80.4 frames per second (fps) on the MLITcctv dataset, an mAP of 97.0<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="40747_2025_2073_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> at 55.7 fps on the BitVehicle dataset, and an mAP of 71.8<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="40747_2025_2073_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>%</mo> </math></EquationSource> </InlineEquation> on the i2 dataset. These results showcase the feasibility of our proposed model in the field of vehicle detection in actual CCTV environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vehicle detection in CCTV with global-guided self-attention and convolution

  • Yupei Guo,
  • Yota Yamamoto,
  • Hideki Yaginuma,
  • Yukinobu Taniguchi

摘要

Vehicle detection in closed-circuit television (CCTV) footage has long been a significant area of research due to the unique challenges posed by the nature of camera placement. CCTV cameras are typically mounted at oblique angles, leading to frequent occlusion in dense traffic scenes. Additionally, the varying distances of vehicles from the cameras, combined with the different angles at which the cameras are positioned, introduce the problem of scale variation. To address these challenges and enable accurate and efficient vehicle detection, we propose Global-Guided Mixing (GGMix), a novel module that combines multi-head self-attention (MHSA) and depthwise convolution in a sequential structure designed to enhance feature representation. It integrates multi-attention mechanisms to better capture vehicle-related features while suppressing irrelevant background information. We introduce GGMix blocks to the backbone of YOLOv8 in order to propose VDC-YOLO, a real-time vehicle detection model. Extensive experiments on a proprietary dataset, MLITcctv, and two public datasets, BitVehicle and i2, demonstrate that our model achieves superior performance, in terms of mAP while balancing accuracy and efficiency. Specifically, VDC-YOLO achieved an mAP of 46.9 \(\%\) % at 80.4 frames per second (fps) on the MLITcctv dataset, an mAP of 97.0 \(\%\) % at 55.7 fps on the BitVehicle dataset, and an mAP of 71.8 \(\%\) % on the i2 dataset. These results showcase the feasibility of our proposed model in the field of vehicle detection in actual CCTV environments.