<p>Vehicle target detection is a crucial component of autonomous driving systems, as its performance directly impacts the accuracy of perception and decision-making. However, existing vehicle detection algorithms fail to fully exploit the complementary advantages of Convolutional Neural Network (CNN) and Transformer in feature extraction, thereby limiting the model’s generalization capability and detection accuracy in complex scenarios. To address this issue, this paper proposes a dual-branch vehicle detection model that fuses global and local features, referred to as DB-GLF (Dual-Branch Global-Local Fusion). In the feature encoding stage, the model leverages the strengths of CNN in capturing local details and Transformer in modeling global contextual information, designing a dual-branch backbone network for efficient extraction of both local and global semantic features from images. Additionally, a Hybrid Attention Module is introduced to dynamically fuse global and local features, enhancing the model’s focus on target regions and improving the representation of fine-grained details. Furthermore, a detection head based on Transformer encoder layers is constructed to optimize global feature modeling, significantly improving classification and localization accuracy. Experimental results demonstrate that the proposed DB-GLF model achieves superior performance in both ablation and comparative experiments. Compared with mainstream object detection methods, it improves the mean Average Precision (mAP) metric by approximately 3.5%, thereby validating the effectiveness and superiority of the algorithm.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-Branch Vehicle Detection Model with Global-Local Feature Fusion

  • Wei Baoli,
  • Zhang Siwan,
  • Zhang Jia,
  • Gao Pengfei

摘要

Vehicle target detection is a crucial component of autonomous driving systems, as its performance directly impacts the accuracy of perception and decision-making. However, existing vehicle detection algorithms fail to fully exploit the complementary advantages of Convolutional Neural Network (CNN) and Transformer in feature extraction, thereby limiting the model’s generalization capability and detection accuracy in complex scenarios. To address this issue, this paper proposes a dual-branch vehicle detection model that fuses global and local features, referred to as DB-GLF (Dual-Branch Global-Local Fusion). In the feature encoding stage, the model leverages the strengths of CNN in capturing local details and Transformer in modeling global contextual information, designing a dual-branch backbone network for efficient extraction of both local and global semantic features from images. Additionally, a Hybrid Attention Module is introduced to dynamically fuse global and local features, enhancing the model’s focus on target regions and improving the representation of fine-grained details. Furthermore, a detection head based on Transformer encoder layers is constructed to optimize global feature modeling, significantly improving classification and localization accuracy. Experimental results demonstrate that the proposed DB-GLF model achieves superior performance in both ablation and comparative experiments. Compared with mainstream object detection methods, it improves the mean Average Precision (mAP) metric by approximately 3.5%, thereby validating the effectiveness and superiority of the algorithm.