Unmanned Aerial Vehicle (UAV) aerial photography has broad applications and is essential in timely and accurate object detection. To ensure precise detection while minimizing parameters and detection time, the constraints of embedded devices during deployment are taken into consideration. We propose Convolutional Attention Fusion-YOLO (CAF-YOLO). Firstly, the Convolutional Attention Fusion Network (CAFNet) is introduced to enhance feature extraction in the neck network by integrating two distinct branches that leverage diverse information flows. Secondly, the Dilated-Enhanced Residual Block (DERB) is designed in the backbone network, efforts are made to broaden the perception range and gather additional contextual data, improving multi-scale object detection. Finally, the fusion head is developed to reduce parameters and computational cost by merging feature extraction and eliminating redundant branches. The experimental results show that CAF-YOLO effectively improves object detection accuracy in UAV photography. In the VisDrone2019 public dataset, CAF-YOLO precision increased by 3.3% and improved by 2.4% compared to YOLOv8n. Additionally, the model’s parameters were reduced to 2.9M, and detection time decreased to 6.5ms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale Feature Fusion for Unmanned Aerial Vehicle Object Detection

  • Chunlong Fan,
  • Lanxin Li,
  • Luli Zhu

摘要

Unmanned Aerial Vehicle (UAV) aerial photography has broad applications and is essential in timely and accurate object detection. To ensure precise detection while minimizing parameters and detection time, the constraints of embedded devices during deployment are taken into consideration. We propose Convolutional Attention Fusion-YOLO (CAF-YOLO). Firstly, the Convolutional Attention Fusion Network (CAFNet) is introduced to enhance feature extraction in the neck network by integrating two distinct branches that leverage diverse information flows. Secondly, the Dilated-Enhanced Residual Block (DERB) is designed in the backbone network, efforts are made to broaden the perception range and gather additional contextual data, improving multi-scale object detection. Finally, the fusion head is developed to reduce parameters and computational cost by merging feature extraction and eliminating redundant branches. The experimental results show that CAF-YOLO effectively improves object detection accuracy in UAV photography. In the VisDrone2019 public dataset, CAF-YOLO precision increased by 3.3% and improved by 2.4% compared to YOLOv8n. Additionally, the model’s parameters were reduced to 2.9M, and detection time decreased to 6.5ms.