The rapid advancement of Unmanned Aerial Vehicle (UAV) detection technology has led to its widespread use in traffic surveillance and management. However, ground targets in high-altitude aerial images are often small and low in resolution, resulting in blurred details and weak features, which significantly increases the challenge of small target detection. To address this, we propose RMS-YOLO, an improved YOLO11-based small target detection model for UAV images, designed to enhance detection accuracy while reducing model parameters. RMS-YOLO introduces a re-parameterized Vision Transformer Block (RVB) in the C3k2 module, which further enhances the model’s performance in complex environments through its unique re-parameterization technique and EMA attention mechanism. Additionally, an improved Bi-directional Feature Pyramid Network (MBiFPN) is introduced, where feature information at different levels is efficiently fused through top-down and bottom-up links, significantly enhancing small object detection. To further simplify the model, we did not adopt the decoupled head design, but instead employed a lightweight Shared Convolutional Detection Head (SCDH). We evaluated the proposed model on the VisDrone2019 dataset using five performance metrics: P, R, mAP@50, mAP@50–95, and model size. Experimental results show that compared to the YOLO11n baseline model, RMS-YOLO improves P, R, mAP@50, and mAP@50–95 by 3.4%, 4.9%, 5%, and 4%, respectively, while reducing parameters by 29.7%. The model size is reduced to just 4.2 MB.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-scale Feature Fusion Method for Small Target Detection in High-Altitude Aerial Photography

  • Yifei Wang,
  • Yuan Zhu,
  • Dongsheng Wang,
  • Man Li,
  • Yibo Jie

摘要

The rapid advancement of Unmanned Aerial Vehicle (UAV) detection technology has led to its widespread use in traffic surveillance and management. However, ground targets in high-altitude aerial images are often small and low in resolution, resulting in blurred details and weak features, which significantly increases the challenge of small target detection. To address this, we propose RMS-YOLO, an improved YOLO11-based small target detection model for UAV images, designed to enhance detection accuracy while reducing model parameters. RMS-YOLO introduces a re-parameterized Vision Transformer Block (RVB) in the C3k2 module, which further enhances the model’s performance in complex environments through its unique re-parameterization technique and EMA attention mechanism. Additionally, an improved Bi-directional Feature Pyramid Network (MBiFPN) is introduced, where feature information at different levels is efficiently fused through top-down and bottom-up links, significantly enhancing small object detection. To further simplify the model, we did not adopt the decoupled head design, but instead employed a lightweight Shared Convolutional Detection Head (SCDH). We evaluated the proposed model on the VisDrone2019 dataset using five performance metrics: P, R, mAP@50, mAP@50–95, and model size. Experimental results show that compared to the YOLO11n baseline model, RMS-YOLO improves P, R, mAP@50, and mAP@50–95 by 3.4%, 4.9%, 5%, and 4%, respectively, while reducing parameters by 29.7%. The model size is reduced to just 4.2 MB.