Images captured by unmanned aerial vehicles (UAVs) vary in flight altitude, which presents challenges in detection due to significant variations in target scales, sparse distribution, and complex background interference. Traditional detection methods often struggle to balance accuracy and efficiency. To overcome these challenges, we propose PMA-YOLO network. Specifically, we propose a Parallel Multi-scale Depthwise Separable Convolution (PDSC), which adaptively captures features from targets of varying scales. This significantly enhances sensitivity and representation for multi-scale targets while minimizing computational cost. Next, we propose a Multi-Dimensional Attention (MDA) mechanism, which strengthens the feature representation of small targets. The MDA mechanism refines features across three dimensions—spatial, channel, and coordinate—by facilitating cross-dimensional information interaction and dynamic weighting, thereby enhancing key features while suppressing background noise. Finally, we propose a high-resolution small target detection layer. By generating larger feature maps to better capture and represent the detailed characteristics of fine-grained targets. Experimental results in the VisDrone2019 and DOTA public datasets proves that our innovative approach achieved a 5.8% increase in mAP on the VisDrone2019 dataset and a 3.4% increase on the DOTA dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multiscale and Multidimensional Lightweight Network for Small-Target Detection in UAV Images

  • Shuai Li,
  • Boyuan Li,
  • Hongji Ma,
  • Kurban Ubul

摘要

Images captured by unmanned aerial vehicles (UAVs) vary in flight altitude, which presents challenges in detection due to significant variations in target scales, sparse distribution, and complex background interference. Traditional detection methods often struggle to balance accuracy and efficiency. To overcome these challenges, we propose PMA-YOLO network. Specifically, we propose a Parallel Multi-scale Depthwise Separable Convolution (PDSC), which adaptively captures features from targets of varying scales. This significantly enhances sensitivity and representation for multi-scale targets while minimizing computational cost. Next, we propose a Multi-Dimensional Attention (MDA) mechanism, which strengthens the feature representation of small targets. The MDA mechanism refines features across three dimensions—spatial, channel, and coordinate—by facilitating cross-dimensional information interaction and dynamic weighting, thereby enhancing key features while suppressing background noise. Finally, we propose a high-resolution small target detection layer. By generating larger feature maps to better capture and represent the detailed characteristics of fine-grained targets. Experimental results in the VisDrone2019 and DOTA public datasets proves that our innovative approach achieved a 5.8% increase in mAP on the VisDrone2019 dataset and a 3.4% increase on the DOTA dataset.