<p>Detecting small objects in remote sensing images presents a significant challenge due to their limited size, low pixel density, and susceptibility to background interference, significantly heightening the risk of missed detections. In order to overcome these challenges, our research refines the YOLO11s algorithm. We center on improving detection exactness, streamlining computational efficiency, and managing the detection of small objects in complex settings. A key innovation of this method lies in incorporating Adaptive Fine-Grained Channel Attention(AFGC), which enhances the model’s ability to concentrate on critical regions while suppressing the impact of extraneous background elements. Additionally, we refined the core framework of the C3k2 module by using wavelet convolution (WTConv). We intended to heighten the model’s proficiency in extracting features during night-time situations. Additionally, we integrate a lightweight Cross-Attention Fusion Module (CAFM), which enhances feature extraction in challenging scenarios by addressing occlusion and lighting variations. By customizing Focaler-CIoU (FCIoU) loss function to optimize target fitting, we can synergistically enhance both model accuracy and robustness. The recall rate also increased by 6.1% on the VisDrone2019 dataset. Additional evaluation of the DOTA dataset yields promising results, providing further evidence of the proposed method’s effectiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Awcf-yolo11: hierarchical attention fusion and adaptive channel refinement for object detection in remote sensing imagery

  • Jingmin Yang,
  • Hongbin Zhang,
  • Wenjie Zhang,
  • Jinghui Ren

摘要

Detecting small objects in remote sensing images presents a significant challenge due to their limited size, low pixel density, and susceptibility to background interference, significantly heightening the risk of missed detections. In order to overcome these challenges, our research refines the YOLO11s algorithm. We center on improving detection exactness, streamlining computational efficiency, and managing the detection of small objects in complex settings. A key innovation of this method lies in incorporating Adaptive Fine-Grained Channel Attention(AFGC), which enhances the model’s ability to concentrate on critical regions while suppressing the impact of extraneous background elements. Additionally, we refined the core framework of the C3k2 module by using wavelet convolution (WTConv). We intended to heighten the model’s proficiency in extracting features during night-time situations. Additionally, we integrate a lightweight Cross-Attention Fusion Module (CAFM), which enhances feature extraction in challenging scenarios by addressing occlusion and lighting variations. By customizing Focaler-CIoU (FCIoU) loss function to optimize target fitting, we can synergistically enhance both model accuracy and robustness. The recall rate also increased by 6.1% on the VisDrone2019 dataset. Additional evaluation of the DOTA dataset yields promising results, providing further evidence of the proposed method’s effectiveness.