Despite significant advances in object detection, detecting small objects is still a challenge. The most common problems in small object detection include densely distributed objects, complex scale variations and background interference. To address this issue, we propose a network named BiCA-YOLO, which is based on YOLOv7 and advances small object detection with Bidirectional Feature Enhancement (BFE) and Cross Coordinate Attention (CCA). BFE removes salient features of large objects from low-level feature maps to highlight the features of small objects. It also transfers these removed features to high-level feature maps to enhance the features of large objects. CCA eliminates background interference in complex backgrounds by capturing long-range information. CCA not only use coordinate attention along the horizontal and vertical directions but also use convolutions of strip-shape kernels in two directions to retain important spatial information. Additionally, we use a weight function named improved-Slide to solve the imbalance between easy and hard samples. The effectiveness of BiCA-YOLO is verified on VisDrone-DET2019 and PASCAL VOC 2012 datasets, with state-of-the-art performance on both benchmarks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BiCA-YOLO: Bidirectional Feature Enhancement and Cross Coordinate Attention for Small Object Detection

  • Jin Yan Lv,
  • Guo Qiang Xiao

摘要

Despite significant advances in object detection, detecting small objects is still a challenge. The most common problems in small object detection include densely distributed objects, complex scale variations and background interference. To address this issue, we propose a network named BiCA-YOLO, which is based on YOLOv7 and advances small object detection with Bidirectional Feature Enhancement (BFE) and Cross Coordinate Attention (CCA). BFE removes salient features of large objects from low-level feature maps to highlight the features of small objects. It also transfers these removed features to high-level feature maps to enhance the features of large objects. CCA eliminates background interference in complex backgrounds by capturing long-range information. CCA not only use coordinate attention along the horizontal and vertical directions but also use convolutions of strip-shape kernels in two directions to retain important spatial information. Additionally, we use a weight function named improved-Slide to solve the imbalance between easy and hard samples. The effectiveness of BiCA-YOLO is verified on VisDrone-DET2019 and PASCAL VOC 2012 datasets, with state-of-the-art performance on both benchmarks.