MLF-YOLO: a novel multiscale feature fusion network for remote sensing small target detection
摘要
Detecting small targets in aerial and remote sensing imagery remains challenging due to factors such as cluttered backgrounds, dense distributions, and low resolution of small objects. Existing detection networks often rely on simple multi-scale fusion strategies, which cause shallow features to be overwhelmed by high-level semantics, limiting the ability to identify fine-grained targets. Additionally, downsampling operations in deep layers result in the loss of spatial details, further reducing detection accuracy. Traditional loss functions also struggle with error sensitivity from small annotation boxes, making precise localization difficult in dense environments. To address these issues, we propose MLF-YOLO, a novel small target detection model tailored for complex remote sensing scenes. Specifically, we design a multiscale local fusion (MLF) module guided by attention, which effectively integrates deep semantic information into shallow layers while preserving local details. This module enhances the model’s ability to detect small targets in complex backgrounds and dense environments. To model global context and refine spatial associations, we introduce the feature aware contextual encoding transformer (FACET) module, which extracts contextual information and further improves detection accuracy in dense small target scenarios. Additionally, we design an IW-IoU loss function, which enhances localization robustness for small targets under occlusion and dense scenarios. We evaluate MLF-YOLO on the RSOD, NWPU VHR-10, VisDrone2-019, and large-scale DOTAv2.0 datasets. Our model achieves improvements of 2%, 26.4%, and 7.7% in the mAP50 metric, respectively, while maintaining excellent performance and real-time capabilities on the large-scale DOTAv2.0 dataset. In summary, the proposed innovations comprehensively enhance the overall performance of MLF-YOLO in remote sensing small object detection tasks, especially in complex backgrounds and dense small target scenarios.