Enhancing Remote Sensing Object Detection with LL-YOLO: Integrating Multi-modal Data Fusion and Latent Diffusion Models
摘要
Accurate object detection in remote sensing images is crucial for various applications, yet it poses significant challenges due to complex backgrounds and the presence of small targets. This paper introduces LL-LOYO model, a novel multi-modal object detection method that synergistically integrates visual and semantic data through La-YOLO and Ld-YOLO, enhanced by reinforcement learning-based latent diffusion model. Our approach significantly outperforms existing detection methods by effectively handling the intricacies of high-resolution remote sensing data. The experimental results show that LL-LOYO achieves a significant improvement in detection accuracy, especially for small and dense targets, confirming potential of LL-LOYO in improving object detection in remote sensing image data.