Detecting small targets in remote sensing and UAV imagery remains challenging due to limited resolution and ambiguous features. To address these issues, we propose a multi-modal detection framework based on an enhanced YOLO architecture that fuses infrared and visible light data. We introduce a Multi-Modal Local-Global Fusion Block (MLGF) to extract globally normalized features and design a dedicated detection head to exploit cross-modal complementarity. Extensive experiments show that our method achieves 82.48% average accuracy on the VEDAI dataset and 81.6% on DroneVehicle, outperforming state-of-the-art methods. It also improves mAP@0.5 by 5.9% on the VisDrone dataset. These results demonstrate the effectiveness and practical potential of our approach for autonomous driving and intelligent transportation systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Intelligent Transportation Target Recognition Based on Multi-Modal Data Fusion in Remote Sensing and UAV Images

  • Zhixing Wang,
  • Jikai Zhang,
  • Qiankai Xi,
  • Wentao Zhao,
  • Dongliang Guo

摘要

Detecting small targets in remote sensing and UAV imagery remains challenging due to limited resolution and ambiguous features. To address these issues, we propose a multi-modal detection framework based on an enhanced YOLO architecture that fuses infrared and visible light data. We introduce a Multi-Modal Local-Global Fusion Block (MLGF) to extract globally normalized features and design a dedicated detection head to exploit cross-modal complementarity. Extensive experiments show that our method achieves 82.48% average accuracy on the VEDAI dataset and 81.6% on DroneVehicle, outperforming state-of-the-art methods. It also improves mAP@0.5 by 5.9% on the VisDrone dataset. These results demonstrate the effectiveness and practical potential of our approach for autonomous driving and intelligent transportation systems.