<p>To address the challenge of insufficient target detection accuracy for UAVs operating in adverse conditions, such as low illumination, dense fog, and extreme weather, this paper proposes a lightweight multi-modal fusion detection network named MMT-NET, designed to enhance UAV perception capabilities in challenging environments. Built upon the RT-DETR framework, the proposed method employs a dual-branch MobileNetV4 backbone to independently extract features from infrared and visible images. Moreover, a lightweight multi-modal feature interaction module is designed to strengthen the interaction between different modalities, and a lightweight cross-modal attention fusion module is designed to efficiently fuse cross-modal features via a spatial attention mechanism with minimal computational overhead. Extensive experiments on the public multi-modal dataset M3FD demonstrate that MMT-NET achieves 89.9% mAP@50 and 60.3% mAP@50:95, validating its effectiveness in multi-modal detection tasks while maintaining a lightweight architecture. Furthermore, qualitative evaluations under diverse real-world and simulated scenarios—including nighttime, fog, snow, and occlusion—confirm the robustness and generalization capability of the proposed method in complex environments. The source code of this work will be publicly available at: <a href="https://github.com/UAVSwarm/MMT-NET.">https://github.com/UAVSwarm/MMT-NET.</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MMT-NET: a lightweight multi-modal fusion network for UAV target detection in adverse environments

  • Chuanyun Wang,
  • Mingqi Zhou,
  • Dongdong Sun,
  • Qian Gao,
  • Zhaokui Li,
  • Tian Wang

摘要

To address the challenge of insufficient target detection accuracy for UAVs operating in adverse conditions, such as low illumination, dense fog, and extreme weather, this paper proposes a lightweight multi-modal fusion detection network named MMT-NET, designed to enhance UAV perception capabilities in challenging environments. Built upon the RT-DETR framework, the proposed method employs a dual-branch MobileNetV4 backbone to independently extract features from infrared and visible images. Moreover, a lightweight multi-modal feature interaction module is designed to strengthen the interaction between different modalities, and a lightweight cross-modal attention fusion module is designed to efficiently fuse cross-modal features via a spatial attention mechanism with minimal computational overhead. Extensive experiments on the public multi-modal dataset M3FD demonstrate that MMT-NET achieves 89.9% mAP@50 and 60.3% mAP@50:95, validating its effectiveness in multi-modal detection tasks while maintaining a lightweight architecture. Furthermore, qualitative evaluations under diverse real-world and simulated scenarios—including nighttime, fog, snow, and occlusion—confirm the robustness and generalization capability of the proposed method in complex environments. The source code of this work will be publicly available at: https://github.com/UAVSwarm/MMT-NET.