Advancements in 3D object detection are pivotal for the development of autonomous driving technologies, demanding high accuracy, robustness, and real-time processing capabilities. Current state-of-the-art multi-modal 3d object detection frameworks often struggle to balance these demands, particularly under the computational constraints of autonomous vehicles. This study introduces a novel 3D object detection framework that leverages a transformer-based fusion module, employing unique radial and zigzag partitioning techniques to efficiently integrate LiDAR and camera data. Our method, termed CTP-net, is designed to optimize inference speed while maintaining competitive detection accuracy. Tested on the NuScenes validation dataset, CTP-net achieves a NuScenes Detection Score (NDS) of 68.39. Notably, it demonstrates remarkable inference speeds of 8.50 FPS on an NVIDIA RTX 3060 and 20.72 FPS on a Tesla A100, indicating substantial improvements over existing methods making it a viable solution for deployment on edge devices with limited computational resources.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Continuous Token Partitioning for Real-Time Multi-modal 3d Object Detection

  • Nikolay Filatov,
  • Roman Potekhin

摘要

Advancements in 3D object detection are pivotal for the development of autonomous driving technologies, demanding high accuracy, robustness, and real-time processing capabilities. Current state-of-the-art multi-modal 3d object detection frameworks often struggle to balance these demands, particularly under the computational constraints of autonomous vehicles. This study introduces a novel 3D object detection framework that leverages a transformer-based fusion module, employing unique radial and zigzag partitioning techniques to efficiently integrate LiDAR and camera data. Our method, termed CTP-net, is designed to optimize inference speed while maintaining competitive detection accuracy. Tested on the NuScenes validation dataset, CTP-net achieves a NuScenes Detection Score (NDS) of 68.39. Notably, it demonstrates remarkable inference speeds of 8.50 FPS on an NVIDIA RTX 3060 and 20.72 FPS on a Tesla A100, indicating substantial improvements over existing methods making it a viable solution for deployment on edge devices with limited computational resources.