<p>Point cloud object detection plays a critical role in fields such as autonomous driving and robotics. While detection accuracy has advanced, balancing this with efficient inference speed remains a critical challenge, particularly given hardware constraints. Pillar-based methods, owing to their streamlined design, provide notable advantages in computational efficiency. Addressing this, we present XPillars, a 3D point cloud detector focused on cross-pillar feature interaction. Our proposed efficient cross-pillar feature fusion method significantly enhances the model’s ability to capture contextual features. Furthermore, innovations in the backbone network and data augmentation contribute to the robust overall performance. Extensive experiments on the challenging KITTI benchmark dataset, evaluated using standard metrics including mean average precision for both 3D detection and bird’s eye view, demonstrate that XPillars surpasses previous single-stage methods. The extended experiments on the DAIR-V2X-V dataset further demonstrate the generalization capability and robustness of XPillars. Our method achieves a new state-of-the-art balance between high detection accuracy and efficient inference speed, making it highly suitable for real-time deployment in resource-constrained robotic and autonomous systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

XPillars: enhancing 3D object detection through cross-pillar feature fusion

  • Lijuan Zhang,
  • Zihan Fu,
  • Zhiyi Li,
  • Dongming Li

摘要

Point cloud object detection plays a critical role in fields such as autonomous driving and robotics. While detection accuracy has advanced, balancing this with efficient inference speed remains a critical challenge, particularly given hardware constraints. Pillar-based methods, owing to their streamlined design, provide notable advantages in computational efficiency. Addressing this, we present XPillars, a 3D point cloud detector focused on cross-pillar feature interaction. Our proposed efficient cross-pillar feature fusion method significantly enhances the model’s ability to capture contextual features. Furthermore, innovations in the backbone network and data augmentation contribute to the robust overall performance. Extensive experiments on the challenging KITTI benchmark dataset, evaluated using standard metrics including mean average precision for both 3D detection and bird’s eye view, demonstrate that XPillars surpasses previous single-stage methods. The extended experiments on the DAIR-V2X-V dataset further demonstrate the generalization capability and robustness of XPillars. Our method achieves a new state-of-the-art balance between high detection accuracy and efficient inference speed, making it highly suitable for real-time deployment in resource-constrained robotic and autonomous systems.