XPillars: enhancing 3D object detection through cross-pillar feature fusion
摘要
Point cloud object detection plays a critical role in fields such as autonomous driving and robotics. While detection accuracy has advanced, balancing this with efficient inference speed remains a critical challenge, particularly given hardware constraints. Pillar-based methods, owing to their streamlined design, provide notable advantages in computational efficiency. Addressing this, we present XPillars, a 3D point cloud detector focused on cross-pillar feature interaction. Our proposed efficient cross-pillar feature fusion method significantly enhances the model’s ability to capture contextual features. Furthermore, innovations in the backbone network and data augmentation contribute to the robust overall performance. Extensive experiments on the challenging KITTI benchmark dataset, evaluated using standard metrics including mean average precision for both 3D detection and bird’s eye view, demonstrate that XPillars surpasses previous single-stage methods. The extended experiments on the DAIR-V2X-V dataset further demonstrate the generalization capability and robustness of XPillars. Our method achieves a new state-of-the-art balance between high detection accuracy and efficient inference speed, making it highly suitable for real-time deployment in resource-constrained robotic and autonomous systems.