BEVDot: Enhancing Environmental Perception for Autonomous Driving with a Deformable Depth Mechanism
摘要
In the field of autonomous driving, environmental perception is crucial for driving safety. Addressing the limitations of existing visual perception methods in complex scenarios, this study proposes a deformable depth visual perception framework based on a multi-camera system. The framework processes multi-camera data through a feature extraction network to generate and fuse multi-scale features. And a deformable depth prediction mechanism incorporating self-vehicle temporal difference features is introduced to enhance the accuracy of the model in depth prediction. Experimental results show that on the NuScenes dataset, our method achieves a detection accuracy (mAP) of 0.508 using only 5 random cameras out of 6, surpassing existing technologies such as Lift-Splat (0.446), RC-BEVFusion (0.476), and SOGDet-SE (0.474). Future research will focus on improving the prediction accuracy of distant vehicles to further enhance the performance of the model.