Lightweight occupancy network via semantic supervision & back projection enhancement
摘要
Vision-based occupancy prediction is an emerging perception paradigm in autonomous driving, offering the capability to reconstruct 3D scene geometry from 2D images. BEV-based methods, with their low computational cost, have become a focal point in this area of research. However, existing methods are hindered by limited semantic understanding of scenes, and they suffer from deficiencies in both supervision strategies and transformation mechanisms during critical view transformation stages, leading to suboptimal occupancy prediction performance. To address these challenges, we propose a lightweight framework called Semantic Supervision & Back Projection Network (SSBPN). Specifically, SSBPN integrates a 2D Semantic Supervision (2D-SS) module to enhance the semantic understanding of the image encoder, while a Back Projection Enhancement (BPE) module is used to generate high-quality BEV features. Additionally, we introduce an improved Instance-based Depth Soft Label (IDSL) for accurate depth supervision during the crucial 2D-to-3D transformation stage. SSBPN achieves state-of-the-art performance on the Occ3D-nuScenes dataset, with a 37.22 mIoU using single-frame input.