<p>Vision-based occupancy prediction is an emerging perception paradigm in autonomous driving, offering the capability to reconstruct 3D scene geometry from 2D images. BEV-based methods, with their low computational cost, have become a focal point in this area of research. However, existing methods are hindered by limited semantic understanding of scenes, and they suffer from deficiencies in both supervision strategies and transformation mechanisms during critical view transformation stages, leading to suboptimal occupancy prediction performance. To address these challenges, we propose a lightweight framework called Semantic Supervision &amp; Back Projection Network (SSBPN). Specifically, SSBPN integrates a 2D Semantic Supervision (2D-SS) module to enhance the semantic understanding of the image encoder, while a Back Projection Enhancement (BPE) module is used to generate high-quality BEV features. Additionally, we introduce an improved Instance-based Depth Soft Label (IDSL) for accurate depth supervision during the crucial 2D-to-3D transformation stage. SSBPN achieves state-of-the-art performance on the Occ3D-nuScenes dataset, with a 37.22 mIoU using single-frame input.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lightweight occupancy network via semantic supervision & back projection enhancement

  • Xianbao Wang,
  • Shunwen Zuo,
  • Sheng Xiang,
  • Hanyun Zhou

摘要

Vision-based occupancy prediction is an emerging perception paradigm in autonomous driving, offering the capability to reconstruct 3D scene geometry from 2D images. BEV-based methods, with their low computational cost, have become a focal point in this area of research. However, existing methods are hindered by limited semantic understanding of scenes, and they suffer from deficiencies in both supervision strategies and transformation mechanisms during critical view transformation stages, leading to suboptimal occupancy prediction performance. To address these challenges, we propose a lightweight framework called Semantic Supervision & Back Projection Network (SSBPN). Specifically, SSBPN integrates a 2D Semantic Supervision (2D-SS) module to enhance the semantic understanding of the image encoder, while a Back Projection Enhancement (BPE) module is used to generate high-quality BEV features. Additionally, we introduce an improved Instance-based Depth Soft Label (IDSL) for accurate depth supervision during the crucial 2D-to-3D transformation stage. SSBPN achieves state-of-the-art performance on the Occ3D-nuScenes dataset, with a 37.22 mIoU using single-frame input.