Semantic segmentation for LiDAR point clouds plays a crucial role in autonomous driving and robotic navigation. Currently, voxel-based methods have emerged as the predominant type of approaches in this domain. However, such methods inevitably introduces information loss due to quantization. To alleviate this problem, we propose a Multi-attribute perception fusion (MAF) module, that uses 2D convolution to extract scene semantic information from Range View to enhance the original point cloud features, and effectively improve the segmentation accuracy. Meanwhile, a point-based branch is introduced to preserve detail information even in large outdoor scenes. Furthermore, despite advanced segmentation methods widely use voxel-based 3D convolutions to extract multi-scale spatial features, their capacity to effectively integrate multi-scale features is constrained due to their reliance on direct feature summation or concatenation for feature fusion. Therefore, to accurately segment objects of different scales, we propose a Multi-scale adaptive fusion (MSAF) module. Through which, the features of different layers are selectively combined by weighting and adaptive fusion in the decoder part. Experimental results on the SemanticKITTI dataset demonstrate that our method achieves new state-of-the-art performance. The VPFNet codebase will be made publicly available.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VPFNET: A Scale-Adaptive Voxel Point Fusion Network for Semantic Segmentation of Point Clouds

  • Xiaoyang Wang,
  • Kaining Cui,
  • Lu Wang,
  • Zhenfei Liu,
  • Bingxin Yu,
  • Yuhong He,
  • Jun Cheng

摘要

Semantic segmentation for LiDAR point clouds plays a crucial role in autonomous driving and robotic navigation. Currently, voxel-based methods have emerged as the predominant type of approaches in this domain. However, such methods inevitably introduces information loss due to quantization. To alleviate this problem, we propose a Multi-attribute perception fusion (MAF) module, that uses 2D convolution to extract scene semantic information from Range View to enhance the original point cloud features, and effectively improve the segmentation accuracy. Meanwhile, a point-based branch is introduced to preserve detail information even in large outdoor scenes. Furthermore, despite advanced segmentation methods widely use voxel-based 3D convolutions to extract multi-scale spatial features, their capacity to effectively integrate multi-scale features is constrained due to their reliance on direct feature summation or concatenation for feature fusion. Therefore, to accurately segment objects of different scales, we propose a Multi-scale adaptive fusion (MSAF) module. Through which, the features of different layers are selectively combined by weighting and adaptive fusion in the decoder part. Experimental results on the SemanticKITTI dataset demonstrate that our method achieves new state-of-the-art performance. The VPFNet codebase will be made publicly available.