Labeling every point in a scene is a laborious journey for 3D understanding. To achieve annotation-free training, existing works introduce Contrastive Language-Image Pre-training (CLIP) to transfer the pre-trained capability of visual-linguistic correspondence to 3D-linguistic matching. However, directly adopting this CLIP-driven strategy can inevitably introduce bias: The overrated roles of the color and texture from an RGB image could overshadow the geometric nature of the corresponding 3D scene, resulting in a sub-optimal alignment. We note that different from RGB images, a depth map contains rich geometric information. Inspired by this, we propose Depth-Enhanced Alignment (D-EA) for label-free 3D semantic segmentation. D-EA aims to explore the rich geometric cues in depth maps and mitigate the color and texture biases rooted in the original CLIP-driven strategy. Specifically, we first tune a geometry-enhanced CLIP by aligning its depth prediction to the paired RGB prediction given by the original CLIP. Next, the point cloud feature space is matched with the RGB-Depth aggregated CLIP space by aligning point prediction to RGB and depth predictions. Moreover, to mitigate the semantic ambiguity caused by view-specific noise, we propose a View-Integrated Pseudo Label Generation paradigm. Experiments demonstrate the effectiveness of the proposed D-EA on the ScanNet (indoor) and GraspNet-1Billion (desktop) datasets in the label-free setting. Our method is also competitive in limited annotation semantic segmentation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Depth-Enhanced Alignment for Label-Free 3D Semantic Segmentation

  • Shangjin Xie,
  • Jiawei Feng,
  • Zibo Chen,
  • Zhixuan Liu,
  • Wei-Shi Zheng

摘要

Labeling every point in a scene is a laborious journey for 3D understanding. To achieve annotation-free training, existing works introduce Contrastive Language-Image Pre-training (CLIP) to transfer the pre-trained capability of visual-linguistic correspondence to 3D-linguistic matching. However, directly adopting this CLIP-driven strategy can inevitably introduce bias: The overrated roles of the color and texture from an RGB image could overshadow the geometric nature of the corresponding 3D scene, resulting in a sub-optimal alignment. We note that different from RGB images, a depth map contains rich geometric information. Inspired by this, we propose Depth-Enhanced Alignment (D-EA) for label-free 3D semantic segmentation. D-EA aims to explore the rich geometric cues in depth maps and mitigate the color and texture biases rooted in the original CLIP-driven strategy. Specifically, we first tune a geometry-enhanced CLIP by aligning its depth prediction to the paired RGB prediction given by the original CLIP. Next, the point cloud feature space is matched with the RGB-Depth aggregated CLIP space by aligning point prediction to RGB and depth predictions. Moreover, to mitigate the semantic ambiguity caused by view-specific noise, we propose a View-Integrated Pseudo Label Generation paradigm. Experiments demonstrate the effectiveness of the proposed D-EA on the ScanNet (indoor) and GraspNet-1Billion (desktop) datasets in the label-free setting. Our method is also competitive in limited annotation semantic segmentation.