PointCLIP shows that transferring 2D knowledge to zero-shot 3D shape classification tasks can be effective, even without incorporating 3D-specific knowledge. However, the absence of 3D knowledge limits the model’s performance, leading to efforts focused on improving 3D-to-2D projection algorithms. This work advances previous methods by introducing a learnable fine-tuning strategy, called MDT-Net, for Contrastive Language-Image Pre-training (CLIP) on 3D datasets. MDT-Net freezes the CLIP encoder to keep view features aligned with semantic information and incorporates a shape decoder with masked view consistency and attention-guided fusion, enabling more effective fine-tuning of CLIP for 3D tasks. The decoder receives both raw view features and masked view features. In the shallow layers, a mask consistency constraint is utilized between two types of view features to enhance the model’s generalization and robustness. In the deeper layers, attention-guided fusion is introduced to better capture the contextual information between views and compact all view features into one. MDT-Net is evaluated on different benchmarks of 5 datasets for zero-shot classification, achieving a competitive classification accuracy on ModelNet10, ModelNet40, McGill, and ZS3D datasets. These experiment results demonstrate the effectiveness of this approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MDT-Net: A Mask Decoder Tuning Strategy for CLIP-Based Zero-Shot 3D Classification

  • Hao Yan,
  • Jing Bai

摘要

PointCLIP shows that transferring 2D knowledge to zero-shot 3D shape classification tasks can be effective, even without incorporating 3D-specific knowledge. However, the absence of 3D knowledge limits the model’s performance, leading to efforts focused on improving 3D-to-2D projection algorithms. This work advances previous methods by introducing a learnable fine-tuning strategy, called MDT-Net, for Contrastive Language-Image Pre-training (CLIP) on 3D datasets. MDT-Net freezes the CLIP encoder to keep view features aligned with semantic information and incorporates a shape decoder with masked view consistency and attention-guided fusion, enabling more effective fine-tuning of CLIP for 3D tasks. The decoder receives both raw view features and masked view features. In the shallow layers, a mask consistency constraint is utilized between two types of view features to enhance the model’s generalization and robustness. In the deeper layers, attention-guided fusion is introduced to better capture the contextual information between views and compact all view features into one. MDT-Net is evaluated on different benchmarks of 5 datasets for zero-shot classification, achieving a competitive classification accuracy on ModelNet10, ModelNet40, McGill, and ZS3D datasets. These experiment results demonstrate the effectiveness of this approach.