The text-to-3D task aims to generate high-fidelity 3D contents from textual descriptions, with broad applications in virtual reality, game design, and other fields. However, existing methods often struggle with insufficient reconstruction details, low fidelity, and weak semantic alignment. To address these challenges, we propose SC-D3D, a novel text-to-3D framework comprising two key components: the Dual Collaborative Distillation module and the Semantic Calibration module. First, we enhance traditional score distillation by integrating a consistency model, forming a Consistency Distillation mechanism that refines global structure while preserving local details. By combining this with score distillation, our approach achieves a balance between fine-grained texture enhancement and global structural consistency. Furthermore, we introduce the Semantic Calibration module, which employs a human feedback-driven reward mechanism to guide the geometric optimization process, ensuring that the generated 3D models align more accurately with textual descriptions. Ablation studies validate the effectiveness of the Dual Collaborative Distillation module in improving both structural integrity and detail fidelity, while the Semantic Calibration module further enhances semantic correspondence. Comparative experiments with state-of-the-art methods demonstrate that SC-D3D achieves superior performance in both high-quality 3D model generation and text-3D semantic alignment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SC-D3D: A Dual Collaborative Distillation Framework for High-Fidelity Text-To-3D Generation

  • ZiHong Li,
  • Wen Yang,
  • RuiSheng Ran

摘要

The text-to-3D task aims to generate high-fidelity 3D contents from textual descriptions, with broad applications in virtual reality, game design, and other fields. However, existing methods often struggle with insufficient reconstruction details, low fidelity, and weak semantic alignment. To address these challenges, we propose SC-D3D, a novel text-to-3D framework comprising two key components: the Dual Collaborative Distillation module and the Semantic Calibration module. First, we enhance traditional score distillation by integrating a consistency model, forming a Consistency Distillation mechanism that refines global structure while preserving local details. By combining this with score distillation, our approach achieves a balance between fine-grained texture enhancement and global structural consistency. Furthermore, we introduce the Semantic Calibration module, which employs a human feedback-driven reward mechanism to guide the geometric optimization process, ensuring that the generated 3D models align more accurately with textual descriptions. Ablation studies validate the effectiveness of the Dual Collaborative Distillation module in improving both structural integrity and detail fidelity, while the Semantic Calibration module further enhances semantic correspondence. Comparative experiments with state-of-the-art methods demonstrate that SC-D3D achieves superior performance in both high-quality 3D model generation and text-3D semantic alignment.