<p>The burgeoning field of text-to-3D synthesis offers transformative potential in diverse domains such as computer-aided design, gaming, virtual reality, and artistic creation. However, the generation struggles with issues of inconsistency and low resolution, primarily due to the lack of critical visual clues like views and attributes. Furthermore, random constraint in rendering may impair model inference, leading to the Janus problem. In response to these challenges, we introduce HexaDream to produce high-quality 3D content. Hexaview Generation Diffusion Model is designed to merge object types, attributes, and view-specific text into unified latent space. Besides, the feature aggregation attention significantly enhances the detail and consistency of the generated output by mapping point features from orthogonal view into the 3D domain. Another innovation is the Dynamic-weighted HexaConstraint. This module employs a projection matrix to generate projected views and calculates the differential loss between these projections and the hexaviews, ensuring high fidelity. Our comparative experiments show that HexaDream achieves improvements of 8% in CLIP-R, 12% in Keypart Fidelity, and especially 20.6% in Multihead Alleviation compared with existing methods respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HexaDream: hexaview prior and constraint for text to 3D creation

  • Zhi-Chao Zhang,
  • Hui Chen,
  • Jin-Sheng Deng,
  • Ming Xu,
  • Zheng-Bin Pang

摘要

The burgeoning field of text-to-3D synthesis offers transformative potential in diverse domains such as computer-aided design, gaming, virtual reality, and artistic creation. However, the generation struggles with issues of inconsistency and low resolution, primarily due to the lack of critical visual clues like views and attributes. Furthermore, random constraint in rendering may impair model inference, leading to the Janus problem. In response to these challenges, we introduce HexaDream to produce high-quality 3D content. Hexaview Generation Diffusion Model is designed to merge object types, attributes, and view-specific text into unified latent space. Besides, the feature aggregation attention significantly enhances the detail and consistency of the generated output by mapping point features from orthogonal view into the 3D domain. Another innovation is the Dynamic-weighted HexaConstraint. This module employs a projection matrix to generate projected views and calculates the differential loss between these projections and the hexaviews, ensuring high fidelity. Our comparative experiments show that HexaDream achieves improvements of 8% in CLIP-R, 12% in Keypart Fidelity, and especially 20.6% in Multihead Alleviation compared with existing methods respectively.