TVAE-3D: Efficient multi-view 3D shape reconstruction with diffusion models and transformer based VAE
摘要
Recent studies on diffusion models and transformers have made notable progress in computer vision by enhancing 3D image reconstruction. Diffusion models have demonstrated significant efficacy in producing images of exceptional quality. Creating high-quality 3D images has always been difficult, mainly because there is a lack of scalable 3D representations that can accurately capture complex geometric patterns. Despite significant efforts, current methods in 3D creation have challenges in effectively creating large-scale 3D shapes. Real-world shapes exhibit various geometry and texture variations, resulting in complicated appearances. To overcome these challenges, we suggest a more robust approach for 3D shape reconstruction using a transformer-based VAE for point cloud reconstruction. We provide a new time-dependent self-attention module that enables attention layers to adjust their behavior effectively at various phases of the denoising process. We provide a unique triplane-based 3D-aware Diffusion model using Transformer. We show a new 3D-aware knowledgeable transformer that uses a variational autoencoder to get general 3D knowledge from different types of shapes. In the TVAE-3D architecture, a transformer-based variational auto-encoder (VAE) models the whole shape, and a diffusion network that guesses the 3D reconstruction shapes. The VAE encoder utilizes the overall latent shape and the specific point features to understand the distribution of shapes. The 3D-aware transformers’ strong capacity to extract features guarantees excellent quality and smoothness of the resulting point clouds. Several experiments using the ShapeNet, OmniObject3D, MVP, ShapeNet-ViPC, and KITTI datasets have shown that TVAE-3D works better than other state-of-the-art algorithms, both in terms of FD, CD, MMD, and quality of the reconstructed shapes.