MGTR-Avatar: Multi-scale Gaussian Triplane Representation for High-Fidelity 3D Facial Model Reconstruction from a Monocular Video
摘要
Creating head avatar animations from monocular portrait videos is key to bridging virtual and real worlds. Existing methods like Explicit 3D Deformable Mesh (3DMM), neural implicit representations, and point clouds have advanced this field but face limitations: 3DMMs often produce overly smooth shapes due to fixed topologies, and neural implicit models require long training times with limited animation capacity. To address these limitations, we propose MGTR-Avatar, which combines animated 3D Gaussians with a parametric face model for photo-realistic avatars. Using FLAME as the initial point cloud, we convert model points into 3D Gaussian primitives and design a multi-scale triplane feature encoder with hybrid attention to capture detailed facial features. MGTR-Avatar captures diverse expressions and views, supporting real-time rendering and leveraging geometric priors for efficient training. Extensive experiments on the INSTA dataset show that MGTR-Avatar outperforms existing methods in both quality and speed.