Face the Future: Generating Animated 3D Avatars from Monocular Multimedia Instances with Unified Deep Learning Approach
摘要
Incorporating Deep Learning (DL) techniques, this research introduces a groundbreaking method for crafting realistic and easily animatable head avatars from everyday video sequences. The approach leverages a point-based representation that is deformable, effectively disentangling the source color into normal-dependent shading and intrinsic albedo. Unlike prevailing methods relying on either neural implicit representations or explicit 3D morphable meshes. Faf-Avatar strikes a balance between superior look, topological flexibility, deformation simplicity, and efficacy of production. The research showcases the method's prowess in generating animatable 3D avatars from monocular videos from diverse sources, outperforming previous methods in challenging scenarios such as capturing thin hair strands. Significantly, the proposed method exhibits superior training efficiency compared to competing approaches. By disentangling RGB color split into a pose-dependent shading element and a pose-agnostic albedo, the method enables unsupervised albedo disentanglement and rudimentary relighting through shading adjustments. The study offers a comprehensive breakdown of the methodology, encompassing the point-based canonical representation, point deformation, and shading network.