AudioPortrait: A Staged Method for Generating Decoupled Facial Features and Synthesizing Speaker Faces Through Voice
摘要
This study proposes a two-stage talking face generation framework (AudioPortrait) based on audio-driven disentangled motion representation, which can achieve high-synchronization quality facial animation synthesis and cross-personality generalization. Aiming at the deficiencies of existing end-to-end methods in audio-video synchronization, dynamic detail control, and generation efficiency, this method decomposes facial dynamics into identity-independent expression coefficients and head pose parameters through a motion parameter decoupling strategy, which is combined with lightweight timing modeling to achieve efficient mapping. In the first stage, a four-stage encoder is constructed based on Hubert’s audio features, joint Transformer timing modeling predicts the target frame motion parameters, and multi-task loss is used to strengthen the lip synchronization accuracy; The second phase enables personalized adaptation of decoupled motion and static portraits through the LivePortrait rendering framework. Experiments show that the method achieves optimal LSE synchronization metrics (LSE-C = 7.75, LSE-D = 7.42) and realism performance of FVD-25 = 284.80 on the HDTF test set, and also exhibits strong generalization ability on the wild dataset. Compared to baseline models such as SadTalker, this framework has unlimited video length (supports streaming generation) and provides an efficient solution for digital human interaction scenarios.