Pose Preserving Landmark Guided Neural Radiation Fields for Talking Portrait Synthesis
摘要
Talking portrait synthesis is a challenging task of synthesizing an image sequence of portraits with accurate lip synchronization and head pose estimation which correspond to a given audio clip of speech. Recently, the dynamic Neural Radiance Field(NeRF) has achieved real-time synthesizing and high-fidelity 3D modeling of talking portraits, and becomes a hot topic within the research community. In this work, we propose a pose preserving prior and a prior-guided NeRF rendering method for talking portrait synthesis. The prior is a fusion of the facial landmark and the input audio, where the facial landmark is composed of a unified lip-synced landmark and the pose landmark. The unified lip-synced landmarks are generated by a Transformer-like network trained by a large lip-reading corpus, which is suppose to learn the correspondence between the phoneme of an audio and the overall lip motion via its landmark. Lip-synced landmark along with the pose landmark compose the facial landmark for pose preserving. Relying on the uniqueness of the facial landmark and facial action correspondences, it aims to optimize the stability between the synthesized portrait and the driving information. The prior-guided NeRF renderer is to generate the synthesizing portraits according to the input audio and pose preserving prior that learn. Extensive experiments show that our method outperforms existing methods in terms of fidelity and synchronization of speech and video.