Expressive talking face generation via audio visual control
摘要
Generating a talking face that is expressive is an interesting challenge. However, the existing works place more emphasis on lip synchronization or head posture than emotional expressions that make the speaker more realistic and vivid. This paper presents the Audio-Visual Controlled Expressive Talking Face Generation (EFG-AVT), which can leverage another pose source video to compensate for head postures and emotional expressions. Meanwhile, generating lip-synced talking faces is available. On the one hand, face keypoints and their local affine transformations are employed to align the emotional expression of the reference character. On the other hand, a powerful lip synchronization discriminator and a visual discriminator are implemented for lip synchronization. More importantly, the use of deformable convolution and SE attention mechanism makes the feature extraction of face more sufficient. In addition, face restoration is used to generate higher-quality video. Extensive experiments demonstrate that the proposed EFG-AVT outperforms most tasks of audio-driven talking head generation. In addition to producing lip-synced lips and controllable head postures, it can also produce appropriate emotional expressions. Furthermore, to provide a more convenient and faster service, performing a face talking video generation on edge devices is proposed.