<p>Generating a talking face that is expressive is an interesting challenge. However, the existing works place more emphasis on lip synchronization or head posture than emotional expressions that make the speaker more realistic and vivid. This paper presents the Audio-Visual Controlled Expressive Talking Face Generation (EFG-AVT), which can leverage another pose source video to compensate for head postures and emotional expressions. Meanwhile, generating lip-synced talking faces is available. On the one hand, face keypoints and their local affine transformations are employed to align the emotional expression of the reference character. On the other hand, a powerful lip synchronization discriminator and a visual discriminator are implemented for lip synchronization. More importantly, the use of deformable convolution and SE attention mechanism makes the feature extraction of face more sufficient. In addition, face restoration is used to generate higher-quality video. Extensive experiments demonstrate that the proposed EFG-AVT outperforms most tasks of audio-driven talking head generation. In addition to producing lip-synced lips and controllable head postures, it can also produce appropriate emotional expressions. Furthermore, to provide a more convenient and faster service, performing a face talking video generation on edge devices is proposed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Expressive talking face generation via audio visual control

  • Pengfei Li,
  • Huihuang Zhao,
  • Mugang Lin,
  • Qingyun Liu,
  • Peng Tang,
  • Yangfan Zhou

摘要

Generating a talking face that is expressive is an interesting challenge. However, the existing works place more emphasis on lip synchronization or head posture than emotional expressions that make the speaker more realistic and vivid. This paper presents the Audio-Visual Controlled Expressive Talking Face Generation (EFG-AVT), which can leverage another pose source video to compensate for head postures and emotional expressions. Meanwhile, generating lip-synced talking faces is available. On the one hand, face keypoints and their local affine transformations are employed to align the emotional expression of the reference character. On the other hand, a powerful lip synchronization discriminator and a visual discriminator are implemented for lip synchronization. More importantly, the use of deformable convolution and SE attention mechanism makes the feature extraction of face more sufficient. In addition, face restoration is used to generate higher-quality video. Extensive experiments demonstrate that the proposed EFG-AVT outperforms most tasks of audio-driven talking head generation. In addition to producing lip-synced lips and controllable head postures, it can also produce appropriate emotional expressions. Furthermore, to provide a more convenient and faster service, performing a face talking video generation on edge devices is proposed.