<p>One-shot video-driven talking face generation aims to extract facial movements from a driving video and apply them to arbitrary portrait images to synthesize a talking facial video. The core challenge of this process lies in decoupling facial expressions from head poses, as these motion information are coupled together. Although prior researches attempt to achieve this information through self-supervised learning, they have limitations in accuracy. In this paper, we propose a talking face generation model based on Geometric Transformation Supervised Disentanglement of Pose and Expression. In this model, to enhance the model’s learning of facial transformations information, we introduce facial keypoints as inputs and use them as supervision information to improve the fidelity of generated results. By introducing geometric transformations of facial keypoints as supervision, we achieve a more effective decoupling of facial motion information. Specifically, the keypoints of the input driving image are transformed through geometric transformations to approximate the keypoints of the intermediate result generated by the model, then the transformed keypoints supervise the keypoints of the intermediate result. Besides, to address the issue of background blurring in some generated results and improve the authenticity of the generated results, we design a background enhancement module that extracts more background information from sequential images. Experimental results demonstrate that our method outperforms state-of-the-art techniques in separating facial expressions and poses, which achieves high-quality video editing effects.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Geometric transformation supervised disentanglement of pose and expression for talking face generation

  • Mengxiang Wang,
  • Guiyu Xia,
  • Zhedong Jin,
  • Paike Yang,
  • Yubao Sun

摘要

One-shot video-driven talking face generation aims to extract facial movements from a driving video and apply them to arbitrary portrait images to synthesize a talking facial video. The core challenge of this process lies in decoupling facial expressions from head poses, as these motion information are coupled together. Although prior researches attempt to achieve this information through self-supervised learning, they have limitations in accuracy. In this paper, we propose a talking face generation model based on Geometric Transformation Supervised Disentanglement of Pose and Expression. In this model, to enhance the model’s learning of facial transformations information, we introduce facial keypoints as inputs and use them as supervision information to improve the fidelity of generated results. By introducing geometric transformations of facial keypoints as supervision, we achieve a more effective decoupling of facial motion information. Specifically, the keypoints of the input driving image are transformed through geometric transformations to approximate the keypoints of the intermediate result generated by the model, then the transformed keypoints supervise the keypoints of the intermediate result. Besides, to address the issue of background blurring in some generated results and improve the authenticity of the generated results, we design a background enhancement module that extracts more background information from sequential images. Experimental results demonstrate that our method outperforms state-of-the-art techniques in separating facial expressions and poses, which achieves high-quality video editing effects.