Traditional talking head synthesis algorithms decompose the lips and head pose information from the facial landmark points based on the spatial point registration. In this paper, we show that the hypothesis of the traditional point registration methods is too strong to result in unnatural talking head. Instead of registration, we propose a latent lip-head pose coding method. The proposed method performs self-supervised learning and equivariant data augmentation to the facial landmark points. The experimental results show that the proposed latent lip-head pose coding method outperforms the traditional registration-based methods and can generate natural looking talking head with accurate mouth shapes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Controllable Talking Head Synthesis by Equivariant Data Augmentation for Spatial Coordinates

  • Wan Ding,
  • Dong-Yan Huang,
  • Zehong Zheng,
  • Tianyu Wang,
  • Linhuang Yan,
  • Xianjie Yang,
  • Penghui Li

摘要

Traditional talking head synthesis algorithms decompose the lips and head pose information from the facial landmark points based on the spatial point registration. In this paper, we show that the hypothesis of the traditional point registration methods is too strong to result in unnatural talking head. Instead of registration, we propose a latent lip-head pose coding method. The proposed method performs self-supervised learning and equivariant data augmentation to the facial landmark points. The experimental results show that the proposed latent lip-head pose coding method outperforms the traditional registration-based methods and can generate natural looking talking head with accurate mouth shapes.