Guidance-free talking face generation is a complex and challenging task for high-resolution videos. Most guidance-free methods generate the target video by deforming the appearance of reference frames to match the source video. However, the potential pose discrepancy between reference and source frames often leads to appearance distortions in these methods. To address this issue, we propose a Masked Appearance Restoration (MAR) approach for generating high-resolution videos with reasonable facial structure and detailed texture. Specifically, the proposed MAR comprises two main modules: a Masked Structure Reasoning (MSR) module and a Motion Detail Aligning (MDA) module. The MSR module focuses on reasoning multi-scale structure features by combining a masked autoencoder with multi-scale upsampling layers, which ensures that the generated videos exhibit precise structural representations. Subsequently, to preserve detailed textures, the MDA module leverages audio-driven cross-attention to dynamically align the texture and structure features. Finally, the structure and texture features are fused and fed into a decoder, generating the target video frames. Extensive experiments indicate that our method generates accurately lip-synced talking faces in high-resolution video. Moreover, our model demonstrates robustness to extreme pose variations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Masked Appearance Restoration for High Resolution Talking Face Generation

  • Yuqing Wen,
  • Zhijing Cheng,
  • Congcong Zhu,
  • Rui Du

摘要

Guidance-free talking face generation is a complex and challenging task for high-resolution videos. Most guidance-free methods generate the target video by deforming the appearance of reference frames to match the source video. However, the potential pose discrepancy between reference and source frames often leads to appearance distortions in these methods. To address this issue, we propose a Masked Appearance Restoration (MAR) approach for generating high-resolution videos with reasonable facial structure and detailed texture. Specifically, the proposed MAR comprises two main modules: a Masked Structure Reasoning (MSR) module and a Motion Detail Aligning (MDA) module. The MSR module focuses on reasoning multi-scale structure features by combining a masked autoencoder with multi-scale upsampling layers, which ensures that the generated videos exhibit precise structural representations. Subsequently, to preserve detailed textures, the MDA module leverages audio-driven cross-attention to dynamically align the texture and structure features. Finally, the structure and texture features are fused and fed into a decoder, generating the target video frames. Extensive experiments indicate that our method generates accurately lip-synced talking faces in high-resolution video. Moreover, our model demonstrates robustness to extreme pose variations.