Expression Fusion to Enhance Video and Speech-Driven 3D Facial Animation
摘要
Currently, monocular 3D face capture and tracking methods are difficult to ensure the accuracy of identity and expression, and speech-driven 3D facial animation methods are challenging to acquire head pose and upper facial expression. In response, we propose a video- and speech-driven 3D facial animation synthesis method that attempts to combine the benefits of these methods and avoid their weaknesses. Specifically, to facilitate animation and handle the problem of character identity accuracy, we generate character appearance templates by registering a 3D morphable model (3DMM) to a rigid model. To address the limitations of different methods for acquiring expressions and poses, we design an expression fusion network based on the 3DMM space to fuse the expression data acquired by different modalities and output unified facial expression data. Finally, due to the limitation of the dataset ignoring the eye movements, we design an eye movement enhancement network to add eyelid movements by modifying the facial expression data and then replacing the eye region to get the final 3D face mesh animation. Through detailed experiments, we demonstrate that our method can generate speech-visual synchronized 3D face animations while obtaining better performance results than the current concerns of using monocular video images or speech-driven generation of 3D face animation methods independently.