3D facial animation has been widely applied in entertainment, virtual reality, education, and advertising. Achieving coordination between lip movements, facial expressions, and speech content is crucial for enhancing the realism of animations. Existing research often focuses on directly driving 3D facial animations using audio signals, failing to fully process the complex relationship between audio signals and facial expressions, and overlooking the critical role of lip shape constraints in improving lip synchronization accuracy. This paper proposes MambaTalk, a speech-driven 3D lip synchronization method. MambaTalk, based on the Mamba architecture, generates more refined and natural 3D facial animations by separately and independently handling lip movements strongly correlated with speech information and upper facial expressions weakly correlated with speech information. Additionally, the introduction of lip shape constraints further enhances the naturalness and synchronization of lip movements in 3D facial animations. Quantitative evaluation demonstrates that MambaTalk produces animations of higher quality than existing state-of-the-art methods on the VOCASET and MeshTalk dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MambaTalk: Speech-Driven 3D Facial Animation with Mamba

  • Deli Zhu,
  • Zhao Xu,
  • Yunong Yang

摘要

3D facial animation has been widely applied in entertainment, virtual reality, education, and advertising. Achieving coordination between lip movements, facial expressions, and speech content is crucial for enhancing the realism of animations. Existing research often focuses on directly driving 3D facial animations using audio signals, failing to fully process the complex relationship between audio signals and facial expressions, and overlooking the critical role of lip shape constraints in improving lip synchronization accuracy. This paper proposes MambaTalk, a speech-driven 3D lip synchronization method. MambaTalk, based on the Mamba architecture, generates more refined and natural 3D facial animations by separately and independently handling lip movements strongly correlated with speech information and upper facial expressions weakly correlated with speech information. Additionally, the introduction of lip shape constraints further enhances the naturalness and synchronization of lip movements in 3D facial animations. Quantitative evaluation demonstrates that MambaTalk produces animations of higher quality than existing state-of-the-art methods on the VOCASET and MeshTalk dataset.