MambaTalk: Speech-Driven 3D Facial Animation with Mamba
摘要
3D facial animation has been widely applied in entertainment, virtual reality, education, and advertising. Achieving coordination between lip movements, facial expressions, and speech content is crucial for enhancing the realism of animations. Existing research often focuses on directly driving 3D facial animations using audio signals, failing to fully process the complex relationship between audio signals and facial expressions, and overlooking the critical role of lip shape constraints in improving lip synchronization accuracy. This paper proposes MambaTalk, a speech-driven 3D lip synchronization method. MambaTalk, based on the Mamba architecture, generates more refined and natural 3D facial animations by separately and independently handling lip movements strongly correlated with speech information and upper facial expressions weakly correlated with speech information. Additionally, the introduction of lip shape constraints further enhances the naturalness and synchronization of lip movements in 3D facial animations. Quantitative evaluation demonstrates that MambaTalk produces animations of higher quality than existing state-of-the-art methods on the VOCASET and MeshTalk dataset.