EC Speaker: Speech-Driven 3D Facial Animation with Latent Emotion Constraints
摘要
Recent advancements in artificial intelligence have significantly propelled speech-driven 3D facial animation technologies. However, the existing facial animation generation methods often overlook the critical interplay between latent emotional cues in audio signals and their manifestation in facial dynamics, consequently restricting the fidelity of generated animations. To bridge this gap, we present EC Speaker, a novel conditional diffusion framework that explicitly models emotional speech characteristics to synthesize realistic facial animations. We design a novel emotion-aware extractor that enables effective extraction of latent emotional features from raw audio through the perceptual fusion of time-frequency domain features. In addition, we introduce an emotion-constrained loss function (EmoLoss) to regulate the emotional expression of generated facial animations. Extensive experimental results on the VOCASET dataset demonstrate that the proposed model achieves higher accuracy in facial animation generation compared to other state-of-the-art methods, reducing vertex error and lip error by 41.28% and 75.84%.