CMATalk: Cross modality alignment for talking head generation
摘要
Talking head generation has wide applications in virtual assistants, education, and entertainment. Indirect approaches, which leverage facial landmarks as intermediates, offer flexibility in output identity and resolution but often fail to effectively integrate multi-domain information. In this paper, we propose CMATalk, a novel framework that enhances the alignment between audio and facial landmarks within a shared latent space. CMATalk incorporates audio, landmark, and emotional cues through a cross-modal attention mechanism to generate expressive and synchronized facial landmarks. We further introduce a VAE-based image synthesis module that distills knowledge from a pre-trained model to reconstruct high-quality facial images from the predicted landmarks. Extensive experiments demonstrate that our method achieves competitive performance compared to state-of-the-art techniques. Our key contributions include: (1) a latent space alignment strategy for improved audio-landmark correlation, (2) a cross-modal architecture that fuses emotion with audio and landmark features, and (3) a VAE-enhanced decoder for high-fidelity facial reconstruction.