<p>Talking head generation has wide applications in virtual assistants, education, and entertainment. Indirect approaches, which leverage facial landmarks as intermediates, offer flexibility in output identity and resolution but often fail to effectively integrate multi-domain information. In this paper, we propose CMATalk, a novel framework that enhances the alignment between audio and facial landmarks within a shared latent space. CMATalk incorporates audio, landmark, and emotional cues through a cross-modal attention mechanism to generate expressive and synchronized facial landmarks. We further introduce a VAE-based image synthesis module that distills knowledge from a pre-trained model to reconstruct high-quality facial images from the predicted landmarks. Extensive experiments demonstrate that our method achieves competitive performance compared to state-of-the-art techniques. Our key contributions include: (1) a latent space alignment strategy for improved audio-landmark correlation, (2) a cross-modal architecture that fuses emotion with audio and landmark features, and (3) a VAE-enhanced decoder for high-fidelity facial reconstruction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CMATalk: Cross modality alignment for talking head generation

  • Xuan-Nam Cao,
  • Quoc-Huy Trinh,
  • Minh-Triet Tran

摘要

Talking head generation has wide applications in virtual assistants, education, and entertainment. Indirect approaches, which leverage facial landmarks as intermediates, offer flexibility in output identity and resolution but often fail to effectively integrate multi-domain information. In this paper, we propose CMATalk, a novel framework that enhances the alignment between audio and facial landmarks within a shared latent space. CMATalk incorporates audio, landmark, and emotional cues through a cross-modal attention mechanism to generate expressive and synchronized facial landmarks. We further introduce a VAE-based image synthesis module that distills knowledge from a pre-trained model to reconstruct high-quality facial images from the predicted landmarks. Extensive experiments demonstrate that our method achieves competitive performance compared to state-of-the-art techniques. Our key contributions include: (1) a latent space alignment strategy for improved audio-landmark correlation, (2) a cross-modal architecture that fuses emotion with audio and landmark features, and (3) a VAE-enhanced decoder for high-fidelity facial reconstruction.