<p>With the advancement of deep-learning technology and the booming of the virtual-human industry, speech-driven generation of talking-face videos has emerged as a crucial research area in computer vision. Nevertheless, current research confronts several key challenges. Existing methods frequently struggle to guarantee the congruence between speech and facial emotional expressions during video generation, leading to a dearth of emotional expressiveness in the generated videos. Moreover, the lip-speech synchronization and image quality of these videos remain suboptimal. Specifically, when dealing with complex speech inputs, issues like lip-mismatch or image blurring are prone to occur. Additionally, the discriminator models in use have limitations in feature discrimination capabilities. They are unable to effectively discern the subtle disparities between generated and real-world videos, thereby impeding further enhancement of the generation effect. Consequently, inspired by the concept of Generative Adversarial Networks (GANs), this paper constructs a network for generating talking face videos with emotional expression. To address the problem of emotional consistency in speech-driven generated videos, a speech emotion recognition model is constructed. This model automatically extracts emotional information from the input audio and video. Experiments indicate that the Weighted Accuracy (WA) and Unweighted Accuracy (UA) of this model on the IEMOCAP (Interactive Emotional Dyadic Motion Capture), CASIA (Chinese Academy of Sciences Institute of Automation Speech Emotion), and CREMA-D (Crowd-sourced Emotional Multimodal Actors) datasets outperform those of existing approaches. Furthermore, this paper presents a generator model based on the Residual Dense Network (RDN) and Squeeze-and-Excitation Networks (SENet), as well as a discriminator model based on depthwise separable convolution and the Convolutional Block Attention Module (CBAM). Experimental results demonstrate that the improved generator and discriminator outperform the baseline models in terms of Lip Synchronization Error – Distance (LSE-D), Lip Synchronization Error-Confidence (LSE-C), Fréchet Inception Distance (FID), and Emotional Accuracy (EmoAcc). This verifies the enhanced performance of the proposed model in lip synchronization, facial movement naturalness, emotional expression authenticity, and generated - image quality.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech-driven talking face video generation

  • Zhao Wei,
  • Tang Zhuoran,
  • Yang Qihan

摘要

With the advancement of deep-learning technology and the booming of the virtual-human industry, speech-driven generation of talking-face videos has emerged as a crucial research area in computer vision. Nevertheless, current research confronts several key challenges. Existing methods frequently struggle to guarantee the congruence between speech and facial emotional expressions during video generation, leading to a dearth of emotional expressiveness in the generated videos. Moreover, the lip-speech synchronization and image quality of these videos remain suboptimal. Specifically, when dealing with complex speech inputs, issues like lip-mismatch or image blurring are prone to occur. Additionally, the discriminator models in use have limitations in feature discrimination capabilities. They are unable to effectively discern the subtle disparities between generated and real-world videos, thereby impeding further enhancement of the generation effect. Consequently, inspired by the concept of Generative Adversarial Networks (GANs), this paper constructs a network for generating talking face videos with emotional expression. To address the problem of emotional consistency in speech-driven generated videos, a speech emotion recognition model is constructed. This model automatically extracts emotional information from the input audio and video. Experiments indicate that the Weighted Accuracy (WA) and Unweighted Accuracy (UA) of this model on the IEMOCAP (Interactive Emotional Dyadic Motion Capture), CASIA (Chinese Academy of Sciences Institute of Automation Speech Emotion), and CREMA-D (Crowd-sourced Emotional Multimodal Actors) datasets outperform those of existing approaches. Furthermore, this paper presents a generator model based on the Residual Dense Network (RDN) and Squeeze-and-Excitation Networks (SENet), as well as a discriminator model based on depthwise separable convolution and the Convolutional Block Attention Module (CBAM). Experimental results demonstrate that the improved generator and discriminator outperform the baseline models in terms of Lip Synchronization Error – Distance (LSE-D), Lip Synchronization Error-Confidence (LSE-C), Fréchet Inception Distance (FID), and Emotional Accuracy (EmoAcc). This verifies the enhanced performance of the proposed model in lip synchronization, facial movement naturalness, emotional expression authenticity, and generated - image quality.