Speech enhancement (SE) is essential for improving the quality and intelligibility of speech signals, particularly in noisy environments. In this paper, we propose an innovative approach to Audio-Visual Speech Enhancement (AVSE) by modifying an audio-visual variational autoencoder (AV-VAE) framework to integrate both lip movements and hand gestures from Cued Speech (CS) as visual cues. This is the first work to incorporate hand gestures, in addition to lip movements, within the AVSE task. By introducing hand cues, our approach aims to address the inherent challenges of lip reading, such as the high ambiguity in interpreting lip movements, which can limit the effectiveness of traditional AVSE methods. Leveraging deep learning and computer vision techniques, our method offers a more comprehensive representation of spoken content. Through empirical evaluation, we demonstrate the effectiveness of our approach in enhancing the clarity and quality of speech signals, even in challenging acoustic conditions. The results indicate that the integration of hand cues significantly improves speech quality, providing a promising solution for AVSE in noisy environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cued Speech-Integrated Audio-Visual Variational Autoencoder for Speech Enhancement

  • Lufei Gao,
  • Yan Rong,
  • Li Liu

摘要

Speech enhancement (SE) is essential for improving the quality and intelligibility of speech signals, particularly in noisy environments. In this paper, we propose an innovative approach to Audio-Visual Speech Enhancement (AVSE) by modifying an audio-visual variational autoencoder (AV-VAE) framework to integrate both lip movements and hand gestures from Cued Speech (CS) as visual cues. This is the first work to incorporate hand gestures, in addition to lip movements, within the AVSE task. By introducing hand cues, our approach aims to address the inherent challenges of lip reading, such as the high ambiguity in interpreting lip movements, which can limit the effectiveness of traditional AVSE methods. Leveraging deep learning and computer vision techniques, our method offers a more comprehensive representation of spoken content. Through empirical evaluation, we demonstrate the effectiveness of our approach in enhancing the clarity and quality of speech signals, even in challenging acoustic conditions. The results indicate that the integration of hand cues significantly improves speech quality, providing a promising solution for AVSE in noisy environments.