Silent Listener project addresses the limitations of traditional audio-centric speech detection by introducing a Visual Speech Recognition (VSR) system that leverages deep learning techniques. The objective is to convert spoken words into text by analyzing lip movements, with a specific focus on applications in noisy environments and for individuals with hearing impairments. The proposed model combines a 3D Convolutional Neural Network (CNN) trained on a GRID video dataset. For improved spoken word prediction accuracy, this workflow incorporates visual features extracted from the mouth region, alongside analysis of the corresponding audio signal. The implementation process acknowledges potential limitations, including poor input video quality, lighting variations, and other environmental factors that may impact the system's effectiveness. This project contributes to the exploration of solutions for speech recognition, specifically emphasizing the integration of visual information to improve accuracy and robustness. This project's findings hold promise for advancing real-world speech recognition applications. By incorporating these insights, future systems can become more accessible and robust across various environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Silent Listener: Deep Learning-Based Visual Speech Recognition

  • Aiswarya Mohan K,
  • T Allwin Kingstan,
  • Don Philip,
  • Anandhu Santhosh

摘要

Silent Listener project addresses the limitations of traditional audio-centric speech detection by introducing a Visual Speech Recognition (VSR) system that leverages deep learning techniques. The objective is to convert spoken words into text by analyzing lip movements, with a specific focus on applications in noisy environments and for individuals with hearing impairments. The proposed model combines a 3D Convolutional Neural Network (CNN) trained on a GRID video dataset. For improved spoken word prediction accuracy, this workflow incorporates visual features extracted from the mouth region, alongside analysis of the corresponding audio signal. The implementation process acknowledges potential limitations, including poor input video quality, lighting variations, and other environmental factors that may impact the system's effectiveness. This project contributes to the exploration of solutions for speech recognition, specifically emphasizing the integration of visual information to improve accuracy and robustness. This project's findings hold promise for advancing real-world speech recognition applications. By incorporating these insights, future systems can become more accessible and robust across various environments.