The comparative analysis of recognition techniques using the GRID dataset enables us to evaluate the performance of models on visual lip movements and auditory signals. Therefore, we construct a common architecture for video and audio tasks. Similar model construction blocks are utilized to perform different tasks such as lip-reading from a sequence of video frames and speech recognition from audio waveforms using Mel-Frequency Cepstral Coefficients (MFCCs). In this study, we aim at determining how well these models can process the distinct characteristics of visual lip movements and auditory signals by training both with the GRID dataset. This approach also provides an opportunity to assess the effectiveness and adaptability of our models across various modalities directly. As such, this paper discusses how well the shared architecture deals with temporal and spatial characteristics, impacting advancements in video-based recognition, audio-based technologies, and analyzing cross-domain applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparing Audio Speech Recognition and Video Speech Recognition Using Neural Networks

  • Atharva Swami,
  • Neetigya Bisen,
  • Darshak Savant

摘要

The comparative analysis of recognition techniques using the GRID dataset enables us to evaluate the performance of models on visual lip movements and auditory signals. Therefore, we construct a common architecture for video and audio tasks. Similar model construction blocks are utilized to perform different tasks such as lip-reading from a sequence of video frames and speech recognition from audio waveforms using Mel-Frequency Cepstral Coefficients (MFCCs). In this study, we aim at determining how well these models can process the distinct characteristics of visual lip movements and auditory signals by training both with the GRID dataset. This approach also provides an opportunity to assess the effectiveness and adaptability of our models across various modalities directly. As such, this paper discusses how well the shared architecture deals with temporal and spatial characteristics, impacting advancements in video-based recognition, audio-based technologies, and analyzing cross-domain applications.