Comparing Audio Speech Recognition and Video Speech Recognition Using Neural Networks
摘要
The comparative analysis of recognition techniques using the GRID dataset enables us to evaluate the performance of models on visual lip movements and auditory signals. Therefore, we construct a common architecture for video and audio tasks. Similar model construction blocks are utilized to perform different tasks such as lip-reading from a sequence of video frames and speech recognition from audio waveforms using Mel-Frequency Cepstral Coefficients (MFCCs). In this study, we aim at determining how well these models can process the distinct characteristics of visual lip movements and auditory signals by training both with the GRID dataset. This approach also provides an opportunity to assess the effectiveness and adaptability of our models across various modalities directly. As such, this paper discusses how well the shared architecture deals with temporal and spatial characteristics, impacting advancements in video-based recognition, audio-based technologies, and analyzing cross-domain applications.