Noise Robust E2E Continuous Kannada ASR System Under Real Time Conditions
摘要
In this study, we showcase the latest enhancements integrated into the previously developed end-to-end (E2E) continuous Kannada automatic speech recognition (ASR) system. The objective is to enhance the accuracy of the existing E2E continuous Kannada ASR model by integrating advanced speech enhancement and acoustic modeling techniques. To achieve this, we propose a novel method that combines magnitude and phase features using a deep neural network (DNN) to reconstruct enhanced speech signals. During the training stage, we treat the phase as a target, transforming unstructured spectrogram information into its derivative along a time axis, referred to as sudden frequency deviation (SFD). In the decoding process, the phase information (spectrogram) is recovered from the estimated SFD. The reconstruction of an enhanced speech signal involves amalgamating the features of magnitude and SFD within a learning framework. Our proposed noise elimination algorithm is applied to a degraded continuous Kannada speech database for enhancement. It is integrated into the real-time spoken query system before the speech feature extraction stage. Additionally, we explore the efficacy of time delay neural network (TDNN) and long short-term memory (LSTM) acoustic modeling techniques. The experiments conducted on both tasks, namely speech enhancement and ASR, demonstrate that the proposed algorithm, along with recent acoustic modeling techniques such as TDNN and LSTM, significantly reduces the word error rate for a noise-robust ASR system. The main findings of this work indicate that the integration of magnitude and phase features using a novel DNN-based technique significantly improves the performance of the E2E continuous Kannada ASR system under real-time conditions. To the best of our knowledge, these results represent the best performance reported for an E2E continuous Kannada ASR system.