Robust Multi-Dialect End-to-End ASR Model Jointly with Beam Search Threshold Pruning and LLM
摘要
This paper aims to develop a novel robust multi-dialect end-to-end ASR system with beam search threshold pruning. The efficacy of our proposed model is evaluated using word error rate (WER). Our key contributions are: (1) To develop an end-to-end ASR system using attention-based neural network architecture and analyze the effectiveness of two features such as MFCC and log mel filter bank energies on multiple speech dialect corpora including American, Britain, and Indian accents; (2) To integrate beam search threshold pruning as a decoding mechanism to reduce the decoding time (3) To conduct an experimental analysis to test the model performance and compare the results against baseline system. (4) Post processing analysis are carried out using Llama2-7B based large language model(LLM) for enhancing the performance of proposed ASR system. The proposed model significantly improves performance by 1.91% and 4.29% over clean and noisy speech in librispeech corpus. Similarly, for the Indian accented speech, the model attains an average WER of about 6.6%.