Deep Neural Network-Based Phoneme Recognition Techniques for Non-native English Speakers
摘要
Phoneme recognition is one of the important real-life applications of automatic speech recognition. It is a speech-processing technology that transcribes and analyzes spoken language by breaking it into small distinct sounds, known as phonemes. In recent years, DNN has seen significant improvement in ASR. In this paper, transformer-based architecture is implemented to perform phoneme recognition. Additionally, GMM-HMM and DNN are deployed for this task and their results are compared. The paper discusses the results obtained by using these models on two datasets, namely TIMIT and L2 ARCTIC. Experimental results show that transformer architecture gave 26% less phoneme error rate when compared with the base-like model of GMM-HMM. Moreover, the paper provides a fresh perspective on phoneme recognition for a more diverse range of people with a special focus on non-native English speakers.