Phoneme recognition is one of the important real-life applications of automatic speech recognition. It is a speech-processing technology that transcribes and analyzes spoken language by breaking it into small distinct sounds, known as phonemes. In recent years, DNN has seen significant improvement in ASR. In this paper, transformer-based architecture is implemented to perform phoneme recognition. Additionally, GMM-HMM and DNN are deployed for this task and their results are compared. The paper discusses the results obtained by using these models on two datasets, namely TIMIT and L2 ARCTIC. Experimental results show that transformer architecture gave 26% less phoneme error rate when compared with the base-like model of GMM-HMM. Moreover, the paper provides a fresh perspective on phoneme recognition for a more diverse range of people with a special focus on non-native English speakers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Neural Network-Based Phoneme Recognition Techniques for Non-native English Speakers

  • Adyatoni Biswas,
  • Puja Bharati,
  • Aniket Aitawade,
  • Sabyasachi Chandra,
  • Debolina Pramanik,
  • Shyamal Kumar Das Mandal

摘要

Phoneme recognition is one of the important real-life applications of automatic speech recognition. It is a speech-processing technology that transcribes and analyzes spoken language by breaking it into small distinct sounds, known as phonemes. In recent years, DNN has seen significant improvement in ASR. In this paper, transformer-based architecture is implemented to perform phoneme recognition. Additionally, GMM-HMM and DNN are deployed for this task and their results are compared. The paper discusses the results obtained by using these models on two datasets, namely TIMIT and L2 ARCTIC. Experimental results show that transformer architecture gave 26% less phoneme error rate when compared with the base-like model of GMM-HMM. Moreover, the paper provides a fresh perspective on phoneme recognition for a more diverse range of people with a special focus on non-native English speakers.