<p>Deep learning in combination with multiwavelet (MW) techniques is introduced to phoneme recognition in the under-resourced Nyishi language in this study. On both Nyishi and English phonemes across datasets, the proposed model is demonstrated to be high accurate, This research begins by the initial acoustic spectrum analysis of Nyishi word syllables which revealed the range of syllable structures. The speaker dependent factors, noise interference and speech variability are addressed in this paper in terms of challenges in Nyishi phoneme recognition. A comprehensive acoustic analysis of the Nyishi speech corpus is presented describing the syllable composition, vowels, consonants and voice onset time. The corpus, being collected from native speakers, is manually transcribed and annotated at phoneme and syllable levels. It is found that the phoneme error rate of the proposed DNN-Multiwavelet model outperforms the currently available methods against Nyishi datasets. This work represents an important use of deep learning and multiwavelet methods to phoneme recognition in low resource languages, achieving accuracy of 92%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep learning and multiwavelet approach for Nyishi phoneme recognition: acoustic analysis and model development

  • Likha Ganu,
  • Biri Arun

摘要

Deep learning in combination with multiwavelet (MW) techniques is introduced to phoneme recognition in the under-resourced Nyishi language in this study. On both Nyishi and English phonemes across datasets, the proposed model is demonstrated to be high accurate, This research begins by the initial acoustic spectrum analysis of Nyishi word syllables which revealed the range of syllable structures. The speaker dependent factors, noise interference and speech variability are addressed in this paper in terms of challenges in Nyishi phoneme recognition. A comprehensive acoustic analysis of the Nyishi speech corpus is presented describing the syllable composition, vowels, consonants and voice onset time. The corpus, being collected from native speakers, is manually transcribed and annotated at phoneme and syllable levels. It is found that the phoneme error rate of the proposed DNN-Multiwavelet model outperforms the currently available methods against Nyishi datasets. This work represents an important use of deep learning and multiwavelet methods to phoneme recognition in low resource languages, achieving accuracy of 92%.