Mispronunciation detection (MD) for children is a crucial Natural Language Processing problem to teach them languages and develop proper pronunciation skills. Current computer-assisted pronunciation training (CAPT) systems studied minor noise types, with a wide age range. However, the unique challenges of children’s behaviors were not addressed, in terms of the acoustic variations caused in uncontrolled environments with noises, moving around, changing distance from microphones, which introduce additional speech variations that complicate MD especially with Arabic language. In this paper, we propose the Noise-Robust Arabic Utterance Mispronunciation Detector for Children (NR-AUMDChild) that detects mispronunciation from the 2D spectrogram representation of Arabic words in noisy environments at real-time. The proposed system utilizes vision transformers (ViT) to address the various children’s behaviors and encoder-decoder codec (EnCodec) to clean numerous types of noises, including background, structure, and occlusion noises. The experimental results are obtained on a real dataset of 16 Arabic words collected from children aged between 7 and 12 years old. Our proposed NR-AUMDChild system achieved an average accuracy of 78.88% for both noise and children’s behaviors handling, improving ViT with 11.6% compared to ViT when trained on clean data augmented using traditional data augmentation techniques with an average processing time of 25ms. This supports our objective towards a reliable real-time system for children that is robust against different noise types and children’s behaviors.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Arabic Mispronunciation Detection for Children in Noisy Environments via Vision Transformers

  • Mona Sadik,
  • Ahmed ElSayed,
  • Sherin Moussa,
  • Zaki T. Fayed

摘要

Mispronunciation detection (MD) for children is a crucial Natural Language Processing problem to teach them languages and develop proper pronunciation skills. Current computer-assisted pronunciation training (CAPT) systems studied minor noise types, with a wide age range. However, the unique challenges of children’s behaviors were not addressed, in terms of the acoustic variations caused in uncontrolled environments with noises, moving around, changing distance from microphones, which introduce additional speech variations that complicate MD especially with Arabic language. In this paper, we propose the Noise-Robust Arabic Utterance Mispronunciation Detector for Children (NR-AUMDChild) that detects mispronunciation from the 2D spectrogram representation of Arabic words in noisy environments at real-time. The proposed system utilizes vision transformers (ViT) to address the various children’s behaviors and encoder-decoder codec (EnCodec) to clean numerous types of noises, including background, structure, and occlusion noises. The experimental results are obtained on a real dataset of 16 Arabic words collected from children aged between 7 and 12 years old. Our proposed NR-AUMDChild system achieved an average accuracy of 78.88% for both noise and children’s behaviors handling, improving ViT with 11.6% compared to ViT when trained on clean data augmented using traditional data augmentation techniques with an average processing time of 25ms. This supports our objective towards a reliable real-time system for children that is robust against different noise types and children’s behaviors.