Stuttering is a speech disorder causing difficulties in expressing words or sentences. Although stuttering is widely acknowledged and is the focus of current research, there are no effective cures. Research centres on easing the lives of people who stutter (PWS). This paper introduces Ebana, a stuttering correction system. Ebana allows PWS to share their stuttered voice virtually as both stutter-free text and voice. In this study, an Arabic speech dataset was recorded and collected in collaboration with King Fahd Hospital of the University (KFHU). The proposed methodology is based on the automatic speech recognition (ASR) system NeuralSpace and the large language model (LLM). As the process of building Ebana includes multiple stages, each stage was evaluated independently, both objectively and subjectively. The average of the overall subjective evaluation had an MOS value of 86.1%. The objective evaluation of the ASR indicated a word error rate (WER) of 33.11%, while for the LLM this value was 12.74%. Voice cloning technology was evaluated using MCD with a resulting value of 19.32. To conclude, Ebana is a user-friendly application through which PWS can easily express themselves without the fear of being misinterpreted.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Stuttering Disfluency Correction System Using Artificial Intelligence

  • Sumayh S. Aljameel,
  • Deemah Alqahtani,
  • Danah A. Algarni,
  • Dlayel A. Aluhideb,
  • Fatema A. Alamoodi,
  • Shahad F. Aljafaari,
  • Zainab A. Alsafwani

摘要

Stuttering is a speech disorder causing difficulties in expressing words or sentences. Although stuttering is widely acknowledged and is the focus of current research, there are no effective cures. Research centres on easing the lives of people who stutter (PWS). This paper introduces Ebana, a stuttering correction system. Ebana allows PWS to share their stuttered voice virtually as both stutter-free text and voice. In this study, an Arabic speech dataset was recorded and collected in collaboration with King Fahd Hospital of the University (KFHU). The proposed methodology is based on the automatic speech recognition (ASR) system NeuralSpace and the large language model (LLM). As the process of building Ebana includes multiple stages, each stage was evaluated independently, both objectively and subjectively. The average of the overall subjective evaluation had an MOS value of 86.1%. The objective evaluation of the ASR indicated a word error rate (WER) of 33.11%, while for the LLM this value was 12.74%. Voice cloning technology was evaluated using MCD with a resulting value of 19.32. To conclude, Ebana is a user-friendly application through which PWS can easily express themselves without the fear of being misinterpreted.