<p>Dysarthria is a speech disorder caused by damage to the motor areas of the brain, resulting in speech that is effortful, slow, slurred, or disordered in rhythm and intonation. Assessing the severity of dysarthria helps pathologists plan therapy, track patient progress, and support automatic speech recognition systems. We approached the task of classifying dysarthria severity using deep learning models, including deep neural networks (DNN), convolutional neural networks (CNN), and long short-term memory networks (LSTM), applied to the standard UA Speech and TORGO databases with Mel-frequency cepstral coefficients (MFCCs). We also propose a data augmentation strategy based on multi-voice synthesis from lyrics. In the front end, a text-to-speech (TTS) converter generates synthetic speech in the target speaker’s voice. This is then processed by a style-transfer module to adjust the style of the synthesized speech. The best-performing DNN-MFCC framework achieved 97.68% and 97.16% accuracy in speaker-dependent scenarios for the UA Speech and TORGO databases, respectively. In speaker-independent scenarios, the accuracies were 40.12% and 48.07%, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker independent dysarthria severity classification using synthesis-based augmentation

  • Manju Suresh,
  • Rajeev Rajan,
  • Joshua Thomas

摘要

Dysarthria is a speech disorder caused by damage to the motor areas of the brain, resulting in speech that is effortful, slow, slurred, or disordered in rhythm and intonation. Assessing the severity of dysarthria helps pathologists plan therapy, track patient progress, and support automatic speech recognition systems. We approached the task of classifying dysarthria severity using deep learning models, including deep neural networks (DNN), convolutional neural networks (CNN), and long short-term memory networks (LSTM), applied to the standard UA Speech and TORGO databases with Mel-frequency cepstral coefficients (MFCCs). We also propose a data augmentation strategy based on multi-voice synthesis from lyrics. In the front end, a text-to-speech (TTS) converter generates synthetic speech in the target speaker’s voice. This is then processed by a style-transfer module to adjust the style of the synthesized speech. The best-performing DNN-MFCC framework achieved 97.68% and 97.16% accuracy in speaker-dependent scenarios for the UA Speech and TORGO databases, respectively. In speaker-independent scenarios, the accuracies were 40.12% and 48.07%, respectively.