The dysarthric severity-level classification system serves as a valuable diagnostic tool, which enables the assessment and monitoring of the condition progression in patients, and a selection of appropriate severity-specific models for recognizing dysarthric speech-an important assistive technology. Determining the severity of dysarthria presents a considerable challenge in clinical practice, given the heterogeneous nature of speech impairments associated with this motor speech disorder. This study investigates the application of Linear Frequency Residual Cepstral Coefficients (LFRCC), which are derived from the excitation source information captured via the Linear Prediction (LP) residual signal, for classification of dysarthria severity-levels. To our knowledge, this is the first work to utilize LFRCC for this purpose. Experimental assessments were conducted on two extensively employed datasets, namely, UA-Speech and TORGO. Validation of the results was carried out using a Convolutional Neural Network (CNN) with 5-fold cross-validation and test accuracies with MFCC, LFCC, and web-scale Supervised Pretraining for Speech Recognition (WSPSR), also known as Whisper, encoder module as the baseline features. Additionally, to ensure speaker-independence, Leave-One-Speaker-Out (LOSO) experiments were conducted. Furthermore, the robustness of LFRCC features against noise was explored, encompassing both stationary and non-stationary noises at varying Signal-to-Noise Ratio (SNR)-levels. Lastly, comparative analysis of latency periods with baseline feature sets suggests the potential applicability of LFRCC in real-world scenarios for severity-level classification systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Linear Frequency Residual Cepstral Features for Dysarthria Severity Classification

  • Aditya Pusuluri,
  • Hemant A. Patil

摘要

The dysarthric severity-level classification system serves as a valuable diagnostic tool, which enables the assessment and monitoring of the condition progression in patients, and a selection of appropriate severity-specific models for recognizing dysarthric speech-an important assistive technology. Determining the severity of dysarthria presents a considerable challenge in clinical practice, given the heterogeneous nature of speech impairments associated with this motor speech disorder. This study investigates the application of Linear Frequency Residual Cepstral Coefficients (LFRCC), which are derived from the excitation source information captured via the Linear Prediction (LP) residual signal, for classification of dysarthria severity-levels. To our knowledge, this is the first work to utilize LFRCC for this purpose. Experimental assessments were conducted on two extensively employed datasets, namely, UA-Speech and TORGO. Validation of the results was carried out using a Convolutional Neural Network (CNN) with 5-fold cross-validation and test accuracies with MFCC, LFCC, and web-scale Supervised Pretraining for Speech Recognition (WSPSR), also known as Whisper, encoder module as the baseline features. Additionally, to ensure speaker-independence, Leave-One-Speaker-Out (LOSO) experiments were conducted. Furthermore, the robustness of LFRCC features against noise was explored, encompassing both stationary and non-stationary noises at varying Signal-to-Noise Ratio (SNR)-levels. Lastly, comparative analysis of latency periods with baseline feature sets suggests the potential applicability of LFRCC in real-world scenarios for severity-level classification systems.