<p>Dysarthria, a motor speech disorder marked by irregular acoustic patterns and rapid articulatory shifts, poses significant challenges for Dysarthric Speech Recognition (DSR) systems. Although spectrogram-based methods are extensively used in current research and have demonstrated strong performance, their fixed windowing approach may not fully capture the transient dynamics of dysarthric speech. To overcome this limitation, the Hilbert Spectrum is introduced as a novel feature extraction approach that offers adaptive time–frequency analysis to capture instantaneous amplitude and frequency variations. An end-to-end system was designed to integrate this approach. Classification evaluation using Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) revealed that ViTs achieve 91% accuracy with spectrograms, which drops to 81% with Hilbert Spectrum features, while CNNs attain 95% accuracy with the Hilbert Spectrum, closely matching spectrogram-based CNN’s accuracy (96–97%). Additionally, Hidden-Unit BERT (HuBERT), which processes raw audio signals, reaches 97.27% accuracy, demonstrating the strength of direct waveform modeling. The end-to-end assistive system integrating a fine-tuned Whisper Automatic Speech Recognition (ASR) system for transcription with Google Text-to-Speech (gTTS) for synthesis further supports the approach. These results highlight the Hilbert Spectrum as a competitive alternative, offering enhanced adaptability for dysarthric speech analysis and promising improvements in assistive communication technologies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigating the Impact of Feature Extraction Techniques for Classification and Synthesis of Dysarthric Speech

  • Daniel Prem,
  • Kumudha Raimond,
  • Andrew Jeyabose

摘要

Dysarthria, a motor speech disorder marked by irregular acoustic patterns and rapid articulatory shifts, poses significant challenges for Dysarthric Speech Recognition (DSR) systems. Although spectrogram-based methods are extensively used in current research and have demonstrated strong performance, their fixed windowing approach may not fully capture the transient dynamics of dysarthric speech. To overcome this limitation, the Hilbert Spectrum is introduced as a novel feature extraction approach that offers adaptive time–frequency analysis to capture instantaneous amplitude and frequency variations. An end-to-end system was designed to integrate this approach. Classification evaluation using Vision Transformers (ViTs) and Convolutional Neural Networks (CNNs) revealed that ViTs achieve 91% accuracy with spectrograms, which drops to 81% with Hilbert Spectrum features, while CNNs attain 95% accuracy with the Hilbert Spectrum, closely matching spectrogram-based CNN’s accuracy (96–97%). Additionally, Hidden-Unit BERT (HuBERT), which processes raw audio signals, reaches 97.27% accuracy, demonstrating the strength of direct waveform modeling. The end-to-end assistive system integrating a fine-tuned Whisper Automatic Speech Recognition (ASR) system for transcription with Google Text-to-Speech (gTTS) for synthesis further supports the approach. These results highlight the Hilbert Spectrum as a competitive alternative, offering enhanced adaptability for dysarthric speech analysis and promising improvements in assistive communication technologies.