Dysarthria is a motor speech disorder frequently associated with Autism Spectrum Disorder (ASD) and other neurological conditions, resulting in impaired articulation and reduced speech clarity. This paper presents an AI-driven speech analysis and feedback system designed to detect, classify, and assist in the correction of dysarthric speech patterns. The system utilizes acoustic signal processing techniques to extract key speech features such as pitch, pauses, speech rate, and spectral properties. A rule-based classification approach evaluates these features against clinically-informed thresholds, while a pronunciation similarity module offers interactive, user-centered feedback. The framework supports both real-time and pre-recorded audio inputs, ensuring adaptability for therapeutic and educational use. It provides intuitive visualizations—such as spectrograms, waveform plots, and pitch contours—to help users understand and improve their speech. The system was evaluated through multiple functional test cases including speech recording, feature extraction, classification, pronunciation assessment, and text-to-speech synthesis. Designed for early diagnosis and guided rehabilitation, the system bridges conventional speech therapy and intelligent, technology-enabled speech correction, offering a scalable and accessible tool for assistive communication.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive Dysarthria Diagnostic System: An AI-Based Real-Time Speech Performance Tracking Approach

  • C. S. Ashwini,
  • N. Anushree Rai,
  • K. J. Chaithanya,
  • Punya,
  • M. C. Thanvi

摘要

Dysarthria is a motor speech disorder frequently associated with Autism Spectrum Disorder (ASD) and other neurological conditions, resulting in impaired articulation and reduced speech clarity. This paper presents an AI-driven speech analysis and feedback system designed to detect, classify, and assist in the correction of dysarthric speech patterns. The system utilizes acoustic signal processing techniques to extract key speech features such as pitch, pauses, speech rate, and spectral properties. A rule-based classification approach evaluates these features against clinically-informed thresholds, while a pronunciation similarity module offers interactive, user-centered feedback. The framework supports both real-time and pre-recorded audio inputs, ensuring adaptability for therapeutic and educational use. It provides intuitive visualizations—such as spectrograms, waveform plots, and pitch contours—to help users understand and improve their speech. The system was evaluated through multiple functional test cases including speech recording, feature extraction, classification, pronunciation assessment, and text-to-speech synthesis. Designed for early diagnosis and guided rehabilitation, the system bridges conventional speech therapy and intelligent, technology-enabled speech correction, offering a scalable and accessible tool for assistive communication.