<p>Parkinson’s disease (PD) is a progressive neurological disorder that often leads to impairments in speech (dysarthria) and facial expressiveness (hypomimia). These signs are frequently used in clinical assessments, yet their evaluation remains largely observational and subject to specialist interpretation. Furthermore, current approaches typically address these impairments independently, despite their shared neurophysiological origin. This study introduces a multimodal deep learning framework that jointly analyzes audio and video data to detect patterns associated with dysarthria and hypomimia. Two independent neural networks are trained to extract deep representations from facial gestures and voice signals during communication exercises. These representations are then integrated to enhance the detection of PD-related alterations. The approach was evaluated in a retrospective setting with synchronized audio-visual recordings from 14 subjects (7 with PD and 7 controls), who performed standardized pronunciation tasks. The proposed approach achieved an area under the receiver operating characteristic (ROC) curve—commonly referred to as AUC—of 85.39% in the phoneme pronunciation task, outperforming unimodal systems by up to 17.67%. These findings suggest that combining neurologically synchronized audio and visual cues can improve the sensitivity of automated PD screening tools.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating synchronized dysarthria and hypomimia deep patterns to quantify Parkinson disease

  • Brayan Valenzuela,
  • John Archila,
  • John Arevalo,
  • William Contreras,
  • Fabio Martinez

摘要

Parkinson’s disease (PD) is a progressive neurological disorder that often leads to impairments in speech (dysarthria) and facial expressiveness (hypomimia). These signs are frequently used in clinical assessments, yet their evaluation remains largely observational and subject to specialist interpretation. Furthermore, current approaches typically address these impairments independently, despite their shared neurophysiological origin. This study introduces a multimodal deep learning framework that jointly analyzes audio and video data to detect patterns associated with dysarthria and hypomimia. Two independent neural networks are trained to extract deep representations from facial gestures and voice signals during communication exercises. These representations are then integrated to enhance the detection of PD-related alterations. The approach was evaluated in a retrospective setting with synchronized audio-visual recordings from 14 subjects (7 with PD and 7 controls), who performed standardized pronunciation tasks. The proposed approach achieved an area under the receiver operating characteristic (ROC) curve—commonly referred to as AUC—of 85.39% in the phoneme pronunciation task, outperforming unimodal systems by up to 17.67%. These findings suggest that combining neurologically synchronized audio and visual cues can improve the sensitivity of automated PD screening tools.