Integrating synchronized dysarthria and hypomimia deep patterns to quantify Parkinson disease
摘要
Parkinson’s disease (PD) is a progressive neurological disorder that often leads to impairments in speech (dysarthria) and facial expressiveness (hypomimia). These signs are frequently used in clinical assessments, yet their evaluation remains largely observational and subject to specialist interpretation. Furthermore, current approaches typically address these impairments independently, despite their shared neurophysiological origin. This study introduces a multimodal deep learning framework that jointly analyzes audio and video data to detect patterns associated with dysarthria and hypomimia. Two independent neural networks are trained to extract deep representations from facial gestures and voice signals during communication exercises. These representations are then integrated to enhance the detection of PD-related alterations. The approach was evaluated in a retrospective setting with synchronized audio-visual recordings from 14 subjects (7 with PD and 7 controls), who performed standardized pronunciation tasks. The proposed approach achieved an area under the receiver operating characteristic (ROC) curve—commonly referred to as AUC—of 85.39% in the phoneme pronunciation task, outperforming unimodal systems by up to 17.67%. These findings suggest that combining neurologically synchronized audio and visual cues can improve the sensitivity of automated PD screening tools.