Multimodal machine learning for Parkinson’s disease diagnosis and genetic subtyping
摘要
Parkinson’s disease (PD) is a complex neurodegenerative disorder influenced by genetic, clinical, and lifestyle factors. While machine learning (ML) has demonstrated potential in supporting PD diagnosis, existing approaches often rely on single-modality data and lack interpretability for individualized risk estimation. This study aimed to develop and validate an interpretable multimodal ML framework that enhances PD diagnosis and genetic subtype classification, supporting precision health and clinical decision-making.
MethodsWe integrated clinical assessments, voice-derived acoustic features, and genetic variant data to construct a multimodal prediction framework. Random Forest, XGBoost, and LightGBM models were trained on clinical and voice data, while a separate classifier was trained on genetic variants for subtype identification. Stacked generalization was employed using logistic regression and support vector machine (SVM) as meta-learners. Final predictions were aggregated using weighted ensemble stacking, with AUC-informed contributions. Hierarchical clustering was applied to genetic data to explore subtype patterns. Model interpretability was examined using SHAP analysis.
ResultsThe proposed ensemble model achieved strong predictive performance, with AUCs exceeding 0.96 for PD diagnosis. SHAP analysis highlighted clinically relevant predictors including UPDRS scores, tremor severity, and voice frequency features. Genetic clustering revealed subtype-specific variants, including GBA1 and SNCA (linked to cognitive decline) and LRRK2 (associated with early-onset PD).
ConclusionsThis interpretable and integrative framework provides accurate PD prediction and meaningful genetic subtype classification. It offers a practical foundation for AI-driven clinical decision support, aligning with personalized care initiatives and advancing the role of health technology in neurodegenerative disease management.