<p>Parkinson’s Disease (PD) is a progressive neurodegenerative disorder that significantly impacts motor and speech functions, substantially reducing quality of life. Early and accurate diagnosis remains challenging due to subtle initial symptoms and the absence of definitive diagnostic tests. This exploratory study investigates the application of machine learning (ML) techniques for non-invasive PD detection through comprehensive voice data analysis. A dataset comprising 195 voice recordings from 31 individuals (23 with PD, 8 healthy controls) was systematically analyzed, incorporating diverse vocal metrics including fundamental frequency variations, amplitude perturbations, noise-to-harmonic ratios, and nonlinear dynamical features. Eight ML algorithms were rigorously evaluated: XGBoost, Random Forest, KNN, Gradient Boosting, AdaBoost, Logistic Regression, SVM, and Decision Tree. A strategic preprocessing pipeline was implemented, applying the Synthetic Minority Oversampling Technique (SMOTE) first to address inherent class imbalance while preserving original feature relationships, followed by Principal Component Analysis (PCA) for dimensionality reduction. This SMOTE then PCA sequence ensured synthetic sample generation in the original feature space, maintaining biological relevance of vocal biomarkers. XGBoost demonstrated superior performance, achieving 94.92% accuracy with the combined preprocessing approach, followed by Random Forest (93.22%) and KNN (89.83%). Comprehensive feature importance analysis identified Pitch Period Entropy (PPE), Recurrence Period Density Entropy (RPDE), and spread1 as critical biomarkers for PD differentiation. The SMOTE first methodology enhanced model sensitivity for PD case detection while maintaining high specificity through preservation of natural vocal feature relationships. These preliminary findings demonstrate the potential of integrating voice analysis with sophisticated ML preprocessing techniques for developing cost-effective, non-invasive PD diagnostic tools. However, the limited sample size necessitates validation on larger, more diverse datasets before clinical translation, emphasizing the exploratory nature of this foundational research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Non-invasive detection of Parkinson’s disease using voice analysis and machine learning techniques

  • Majdoubi Oumaima,
  • Benba Achraf,
  • Jilbab Abdelilah,
  • Hammouch Ahmed

摘要

Parkinson’s Disease (PD) is a progressive neurodegenerative disorder that significantly impacts motor and speech functions, substantially reducing quality of life. Early and accurate diagnosis remains challenging due to subtle initial symptoms and the absence of definitive diagnostic tests. This exploratory study investigates the application of machine learning (ML) techniques for non-invasive PD detection through comprehensive voice data analysis. A dataset comprising 195 voice recordings from 31 individuals (23 with PD, 8 healthy controls) was systematically analyzed, incorporating diverse vocal metrics including fundamental frequency variations, amplitude perturbations, noise-to-harmonic ratios, and nonlinear dynamical features. Eight ML algorithms were rigorously evaluated: XGBoost, Random Forest, KNN, Gradient Boosting, AdaBoost, Logistic Regression, SVM, and Decision Tree. A strategic preprocessing pipeline was implemented, applying the Synthetic Minority Oversampling Technique (SMOTE) first to address inherent class imbalance while preserving original feature relationships, followed by Principal Component Analysis (PCA) for dimensionality reduction. This SMOTE then PCA sequence ensured synthetic sample generation in the original feature space, maintaining biological relevance of vocal biomarkers. XGBoost demonstrated superior performance, achieving 94.92% accuracy with the combined preprocessing approach, followed by Random Forest (93.22%) and KNN (89.83%). Comprehensive feature importance analysis identified Pitch Period Entropy (PPE), Recurrence Period Density Entropy (RPDE), and spread1 as critical biomarkers for PD differentiation. The SMOTE first methodology enhanced model sensitivity for PD case detection while maintaining high specificity through preservation of natural vocal feature relationships. These preliminary findings demonstrate the potential of integrating voice analysis with sophisticated ML preprocessing techniques for developing cost-effective, non-invasive PD diagnostic tools. However, the limited sample size necessitates validation on larger, more diverse datasets before clinical translation, emphasizing the exploratory nature of this foundational research.