MPD-MFF: A Multimodal Parkinson’s Disease Detection Method Based on Multi-feature Fusion
摘要
This paper proposes a Parkinson’s disease detection method based on multimodal feature fusion (Multimodal Parkinson’s Detection Multi-feature Fusion, MPD-MFF), which improves the accuracy of early Parkinson’s detection by leveraging both speech and gait data. First, we introduce the HuBERT-Speech Encoder (HSE) to extract key features from Parkinson’s speech. Furthermore, the extracted speech features are fused with Mel-frequency Cepstral Coefficients (MFCC) to form a multi-feature speech fusion dataset, enhancing the accuracy of Parkinson’s speech recognition. Second, we propose a Point-wise Convolutional Depth Gait Encoder (PDGE) to extract critical features from gait data. Finally, we introduce an Enhanced Bi-LSTM (EBi-LSTM) to capture the temporal dependencies between the multi-feature speech fusion data and gait feature data. Experiments conducted on the mPower, GAIT-IT, and MaxLittle datasets show that MPD-MFF outperforms existing methods in multiple metrics such as detection accuracy, recall, and F1-score. Ablation experiments confirm the contribution of each module to the overall model performance, with EBi-LSTM demonstrating significant effectiveness in handling long sequential data, further proving the validity of the proposed method.