LiteCrossFusion: Lightweight Transformer-Based Multimodal Fusion with Explainability and Uncertainty Quantification for Parkinson’s Disease Detection
摘要
Parkinson’s disease is the second most common neurodegenerative disorder, and its early diagnosis remains challenging due to heterogeneous clinical and imaging manifestations. Single-modality approaches based on either neuroimaging or clinical data often fail to capture complementary information, limiting diagnostic reliability, particularly in early or atypical cases. Existing multimodal methods, while effective, are typically computationally intensive and lack interpretability and uncertainty awareness. To address these limitations, we propose LiteCrossFusion, a lightweight Transformer-based multimodal framework that integrates 3D MRI scans with a compact set of clinically relevant features. The architecture combines a 3D convolutional encoder and a parallel clinical network through a cross-attention fusion mechanism, enabling efficient representation learning with only 1.3M parameters. The framework incorporates uncertainty quantification and supports explainability through Grad-CAM for imaging and SHAP for clinical features. Experimental results on the Parkinson’s Progression Markers Initiative (PPMI) dataset demonstrate that the proposed model achieves superior performance compared to unimodal baselines, with improved class balance and discrimination capability. These findings highlight the effectiveness of lightweight multimodal fusion for accurate and interpretable Parkinson’s disease detection.