A calibrated multimodal stacking framework integrating machine learning and deep learning for mental health analysis
摘要
Mental health conditions are increasing globally; therefore, novel computational tools with greater capabilities to screen earlier, more accurately, and fairly are to be developed. Traditional diagnostic techniques are heavily dependent on subjective self-reports and clinical interviews, which introduce subjectivities and delays to the assessment process. In this study, a Calibrated multimodal stacking framework is introduced that integrates behavioral, physiological, demographic, and self-reported attributes to improve predictive modeling for mental health analysis. A framework is proposed with both machine learning and deep learning models as the base learners, including Support Vector Machines, Random Forest, Convolutional Neural Networks, and Long Short-Term Memory networks. Our framework is validated using stratified internal testing and domain-wise external-like evaluation with GroupKFold and Leave-One-Country-Out strategies to reduce the spatio-demographic bias. The proposed framework attained excellent performance with 96.7% validation accuracy and 96.5% test accuracy, yielding an area under the receiver operating characteristic curve of 0.982, outperforming all baseline models. The external-like validation provided the accuracy of 89.8% and area under roc curve of 0.92, which indicates strong generalized model to the real-world scenario. The proposed framework attained high visibility made it easier to reason with SHAP-LIME and also performed subgroup fairness where the demographic bias was minimal. The proposed model offers high screening capability and interpretability while aligning with the clinical decision support system, providing upscaling ability with the digital ecosystem.