Risk prediction of multiple cardiometabolic diseases in middle-aged and older adults using lipid metabolism biomarkers and derived composite measures based on machine learning: a prospective cohort study of CHARLS
摘要
Cardiometabolic multimorbidity (CMM) has become an increasing global public health challenge. In China, the prevalence of CMM is rising rapidly among middle-aged and older adults, with estimates ranging from 11.6% to 16.9%, posing a substantial burden on both individuals and healthcare systems. However, effective tools for predicting individual risk of CMM remain limited, hindering timely prevention and intervention.
MethodsThis study used data from the China Health and Retirement Longitudinal Study (CHARLS) between 2011 and 2015, including 7,913 participants aged ≥ 45 years without CMM at baseline. Incident CMM events were identified during the 2015 follow-up based on self-reported diagnoses of cardiometabolic diseases. Ten lipid metabolism biomarkers and derived composite indices (TC, TG, LDL-C, HDL-C, TyG, TyG-BMI, LAP, CTI, non-HDL-C, and RC) were evaluated. Predictive models were estimated using logistic regression, random forest, gradient boosting machine, eXtreme Gradient Boosting (XGBoost), support vector machine, naïve Bayes, deep learning (DL), and an ensemble model. The dataset was randomly split into training (75%) and validation (25%) subsets. Model discrimination was assessed using ROC curves and Area Under the Curve (AUC); calibration was evaluated with calibration plots and Brier scores; classification performance was examined using confusion matrices. Decision curve analysis (DCA) and clinical impact curves (CIC) were applied to assess clinical utility across risk thresholds. Feature importance ranking and SHapley Additive exPlanations (SHAP) were used to quantify variable contributions, marginal effects, and feature interactions. In addition, regional variations in CMM incidence were illustrated using choropleth maps, and correlations between lipid markers and CMM prevalence were analyzed with Pearson coefficients and heatmaps.
ResultsOver the four-year follow-up, 1,355 participants (17.1%) developed CMM. Compared with controls, incident cases were older, had a higher proportion of women and urban residents, and showed higher BMI. They also had significantly elevated triglycerides (126.6 vs. 101.8 mg/dL), reduced HDL-C (45.2 vs. 50.3 mg/dL, P < 0.001), and increased TyG-BMI and LAP (P < 0.001). Geographical analysis revealed markedly higher CMM incidence in northern cold regions (> 40%) than in southern regions (< 20%). The ensemble model achieved robust predictive performance (AUC = 0.715), followed closely by the DL model (AUC = 0.716) and GBM (AUC = 0.714). These non-linear models consistently outperformed GLM (AUC = 0.696), SVM (AUC = 0.696), and XGBoost (AUC = 0.683). Ensemble, DL, and RF models also demonstrated the best calibration (lowest Brier score, 0.125) and provided the greatest net benefit across risk thresholds. SHAP analysis indicated that composite indices, particularly TyG-BMI, LAP, and TyG, contributed most to risk prediction, whereas HDL-C exerted a protective effect. In contrast, traditional single lipid markers such as LDL-C and TC ranked lower in predictive importance.
ConclusionsThis study demonstrates that machine learning models incorporating lipid metabolism biomarkers and derived indices can predict the risk of CMM. Composite indicators such as TyG and LAP, which capture insulin resistance and visceral adiposity, showed superior predictive value. DL and ensemble models provided higher discrimination and clinical utility compared with traditional approaches. These models may enable early identification of high-risk individuals, underscoring the importance of lipid and metabolic management in CMM prevention, with potential implications for clinical decision-making and public health strategies.