Machine learning‑based risk stratification for moderate‑to‑severe anxiety‑depression in patients with hypertension and/or coronary heart disease: a multicenter cross‑sectional study
摘要
Anxiety and depression in patients with hypertension and/or coronary heart disease (CHD) are associated with severe outcomes. However, there is a lack of objective risk stratification tools. This proof‑of‑concept study developed eight machine learning (ML) models to assess the feasibility of risk stratification for moderate‑to‑severe anxiety‑depression in this population and to identify key associated factors.
MethodsThe study retrospectively included patients with hypertension and/or CHD from the 2023 Psychology and Behavior Investigation of Chinese Residents (PBICR) database, a cross‑sectional, self‑report survey. Patients were randomly split into training (70%) and test (30%) sets for internal validation. In the training set, least absolute shrinkage and selection operator (LASSO) regression with cross-validation was employed to select features for constructing the model. Eight ML models were developed for risk stratification of moderate‑to‑severe anxiety‑depression (PHQ‑9 ≥ 10 and/or GAD‑7 ≥ 10). Model performance was assessed using the area under the receiver operating characteristic curve (AUROC), calibration, and decision curve analysis (DCA).
ResultsThe final cohort comprised 4,006 eligible patients, allocated to training (n = 2,805) and test (n = 1,201) sets. LASSO regression selected 15 predictors from 43 initial features. This study established eight ML models for risk stratification of moderate-to-severe anxiety-depression in patients with hypertension and/or CHD. Random Forest (RF) demonstrated optimal discrimination in the training set (AUROC = 0.837; 95% CI: 0.812–0.862). In the test set, it maintained acceptable discrimination (AUROC = 0.838; 95% CI: 0.814–0.862; Brier score = 0.134). Family health, risk taking, and quality of life (QoL) were identified as the strongest correlates via SHAP analysis. DCA suggested net benefit across risk thresholds of 10–40%.
ConclusionsThe RF model achieved the best performance, with family health, risk taking, and QoL as the strongest correlates. The proposed two‑stage screening workflow is hypothetical and has not been prospectively validated. Due to a lack of external validation and the cross‑sectional, self‑report design, this study is hypothesis‑generating and proof‑of‑concept. Prospective validation in independent cohorts is required before any clinical application.