Identification of key clinical features and development of a machine learning model for screening-based detection of diabetic retinopathy
摘要
Diabetic retinopathy (DR) is a major cause of preventable visual impairment. Routinely collected clinical information may help prioritize patients for retinal assessment where immediate fundus examination is not universally available. Cross-sectional identification of DR at screening must, however, be distinguished from prediction of future disease progression. To develop and internally validate machine-learning models based on routine tabular clinical variables for identifying prevalent DR at community screening. This cross-sectional study included 1,741 adults with type 2 diabetes from a community screening program in Changzhi, China. The primary outcome was prevalent DR at screening, defined as mild non-proliferative DR or worse in one prespecified study eye. Fundus photographs were used only to establish the ophthalmologist-graded reference outcome; raw images and image-derived features were not used as predictors. Twenty-two routine clinical predictors were analyzed. Participants were divided at the patient level into a training set (n = 1,219) and a held-out test set (n = 522). Elastic-net logistic regression, random forest, and gradient boosting were compared using 10-fold cross-validation within the training set. Test-set evaluation included the area under the receiver operating characteristic curve (AUC), area under the precision-recall curve (AUPRC), Brier score, calibration, threshold-dependent performance, permutation importance, and a complete-known-case sensitivity analysis. DR was identified in 470 of 1,741 participants (27.0%). The random forest showed the highest training-set cross-validated AUC (mean 0.896, standard deviation 0.031) and achieved a test-set AUC of 0.889 (95% confidence interval 0.852–0.925), AUPRC of 0.834, and Brier score of 0.101. Elastic-net logistic regression and gradient boosting achieved test-set AUCs of 0.872 and 0.887, respectively, compared with 0.666 for an HbA1c-only logistic model. Observed DR prevalence was 9.5%, 45.6%, and 93.3% in the exploratory probability strata < 0.3, 0.3 to < 0.6, and > = 0.6. Hypertension, diabetic nephropathy, and diabetes duration ranked highest by test-set permutation importance. In the complete-known-case sensitivity analysis, the random forest AUC was 0.865. A random forest model based on routinely available clinical variables showed good internal discrimination for identifying prevalent DR at screening. Calibration was not perfect, and the model does not predict future DR progression. External validation, recalibration, and prospective evaluation are required before clinical implementation.