Identifying Patients at High Risk for Colorectal Carcinoma Using the Electronic Health Record
摘要
Colorectal cancer (CRC) is the fourth most common and second deadliest cancer in the US. Screening is effective at reducing CRC incidence and mortality, but rates of screening remain suboptimal. There are no sensitive machine learning models for accurately identifying individuals at risk for colorectal cancer or precancerous polyps.
ObjectivesThe aim of our study was to develop and validate a novel machine learning model that uses multimodal Electronic Health Record (EHR) data, including the most recent complete blood count (CBC), basic metabolic panel (BMP), ICD codes, and medications, to estimate a patient’s likelihood of having CRC or an advanced precancerous lesion.
MethodsWe developed ColAI, an L1-regularized logistic regression model trained on 1-year trailing EHR data, to predict CRC or advanced adenoma at screening colonoscopy. Labs are treated as continuous variables, while ICD codes and medications are represented as binary indicators of presence. ColAI was trained using 87,825 screening colonoscopies and validated using 21,957 independent colonoscopies between August 1, 2020, and March 31, 2024, from the NYU Langone Health system.
ResultsColAI achieved an AUROC of 0.93 for CRC and 0.98 for CRC or advanced adenoma. Performance remained consistent across different hospitals and time periods within NYU Langone, demonstrating strong generalizability. Performance also remained consistent between first and follow-up colonoscopies, decreasing concern for selection bias.
ConclusionsColAI accurately identifies patients at elevated risk for CRC using only routine EHR data. It has the potential to enhance targeted outreach to high-risk, unscreened individuals and improve early cancer detection at the population level.