Construction and validation of a machine learning model to predict the risk of nasopharyngeal carcinoma using multimodal clinical data: a single-center, retrospective study
摘要
Early detection and treatment of nasopharyngeal carcinoma (NPC) are critical for improving patient prognosis. The aim of this study is to develop and compare multiple machine learning (ML) models using multimodal clinical data to identify a predictive model for NPC risk, increase diagnostic accuracy, and guide personalized treatment strategies.
MethodsClinical data were retrospectively collected from 1337 patients suspected of having NPC at the First People’s Hospital of Yulin. Feature selection was performed using the least absolute shrinkage and selection operator (LASSO) regression. Patients were divided into training and test sets (80:20 ratio), and seven ML models were developed based on the training set. Model performance was assessed using metrics such as the area under the receiver operating characteristic curve (AUC), sensitivity, and specificity. The best-performing model was further evaluated through decision curve analysis (DCA), calibration, and learning curves. SHapley Additive exPlanations (SHAP) were used to interpret key clinical features.
ResultsSeven models were developed using 17 clinical features selected from 53 parameters. The gradient boosting decision tree (GBDT) model demonstrated superior performance (AUC of 0.95 in the training cohort and 0.82 in the validation cohort). Calibration curves and DCA confirmed the model’s strong accuracy and clinical benefit. SHAP analysis revealed that age, lymphocyte percentage, serum albumin, sex, and EBV IgM were the five most significant predictors of NPC risk.
ConclusionThe GBDT-based ML model, using multimodal clinical data, accurately identifies patients at high risk for NPC, providing a valuable tool for early screening and personalized treatment strategies.