Developing an interpretable machine learning model for easily detecting insulin resistance among breast cancer survivors: a cross-sectional study
摘要
To develop and validate a classification model for insulin resistance in female individuals who have survived breast cancer using easily obtainable clinical and demographic features.
MethodsData were obtained from the U.S. National Health and Nutrition Examination Survey (NHANES) spanning 1999 to March 2020. A total of 340 female individuals who have survived breast cancer were included, and participants were randomly assigned to a training set (n = 239) and a testing set (n = 101). Multiple machine learning algorithms were trained, including Logistic Regression, Random Forest, and Support Vector Machine. Model performance was evaluated using area under the receiver operating characteristic curve (AUC) and decision curve analysis (DCA).
ResultsAll models demonstrated strong classification performance in the testing set, with AUC values exceeding 0.87. Among them, the Random Forest and Support Vector Machine models showed superior performance in DCA. Of the seven input features—body mass index, fasting blood glucose, triglyceride, HDL cholesterol, poverty income ratio, race, and education—fasting blood glucose had the highest positive feature importance for classifying insulin resistance.
ConclusionsThis study demonstrates the feasibility of using machine learning algorithms to accurately predict insulin resistance in individuals who have survived breast cancer with a limited set of clinical and demographic variables. The Random Forest and Support Vector Machine models, in particular, offer strong classification performance and may support clinicians in early identification and management of insulin resistance among individuals in this high-risk population.