Purpose <p> This study aimed to establish and validate a machine learning model for predicting moderate-to-severe cancer-related fatigue (CRF) 2 years after completion of anti-tumor therapy in breast cancer patients.</p> Methods <p>Clinical and laboratory data from 183 patients were retrospectively collected. Candidate predictors were screened using multivariate logistic regression, and seven algorithms—logistic regression, decision tree, random forest, support vector machine, extreme gradient boosting (XGBoost), k-nearest neighbor, and naïve Bayes—were constructed in the training cohort and validated in the testing cohort. Model performance was assessed by discrimination, calibration, and decision curve analysis.</p> Results <p>The 2-year incidence of moderate-to-severe CRF was 54.0%. Eight independent predictors were identified, including age, body mass index, histological grade, menopausal status, hemoglobin, platelet count, neutrophil count, and systemic immune-inflammation index. Models built on these features demonstrated variable performance, with XGBoost showing the most favorable balance. It achieved an AUC of 0.983 in the training set and 0.766 in the validation set, with robust accuracy, sensitivity, and specificity. Calibration plots indicated good agreement between predicted and observed risks, while decision curve analysis confirmed higher net clinical benefit across a wide range of thresholds.</p> Conclusion <p>The XGBoost-based model provided reliable long-term CRF risk prediction, supporting early identification of high-risk patients and informing personalized survivorship care.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development and validation of a machine learning model to predict moderate-to-severe cancer-related fatigue in breast cancer

  • Zhen Liu,
  • Guoshuang Shen,
  • Miaozhou Wang,
  • Jiabin Wang,
  • Yongxin Li,
  • Jiuda Zhao

摘要

Purpose

This study aimed to establish and validate a machine learning model for predicting moderate-to-severe cancer-related fatigue (CRF) 2 years after completion of anti-tumor therapy in breast cancer patients.

Methods

Clinical and laboratory data from 183 patients were retrospectively collected. Candidate predictors were screened using multivariate logistic regression, and seven algorithms—logistic regression, decision tree, random forest, support vector machine, extreme gradient boosting (XGBoost), k-nearest neighbor, and naïve Bayes—were constructed in the training cohort and validated in the testing cohort. Model performance was assessed by discrimination, calibration, and decision curve analysis.

Results

The 2-year incidence of moderate-to-severe CRF was 54.0%. Eight independent predictors were identified, including age, body mass index, histological grade, menopausal status, hemoglobin, platelet count, neutrophil count, and systemic immune-inflammation index. Models built on these features demonstrated variable performance, with XGBoost showing the most favorable balance. It achieved an AUC of 0.983 in the training set and 0.766 in the validation set, with robust accuracy, sensitivity, and specificity. Calibration plots indicated good agreement between predicted and observed risks, while decision curve analysis confirmed higher net clinical benefit across a wide range of thresholds.

Conclusion

The XGBoost-based model provided reliable long-term CRF risk prediction, supporting early identification of high-risk patients and informing personalized survivorship care.