<p>Our study evaluated nine machine learning algorithms using a Leprosy Management Information System dataset of 3,316 newly diagnosed leprosy patients admitted between 1985 and 2023 in Yunan Province, China. The model compared performance in a train-test split of 70 − 30 using the area under receiver operating characteristic curve (AUC-ROC), sensitivity, specificity, and F1 score. The SHapley Additive exPlanation technique ranked feature importance. Leprosy reaction was 12.85% (95% CI: 11.71–13.99%). Naive Bayesian Networks (NBN) achieved an AUC of 0.753 in training and 0.761 in testing, offering the highest sensitivity of 0.418 in training and 0.441 in testing among all models to minimize missed leprosy reaction cases. The model maintained specificity of 0.92 and F1 score of 0.4, indicating reliable performance across metrics. Extreme gradient boosting (XGB) showed superior overall discrimination (AUC: 0.824 training, 0.772 testing); however, lower sensitivity (0.251) limited clinical utility. The remaining seven models had competitive AUC with NBN. The potential key predictors revealed by interpretability analysis were sequentially ranked: family history of leprosy, education level, detection mode, source of infections, treatment regimen, bacterial index, gender, age, marital status, treatment duration, and time to initiate treatment. Bayesian model provided intuitive risk assessments for resource-limited settings, providing that model selection must weigh statistical precision against clinical needs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting leprosy reactions: a machine learning framework incorporating clinical, demographic, and healthcare system factors in China

  • Yingwu Guo,
  • Lijiao Yin,
  • Hailong Yang,
  • Xi Yang,
  • Xiufeng Yu,
  • Chunyu Zhang,
  • Lijuan Zhou,
  • Fushuai Zhao,
  • Shengqing Lu,
  • Qiaojing He,
  • Lin Han,
  • Weiwei Wang,
  • Youhong Liu,
  • Yu-Ye Li

摘要

Our study evaluated nine machine learning algorithms using a Leprosy Management Information System dataset of 3,316 newly diagnosed leprosy patients admitted between 1985 and 2023 in Yunan Province, China. The model compared performance in a train-test split of 70 − 30 using the area under receiver operating characteristic curve (AUC-ROC), sensitivity, specificity, and F1 score. The SHapley Additive exPlanation technique ranked feature importance. Leprosy reaction was 12.85% (95% CI: 11.71–13.99%). Naive Bayesian Networks (NBN) achieved an AUC of 0.753 in training and 0.761 in testing, offering the highest sensitivity of 0.418 in training and 0.441 in testing among all models to minimize missed leprosy reaction cases. The model maintained specificity of 0.92 and F1 score of 0.4, indicating reliable performance across metrics. Extreme gradient boosting (XGB) showed superior overall discrimination (AUC: 0.824 training, 0.772 testing); however, lower sensitivity (0.251) limited clinical utility. The remaining seven models had competitive AUC with NBN. The potential key predictors revealed by interpretability analysis were sequentially ranked: family history of leprosy, education level, detection mode, source of infections, treatment regimen, bacterial index, gender, age, marital status, treatment duration, and time to initiate treatment. Bayesian model provided intuitive risk assessments for resource-limited settings, providing that model selection must weigh statistical precision against clinical needs.