A data-driven machine learning model for effective diabetes diagnosis
摘要
Diabetes is a prevalent long-lasting disease marked by high blood glucose due to inadequate insulin secretion and insulin resistance may result in serious life-threatening complications. Diabetes global prevalence has raised by fourfold over the past thirty years, and is the ninth foremost disease leads to death across the globe. Meanwhile, developments in machine learning presents new opportunities for prediction and classification of disease. However, despite of numerous existing models, a need for classifying types of diabetes still remains. The objective of this study is to develop an integrated, data-driven machine learning model for predicting the occurrence and classification of diabetes in an effort to improve on these limitations and investigate the ability to differentiate between types of diabetes through machine learning approach. To evaluate the proposed framework for classification of diabetes and its subtypes, publicly accessible diabetes-related datasets were employed for model development and evaluation. Binary classification is used for detecting occurrence of diabetes, while multiclass classification was employed for subtype classification, namely Prediabetes(PD), Type 1 Diabetes(T1D), Type 2 Diabetes(T2D), and Pancreatogenic (Type 3c-T3cD) Diabetes. K-Nearest Neighbors (KNN), Logistic Regression, Naive Bayes, Random Forest, and XGBoost machine learning algorithms were implemented. The XGBoost demonstrated the highest performance among all models with an accuracy of 0.97. Its feature importance scores validated predictive accuracy and to identify key factors that distinguish types of diabetes. The proposed model is intended to serve as a decision-support system for screening and classification tool using routine clinical data to facilitate early diagnosis and treatment planning.