Diabetes is one of the chronic diseases that is required to be diagnosed and treated immediately because of its health hazards. The enhancement of the need for accurate and efficient prediction models has however been occasioned by increasing incidences of diabetes in the global market. In this research, patterns of different machine learning techniques as applied in predicting the likelihood of a person being a diabetic are determined. We utilized a data set that is taken from the patient’s electronic health records; age, blood pressure, glucose level, BMI were the predictors. When the data was being pre-processed, features had been selected, features normalized while missing values were also dealt with. Thus, in the context of performance comparison of different ML, we trained together with the contrast of prognosis models of Diabetes as Decision Trees, RandomForest, SVM, NeuralNet, and LogReg. The analytical models were developed using cross validation technique and hyper parameters were tuned for best results of all the models. The models were assessed based on metrics such as accuracy, precision, recall, F1-measure, and AUC-ROC. The results found out that with an ensemble of learning models like the Random Forest and boosting techniques, the method predicted higher accuracy than the classical approaches. Consequently, it was also revealed that such factors as BMI and glucose protected the model from being over predicted. So, in the framework of the specified research, it is concluded that machine learning models can benefit in the assessment of diabetes so as to make timely diagnostics and individualised therapy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning Approches for the Prediction of Diabetes

  • Venkata Bhujangarao Madamanchi,
  • A. Nagamuruganandam,
  • C. P. Chandran,
  • S. Rajathi

摘要

Diabetes is one of the chronic diseases that is required to be diagnosed and treated immediately because of its health hazards. The enhancement of the need for accurate and efficient prediction models has however been occasioned by increasing incidences of diabetes in the global market. In this research, patterns of different machine learning techniques as applied in predicting the likelihood of a person being a diabetic are determined. We utilized a data set that is taken from the patient’s electronic health records; age, blood pressure, glucose level, BMI were the predictors. When the data was being pre-processed, features had been selected, features normalized while missing values were also dealt with. Thus, in the context of performance comparison of different ML, we trained together with the contrast of prognosis models of Diabetes as Decision Trees, RandomForest, SVM, NeuralNet, and LogReg. The analytical models were developed using cross validation technique and hyper parameters were tuned for best results of all the models. The models were assessed based on metrics such as accuracy, precision, recall, F1-measure, and AUC-ROC. The results found out that with an ensemble of learning models like the Random Forest and boosting techniques, the method predicted higher accuracy than the classical approaches. Consequently, it was also revealed that such factors as BMI and glucose protected the model from being over predicted. So, in the framework of the specified research, it is concluded that machine learning models can benefit in the assessment of diabetes so as to make timely diagnostics and individualised therapy.