Diabetes is a long-term metabolic condition that arises when the body is unable to regulate blood sugar levels effectively. The need for research, particularly utilizing machine learning (ML), is critical for several reasons. ML can analyze vast datasets to uncover patterns, leading to a more comprehensive understanding of the various factors that contribute to diabetes development. This research explores the application of four interpretable supervised machine learning models: the naïve Bayes classifier, ensemble methods, K-Nearest Neighbors (K-NN) classifier, and Support Vector Machine (SVM) classifier. This paper utilizes machine learning techniques to analyze the Pima Indian Diabetes dataset. The performance of each algorithm is evaluated to identify the one that achieves the highest accuracy, recall, precision, specificity, and F1-score. The highest accuracy of 79.17% was observed on the dataset used. The focus is on uncovering patterns and factors influencing diabetes prevalence within the Pima Indian population. The analysis of this data illustrates the real-world use of machine learning in improving diagnostics and refining treatment strategies for improved diabetes management within this specific population.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Machine Learning Techniques in Diabetes Prediction

  • Sneha Sharma,
  • Jayaprakash Vemuri,
  • Jyoti Kainthola

摘要

Diabetes is a long-term metabolic condition that arises when the body is unable to regulate blood sugar levels effectively. The need for research, particularly utilizing machine learning (ML), is critical for several reasons. ML can analyze vast datasets to uncover patterns, leading to a more comprehensive understanding of the various factors that contribute to diabetes development. This research explores the application of four interpretable supervised machine learning models: the naïve Bayes classifier, ensemble methods, K-Nearest Neighbors (K-NN) classifier, and Support Vector Machine (SVM) classifier. This paper utilizes machine learning techniques to analyze the Pima Indian Diabetes dataset. The performance of each algorithm is evaluated to identify the one that achieves the highest accuracy, recall, precision, specificity, and F1-score. The highest accuracy of 79.17% was observed on the dataset used. The focus is on uncovering patterns and factors influencing diabetes prevalence within the Pima Indian population. The analysis of this data illustrates the real-world use of machine learning in improving diagnostics and refining treatment strategies for improved diabetes management within this specific population.