Unfolding of Feature Contributions for Diabetes and Non-diabetes Predictions: A Multi-classifier Approach Using LIME
摘要
In the early-stage diagnosis of diabetes for a patient, it is important to construct a significant framework for accurate prediction of an instance and improve patient outcomes. In this study, we craft a novel framework where it analyzes the significant contribution of features viz. level of Glucose, Insulin, BMI in the body, no of times Pregnancies, Age, and others through the usage of Expandable Artificial Intelligence (XAI) techniques such as Local Interpretable Model-Agnostic Explanations (LIME). We utilize the Pima Indians Diabetes dataset to examine the feature contributions on predictions of diabetes and non-diabetes instances across widely used machine learning classifiers: Multi-layer Perceptron (MLP), K-Nearest Neighbors (KNN), Random Forest (RF), Support Vector Machine (SVM), Gaussian Naive Bayes (GNB), Decision Tree (DT), and Linear Discriminant Analysis (LDA). The results expose consistent influential patterns with the following features: Glucose, BMI, Age, and Pregnancy to become the most crucial factors in almost all Classifiers. The LDA and GNB model, for instance, illustrates stronger dependence on fewer important features compared to others, which share feature significance more consistently. Our findings from this comparative analysis provide potential insights into the selection of precise models for diabetes prediction and optimizing their utilities in medical decision-making.