Non-alcoholic fatty liver disease (NAFLD) is a public health challenge and a significant cause of morbidity and mortality worldwide. Machine learning models can predict NAFLD, aiding physicians in assessing the severity of patients and offering a fresh approach to diagnosis, prevention, and management of NAFLD. Classification models such as logistic regression (LR), random forest (RF), support vector machine (SVM), Naïve Bayes (NB), K-Nearest Neighbor (KNN), Multilayer Perceptron (MLP) are implemented on dataset. In the proposed method, the model agonistic approach SHAP is used for feature selection technique. This dataset is highly imbalanced, and two oversampling techniques, SMOTE and Borderline SMOTE, were implemented to balance the dataset. The AUROC score for classification models with class balance technique borderline SMOTE was LR (0.93), RF (1.00), SVM (0.97), NB (0.89), KNN (0.95), and MLP (1.00). Additionally, the accuracy of classification models is LR (98.35%), RF (100.00%), SVM (73.55%), NB (98.35%), KNN (84.30%), and MLP (100.00%). However, the random forest model and multilayer perceptron showed higher performance than other classification models. Implementation of a random forest model, multilayer perceptron in the clinical setting could help physicians to identify the severity of fatty liver patients for primary prevention, surveillance, early treatment, and management.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identification of Severity in Non-Alcoholic Fatty Liver Disease Using Machine Learning Algorithms

  • Koteswara Rao Makkena,
  • Karthika Natarajan

摘要

Non-alcoholic fatty liver disease (NAFLD) is a public health challenge and a significant cause of morbidity and mortality worldwide. Machine learning models can predict NAFLD, aiding physicians in assessing the severity of patients and offering a fresh approach to diagnosis, prevention, and management of NAFLD. Classification models such as logistic regression (LR), random forest (RF), support vector machine (SVM), Naïve Bayes (NB), K-Nearest Neighbor (KNN), Multilayer Perceptron (MLP) are implemented on dataset. In the proposed method, the model agonistic approach SHAP is used for feature selection technique. This dataset is highly imbalanced, and two oversampling techniques, SMOTE and Borderline SMOTE, were implemented to balance the dataset. The AUROC score for classification models with class balance technique borderline SMOTE was LR (0.93), RF (1.00), SVM (0.97), NB (0.89), KNN (0.95), and MLP (1.00). Additionally, the accuracy of classification models is LR (98.35%), RF (100.00%), SVM (73.55%), NB (98.35%), KNN (84.30%), and MLP (100.00%). However, the random forest model and multilayer perceptron showed higher performance than other classification models. Implementation of a random forest model, multilayer perceptron in the clinical setting could help physicians to identify the severity of fatty liver patients for primary prevention, surveillance, early treatment, and management.