Liver diseases are a major global health concern, contributing to approximately 2 million deaths annually. Early detection is critical to improving treatment outcomes and preventing severe conditions such as cirrhosis and liver cancer. This study leverages a variety of Machine Learning (ML) techniques, including Support Vector Machine (SVM), Decision Tree (DT), and ensemble methods such as XGBoost, stacking, and voting, to enhance the accuracy of liver disease classification. Two datasets were used: the Indian Liver Patient Dataset (ILPD) and the Liver Disease Patient Dataset (LDPD). Feature selection methods, Principal Component Analysis (PCA), and Recursive Feature Elimination (RFE) were employed to improve model performance. Experimental results demonstrated that ensemble stacking (DT + XGBoost) with RFE achieved the highest accuracy of 85.70% and an F1 score of 86.49% on the ILPD. On the LDPD, most models, except SVM, achieved nearly 99% accuracy and F1-score with all features and RFE. The findings suggest that the size of the dataset and feature selection methods significantly impact the performance of liver disease classification models. This study underscores the importance of combining ML techniques and feature selection to optimize classification accuracy in medical diagnostics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Liver Disease Classification Using Ensemble Learning and Feature Selection

  • Nurul Asmi Amalia,
  • Fadhilah Syafa,
  • Hafizha Dini Giandra,
  • Taufik Fuadi Abidin,
  • Rumaisa Kruba

摘要

Liver diseases are a major global health concern, contributing to approximately 2 million deaths annually. Early detection is critical to improving treatment outcomes and preventing severe conditions such as cirrhosis and liver cancer. This study leverages a variety of Machine Learning (ML) techniques, including Support Vector Machine (SVM), Decision Tree (DT), and ensemble methods such as XGBoost, stacking, and voting, to enhance the accuracy of liver disease classification. Two datasets were used: the Indian Liver Patient Dataset (ILPD) and the Liver Disease Patient Dataset (LDPD). Feature selection methods, Principal Component Analysis (PCA), and Recursive Feature Elimination (RFE) were employed to improve model performance. Experimental results demonstrated that ensemble stacking (DT + XGBoost) with RFE achieved the highest accuracy of 85.70% and an F1 score of 86.49% on the ILPD. On the LDPD, most models, except SVM, achieved nearly 99% accuracy and F1-score with all features and RFE. The findings suggest that the size of the dataset and feature selection methods significantly impact the performance of liver disease classification models. This study underscores the importance of combining ML techniques and feature selection to optimize classification accuracy in medical diagnostics.