Comprehensive Review of Machine Learning Applications on the DHS Dataset Across Multiple Countries
摘要
This review explores the application of machine learning (ML) techniques to the Demographic and Health Surveys (DHS) dataset from countries like Bangladesh, Indonesia, Ethiopia, and others, providing critical insights for effective health policy development. It assesses how ML enhances the analysis of varied health indicators—ranging from fertility to HIV/AIDS—by employing a broad spectrum of techniques, including logistic regression, random forests, and support vector machines for prediction and classification, as well as SMOTE for data balancing and the Boruta algorithm for feature selection. The findings demonstrate that ML algorithms not only outperform traditional statistical methods in terms of accuracy and computational efficiency but also adeptly handle missing data. The review underscores the transformative potential of ML in public health analytics, suggesting further exploration of advanced techniques like deep learning to fully leverage the high-dimensional nature of DHS data, thereby enhancing research and policy-making in public health.