The health industry usually faces challenges in collecting data due to the sensitive nature of the information involved. Missing data can occur when patients do not provide complete responses to questionnaires or interviews. This can make it difficult to analyze the data using machine learning techniques, as these methods often require complete data. To address this issue, researchers use imputation techniques to fill in missing values. However, these techniques can introduce bias and noise into the data. In this study, we examine the effects of four major imputation techniques (MICE, HMISC, Amelia, and missForest) on various machine learning algorithms using three missing data mechanisms (MCAR, MAR, and MNAR). Our results show that the MICE technique performs better than the other techniques for most machine learning algorithms, followed by HMISC and Amelia. The missForest technique has the highest average F1 score for all three missing data mechanisms. It can be deduced that the choice of an imputation technique might depend on the missing data mechanism and the machine learning model(s) used.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Effects of Imputation Techniques on Predictive Performance of Supervised Machine Learning Algorithms

  • Faustus Domebale Maale,
  • Gabriel Asaare Okyere,
  • O. Olawale Awe

摘要

The health industry usually faces challenges in collecting data due to the sensitive nature of the information involved. Missing data can occur when patients do not provide complete responses to questionnaires or interviews. This can make it difficult to analyze the data using machine learning techniques, as these methods often require complete data. To address this issue, researchers use imputation techniques to fill in missing values. However, these techniques can introduce bias and noise into the data. In this study, we examine the effects of four major imputation techniques (MICE, HMISC, Amelia, and missForest) on various machine learning algorithms using three missing data mechanisms (MCAR, MAR, and MNAR). Our results show that the MICE technique performs better than the other techniques for most machine learning algorithms, followed by HMISC and Amelia. The missForest technique has the highest average F1 score for all three missing data mechanisms. It can be deduced that the choice of an imputation technique might depend on the missing data mechanism and the machine learning model(s) used.