Missing values are common in real-world data sets and represent a challenging problem in performing most data analytics tasks. For that reason, many data imputation techniques have been proposed in the past to fill the missing values. However, these existing techniques may not capture the characteristics of the data and mislead the data analytics techniques, resulting in inaccurate conclusions. Generative Adversarial Networks (GANs) proved to be a good technique for generating synthetic data; using GANs, synthetic examples are generated that preserve the existing values in the record. Then, these synthetic examples can be utilized to fill the missing values and capture the data characteristics better than other data imputation techniques. In this paper, we propose a framework based on Generative Adversarial Networks to impute the missing values for incomplete datasets. The performance of the framework is evaluated using two different methodologies: 1) determining the prediction error of the imputed values after introducing missing values in an otherwise complete data set, and 2) comparing the performance of a classifier trained on a post-imputed data set, which has been imputed using our proposed framework and other imputation frameworks. The proposed framework outperformed other state-of-the-art tools at high missing rates (50% and beyond) while achieving comparable results at lower missing rates. In addition, classifiers trained on the imputed data using this proposed framework lead to higher accuracy compared with some of the other baseline methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TGAIN: Missing Data Imputation for Mixed-Type Relational Datasets Using Generative Adversarial Networks

  • Ouassim Bannany,
  • Abdulhakim A. Qahtan

摘要

Missing values are common in real-world data sets and represent a challenging problem in performing most data analytics tasks. For that reason, many data imputation techniques have been proposed in the past to fill the missing values. However, these existing techniques may not capture the characteristics of the data and mislead the data analytics techniques, resulting in inaccurate conclusions. Generative Adversarial Networks (GANs) proved to be a good technique for generating synthetic data; using GANs, synthetic examples are generated that preserve the existing values in the record. Then, these synthetic examples can be utilized to fill the missing values and capture the data characteristics better than other data imputation techniques. In this paper, we propose a framework based on Generative Adversarial Networks to impute the missing values for incomplete datasets. The performance of the framework is evaluated using two different methodologies: 1) determining the prediction error of the imputed values after introducing missing values in an otherwise complete data set, and 2) comparing the performance of a classifier trained on a post-imputed data set, which has been imputed using our proposed framework and other imputation frameworks. The proposed framework outperformed other state-of-the-art tools at high missing rates (50% and beyond) while achieving comparable results at lower missing rates. In addition, classifiers trained on the imputed data using this proposed framework lead to higher accuracy compared with some of the other baseline methods.