Clustering is a fundamental and essential technique in data mining that has been widely applied in many fields. Recently, incomplete data clustering has received significant attention due to the issue of missing data in many practical applications. However, traditional incomplete data clustering methods generally rely on single imputation techniques, which neglect the inestimable uncertainty of missing data, resulting in suboptimal imputation results. In addition, existing algorithms typically treat data imputation and clustering as separate processes. This separation prevents the optimization of the imputation process to better align with the clustering task, leading to suboptimal clustering performance. To address these issues, this paper proposes a novel one-stage incomplete data clustering algorithm called IDCMIA. First, the clustering algorithm performs multiple imputation to generate multiple imputation estimates, effectively alleviating the inestimable uncertainty of missing data. Next, multiple autoencoders are used to extract the latent features of the pre-imputed data, converting the incomplete data clustering problem into a multi-view clustering problem. Finally, a KL-divergence-based clustering loss function is adopted to optimize the cluster assignment derived from the fused latent features. The experimental results show superior performance compared to state-of-the-art incomplete data clustering methods across different missing ratios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Incomplete Data Clustering Based on Multiple Imputation and Autoencoders

  • Xinyu Han,
  • Jin Zhou,
  • Shiyuan Han,
  • Tao Du,
  • Cheng Yang,
  • Bowen Liu

摘要

Clustering is a fundamental and essential technique in data mining that has been widely applied in many fields. Recently, incomplete data clustering has received significant attention due to the issue of missing data in many practical applications. However, traditional incomplete data clustering methods generally rely on single imputation techniques, which neglect the inestimable uncertainty of missing data, resulting in suboptimal imputation results. In addition, existing algorithms typically treat data imputation and clustering as separate processes. This separation prevents the optimization of the imputation process to better align with the clustering task, leading to suboptimal clustering performance. To address these issues, this paper proposes a novel one-stage incomplete data clustering algorithm called IDCMIA. First, the clustering algorithm performs multiple imputation to generate multiple imputation estimates, effectively alleviating the inestimable uncertainty of missing data. Next, multiple autoencoders are used to extract the latent features of the pre-imputed data, converting the incomplete data clustering problem into a multi-view clustering problem. Finally, a KL-divergence-based clustering loss function is adopted to optimize the cluster assignment derived from the fused latent features. The experimental results show superior performance compared to state-of-the-art incomplete data clustering methods across different missing ratios.