<p>It is imperative to handle missing data attentively in the preprocessing stage as it may affects the integrity and quality of real-world datasets. However, existing soft clustering-based imputation neglect the underlying non-spherical separability of the data in feature space. This study proposes two robust missing data imputation (MDI) algorithms: Linear Interpolation-based Iterative Intuitionistic Fuzzy C-Means with Euclidean distance (LI-IIFCM) and its weighted variant LI-IIFCM-σ. LI-IIFCM and LI-IIFCM-σ uses linear interpolation for initial imputation followed by iterative IFCM and IFCM-σ, respectively. The approach leverages the soft Davies–Bouldin index to determine the optimal number of clusters and then iteratively refines imputations by minimizing average variation. Experimental analysis and statistical analysis (Friedman Test) on four UCI datasets, using two performance metrics, Mean Absolute Error (MAE) and Root Mean Square Error (RMSE), demonstrate that the proposed algorithms consistently outperform eight existing fuzzy clustering-based MDI algorithms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An effective imputation approach for handling missing data using intuitionistic fuzzy clustering algorithms

  • Kavita Sethia,
  • Jaspreeti Singh,
  • Anjana Gosain

摘要

It is imperative to handle missing data attentively in the preprocessing stage as it may affects the integrity and quality of real-world datasets. However, existing soft clustering-based imputation neglect the underlying non-spherical separability of the data in feature space. This study proposes two robust missing data imputation (MDI) algorithms: Linear Interpolation-based Iterative Intuitionistic Fuzzy C-Means with Euclidean distance (LI-IIFCM) and its weighted variant LI-IIFCM-σ. LI-IIFCM and LI-IIFCM-σ uses linear interpolation for initial imputation followed by iterative IFCM and IFCM-σ, respectively. The approach leverages the soft Davies–Bouldin index to determine the optimal number of clusters and then iteratively refines imputations by minimizing average variation. Experimental analysis and statistical analysis (Friedman Test) on four UCI datasets, using two performance metrics, Mean Absolute Error (MAE) and Root Mean Square Error (RMSE), demonstrate that the proposed algorithms consistently outperform eight existing fuzzy clustering-based MDI algorithms.