Feature selection for heterogeneous data based on weighted KNN-neighborhood rough set model and self information
摘要
This paper proposes a heterogeneous data feature selection method based on weighted KNN-neighborhood rough set model and self-information. First, weighted KNN-neighborhood rough set model is constructed for heterogeneous data. This model integrates the KNN-rule and rough set theory, enabling effective handling of heterogeneous information systems that include continuous, ordinal, and nominal data. Next, four decision self-information measures are developed as feature evaluation functions, and their performance is verified through numerical experiments. Experimental results show that possible decision self-information performs best in feature selection, as it considers both the upper and lower approximations of the decision and exhibits the largest variation with changes in the feature subset, which helps identify the optimal feature subset. Based on possible decision self-information, a feature selection algorithm is designed and evaluated on multiple real-world datasets. The experimental results show that the algorithm is statistically significantly better than the other six feature selection algorithms. Statistical analysis indicates that RSIFS performs better than the other six algorithms on both classifiers, with average improvements of 18.86% in accuracy and 13.28% in F1 score.