The ensemble of self-information-based feature selection for heterogeneous data via k-nearest neighborhood rough set model
摘要
The kernel step of feature selection using rough set theory is to establish feature evaluation function to assess the classification ability of feature subset. Dependency in rough set theory is a feature evaluation function. However, this function considers only the classification information contained in the lower approximation of the decision while ignoring the upper approximation.k-nearest-neighbor rule (KNN-rule) is an important classification technique. By combining neighborhood rough set model with KNN-rule, KNN-neighborhood rough set model can be obtained. This model has a strong ability to approximate decision. This paper studies feature selection for heterogeneous data based on the ensemble of self-information and KNN-neighborhood rough set model. First, KNN-neighborhood rough set model for heterogeneous data is established. Using this model, decision self-information for four types of heterogeneous data is constructed. By applying linear fusion approach, the ensemble operator of self-information is derived, which can be used as feature evaluation function. This ensemble operator integrates multiple self-information measures, comprehensively considering both the upper and lower approximations of the decision boundary. Moreover, this ensemble operator balances the contributions of different self-information measures in evaluating feature subsets, accurately determines feature importance, effectively avoids feature redundancy, and significantly enhances classification accuracy and efficiency. It provides a novel pathway for the efficient processing and precise analysis of heterogeneous data. Next, a feature selection algorithm for heterogeneous data is designed based on this ensemble operator. Finally, a series of numerical experiments conducted on real-world datasets is used to evaluate the designed algorithm. The experimental results demonstrated that this algorithm offers statistically significant advantages over six other feature selection algorithms. Statistical analysis shows that SI-FFS achieves superior performance across four classifiers compared to the other six algorithms, with an average improvement of 15.95% in F1 score.