This paper proposes a composite membership model based on information sets and fuzzy rough sets to address the unsupervised feature selection problem in data with uncertainty and fuzziness. First, by defining information source values, information values, and Shannon source transform entropy, an information valuation metric (IVM) is constructed to quantify the information consistency between features, effectively balance feature information and redundancy. Secondly, combined with fuzzy rough set theory, a composite membership model is designed. The membership degree of the feature subset is calculated dynamically in the model to ensure that the candidate features are highly correlated with the selected features. The algorithm is divided into two steps: candidate feature screening based on information sets, and feature determination based on the composite membership model. Experiments compare the proposed algorithm with other classic or advanced unsupervised feature selection algorithms on eight datasets. The results show that the proposed algorithm performs well in clustering and can select fewer features.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Dynamic Unsupervised Feature Selection Method Based on Information Sets and Fuzzy Rough Sets

  • Yuxin Zhao,
  • Pengfei Zhang,
  • Dexian Wang,
  • Tianrui Li

摘要

This paper proposes a composite membership model based on information sets and fuzzy rough sets to address the unsupervised feature selection problem in data with uncertainty and fuzziness. First, by defining information source values, information values, and Shannon source transform entropy, an information valuation metric (IVM) is constructed to quantify the information consistency between features, effectively balance feature information and redundancy. Secondly, combined with fuzzy rough set theory, a composite membership model is designed. The membership degree of the feature subset is calculated dynamically in the model to ensure that the candidate features are highly correlated with the selected features. The algorithm is divided into two steps: candidate feature screening based on information sets, and feature determination based on the composite membership model. Experiments compare the proposed algorithm with other classic or advanced unsupervised feature selection algorithms on eight datasets. The results show that the proposed algorithm performs well in clustering and can select fewer features.