<p>With the continuous growth of data scale and dimensionality, effective feature selection from massive datasets has become crucial in data mining and machine learning. As an effective mathematical tool for processing uncertain data, rough set theory provides theoretical support for feature selection. However, when dealing with continuous data, traditional rough set methods mainly rely on single-scale discretization strategies and single-scale indicators to evaluate attribute importance, which can easily lead to information loss and reduced reduction accuracy. To overcome these limitations, this paper proposes a multi-scale information-fusion feature selection algorithm. Continuous data are discretized at multiple scales, with conditional entropy and dependency calculated for each scale. Then, a joint evaluation index is designed, and a dynamic weight adaptive-adjustment mechanism is introduced to achieve effective reduction of continuous data. Experiments on UCI datasets demonstrate that the proposed method outperforms existing algorithms in feature reduction performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A feature selection method based on multi-scale information fusion

  • Chunyuan Chen,
  • LiangKun Wu,
  • Ming Yin,
  • Feng Yin

摘要

With the continuous growth of data scale and dimensionality, effective feature selection from massive datasets has become crucial in data mining and machine learning. As an effective mathematical tool for processing uncertain data, rough set theory provides theoretical support for feature selection. However, when dealing with continuous data, traditional rough set methods mainly rely on single-scale discretization strategies and single-scale indicators to evaluate attribute importance, which can easily lead to information loss and reduced reduction accuracy. To overcome these limitations, this paper proposes a multi-scale information-fusion feature selection algorithm. Continuous data are discretized at multiple scales, with conditional entropy and dependency calculated for each scale. Then, a joint evaluation index is designed, and a dynamic weight adaptive-adjustment mechanism is introduced to achieve effective reduction of continuous data. Experiments on UCI datasets demonstrate that the proposed method outperforms existing algorithms in feature reduction performance.