This paper introduces a novel approach to ranking-based feature selection. Numerous metrics exist to assess univariate filters. This contribution identifies the top best features by employing two distinct methods, computes the global weights of each feature, and sorts these weights in decreasing order for those which are common to both procedures. According to the first part, this work is a feature selection proposal and concerning the last, it is besides a feature sorting approach. The benchmark data sets originate from NIPS 2003. This approach has been tested on problems with up to 20000 features and up to 7000 samples. The test results are very promising since, in some cases the outcomes significantly outperform the raw scenario, and in others, they surpass those of an approach which may be considered to be related to the current contribution. Finally, this study opens new avenues for research, particularly in examining the importance of specific feature sorting methods for certain supervised machine learning algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Weighted Feature Ranking Merging for Supervised Machine Learning

  • Jessica Coto-Palacio,
  • Daniel Alejandro Ortiz-Tandazo,
  • Alejandro Bautista-Juárez,
  • Agustina Grangetto,
  • Kelsy Cabello-Solorzano,
  • Diana León-Castro,
  • Paola Santana-Morales,
  • Antonio J. Tallón-Ballesteros

摘要

This paper introduces a novel approach to ranking-based feature selection. Numerous metrics exist to assess univariate filters. This contribution identifies the top best features by employing two distinct methods, computes the global weights of each feature, and sorts these weights in decreasing order for those which are common to both procedures. According to the first part, this work is a feature selection proposal and concerning the last, it is besides a feature sorting approach. The benchmark data sets originate from NIPS 2003. This approach has been tested on problems with up to 20000 features and up to 7000 samples. The test results are very promising since, in some cases the outcomes significantly outperform the raw scenario, and in others, they surpass those of an approach which may be considered to be related to the current contribution. Finally, this study opens new avenues for research, particularly in examining the importance of specific feature sorting methods for certain supervised machine learning algorithms.