Feature selection is a key step before using any classifiers to handle the problem of high-dimensional data in order to gain better results. This paper develops a hybrid feature selection paradigm to attain strong subsets of features from a high-dimensional dataset to build robust predictive models. The paradigm first constructs a random forest model for a given high-dimensional dataset. Then, all features and their frequency are taken from the forest model. Next, we select the top 5% features with the highest frequency. Finally, we use the list of top 5% features with respect to seven traditional feature selection techniques to retrieve seven different lists of top 1% features based on techniques’ criteria, respectively. These seven retrieved lists are evaluated by four classifiers: SVM, KNN, DT, and NB. We conducted experiments on eight high-dimensional datasets to evaluate these lists. The empirical results have indicated that classifiers with top 1% features are more robust than those with top 5% features and the whole features. Moreover, feature lists selected by wrapper methods are better than those selected by filter methods in our paradigm.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hybrid Feature Selection Paradigm for High-Dimensional Data

  • Hai NguyenDuc,
  • KieuAnh VuThi,
  • Dai NguyenDinhVi,
  • Tamer Z. Emara,
  • Thanh Trinh

摘要

Feature selection is a key step before using any classifiers to handle the problem of high-dimensional data in order to gain better results. This paper develops a hybrid feature selection paradigm to attain strong subsets of features from a high-dimensional dataset to build robust predictive models. The paradigm first constructs a random forest model for a given high-dimensional dataset. Then, all features and their frequency are taken from the forest model. Next, we select the top 5% features with the highest frequency. Finally, we use the list of top 5% features with respect to seven traditional feature selection techniques to retrieve seven different lists of top 1% features based on techniques’ criteria, respectively. These seven retrieved lists are evaluated by four classifiers: SVM, KNN, DT, and NB. We conducted experiments on eight high-dimensional datasets to evaluate these lists. The empirical results have indicated that classifiers with top 1% features are more robust than those with top 5% features and the whole features. Moreover, feature lists selected by wrapper methods are better than those selected by filter methods in our paradigm.