Feature (variable) selection methods are used to detect the most important features (variables) within high-dimensional data. Tree-based models, such as Random Forests, are often exploited for this purpose, as they provide a built-in mechanism to quantify feature importance. However, the stochastic sampling strategies used in these models can lead to unstable feature importance rankings, particularly when the number of trees is low. In our study, we investigate the extent to which these unstable feature rankings can be consolidated through rank aggregation and consensus signal techniques. We propose to compute consensus values from multiple feature selection runs, where each run generates a ranked list of features. We have evaluated our approach while varying a spectrum of hyperparameters such as the number of trees and the number of features available for splitting a node. Our results suggest that consensus ranks provide a more accurate and robust selection of features compared to single-run feature selection procedures. The proposed approach is especially relevant for biomarker discovery, as it can improve the accuracy and reliability of feature selection, leading to the identification of the most informative and relevant biomarkers associated with a particular disease or condition. By consolidating the results from multiple feature selection runs, our approach may help to overcome problems associated with noisy or complex data because it can provide more robust and accurate estimates of the feature importance rankings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving the Reliability of Tree-Based Feature Importance via Consensus Signals

  • Bastian Pfeifer,
  • Michael G. Schimek

摘要

Feature (variable) selection methods are used to detect the most important features (variables) within high-dimensional data. Tree-based models, such as Random Forests, are often exploited for this purpose, as they provide a built-in mechanism to quantify feature importance. However, the stochastic sampling strategies used in these models can lead to unstable feature importance rankings, particularly when the number of trees is low. In our study, we investigate the extent to which these unstable feature rankings can be consolidated through rank aggregation and consensus signal techniques. We propose to compute consensus values from multiple feature selection runs, where each run generates a ranked list of features. We have evaluated our approach while varying a spectrum of hyperparameters such as the number of trees and the number of features available for splitting a node. Our results suggest that consensus ranks provide a more accurate and robust selection of features compared to single-run feature selection procedures. The proposed approach is especially relevant for biomarker discovery, as it can improve the accuracy and reliability of feature selection, leading to the identification of the most informative and relevant biomarkers associated with a particular disease or condition. By consolidating the results from multiple feature selection runs, our approach may help to overcome problems associated with noisy or complex data because it can provide more robust and accurate estimates of the feature importance rankings.