<p>Feature selection in high-dimensional data is an important part of the data mining process and is widely used in bioinformatics, statistics and image processing fields. Successfully selecting informative features can significantly improve learning accuracy and improve result comprehensibility. However, it is a challenging problem to select features accurately and efficiently from high-dimensional data. In this paper, we propose a Weighted Sparse Regression with Mutual Information (WSRMI) for selecting structural features. Differing from traditional sparse feature selection models that focus solely on either feature correlations or feature importance, the proposed model integrates both aspects through a mutual-information-based weighting mechanism. The proposed model can be effectively applied to regression and binary classification tasks, making it more general and practical for real-world applications. The proposed model is statistically compared with several existing classical models over randomly generated classification and benchmark datasets. Experimental results show that the proposed model is more effective at selecting the informative features with a superior prediction performance than the comparative ones.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Structural feature selection via weighted sparse regression with mutual information

  • Yadi Wang,
  • Yulin Xie,
  • Sufang Zhou,
  • Bingbing Jiang,
  • Hangjun Che

摘要

Feature selection in high-dimensional data is an important part of the data mining process and is widely used in bioinformatics, statistics and image processing fields. Successfully selecting informative features can significantly improve learning accuracy and improve result comprehensibility. However, it is a challenging problem to select features accurately and efficiently from high-dimensional data. In this paper, we propose a Weighted Sparse Regression with Mutual Information (WSRMI) for selecting structural features. Differing from traditional sparse feature selection models that focus solely on either feature correlations or feature importance, the proposed model integrates both aspects through a mutual-information-based weighting mechanism. The proposed model can be effectively applied to regression and binary classification tasks, making it more general and practical for real-world applications. The proposed model is statistically compared with several existing classical models over randomly generated classification and benchmark datasets. Experimental results show that the proposed model is more effective at selecting the informative features with a superior prediction performance than the comparative ones.