<p>Among the existing feature selection (FS) methods in machine learning (ML), those known as wrapper methods often produce the most effective subset of features. However, their high computational cost and tendency to overfit make them impractical for high-dimensional data, as is frequently encountered in biological research. The Wrapshap framework presents a novel FS method tailored for high-dimensional biological data, addressing challenges such as redundancy, multicollinearity, and interaction effects. By leveraging the local accuracy property of Additive Local Explanation methods of SHAP, Wrapshap translates complex biological data into interpretable insights. The framework, integrating a trained ML model, allows explanations to be viewed as a new exploratory data space containing a comprehensive understanding of the underlying complexity of the raw data structure. Benchmarks against 84 regression and 106 classification datasets demonstrate its superiority in computational speed, feature ranking quality, and predictive performance compared to state-of-the-art methods like Recursive Feature Elimination. The method’s adaptability extends to various ML models, explanation methods and tasks. By requiring only a single training of the complex model, Wrapshap mitigates the computational inefficiencies of traditional wrapper methods while maintaining performance monitoring thus opening new possibilities for real-time, interactive FS with user involvement. This innovative approach addresses a notable gap in existing methods, and promotes informed biological hypothesis generation, enhancing the interpretability and usability of ML in biomedical research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Wrapshap a high dimension feature selection framework to explore biological systems through machine learning explanations

  • Félix Furger,
  • Miguel Thomas,
  • Colas Foulon,
  • Haomio Wang,
  • Julien Aligon,
  • Emmanuel Doumard,
  • Chantal Soulé-Dupuy,
  • Cyrille Delpierre,
  • Louis Casteilla,
  • Paul Monsarrat

摘要

Among the existing feature selection (FS) methods in machine learning (ML), those known as wrapper methods often produce the most effective subset of features. However, their high computational cost and tendency to overfit make them impractical for high-dimensional data, as is frequently encountered in biological research. The Wrapshap framework presents a novel FS method tailored for high-dimensional biological data, addressing challenges such as redundancy, multicollinearity, and interaction effects. By leveraging the local accuracy property of Additive Local Explanation methods of SHAP, Wrapshap translates complex biological data into interpretable insights. The framework, integrating a trained ML model, allows explanations to be viewed as a new exploratory data space containing a comprehensive understanding of the underlying complexity of the raw data structure. Benchmarks against 84 regression and 106 classification datasets demonstrate its superiority in computational speed, feature ranking quality, and predictive performance compared to state-of-the-art methods like Recursive Feature Elimination. The method’s adaptability extends to various ML models, explanation methods and tasks. By requiring only a single training of the complex model, Wrapshap mitigates the computational inefficiencies of traditional wrapper methods while maintaining performance monitoring thus opening new possibilities for real-time, interactive FS with user involvement. This innovative approach addresses a notable gap in existing methods, and promotes informed biological hypothesis generation, enhancing the interpretability and usability of ML in biomedical research.