<p>Identifying discriminative gene signatures from high-dimensional microarray data remains a major challenge in precision cancer diagnostics, where datasets containing thousands of genes but limited samples challenge traditional feature selection approaches. This study introduces a hybrid symbolic regression–guided knowledge distillation framework that integrates a neural teacher–student architecture with genetic programming for mathematically transparent and compact gene selection. A Multi-Layer Perceptron (MLP) teacher model generates confidence-weighted soft labels that guide a symbolic regression process to evolve algebraic expressions identifying the most discriminative genes. The resulting signatures allow lightweight student and classical models to achieve competitive accuracy using only a few features. All training, tuning, and evaluation steps were performed within a rigorous nested 5-fold cross-validation framework to ensure unbiased estimation of generalization performance. Extensive experiments across five benchmark microarray datasets demonstrate extreme dimensionality reduction (99.85–99.98%), selecting on average 4 to 7 genes while maintaining high classification accuracy (86.7–95.6%) across nine machine learning algorithms. Compared to SVM-RFE, Elastic Net, and ANOVA F-test baselines, the proposed approach achieved up to 7.4% higher accuracy, demonstrating superior feature selection quality. Calibration analysis further showed excellent reliability (ECE &lt; 0.05) for key classifiers after isotonic correction. These results highlight the framework’s ability to combine mathematical transparency, robustness, and computational efficiency for microarray-based cancer classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Symbolic regression-guided knowledge distillation for interpretable gene selection in cancer classification

  • Sara Sfaksi,
  • Leila Djerou

摘要

Identifying discriminative gene signatures from high-dimensional microarray data remains a major challenge in precision cancer diagnostics, where datasets containing thousands of genes but limited samples challenge traditional feature selection approaches. This study introduces a hybrid symbolic regression–guided knowledge distillation framework that integrates a neural teacher–student architecture with genetic programming for mathematically transparent and compact gene selection. A Multi-Layer Perceptron (MLP) teacher model generates confidence-weighted soft labels that guide a symbolic regression process to evolve algebraic expressions identifying the most discriminative genes. The resulting signatures allow lightweight student and classical models to achieve competitive accuracy using only a few features. All training, tuning, and evaluation steps were performed within a rigorous nested 5-fold cross-validation framework to ensure unbiased estimation of generalization performance. Extensive experiments across five benchmark microarray datasets demonstrate extreme dimensionality reduction (99.85–99.98%), selecting on average 4 to 7 genes while maintaining high classification accuracy (86.7–95.6%) across nine machine learning algorithms. Compared to SVM-RFE, Elastic Net, and ANOVA F-test baselines, the proposed approach achieved up to 7.4% higher accuracy, demonstrating superior feature selection quality. Calibration analysis further showed excellent reliability (ECE < 0.05) for key classifiers after isotonic correction. These results highlight the framework’s ability to combine mathematical transparency, robustness, and computational efficiency for microarray-based cancer classification.