Explainable machine learning and multi-layer transcriptomic analysis identify and validate prognostic biomarkers for breast cancer survival
摘要
The inherent biological heterogeneity of tumors continues to pose a challenge for achieving accurate survival status prediction in breast cancer. To address this, the current study seeks to improve the prognostic accuracy by developing an integrated framework that synthesizes clinical parameters with specific genetic variables derived from the comprehensive METABRIC (Molecular Taxonomy of Breast Cancer International Consortium) dataset.
MethodsSeveral machine learning (ML) models were applied under overall survival month (OSM)-dropped scenarios to prevent information leakage and develop clinically applicable predictive models. A diverse array of feature selection strategies, including SelectKBest and an integrated SHapley Additive exPlanations-XGBoost (SHAP-XGB) framework, were implemented. The limma package was used to identify differentially expressed genes (DEGs) with thresholds of |logFC|≥ 0.25 and adjusted p-value < 0.05. High-contribution genes associated with survival were further determined using a Leave-One-Out AUC (LOO-AUC) reduction approach. Gene Ontology (GO) enrichment analysis was performed to investigate biological pathways associated with the identified genes. Monte Carlo permutation testing, Kaplan–Meier survival analysis, multivariate Cox proportional hazards regression, and external validation using the TCGA-BRCA cohort were implemented to validate consensus biomarkers.
ResultsIntegration of clinical data with a 10-gene subset selected by the SHAP-XGB model achieved the highest predictive accuracy of 0.7133 and a balanced accuracy of 0.7217 by the Random Forest model. Integrated analyses of SHAP-XGB/RF, SelectKBest, DEG, and high-contribution gene results consistently identified six prognostic genes: JAK1, JAK2, CASP8, KIT, STAT5A, and GSK3B. GO enrichment analysis revealed significant immune-inflammatory pathways, particularly those related to cytokine production and regulation.
ConclusionThis study presents a biologically interpretable framework that integrates explainable machine learning, transcriptomic analyses, and survival modeling for breast cancer prognosis. The identified six-gene consensus signature was consistently validated through multiple independent analytical approaches, external cohort validation, and survival analyses, highlighting its potential as a promising prognostic biomarker panel for breast cancer risk stratification.