Enabling Transcriptome Classification in Micro-cohorts with Pathway-Anchoring and Single-Subject Studies
摘要
Ninety percent of the 65,000 known diseases are infrequent, collectively affecting over 400 million people. However, low prevalence limits patient accrual, hindering transcriptome classification. Micro-cohorts thus pose a transcriptomic classification challenge due to high dimensionality and small sample sizes, making machine learning prone to overfitting and requiring large datasets (>100 subjects/group; one sample per subject). We hypothesize that two methods can enable classification in micro-cohorts: i) transcriptome dynamics between two paired samples (e.g., cancer vs. unaffected tissue), and ii) single-subject (N-of-1-pathways) analytics that yield more informative features: ~4,000 pathways, their effect size, and their significance. Applying these methods to breast cancer patients with either TP53 or PIK3CA mutations (vs. unaffected paired tissue) identifies 21 pathways, achieving an unseen-set precision of 90% and recall of 90% (n = 27 training; 9 tests): increases of 8.8% and 6% respectively as compared to one transcriptome per subject. These results underscore the potential of integrating human-interpretable transcriptome pathway dynamics to enhance signal-to-noise in rare diseases, overcoming micro-cohort-size limits.