Asymmetric dual pathway fusion for histology driven spatial transcriptomic prediction
摘要
Spatial transcriptomics (ST) measures gene expression at spatial resolution across a tissue section, but its routine clinical use is constrained by high cost, low throughput and tissue consumption. Predicting spatially resolved gene expression directly from the haematoxylin and eosin (H&E) slides already produced by routine pathology offers a scalable alternative. Existing histology-to-ST predictors, however, rely on a single visual encoder that struggles to resolve fine-grained cellular morphology and broader tissue architecture simultaneously—two complementary scales that jointly shape transcriptional state. We introduce AsymST, a dual-pathway framework in which a lightweight convolutional encoder and a pathology foundation model are coupled by an asymmetric cross-attention block: the foundation-model tokens selectively absorb local texture cues from the convolutional branch, but the convolutional tokens are not updated in return, so neither branch homogenises the other. A deep latent information bottleneck provides the main regression pathway, while shallow per-branch auxiliary heads provide explicit branch-level supervision. Across 10 cancer cohorts and 9 spot- and slide-based baselines, AsymST attains the highest mean per-gene Pearson correlation (0.429), ranks first on 9 of 10 cohorts, and remains significantly better than each of the four strongest baselines in a cohort-level Wilcoxon test that treats cancer cohort as the independent unit (BH-adjusted