<p>Single-cell foundation models such as scGPT and Geneformer are large neural networks trained on human single-cell RNA-seq data. They were never shown chronological age during training. Do their internal representations nevertheless encode aging biology in a way that can be interpreted, and how should we test whether an apparent aging signal is real biology rather than an artifact of which donors and cell types happened to be sampled?. We applied a nine-step evaluation pipeline to two foundation models (frozen, no fine-tuning) and five PBMC datasets containing 4 to 5 million cells from <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\sim\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>∼</mo> </math></EquationSource> </InlineEquation>2,000 donors with chronological age. Each step is one specific test: can we read age out of the model’s representation; does the representation place age along a clean axis; do sparse-feature decompositions surface aging-related programs; do the two models agree at the pathway level; do targeted perturbations of those features change predicted age in the expected direction; and finally, does the signal survive when we resample cells so that young and old donors have matched cell-type composition (removing the most obvious confound). (1) The foundation models encode age but do not predict it better than a 50-component PCA of gene expression: in all five cohorts the PCA baseline matches or exceeds the best foundation-model probe. What they add is a complementary interpretability mode—sparse-feature decomposition and activation-level intervention—rather than predictive power; a PCA of gene expression is itself interpretable through its loadings, so the contribution here is the evaluation framework that adjudicates such signals, not a claim that foundation models predict age better. Randomly reinitialising Geneformer’s weights destroys most of its age signal (<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(-0.107\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>-</mo> <mn>0.107</mn> </mrow> </math></EquationSource> </InlineEquation> balanced-accuracy points), while doing the same to scGPT’s layer 9 changes essentially nothing—so the two models encode age asymmetrically. (2) Sparse autoencoders surface 132 robust aging-related features across the two models, of which 193 cross-model pairs match each other at pathway level, concentrated in inflammation. The shared inflammation signal resolves into specific submodules: TNF / NF-<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\kappa\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>κ</mi> </math></EquationSource> </InlineEquation>B classical and type-II IFN-<InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\gamma\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>γ</mi> </math></EquationSource> </InlineEquation> (both models agree), complement (scGPT-specific). (3) The strongest aging signal is Geneformer’s NF-<InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(\kappa\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>κ</mi> </math></EquationSource> </InlineEquation>B program in the AIDA phase 1 v2 cohort. Pushing those features in the “older” direction increases predicted age by 0.15 expected-age units; pushing them the opposite way decreases it; pushing along random unrelated directions does neither—a three-way directional check we call the “strict gate”. When cells are resampled so that the age groups have matched cell-type composition (the strictest control), the directional effect shrinks <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(\sim\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>∼</mo> </math></EquationSource> </InlineEquation>3<InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(\times\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> but <i>7 of 8 resampling seeds still pass</i> the strict gate. One in eight resamplings fully nullifies the effect. The directional aging signal therefore survives confounder removal on most realizations, at attenuated magnitude. An external check on the Yazar OneK1K cohort (981 donors, fully separate from AIDA) reproduces the workflow on a known-strong biological axis (sex), with results within 10% of the AIDA contrast—evidence that the test is calibrated and transfers off-cohort. The paper’s primary contribution is an evaluation framework for deciding when an apparent aging signal in a single-cell foundation model is biology rather than sampling structure. Applied here, it shows that frozen foundation models carry a recoverable aging signal concentrated in NF-<InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(\kappa\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>κ</mi> </math></EquationSource> </InlineEquation>B and IFN-<InlineEquation ID="IEq9"> <EquationSource Format="TEX">\(\gamma\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>γ</mi> </math></EquationSource> </InlineEquation> inflammation submodules—biology that is already established at the gene-expression level, recovered zero-shot from models never trained on age. Reporting both an unrestricted contrast and a composition-matched contrast as side-by-side specificity tests—not just the headline number—is the framework’s central recommendation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inflammation-linked aging signals in frozen single-cell foundation models: donor-aware detection and robustness testing

  • Ihor Kendiukhov

摘要

Single-cell foundation models such as scGPT and Geneformer are large neural networks trained on human single-cell RNA-seq data. They were never shown chronological age during training. Do their internal representations nevertheless encode aging biology in a way that can be interpreted, and how should we test whether an apparent aging signal is real biology rather than an artifact of which donors and cell types happened to be sampled?. We applied a nine-step evaluation pipeline to two foundation models (frozen, no fine-tuning) and five PBMC datasets containing 4 to 5 million cells from \(\sim\) 2,000 donors with chronological age. Each step is one specific test: can we read age out of the model’s representation; does the representation place age along a clean axis; do sparse-feature decompositions surface aging-related programs; do the two models agree at the pathway level; do targeted perturbations of those features change predicted age in the expected direction; and finally, does the signal survive when we resample cells so that young and old donors have matched cell-type composition (removing the most obvious confound). (1) The foundation models encode age but do not predict it better than a 50-component PCA of gene expression: in all five cohorts the PCA baseline matches or exceeds the best foundation-model probe. What they add is a complementary interpretability mode—sparse-feature decomposition and activation-level intervention—rather than predictive power; a PCA of gene expression is itself interpretable through its loadings, so the contribution here is the evaluation framework that adjudicates such signals, not a claim that foundation models predict age better. Randomly reinitialising Geneformer’s weights destroys most of its age signal ( \(-0.107\) - 0.107 balanced-accuracy points), while doing the same to scGPT’s layer 9 changes essentially nothing—so the two models encode age asymmetrically. (2) Sparse autoencoders surface 132 robust aging-related features across the two models, of which 193 cross-model pairs match each other at pathway level, concentrated in inflammation. The shared inflammation signal resolves into specific submodules: TNF / NF- \(\kappa\) κ B classical and type-II IFN- \(\gamma\) γ (both models agree), complement (scGPT-specific). (3) The strongest aging signal is Geneformer’s NF- \(\kappa\) κ B program in the AIDA phase 1 v2 cohort. Pushing those features in the “older” direction increases predicted age by 0.15 expected-age units; pushing them the opposite way decreases it; pushing along random unrelated directions does neither—a three-way directional check we call the “strict gate”. When cells are resampled so that the age groups have matched cell-type composition (the strictest control), the directional effect shrinks \(\sim\) 3 \(\times\) × but 7 of 8 resampling seeds still pass the strict gate. One in eight resamplings fully nullifies the effect. The directional aging signal therefore survives confounder removal on most realizations, at attenuated magnitude. An external check on the Yazar OneK1K cohort (981 donors, fully separate from AIDA) reproduces the workflow on a known-strong biological axis (sex), with results within 10% of the AIDA contrast—evidence that the test is calibrated and transfers off-cohort. The paper’s primary contribution is an evaluation framework for deciding when an apparent aging signal in a single-cell foundation model is biology rather than sampling structure. Applied here, it shows that frozen foundation models carry a recoverable aging signal concentrated in NF- \(\kappa\) κ B and IFN- \(\gamma\) γ inflammation submodules—biology that is already established at the gene-expression level, recovered zero-shot from models never trained on age. Reporting both an unrestricted contrast and a composition-matched contrast as side-by-side specificity tests—not just the headline number—is the framework’s central recommendation.