<p>Large-scale genomic language models (gLMs) hold promise for modeling gene regulation, yet their ability to capture personal gene expression variations remains unresolved. We developed xDecoder, a unified decoding framework that utilizes gLMs and sequence-to-function (S2F) embeddings to learn how personal genetic variation shapes gene expression from paired genome-transcriptome data. Compared to the pretrained genomic models, xDecoder with personalized DNA–RNA training makes cross-individual prediction tractable for seen genes in a few-shot setting. However, zero-shot prediction at unseen loci remains unreliable and gene-dependent, revealing a cross-locus transfer bottleneck of current sequence models. Experiments incorporating individual-level chromatin accessibility suggested that regulatory-state information important for unseen-locus prediction is not fully captured by current DNA-only models. Overall, these results highlight the potential utility of the few-shot setting, the limitations of DNA-only models, and point toward multi-omic, variant-aware frameworks as a promising direction for building personalized regulatory models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

xDecoder unlocks the potential of genomic foundation models for few-shot personal gene expression prediction

  • Shumin Li,
  • Ruibang Luo,
  • Yuanhua Huang

摘要

Large-scale genomic language models (gLMs) hold promise for modeling gene regulation, yet their ability to capture personal gene expression variations remains unresolved. We developed xDecoder, a unified decoding framework that utilizes gLMs and sequence-to-function (S2F) embeddings to learn how personal genetic variation shapes gene expression from paired genome-transcriptome data. Compared to the pretrained genomic models, xDecoder with personalized DNA–RNA training makes cross-individual prediction tractable for seen genes in a few-shot setting. However, zero-shot prediction at unseen loci remains unreliable and gene-dependent, revealing a cross-locus transfer bottleneck of current sequence models. Experiments incorporating individual-level chromatin accessibility suggested that regulatory-state information important for unseen-locus prediction is not fully captured by current DNA-only models. Overall, these results highlight the potential utility of the few-shot setting, the limitations of DNA-only models, and point toward multi-omic, variant-aware frameworks as a promising direction for building personalized regulatory models.