<p>Transformer-based models are rapidly becoming foundational tools for analyzing and integrating multiscale biological data. This Perspective examines recent advances in transformer architectures, tracing their evolution from unimodal and augmented unimodal models to large-scale multimodal foundation models operating across genomic sequences, single-cell transcriptomics and spatial data. We categorize these models into three tiers and evaluate their capabilities for structural learning, representation transfer and tasks such as cell annotation, prediction and imputation. While discussing tokenization, interpretability and scalability challenges, we highlight emerging approaches that leverage masked modeling, contrastive learning and large language models. To support broader adoption, we provide practical guidance through code-based primers, using public datasets and open-source implementations. Finally, we propose designing a modular ‘Super Transformer’ architecture using cross-attention mechanisms to integrate heterogeneous modalities. This Perspective serves as a resource and roadmap for leveraging transformer models in multiscale, multimodal genomics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal foundation transformer models for multiscale genomics

  • Sumeer Ahmad Khan,
  • Xabier Martínez-de-Morentin,
  • Abdel Rahman Alsabbagh,
  • Alberto Maillo,
  • Vincenzo Lagani,
  • David Gomez-Cabrero,
  • Robert Lehmann,
  • Jesper Tegner

摘要

Transformer-based models are rapidly becoming foundational tools for analyzing and integrating multiscale biological data. This Perspective examines recent advances in transformer architectures, tracing their evolution from unimodal and augmented unimodal models to large-scale multimodal foundation models operating across genomic sequences, single-cell transcriptomics and spatial data. We categorize these models into three tiers and evaluate their capabilities for structural learning, representation transfer and tasks such as cell annotation, prediction and imputation. While discussing tokenization, interpretability and scalability challenges, we highlight emerging approaches that leverage masked modeling, contrastive learning and large language models. To support broader adoption, we provide practical guidance through code-based primers, using public datasets and open-source implementations. Finally, we propose designing a modular ‘Super Transformer’ architecture using cross-attention mechanisms to integrate heterogeneous modalities. This Perspective serves as a resource and roadmap for leveraging transformer models in multiscale, multimodal genomics.