<p>High-throughput sequencing has generated large-scale genomic, transcriptomic, and multimodal datasets, but converting these data into biological interpretation remains difficult because of dimensionality, sparsity, noise, and incomplete ground truth. Deep learning offers flexible representation-learning frameworks for modeling local sequence motifs, long-range dependencies, relational structures, and latent cellular states. In this narrative review, we summarize convolutional, recurrent, graph-based, generative, and transformer architectures and examine representative applications in regulatory genomics, RNA structure prediction, single-cell and spatial transcriptomics, pathology-linked multimodal inference, and proteome-related prediction tasks. We emphasize that these models mainly produce statistical predictions, learned representations, and candidate regulatory signals; they can support functional hypotheses but do not establish biological function without independent validation. We also discuss recurring limitations, including dataset and annotation bias, domain shift, tokenization choices, computational cost, limited reproducibility, incomplete uncertainty quantification, and weak causal identifiability. Finally, we discuss artificial-intelligence virtual cells as an emerging conceptual objective rather than an established capability. Overall, deep learning is framed as a powerful tool for pattern discovery and hypothesis generation, whose biological interpretation requires careful benchmarking, external validation, and experimental follow-up.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep learning as a bridge between sequencing-derived biological data and functional interpretation

  • Haopeng Hong,
  • Lunhui Liu,
  • Wanzhe Liao,
  • Jiaxu Zhao,
  • Yongzhuo Liu,
  • Xuan He,
  • Jichang Li,
  • Runzhu He,
  • Yongcheng Jin,
  • Lei Sha,
  • Shanwen Chen,
  • Shuai Zuo,
  • Pengyuan Wang

摘要

High-throughput sequencing has generated large-scale genomic, transcriptomic, and multimodal datasets, but converting these data into biological interpretation remains difficult because of dimensionality, sparsity, noise, and incomplete ground truth. Deep learning offers flexible representation-learning frameworks for modeling local sequence motifs, long-range dependencies, relational structures, and latent cellular states. In this narrative review, we summarize convolutional, recurrent, graph-based, generative, and transformer architectures and examine representative applications in regulatory genomics, RNA structure prediction, single-cell and spatial transcriptomics, pathology-linked multimodal inference, and proteome-related prediction tasks. We emphasize that these models mainly produce statistical predictions, learned representations, and candidate regulatory signals; they can support functional hypotheses but do not establish biological function without independent validation. We also discuss recurring limitations, including dataset and annotation bias, domain shift, tokenization choices, computational cost, limited reproducibility, incomplete uncertainty quantification, and weak causal identifiability. Finally, we discuss artificial-intelligence virtual cells as an emerging conceptual objective rather than an established capability. Overall, deep learning is framed as a powerful tool for pattern discovery and hypothesis generation, whose biological interpretation requires careful benchmarking, external validation, and experimental follow-up.