<p>Existing RNA language models (RLMs) largely overlook structural information in RNA sequences, leading to incomplete feature extraction and suboptimal performance on downstream tasks. In this study, we present ERNIE-RNA (Enhanced Representations with Base-Pairing Restriction for RNA Modeling), an RNA pre-trained language model based on a modified BERT (Bidirectional Encoder Representations from Transformers). Notably, ERNIE-RNA’s attention maps exhibit superior ability to capture RNA structural features through zero-shot prediction, outperforming conventional methods like RNAfold and RNAstructure, suggesting that ERNIE-RNA naturally develops comprehensive representations of RNA architecture during pre-training. Moreover, after fine-tuning, ERNIE-RNA achieves state-of-the-art (SOTA) performance across various downstream tasks, including RNA structure and function predictions. In summary, ERNIE-RNA provides versatile features that can be effectively applied to a wide range of research tasks. Our findings highlight that integrating key knowledge-based priors into the BERT framework may enhance the performance of other language models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ERNIE-RNA: an RNA language model with structure-enhanced representations

  • Weijie Yin,
  • Zhaoyu Zhang,
  • Shuo Zhang,
  • Liang He,
  • Ruiyang Zhang,
  • Rui Jiang,
  • Gan Liu,
  • Jingyi Wang,
  • Xuegong Zhang,
  • Tao Qin,
  • Zhen Xie

摘要

Existing RNA language models (RLMs) largely overlook structural information in RNA sequences, leading to incomplete feature extraction and suboptimal performance on downstream tasks. In this study, we present ERNIE-RNA (Enhanced Representations with Base-Pairing Restriction for RNA Modeling), an RNA pre-trained language model based on a modified BERT (Bidirectional Encoder Representations from Transformers). Notably, ERNIE-RNA’s attention maps exhibit superior ability to capture RNA structural features through zero-shot prediction, outperforming conventional methods like RNAfold and RNAstructure, suggesting that ERNIE-RNA naturally develops comprehensive representations of RNA architecture during pre-training. Moreover, after fine-tuning, ERNIE-RNA achieves state-of-the-art (SOTA) performance across various downstream tasks, including RNA structure and function predictions. In summary, ERNIE-RNA provides versatile features that can be effectively applied to a wide range of research tasks. Our findings highlight that integrating key knowledge-based priors into the BERT framework may enhance the performance of other language models.