Scilinkbert: a BERT-style language model for understanding scientific texts with citations
摘要
The advent of AI for Science signals a major shift of research paradigm, where artificial intelligence can accelerate scientific discovery by automating hypothesis generation, experimental planning, and large-scale simulations on high-performance computing (HPC) infrastructure. However, handling the exponentially increasing volume of scientific literature and complex citation networks requires advanced computational infrastructures. We present SciLinkBERT, a BERT-style language model trained on PubMed articles that incorporates cross-document citation information to enhance performance on scientific NLP tasks. Leveraging high-performance computing (HPC) resources for parallel and distributed model training, SciLinkBERT effectively processes large-scale citation networks. Our model outperforms state-of-the-art baselines, achieving a 0.54% improvement over BioLinkBERT on BLURB, 1.47% on MedQA, and gains of 1.89%, 0.44%, and 1.89% on named entity recognition, relation extraction, and coreference resolution, respectively, with an additional 0.86% improvement on GENIA. These results indicate that our model can offer a foundational solution for scientific literature analysis, thereby contributing to the realization of AI for Science.