Biological Large Language Models
摘要
Beyond traditional NLP applications, LLMs are now being applied to healthcare, biological modeling, and the biopharmaceutical industry, where their ability to maintain long context and understand complex relationships within datasets is proving immensely valuable. In a clinical context, LLMs are driving personalized medicine and medical education. By training on vast amounts of clinical data, these models can assist in identifying patterns and providing decision support, ultimately enhancing patient care. From clinical note generation to conversational agents for mental health, LLMs are reshaping how language-based interactions occur in healthcare and other critical fields. Their role in synthesizing and summarizing medical literature is also becoming indispensable for clinicians and researchers. Efficient extraction of insights from unstructured text (for instance, electronic health records) will become a significant part of daily clinical practice in the near future. In biological modeling, LLMs are being applied to DNA, RNA, and protein sequences, treating these biological molecules as a form of language. Models trained on genomic and proteomic datasets are uncovering intricate gene regulation patterns, predicting protein structures, and designing novel biomolecules, showcasing the potential of LLMs in computational biology. By integrating molecular, genomic, and clinical data, LLMs facilitate the identification of drug targets, prediction of molecular interactions, and accelerating drug discovery. This chapter explores the multifaceted applications of LLMs, with a focus on their impact in healthcare, biological modeling, and biopharmaceutical research. We will review the technical foundations of biological language models, the practical applications, and lessons learned in the development of such models. Finally, we will conclude the chapter by presenting topics that will be the focus of upcoming research in the field.