To cope with ever increasing scientific literature, Literature-Based Discovery (LBD) automatically identifies previously unknown connections within research papers, based on co-occurrence patterns. Word Embeddings (WE), the foundational technology that underlies Large Language Models, convert text into continuous vector representations that encode semantic relationships between entities in the text, based on context and co-occurrence. This work is an initial step in studying the effectiveness of WE in realizing LBD. We (a) measure the ability of different types of WE models in capturing existing functional relations (FR) between genes, diseases and chemicals in the medical literature; and then (b) use the best performing model to discover previously unknown FR from the literature (LBD). Preliminary results on PubMed abstracts for (a) show that WE from PubMedBERT yields the highest average precision (0.92) but exceedingly low recall rates (max. 0.1), which will be addressed in future work. Regarding LBD (b), using a time-sliced set of PubMed abstracts up to 2022, WE were able to capture 37% of FR that were unknown at that point (i.e., not found in a curated set of FRs).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Using Word Embeddings to Extract Semantic Relations from Biomedical Texts: Towards Literature-Based Discovery

  • William Van Woensel,
  • Sushumna S. Pradeep,
  • Ali Daowd,
  • Samina Abidi,
  • Syed Sibte Raza Abidi

摘要

To cope with ever increasing scientific literature, Literature-Based Discovery (LBD) automatically identifies previously unknown connections within research papers, based on co-occurrence patterns. Word Embeddings (WE), the foundational technology that underlies Large Language Models, convert text into continuous vector representations that encode semantic relationships between entities in the text, based on context and co-occurrence. This work is an initial step in studying the effectiveness of WE in realizing LBD. We (a) measure the ability of different types of WE models in capturing existing functional relations (FR) between genes, diseases and chemicals in the medical literature; and then (b) use the best performing model to discover previously unknown FR from the literature (LBD). Preliminary results on PubMed abstracts for (a) show that WE from PubMedBERT yields the highest average precision (0.92) but exceedingly low recall rates (max. 0.1), which will be addressed in future work. Regarding LBD (b), using a time-sliced set of PubMed abstracts up to 2022, WE were able to capture 37% of FR that were unknown at that point (i.e., not found in a curated set of FRs).