Introduction <p>Adverse drug reactions&#xa0;(ADRs), including those resulting from drug interactions, remain a leading cause of morbidity and mortality. Structured product labels (SPLs) serve as a primary source for drug safety information. Having machine-readable product labels, including adverse reactions (ARs) and drug interactions, readily available would allow researchers to streamline medication safety studies. However, extracting this information is complex and requires the use of natural language processing (NLP) methods.</p> Objective <p>In this study, we explored the application of generative language models in the extraction of drug safety information from SPLs.</p> Methods <p>We compared multiple generative LLMs (GPT, Llama, and Mixtral) to two baseline methods in the task of extracting adverse reactions (ARs) from SPLs. We explored various factors, such as prompting strategies and term complexity, that impact the performance of these models&#xa0;in the extraction of ARs. Finally, we explored the generative models' capacity to extract drug interactions from a separate section of SPLs without additional fine-tuning or training, demonstrating their flexibility and adaptability for information retrieval.</p> Results <p>We found that generative language models, specifically GPT-4, are able to match or exceed the performance of previous state-of-the-art models without additional training or fine-tuning. Additionally, we found that the specific SPL section, surrounding&#xa0;context, and complexity of the AR term impacted the extraction performance. Finally, we demonstrated the generalizability of these models by applying them to a separate task of extracting drug&#xa0;names from the drug interaction section where curated training data are not available.</p> Conclusion <p>Generative language models demonstrate significant potential for automating drug safety information extraction from SPLs, offering a promising avenue for improving post-market surveillance and reducing ADRs. Future work should focus on refining prompting strategies and expanding the models’ capabilities to handle increasingly complex and nuanced drug safety information.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Large Language Models in Extracting Drug Safety Information from Prescription Drug Labels

  • Undina Gisladottir,
  • Michael Zietz,
  • Sophia Kivelson,
  • Yutaro Tanaka,
  • Gaurav Sirdeshmukh,
  • Kathleen LaRow Brown,
  • Nicholas P. Tatonetti

摘要

Introduction

Adverse drug reactions (ADRs), including those resulting from drug interactions, remain a leading cause of morbidity and mortality. Structured product labels (SPLs) serve as a primary source for drug safety information. Having machine-readable product labels, including adverse reactions (ARs) and drug interactions, readily available would allow researchers to streamline medication safety studies. However, extracting this information is complex and requires the use of natural language processing (NLP) methods.

Objective

In this study, we explored the application of generative language models in the extraction of drug safety information from SPLs.

Methods

We compared multiple generative LLMs (GPT, Llama, and Mixtral) to two baseline methods in the task of extracting adverse reactions (ARs) from SPLs. We explored various factors, such as prompting strategies and term complexity, that impact the performance of these models in the extraction of ARs. Finally, we explored the generative models' capacity to extract drug interactions from a separate section of SPLs without additional fine-tuning or training, demonstrating their flexibility and adaptability for information retrieval.

Results

We found that generative language models, specifically GPT-4, are able to match or exceed the performance of previous state-of-the-art models without additional training or fine-tuning. Additionally, we found that the specific SPL section, surrounding context, and complexity of the AR term impacted the extraction performance. Finally, we demonstrated the generalizability of these models by applying them to a separate task of extracting drug names from the drug interaction section where curated training data are not available.

Conclusion

Generative language models demonstrate significant potential for automating drug safety information extraction from SPLs, offering a promising avenue for improving post-market surveillance and reducing ADRs. Future work should focus on refining prompting strategies and expanding the models’ capabilities to handle increasingly complex and nuanced drug safety information.