With the recent developments in AI, large language models (LLMs) have been applied to various natural language processing (NLP) tasks, including text summarization. Text summarization is defined as retrieving meaningful information from long texts. In the biomedical domain, text summarization is crucial for physicians to optimize their time and retrieve relevant information. Despite the rapid advancements in LLMs, a systematic comparative analysis of their performance in biomedical text summarization remains limited. This study addresses this gap by evaluating the effectiveness of general-purpose and domain-specific LLMs on a large-scale biomedical dataset. A dataset containing 25,000 records is acquired from PubMed in diverse medical domains, including cardiovascular diseases, gynecology, mental health, neurological disorders, and breast cancer, to conduct this study. An open-source biomedical toolkit, Ascle, is utilized to perform biomedical text summarization on the obtained dataset. Within this framework, the BART, BioBART, BigBird-Pegasus, and T5 Large models are applied, and biomedical text summarization performances are evaluated based on ROUGE, BLEU, and BERTScore. The T5 Large model outperformed the other LLMs, achieving a BERTScore of 0.86, a BLEU score of 2.95, and the highest ROUGE score(s), demonstrating superior performance in biomedical text summarization. The findings provide valuable insights into the suitability of different LLMs for biomedical text summarization, aiding researchers and healthcare professionals in selecting the most appropriate model for their needs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Intelligent Large Language Models for Biomedical Text Summarization: A Performance Evaluation

  • Amine Gonca Toprak,
  • Aytuğ Onan

摘要

With the recent developments in AI, large language models (LLMs) have been applied to various natural language processing (NLP) tasks, including text summarization. Text summarization is defined as retrieving meaningful information from long texts. In the biomedical domain, text summarization is crucial for physicians to optimize their time and retrieve relevant information. Despite the rapid advancements in LLMs, a systematic comparative analysis of their performance in biomedical text summarization remains limited. This study addresses this gap by evaluating the effectiveness of general-purpose and domain-specific LLMs on a large-scale biomedical dataset. A dataset containing 25,000 records is acquired from PubMed in diverse medical domains, including cardiovascular diseases, gynecology, mental health, neurological disorders, and breast cancer, to conduct this study. An open-source biomedical toolkit, Ascle, is utilized to perform biomedical text summarization on the obtained dataset. Within this framework, the BART, BioBART, BigBird-Pegasus, and T5 Large models are applied, and biomedical text summarization performances are evaluated based on ROUGE, BLEU, and BERTScore. The T5 Large model outperformed the other LLMs, achieving a BERTScore of 0.86, a BLEU score of 2.95, and the highest ROUGE score(s), demonstrating superior performance in biomedical text summarization. The findings provide valuable insights into the suitability of different LLMs for biomedical text summarization, aiding researchers and healthcare professionals in selecting the most appropriate model for their needs.