Harnessing Pre-trained Language Models for Efficient Move Recognition in Biomedical Abstracts
摘要
Over the years, the literature in the field of biomedical science has increased exponentially. Structured literature helps readers to grasp the content efficiently. Segmenting abstract sentences into categories like background, objective, method, results, and conclusions will benefit readers and will aid information extraction and information retrieval in complex biomedical literature. This paper focuses on randomized clinical trials (RCTs) and presents a comparative study of fine-tuned BERT-based variants for move recognition on the RCMR_RCT dataset, a subset of RCMR 280k. Among the models used in the experiment, BioBERT has outperformed BERT, SciBERT, PubMedBERT, BioMedBERT, BART, and RoBERTa. Utilizing the RCMR_RCT dataset instead of the PubMed 20k RCT dataset using BERT, BioBERT, SciBERT, and PubMedBERTmodel demonstrates an improvement of 4.95%, 5.71%, 3.95%, and 4.32% in the F1 score metric. This paper contributes to advancing the biomedical literature analysis by developing methods for move recognition.