With the fast evolution of Internet, social media has grown in popularity as a venue for exchanging health-related information about diagnosis, medications, adverse reactions and diseases. Extracting these information from social media is a critical task for patient safety and healthcare excellence. Recently, transformer-based models have yielded competitive performance for dealing with this task. To this end, in this paper, we conduct an empirical analysis by fine-tuning five transformer models including BERT, SciBERT, EndrBERT, BioBERT, and ClinicalBERT. We further explore the ensemble learning techniques, such as majority and weighted voting methods, to combine the predictions of different models. Experimental results performed on the publicly available CADEC dataset, indicate that EndrBERT outperforms other transformer models achieving the best performance with a strict F1-score of 69.88 and a relaxed F1-score of 83.13. Ensemble learning with weighted voting also increase the performance by an average of 1.6% and 0.59% in terms of strict and relaxed F1-score, respectively. These results demonstrate the effectiveness of fine-tuning and ensembling transformer models for extracting health-related information from social media.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ensembling Transformer Models for Medical Information Extraction from Social Media: An Empirical Study

  • Hasna Fartit,
  • Ed-drissiya El-Allaly,
  • Ali Bekri

摘要

With the fast evolution of Internet, social media has grown in popularity as a venue for exchanging health-related information about diagnosis, medications, adverse reactions and diseases. Extracting these information from social media is a critical task for patient safety and healthcare excellence. Recently, transformer-based models have yielded competitive performance for dealing with this task. To this end, in this paper, we conduct an empirical analysis by fine-tuning five transformer models including BERT, SciBERT, EndrBERT, BioBERT, and ClinicalBERT. We further explore the ensemble learning techniques, such as majority and weighted voting methods, to combine the predictions of different models. Experimental results performed on the publicly available CADEC dataset, indicate that EndrBERT outperforms other transformer models achieving the best performance with a strict F1-score of 69.88 and a relaxed F1-score of 83.13. Ensemble learning with weighted voting also increase the performance by an average of 1.6% and 0.59% in terms of strict and relaxed F1-score, respectively. These results demonstrate the effectiveness of fine-tuning and ensembling transformer models for extracting health-related information from social media.