Automatic symptom detection from electronic health records is a valuable source for event-based surveillance systems. In this study, we develop tools to automatically detect symptoms associated with febrile illnesses in electronic health records written in Spanish. Therefore, we use a custom corpus, comprising 6228 expertly labeled and approximately 1 million unlabeled health reports. Our approach involved fine-tuning state-of-the-art named entity recognition models, including BiLSTM-CRF and transformer-based models like RoBERTa. We focused on domain-adaptive and task-adaptive models to enhance performance: the former were pretrained on biomedical corpora, while the latter were further pretrained on our unlabeled health reports. Despite computational constraints, our models demonstrated promising results, with RoBERTa-Clinico, a task-adaptive transformer model pretrained in our unlabeled corpus, showing the best micro recall performance (79.30), and 70.83 micro F1 score, which are comparable to results in similar studies. In this way, we contribute to the limited body of work in BioNLP in Spanish.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Information Extraction from Electronic Health Records Written in Spanish for Epidemic Intelligence

  • Javier Petri,
  • Pilar Barcena Barbeira,
  • Viviana Cotik

摘要

Automatic symptom detection from electronic health records is a valuable source for event-based surveillance systems. In this study, we develop tools to automatically detect symptoms associated with febrile illnesses in electronic health records written in Spanish. Therefore, we use a custom corpus, comprising 6228 expertly labeled and approximately 1 million unlabeled health reports. Our approach involved fine-tuning state-of-the-art named entity recognition models, including BiLSTM-CRF and transformer-based models like RoBERTa. We focused on domain-adaptive and task-adaptive models to enhance performance: the former were pretrained on biomedical corpora, while the latter were further pretrained on our unlabeled health reports. Despite computational constraints, our models demonstrated promising results, with RoBERTa-Clinico, a task-adaptive transformer model pretrained in our unlabeled corpus, showing the best micro recall performance (79.30), and 70.83 micro F1 score, which are comparable to results in similar studies. In this way, we contribute to the limited body of work in BioNLP in Spanish.