The rapid development of artificial intelligence has led to an explosion of literature in the biomedical field, and Biomedical Named Entity Recognition (BioNER) can quickly and accurately identify key information from unstructured text. This task has become an important topic to promote the rapid development of intelligence in the biomedical field. However, in the Named Entity Recognition (NER) of the biomedical field, there are always some problems of unclear boundary recognition, the underutilization of hierarchical information in sentences and the scarcity of training data resources. Based on this, this paper proposes a multi-task BioNER model based on data augmentation, using four data augmentation methods: Mention Replacement (MR), Label-wise token Replacement (LwTR), Shuffle Within Segments (SiS) and Synonym Replacement (SR) to increase the training data. The syntactic information is extracted by incorporating the input sentence into the Graph Convolutional Network (GCN), and then the tag information encoded by BERT is interacted through a co-attention mechanism to obtain an interaction matrix. Subsequently, NER is performed through boundary detection tasks and span classification tasks. Comparative experiments with other methods are conducted on the BC5CDR and JNLPBA datasets, as well as the CCKS2017 dataset. The experimental results demonstrate the effectiveness of the model proposed in this paper.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-task Biomedical Named Entity Recognition Method Based on Data Augmentation

  • Hui Zhao,
  • Di Zhao,
  • Jiana Meng,
  • Shuang Liu,
  • Hongfei Lin

摘要

The rapid development of artificial intelligence has led to an explosion of literature in the biomedical field, and Biomedical Named Entity Recognition (BioNER) can quickly and accurately identify key information from unstructured text. This task has become an important topic to promote the rapid development of intelligence in the biomedical field. However, in the Named Entity Recognition (NER) of the biomedical field, there are always some problems of unclear boundary recognition, the underutilization of hierarchical information in sentences and the scarcity of training data resources. Based on this, this paper proposes a multi-task BioNER model based on data augmentation, using four data augmentation methods: Mention Replacement (MR), Label-wise token Replacement (LwTR), Shuffle Within Segments (SiS) and Synonym Replacement (SR) to increase the training data. The syntactic information is extracted by incorporating the input sentence into the Graph Convolutional Network (GCN), and then the tag information encoded by BERT is interacted through a co-attention mechanism to obtain an interaction matrix. Subsequently, NER is performed through boundary detection tasks and span classification tasks. Comparative experiments with other methods are conducted on the BC5CDR and JNLPBA datasets, as well as the CCKS2017 dataset. The experimental results demonstrate the effectiveness of the model proposed in this paper.