Medical Named Entity Recognition (MedNER) aims to automatically identify medical-related entities from large volumes of unstructured medical text, forming the foundation for the application and development of AI technologies in the medical field. However, Chinese medical text presents significant challenges due to its high level of specialization and complexity, including limited annotated corpora, dense use of domain-specific terminology, and frequent abbreviations. In recent years, Large Language Model (LLM) have achieved remarkable progress across various tasks, offering new approaches for addressing MedNER challenges. This paper proposes an enhanced MedNER method tailored for Chinese medical texts, leveraging the capabilities of LLM. First, a prompt template is designed based on chain-of-thought techniques to guide LLM in extracting preliminary knowledge from medical text. This knowledge is then injected into a MedNER model to strengthen its ability to identify medical terms and abbreviations. Additionally, considering that Chinese medical named entities often exhibit distinctive radical-based features, the proposed model integrates radical information to further improve entity recognition performance. Experiments conducted on the CCKS2019 dataset demonstrate that the proposed method significantly outperforms existing mainstream approaches in terms of precision, recall, and F1-score.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Chinese Medical Named Entity Recognition Enhanced by Large Language Model

  • Jizhao Zhu,
  • Xiaolin Lv,
  • Xinlong Pan,
  • Zhenqiu Zhu,
  • Chunlong Fan

摘要

Medical Named Entity Recognition (MedNER) aims to automatically identify medical-related entities from large volumes of unstructured medical text, forming the foundation for the application and development of AI technologies in the medical field. However, Chinese medical text presents significant challenges due to its high level of specialization and complexity, including limited annotated corpora, dense use of domain-specific terminology, and frequent abbreviations. In recent years, Large Language Model (LLM) have achieved remarkable progress across various tasks, offering new approaches for addressing MedNER challenges. This paper proposes an enhanced MedNER method tailored for Chinese medical texts, leveraging the capabilities of LLM. First, a prompt template is designed based on chain-of-thought techniques to guide LLM in extracting preliminary knowledge from medical text. This knowledge is then injected into a MedNER model to strengthen its ability to identify medical terms and abbreviations. Additionally, considering that Chinese medical named entities often exhibit distinctive radical-based features, the proposed model integrates radical information to further improve entity recognition performance. Experiments conducted on the CCKS2019 dataset demonstrate that the proposed method significantly outperforms existing mainstream approaches in terms of precision, recall, and F1-score.