<p>Named Entity Recognition (NER) plays a crucial role in extracting important information such as treatment methods, symptoms, and herbal prescriptions from Traditional Chinese Medicine (TCM) electronic medical records. However, existing NER methods often struggle with the complexity and variability of TCM language, especially when dealing with overlapping or nested entities. To address these issues, we propose DG-SpanTCM, a novel framework that enhances character-level text understanding using a pre-trained language model and improves entity recognition through lexical-semantic features and robust training strategies. Our method also incorporates techniques to handle label imbalance and better identify complex entity structures. Experiments on a real-world TCM dataset show that DG-SpanTCM achieves superior performance, improving the F1-score over strong baseline models. These findings highlight the potential of DG-SpanTCM in advancing automated information extraction for TCM texts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Nested named entity recognition in traditional Chinese medicine electronic medical records via dual-granularity feature augmentation and span classification

  • Luo Minghao,
  • Ye Qing,
  • Cheng Chunlei,
  • LIn Peng

摘要

Named Entity Recognition (NER) plays a crucial role in extracting important information such as treatment methods, symptoms, and herbal prescriptions from Traditional Chinese Medicine (TCM) electronic medical records. However, existing NER methods often struggle with the complexity and variability of TCM language, especially when dealing with overlapping or nested entities. To address these issues, we propose DG-SpanTCM, a novel framework that enhances character-level text understanding using a pre-trained language model and improves entity recognition through lexical-semantic features and robust training strategies. Our method also incorporates techniques to handle label imbalance and better identify complex entity structures. Experiments on a real-world TCM dataset show that DG-SpanTCM achieves superior performance, improving the F1-score over strong baseline models. These findings highlight the potential of DG-SpanTCM in advancing automated information extraction for TCM texts.