With the widespread adoption of Chinese electronic medical records (CEMR), named entity recognition (NER) has become essential for medical information processing. However, due to the complexity of the Chinese language and diverse medical terminology, existing NER methods often struggle with shallow lexical fusion, limited semantic integration, and ineffective graph structure construction. To address these challenges, we propose LE-MSF, a model based on Lexicon Enhancement and Multi-semantic Fusion. We construct a medical word embedding corpus and incorporate lexicon features into the lower layers of MC-BERT via a lexicon adapter. Part-of-Speech (POS) features, word segmentation boundary features, and radical features are fused through a graph learning algorithm to build a multi-semantic fusion graph, followed by node aggregation using Multi-layer GCN. Cross-gated Feature Fusion is applied to fuse different features, and final entity recognition is performed using Bidirectional Long Short-Term Memory (BiLSTM) and Conditional Random Field (CRF). Comparative ablation experiments on three publicly available benchmark datasets for CEMR show that the LE-MSF model achieves the best performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LE-MSF: A Chinese Medical Named Entity Recognition Method Based on Lexicon Enhancement and Multi-semantic Fusion

  • Yihuan Jiang,
  • Hongyun Huang,
  • Zuohua Ding

摘要

With the widespread adoption of Chinese electronic medical records (CEMR), named entity recognition (NER) has become essential for medical information processing. However, due to the complexity of the Chinese language and diverse medical terminology, existing NER methods often struggle with shallow lexical fusion, limited semantic integration, and ineffective graph structure construction. To address these challenges, we propose LE-MSF, a model based on Lexicon Enhancement and Multi-semantic Fusion. We construct a medical word embedding corpus and incorporate lexicon features into the lower layers of MC-BERT via a lexicon adapter. Part-of-Speech (POS) features, word segmentation boundary features, and radical features are fused through a graph learning algorithm to build a multi-semantic fusion graph, followed by node aggregation using Multi-layer GCN. Cross-gated Feature Fusion is applied to fuse different features, and final entity recognition is performed using Bidirectional Long Short-Term Memory (BiLSTM) and Conditional Random Field (CRF). Comparative ablation experiments on three publicly available benchmark datasets for CEMR show that the LE-MSF model achieves the best performance.