In traditional Chinese medicine (TCM), Large Language Models (LLMs) face challenges due to theoretical differences from modern medicine and a scarcity of specialized data. We address these with a two-stage training: continuous pre-training followed by supervised fine-tuning. Our study organize a 2GB TCM corpus consisting of pre-trained and fine-tuning datasets. In addition, we have developed Qibo-Benchmark, a tool that evaluates the performance of LLM in the TCM on multiple dimensions, including objective, and three TCM NLP tasks. The medical LLM trained with our pipeline, named Qibo, exhibits significant performance boosts. Compared to the baselines, the average objective accuracy improved by 23% to 58%, and the Rouge-L scores for the three TCM NLP tasks are 0.72, 0.61, and 0.55. Finally, we propose a pipline to apply Knowledge Graphs and LLMs to TCM consultation and demonstrate the performance through a case.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating Large Language Models with Knowledge Graphs in Traditional Chinese Medicine Consultation: A Case Study

  • Heyi Zhang,
  • Xin Wang,
  • Zhaopeng Meng,
  • Junhua Zhang,
  • Zhe Chen,
  • Pengwei Zhuang,
  • Yongzhe Jia,
  • Dawei Xu,
  • Wenbin Guo

摘要

In traditional Chinese medicine (TCM), Large Language Models (LLMs) face challenges due to theoretical differences from modern medicine and a scarcity of specialized data. We address these with a two-stage training: continuous pre-training followed by supervised fine-tuning. Our study organize a 2GB TCM corpus consisting of pre-trained and fine-tuning datasets. In addition, we have developed Qibo-Benchmark, a tool that evaluates the performance of LLM in the TCM on multiple dimensions, including objective, and three TCM NLP tasks. The medical LLM trained with our pipeline, named Qibo, exhibits significant performance boosts. Compared to the baselines, the average objective accuracy improved by 23% to 58%, and the Rouge-L scores for the three TCM NLP tasks are 0.72, 0.61, and 0.55. Finally, we propose a pipline to apply Knowledge Graphs and LLMs to TCM consultation and demonstrate the performance through a case.