The evaluation task for syndrome differentiation thought of traditional Chinese medicine (TCM) is a complex and important research task. However, this task presents many challenges due to the complex and specialized nature of TCM texts, as well as the necessity of incorporating extensive expertise for in-depth reasoning and analysis. To address these challenges, we have developed a manually curated and reviewed high-quality TCM syndrome differentiation thought evaluation dataset. This dataset aims to provide a standardized, highly credible, and quantifiable benchmark for assessing the reasoning capabilities of large language models (LLMs) in the field of TCM. It contains 300 TCM medical records, which have been manually cleaned and annotated to ensure the rigor and applicability of the data. Furthermore, the dataset was employed in the evaluation task for syndrome differentiation thought of TCM ( http://cips-chip.org.cn/2024/eval1 ) in CHIP-2024. This task consists of four sub-tasks: clinical information extraction, TCM pathogenesis reasoning, TCM syndrome reasoning, and explanatory summary. In conclusion, this paper reviews and analyzes the methods and the results of the participating teams.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Overview of the Evaluation Task for Syndrome Differentiation Thought in Traditional Chinese Medicine in CHIP 2024

  • Meng Hao,
  • Keyu Yao,
  • Yuyan Huang,
  • Lele Yang,
  • Suyuan Peng,
  • Zhe Wang,
  • Yan Zhu

摘要

The evaluation task for syndrome differentiation thought of traditional Chinese medicine (TCM) is a complex and important research task. However, this task presents many challenges due to the complex and specialized nature of TCM texts, as well as the necessity of incorporating extensive expertise for in-depth reasoning and analysis. To address these challenges, we have developed a manually curated and reviewed high-quality TCM syndrome differentiation thought evaluation dataset. This dataset aims to provide a standardized, highly credible, and quantifiable benchmark for assessing the reasoning capabilities of large language models (LLMs) in the field of TCM. It contains 300 TCM medical records, which have been manually cleaned and annotated to ensure the rigor and applicability of the data. Furthermore, the dataset was employed in the evaluation task for syndrome differentiation thought of TCM ( http://cips-chip.org.cn/2024/eval1 ) in CHIP-2024. This task consists of four sub-tasks: clinical information extraction, TCM pathogenesis reasoning, TCM syndrome reasoning, and explanatory summary. In conclusion, this paper reviews and analyzes the methods and the results of the participating teams.