This article mainly introduces the medical benchmark in the Typical Case Diagnosis Consistency Evaluation Task released at the CHIP 2024 conference (see http://www.cips-chip.org.cn/2024/eval3 for details). Unlike other existing benchmarks, this benchmark is built on the basis of real medical record information screened manually and has been strictly reviewed by medical experts to maximize the diagnostic ability of large models in real clinical environments. In addition, this benchmark not only covers a variety of complex cases, but also fully considers the diversity of cases and the authenticity of diagnostic scenarios, providing an important reference for evaluating the generalization ability and practical application potential of medical large models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmark of the Typical Case Diagnosis Consistency Evaluation Task in CHIP 2024

  • Zehua Wang,
  • Hui Zong,
  • Yang Feng,
  • ZhaoRong Teng,
  • Jun Yan,
  • Shujia Jiang,
  • Buzhou Tang

摘要

This article mainly introduces the medical benchmark in the Typical Case Diagnosis Consistency Evaluation Task released at the CHIP 2024 conference (see http://www.cips-chip.org.cn/2024/eval3 for details). Unlike other existing benchmarks, this benchmark is built on the basis of real medical record information screened manually and has been strictly reviewed by medical experts to maximize the diagnostic ability of large models in real clinical environments. In addition, this benchmark not only covers a variety of complex cases, but also fully considers the diversity of cases and the authenticity of diagnostic scenarios, providing an important reference for evaluating the generalization ability and practical application potential of medical large models.