Benchmark of the Typical Case Diagnosis Consistency Evaluation Task in CHIP 2024
摘要
This article mainly introduces the medical benchmark in the Typical Case Diagnosis Consistency Evaluation Task released at the CHIP 2024 conference (see http://www.cips-chip.org.cn/2024/eval3 for details). Unlike other existing benchmarks, this benchmark is built on the basis of real medical record information screened manually and has been strictly reviewed by medical experts to maximize the diagnostic ability of large models in real clinical environments. In addition, this benchmark not only covers a variety of complex cases, but also fully considers the diversity of cases and the authenticity of diagnostic scenarios, providing an important reference for evaluating the generalization ability and practical application potential of medical large models.