Multi-Axial Analysis of Clinical Reasoning in Large Language Models: Inter-Verifier Disagreement and Its Implications for Automated Evaluation
摘要
Evaluating clinical reasoning in large language models (LLMs) poses two open challenges: reference-oriented semantic metrics do not directly assess whether a model’s stated diagnosis is supported by the evidence in its own justification, and the increasingly popular LLM-as-judge approach rests on a largely untested assumption—that independent verifier LLMs agree with one another. We assess three generator LLMs (HuatuoGPT-o1-8B, Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.3-70B-Instruct) on 1,000 MIMIC-IV hospital-stay cases along four complementary axes (medical concept grounding, semantic similarity, semantic uncertainty, and evidence–conclusion coherence), with coherence judged independently by three frontier verifiers (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.4 mini). Two findings emerge. First, coherence reveals a dissociation that reference-oriented metrics do not capture: a model can score well on those axes yet still produce rationales that do not support its own conclusions. Second, inter-verifier agreement on coherence is consistently low (Fleiss’