<p>Diagnostic pathology depends on complex, structured reasoning to interpret clinical, histologic, and molecular data. Replicating, or even approximating this cognitive process algorithmically remains a significant challenge. As large language models (LLMs) gain traction in medicine, it is critical to assess their clinical utility in specialized domains such as pathology. We evaluated the quality of expressed diagnostic reasoning in the outputs of four LLMs, OpenAI o1, OpenAI o3-mini, Gemini 2.0 Flash Thinking Experimental, and DeepSeek-R1, using 15 constructed-response pathology questions covering board-examination-level content across a range of cognitive demands, from factual recall to integrative knowledge synthesis. Eleven expert pathologists independently scored model-generated responses using a structured rubric encompassing five language quality metrics (accuracy, relevance, coherence, depth, and conciseness) and seven diagnostic reasoning strategies as communicated in the textual output. Scores were normalized and compared using ANOVA and Tukey’s HSD. Gemini significantly outperformed all models in overall reasoning quality, particularly in analytical depth and coherence. While all models achieved comparable factual accuracy, only Gemini and DeepSeek consistently demonstrated expert-like reasoning patterns in their outputs, including algorithmic, inductive, and Bayesian approaches. Performance varied by reasoning type, with the highest scores in algorithmic and deductive reasoning and the lowest in heuristic and pattern recognition. Gemini achieved the highest inter-observer agreement (p &lt; 0.001), suggesting greater interpretability. However, more in-depth reasoning was generally associated with reduced conciseness. Addressing these trade-offs will be key to safe clinical integration.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reasoning beyond accuracy: expert evaluation of expressed diagnostic reasoning in large language model outputs for pathology

  • Asim Waqas,
  • Ehsan Ullah,
  • Asma Khan,
  • Farah Khalil,
  • Zarifa G. Ozturk,
  • Vaibhav Chumbalkar,
  • Daryoush Saeed-Vafa,
  • Zena Jameel,
  • Wei-Shen Chen,
  • Humberto Trejo Bittar,
  • Jasreman Dhillon,
  • Rajendra S. Singh,
  • Andrey Bychkov,
  • Anil V. Parwani,
  • Marilyn M. Bui,
  • Matthew B. Schabath,
  • Ghulam Rasool

摘要

Diagnostic pathology depends on complex, structured reasoning to interpret clinical, histologic, and molecular data. Replicating, or even approximating this cognitive process algorithmically remains a significant challenge. As large language models (LLMs) gain traction in medicine, it is critical to assess their clinical utility in specialized domains such as pathology. We evaluated the quality of expressed diagnostic reasoning in the outputs of four LLMs, OpenAI o1, OpenAI o3-mini, Gemini 2.0 Flash Thinking Experimental, and DeepSeek-R1, using 15 constructed-response pathology questions covering board-examination-level content across a range of cognitive demands, from factual recall to integrative knowledge synthesis. Eleven expert pathologists independently scored model-generated responses using a structured rubric encompassing five language quality metrics (accuracy, relevance, coherence, depth, and conciseness) and seven diagnostic reasoning strategies as communicated in the textual output. Scores were normalized and compared using ANOVA and Tukey’s HSD. Gemini significantly outperformed all models in overall reasoning quality, particularly in analytical depth and coherence. While all models achieved comparable factual accuracy, only Gemini and DeepSeek consistently demonstrated expert-like reasoning patterns in their outputs, including algorithmic, inductive, and Bayesian approaches. Performance varied by reasoning type, with the highest scores in algorithmic and deductive reasoning and the lowest in heuristic and pattern recognition. Gemini achieved the highest inter-observer agreement (p < 0.001), suggesting greater interpretability. However, more in-depth reasoning was generally associated with reduced conciseness. Addressing these trade-offs will be key to safe clinical integration.