Purpose <p>To evaluate differences in performance between two large language models (LLMs), GPT-5.3 and Gemini 2.5 Pro, in diagnostic classification and clinical reasoning for primary open-angle glaucoma (POAG).</p> Methods <p>Forty-eight guideline-based standardized cases were constructed. Using a unified prompt, we fed the cases into both models and obtained outputs on diagnosis, classification, and reasoning. A consensus of three glaucoma specialists served as the reference standard. We compared the diagnostic accuracy and classification consistency (Cohen’s κ) of the two models. Clinical reasoning ability was scored using a Likert scale based on logical coherence, evidence utilization, and conclusion consistency. We further analyzed error patterns and potential safety issues.</p> Results <p>Overall diagnostic accuracy was 85.4% for GPT-5.3 and 75.0% for Gemini (<i>P</i> = 0.306). Both models performed well on typical cases but showed reduced accuracy on borderline cases, including ocular hypertension, suspected glaucoma, and early-stage POAG. For classification consistency, κ values were 0.675 for GPT-5.3 and 0.628 for Gemini. GPT-5.3 scored higher than Gemini in overall clinical reasoning (4.4 ± 0.6 vs. 3.9 ± 0.7, <i>P</i> = 0.011). Error pattern analysis indicated that Gemini was more prone to overdiagnosis and reasoning inconsistency, whereas GPT-5.3 was relatively conservative. Both models had low rates of unsafe outputs, though Gemini showed a slightly higher proportion.</p> Conclusion <p>ChatGPT and Gemini both demonstrate certain capabilities in diagnosing POAG, but their stability on borderline cases remains limited. Comparatively, GPT-5.3 shows higher consistency and more stable reasoning patterns. The application of LLMs in ophthalmic diagnostic support still requires cautious evaluation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative performance of chatgpt and gemini in diagnostic classification and clinical reasoning for open-angle glaucoma: a standardized scenario-based study

  • Zhewen Zhang,
  • Siyu Lu,
  • Zhenqiang Xu,
  • Kangyu Ji,
  • Yan Liang

摘要

Purpose

To evaluate differences in performance between two large language models (LLMs), GPT-5.3 and Gemini 2.5 Pro, in diagnostic classification and clinical reasoning for primary open-angle glaucoma (POAG).

Methods

Forty-eight guideline-based standardized cases were constructed. Using a unified prompt, we fed the cases into both models and obtained outputs on diagnosis, classification, and reasoning. A consensus of three glaucoma specialists served as the reference standard. We compared the diagnostic accuracy and classification consistency (Cohen’s κ) of the two models. Clinical reasoning ability was scored using a Likert scale based on logical coherence, evidence utilization, and conclusion consistency. We further analyzed error patterns and potential safety issues.

Results

Overall diagnostic accuracy was 85.4% for GPT-5.3 and 75.0% for Gemini (P = 0.306). Both models performed well on typical cases but showed reduced accuracy on borderline cases, including ocular hypertension, suspected glaucoma, and early-stage POAG. For classification consistency, κ values were 0.675 for GPT-5.3 and 0.628 for Gemini. GPT-5.3 scored higher than Gemini in overall clinical reasoning (4.4 ± 0.6 vs. 3.9 ± 0.7, P = 0.011). Error pattern analysis indicated that Gemini was more prone to overdiagnosis and reasoning inconsistency, whereas GPT-5.3 was relatively conservative. Both models had low rates of unsafe outputs, though Gemini showed a slightly higher proportion.

Conclusion

ChatGPT and Gemini both demonstrate certain capabilities in diagnosing POAG, but their stability on borderline cases remains limited. Comparatively, GPT-5.3 shows higher consistency and more stable reasoning patterns. The application of LLMs in ophthalmic diagnostic support still requires cautious evaluation.