Evaluation of DeepSeek-R1 for Ophthalmic Diagnosis and Reasoning: A Comparison with OpenAI o1 and o3
摘要
DeepSeek-R1, an open-source reasoning large language model (LLM) clinically deployed in Chinese hospitals, still lacks validation in ophthalmology.
AimsTo compare DeepSeek-R1 against OpenAI’s o1 and upgraded o3 models in diagnostic accuracy and reasoning capability across diverse ophthalmic conditions.
MethodsWe evaluated 98 standardized case vignettes covering13 ophthalmic sub-specialties, each supplied with an expert-validated diagnostic hierarchy, differential list, and reasoning chain. Model performance was assessed with a diagnosis matrix focused on final-diagnosis (FDx) accuracy; incorrect outputs were resubmitted with key diagnostic clues (reasoning-augmented, RA prompt) to test self-correction. Reasoning capacity was quantified by the number/score of diagnostic clues retrieved per case across 13 predefined domains.
ResultsDeepSeek-R1 achieved an 87.8% FDx accuracy, comparable to o3 (91.8%, P = .34) and higher than o1 (58.2%, P < .001). Similar trends were observed for others accuracy (global P < .001). Agreement was moderate–high between R1 and o3 (κ = 0.42–1.00), but slight with o1 (κ = 0.12–0.32). R1 and o3 identified more diagnostic clues than o1 (median count = 4 vs. 3, median score = 100 vs. 80; P < .001). RA prompts corrected 50.0%, 62.5% and 41.5% of FDx errors for R1, o3, and o1, raising FDx accuracy to 93.9%, 96.9%, and 80.6% respectively.
ConclusionsDeepSeek-R1 matched o3 and outperformed o1 in diagnostic accuracy and reasoning, retrieving nearly all expert-defined clues. Its open-source nature, low cost and strong performance support its use as a practical aid for ophthalmic decision-making.