Benchmark evaluation of OpenAI-o1 vs DeepSeek-R1 for junior anesthesiologists in Chinese hospitals: hazards of fast large language models use in high-stakes specialties
摘要
Junior physicians may experience critical competency gaps during the transition to unsupervised practice, particularly in high-risk specialties such as anesthesiology, where errors can compromise patient safety. Although large language models (LLMs) are being rapidly integrated into healthcare, their accuracy and safety in supporting junior physicians—who may lack sufficient experience to identify model errors—remain insufficiently quantified in high-stakes clinical settings. This study evaluated OpenAI-o1 (OA) and DeepSeek-R1 in English (DSE) and Chinese (DSC) to address this evidence gap.
MethodsIn this dual-axis evaluation, 30 anesthesia crisis scenarios were developed through Delphi consensus. Responses generated by OA, DSE, and DSC were assessed by 20 experts for accuracy using a 5-point Likert scale and for clinical logicality using an 8-point Situation–Background–Assessment–Recommendation (SBAR) scoring system, with each component rated from 0 to 2 points. Practicality was assessed by 20 junior physicians using a 5-point Likert scale across three subdimensions: step clarity, guideline applicability, and learning assistance. Accuracy and practicality were the two primary evaluation domains, while clinical logicality, qualitative feedback, and readability were examined as complementary outcomes. Performance was compared across nine pre-specified, order-constrained hypotheses using Bayes Factor Design Analysis (n = 600 per group), with outcomes expressed as posterior probabilities (PPs) and Bayes Factors (BFs).
ResultsOA demonstrated superior accuracy (OA > DSE > DSC; PP = 0.94; BF₄ᵤ = 4.70, strong evidence), whereas DSC showed greater practicality (OA < DSE < DSC; PP = 0.82; BF₇ᵤ = 3.67). This inverse relationship was also observed in the practicality subdimensions, with DSC performing better in step clarity (PP = 0.75) and guideline applicability (PP = 0.82). For high-complexity tasks, such as urgent decision-making, all models failed to establish an effective SBAR Situation-to-Assessment linkage despite showing equivalent overall clinical logicality (PP = 0.59; BF₁ᵤ = 161.05, decisive evidence). Notably, 95.2% of junior physicians reported that DSC alleviated decision-making anxiety, compared with 28.6% for OA, suggesting a preference for actionable guidance despite accuracy limitations.
ConclusionsLLMs exhibit divergent capabilities: OA provides greater accuracy but lacks operational specificity, whereas DeepSeek-R1, particularly the Chinese version, offers more practical scaffolding at the expense of accuracy. This trade-off highlights the urgent need to develop and validate robust, specialty-specific evaluation frameworks and safeguards before such models can be safely deployed to support clinical decision-making in high-stakes environments.