Purpose <p>Junior physicians may experience critical competency gaps during the&#xa0;transition to unsupervised practice, particularly in high-risk specialties such as anesthesiology, where errors can compromise patient safety. Although large language models (LLMs)&#xa0;are being rapidly integrated into healthcare, their accuracy and safety in supporting junior physicians—who may lack sufficient&#xa0;experience to identify model errors—remain insufficiently&#xa0;quantified in high-stakes clinical settings. This study evaluated OpenAI-o1 (OA) and DeepSeek-R1&#xa0;in English (DSE) and&#xa0;Chinese (DSC) to address this evidence gap.</p> Methods <p>In this dual-axis evaluation, 30 anesthesia crisis scenarios were developed through Delphi consensus. Responses generated by OA, DSE, and DSC were assessed by 20 experts for&#xa0;accuracy using a&#xa0;5-point Likert scale and for&#xa0;clinical logicality using an 8-point Situation–Background–Assessment–Recommendation (SBAR)&#xa0;scoring system, with each component rated&#xa0;from&#xa0;0&#xa0;to&#xa0;2 points. Practicality was assessed by&#xa0;20 junior physicians using a 5-point Likert scale across three subdimensions: step clarity,&#xa0;guideline applicability, and&#xa0;learning assistance. Accuracy and practicality were the two primary evaluation domains, while clinical logicality, qualitative feedback, and readability were examined as complementary outcomes. Performance was compared across nine pre-specified, order-constrained hypotheses using Bayes Factor Design Analysis (<i>n</i> = 600 per group), with outcomes expressed as posterior probabilities (PPs) and Bayes Factors (BFs).</p> Results <p>OA demonstrated superior accuracy (OA &gt; DSE &gt; DSC; PP = 0.94; BF₄ᵤ = 4.70,&#xa0;strong evidence), whereas DSC showed greater practicality (OA &lt; DSE &lt; DSC; PP = 0.82; BF₇ᵤ = 3.67). This inverse relationship was also observed in the practicality subdimensions, with&#xa0;DSC performing better&#xa0;in step clarity (PP = 0.75) and guideline applicability (PP = 0.82). For high-complexity tasks, such as&#xa0;urgent decision-making, all models failed to establish an effective&#xa0;SBAR Situation-to-Assessment linkage despite showing&#xa0;equivalent overall&#xa0;clinical logicality (PP = 0.59; BF₁ᵤ = 161.05,&#xa0;decisive evidence). Notably, 95.2% of junior physicians reported that&#xa0;DSC alleviated decision-making anxiety, compared with 28.6% for OA, suggesting a preference for&#xa0;actionable guidance despite accuracy limitations.</p> Conclusions <p>LLMs exhibit divergent capabilities: OA provides greater&#xa0;accuracy but lacks operational specificity, whereas DeepSeek-R1,&#xa0;particularly the Chinese version, offers more&#xa0;practical scaffolding at the expense of accuracy. This trade-off highlights the urgent need to develop and validate robust, specialty-specific evaluation frameworks and safeguards before such models can be safely deployed to support clinical decision-making in high-stakes environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmark evaluation of OpenAI-o1 vs DeepSeek-R1 for junior anesthesiologists in Chinese hospitals: hazards of fast large language models use in high-stakes specialties

  • Zhuoxi Wu,
  • Yuting Tan,
  • Qin Chen,
  • Xinming Ye,
  • Qin Zhang,
  • Jiefeng Bi,
  • Hong Li

摘要

Purpose

Junior physicians may experience critical competency gaps during the transition to unsupervised practice, particularly in high-risk specialties such as anesthesiology, where errors can compromise patient safety. Although large language models (LLMs) are being rapidly integrated into healthcare, their accuracy and safety in supporting junior physicians—who may lack sufficient experience to identify model errors—remain insufficiently quantified in high-stakes clinical settings. This study evaluated OpenAI-o1 (OA) and DeepSeek-R1 in English (DSE) and Chinese (DSC) to address this evidence gap.

Methods

In this dual-axis evaluation, 30 anesthesia crisis scenarios were developed through Delphi consensus. Responses generated by OA, DSE, and DSC were assessed by 20 experts for accuracy using a 5-point Likert scale and for clinical logicality using an 8-point Situation–Background–Assessment–Recommendation (SBAR) scoring system, with each component rated from 0 to 2 points. Practicality was assessed by 20 junior physicians using a 5-point Likert scale across three subdimensions: step clarity, guideline applicability, and learning assistance. Accuracy and practicality were the two primary evaluation domains, while clinical logicality, qualitative feedback, and readability were examined as complementary outcomes. Performance was compared across nine pre-specified, order-constrained hypotheses using Bayes Factor Design Analysis (n = 600 per group), with outcomes expressed as posterior probabilities (PPs) and Bayes Factors (BFs).

Results

OA demonstrated superior accuracy (OA > DSE > DSC; PP = 0.94; BF₄ᵤ = 4.70, strong evidence), whereas DSC showed greater practicality (OA < DSE < DSC; PP = 0.82; BF₇ᵤ = 3.67). This inverse relationship was also observed in the practicality subdimensions, with DSC performing better in step clarity (PP = 0.75) and guideline applicability (PP = 0.82). For high-complexity tasks, such as urgent decision-making, all models failed to establish an effective SBAR Situation-to-Assessment linkage despite showing equivalent overall clinical logicality (PP = 0.59; BF₁ᵤ = 161.05, decisive evidence). Notably, 95.2% of junior physicians reported that DSC alleviated decision-making anxiety, compared with 28.6% for OA, suggesting a preference for actionable guidance despite accuracy limitations.

Conclusions

LLMs exhibit divergent capabilities: OA provides greater accuracy but lacks operational specificity, whereas DeepSeek-R1, particularly the Chinese version, offers more practical scaffolding at the expense of accuracy. This trade-off highlights the urgent need to develop and validate robust, specialty-specific evaluation frameworks and safeguards before such models can be safely deployed to support clinical decision-making in high-stakes environments.