ChatGPT and other large language models in laparoscopic cholecystectomy: a multidimensional audit of reliability, quality, and readability
摘要
The rapid uptake of large language models (LLMs) in surgery demands evidence of their reliability when guiding laparoscopic cholecystectomy (LC).
MethodsAn analytical cross-sectional study (April–June 2025) compared five current LLMs (ChatGPT-o3, Claude-Sonnet-4, DeepSeek-V3.5, Gemini-2.5 Flash, and Grok-3) on 24 guideline-derived questions covering the pre-, intra-, and postoperative phases of LC. Four blinded hepatobiliary surgeons rated 120 answers with the eight-item modified DISCERN (mDISCERN, 8–40) and Global Quality Score (GQS, 1–5). Readability was quantified with FRES, FKGL, SMOG, Fog, CLI, and lexical density indices, and inter-rater agreement assessed by two-way ICC.
ResultsGrok delivered the highest mean mDISCERN (36.3 ± 2.3) and GQS (4.76 ± 0.41), whereas Gemini scored lowest (29.0 ± 2.1; 3.58 ± 0.36). DeepSeek produced the most readable output (FRES ≈ 30.6; FKGL ≈ 12.1), while Claude generated the densest, least readable text (negative FRES; FKGL ≈ 18.3). Quality correlated positively with word count and lexical density (ρ ≈ 0.7) but not with syntactic complexity. Surgeon ratings showed good reliability (ICC(2,k) = 0.775; ICC(3,k) = 0.819).
ConclusionsLLM performance for LC varies markedly; even the best-performing model stops short of full reliability, reinforcing the need for procedure-specific validation before clinical deployment. This multidimensional audit provides a reproducible benchmark for selecting and fine-tuning surgical decision-support LLMs and highlights that terminological richness, rather than sentence complexity, underpins high-quality guidance.