Artificial intelligence chatbots vs. YPUC pediatric urologists: performance on a Campbell Walsh urology hypospadiology questionnaire
摘要
Artificial intelligence (AI) is increasingly used in medicine, with uncertain performance in highly specialized fields, like hypospadiology. This study aimed to compare the performance of AI platforms and pediatric urologists using a structured knowledge assessment.
MethodsA 31-question multiple-choice questionnaire (12th edition of Campbell Walsh Urology) was completed by 23 members of the Young Pediatric Urologist Committee (YPUC) of the European Society for Paediatric Urology (ESPU), and by five AI models (ChatGPT 3.5, ChatGPT 4o, Gemini, Copilot, and Doubao). Scores were analyzed overall, by subgroups (FEAPU-certified (Fellow of the European Academy of Paediatric Urology, n=15), >35 years old (n=16), and self-declared hypospadiology experts (n=12)), and across four thematic categories : basics, clinical and paraclinical evaluation, initial management, and complications.
ResultsAI and human respondents showed comparable overall performance (67.7% vs. 61.2%, p=0.467). However, experienced subgroups outperformed their counterparts: FEAPU-certified vs. non-certified (61.3% vs. 53.2%, p=0.001), >35 vs. ≤35 years (61.3% vs. 54.8%, p=0.039), and self declared experts vs. non-experts (61.3% vs. 54.8%, p=0.039). AI performed better in Basics (83% vs. 67%, p=0.320), Initial management (71% vs. 57%, p=0.450), and Complications (64% vs. 55%, p=0.087), but underperformed significantly in Clinical and paraclinical evaluation (57% vs. 71%, p=0.003). The highly specialized nature of the topic limits generalizability to broader medical contexts.
ConclusionAI platforms achieved scores comparable to pediatric urologists in a structured questionnaire but fell short in domains requiring nuanced clinical reasoning. While AI shows potential as an educational-support tool, it does not replace expert human judgment in hypospadiology.