Background <p>The integration of artificial intelligence (AI) in healthcare has increased rapidly, with large language model-based chatbots emerging as potential tools for education and clinical support. However, their performance in complex medical domains such as sedation and general anesthesia remains underexplored. This study aimed to evaluate the accuracy of AI-supported chatbot models (ChatGPT 4.0 Mini, Gemini 1.5 Pro, and Claude 3 Sonnet) in comparison to human experts (Anesthesiologists, Pediatric Dentists, Oral and Maxillofacial Surgeons) within the context of sedation and general anesthesia.</p> Methods <p>This descriptive and comparative cross-sectional study utilized a 26-item true/false questionnaire based on ASA, AAP, ADA, and AAPD guidelines. The questionnaire was administered to three chatbots and 72 human participants (24 per specialty). For chatbot evaluation, a zero-shot prompting technique with a uniform command (“Is this statement true or false?”) was applied in separate sessions after clearing the cache to ensure standardization. Responses were coded as correct/incorrect by two independent pediatric dentists. Human participants completed the same questionnaire via Google Forms. Descriptive statistics were calculated. Normality of data was assessed using the Shapiro-Wilk test. Non-parametric data were compared using the Kruskal-Wallis test followed by Bonferroni-adjusted post hoc tests. Categorical variables were analyzed using the Pearson Chi-square test. Statistical analyses were performed using IBM SPSS v27.</p> Results <p>Significant differences were found among all groups ($<i>p</i> &lt; 0.001$). In terms of text-based guideline retrieval, Claude 3 Sonnet (92.8%) and Gemini 1.5 Pro (91.7%) demonstrated high factual accuracy, while anesthesiologists achieved the highest performance among human clinicians (87.0%). While the chatbot models showed high proficiency in data retrieval for “General Information” and “Postoperative Monitoring,” anesthesiologists significantly outperformed the AI models in “Patient Evaluation and Preparation” (91.7%), highlighting a critical gap in AI’s clinical reasoning compared to human expertise. Notable knowledge gaps were identified among oral surgeons regarding pharmacological reversal agents.</p> Conclusions <p>Claude and Gemini chatbots exhibit high factual accuracy in retrieving guideline-based pediatric sedation theory, suggesting potential exclusively as supplementary informational resources. However, anesthesiologists’ superior performance in pre-operative assessment underscores the irreplaceable role of clinical judgment. AI integration should therefore follow a synergistic model, positioning language models strictly as adjunct tools under mandatory human oversight.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of chatbot and specialist knowledge on pediatric sedation and general anesthesia: a comparative analysis

  • Dilara Dinç,
  • Aslıhan Ozbilgen

摘要

Background

The integration of artificial intelligence (AI) in healthcare has increased rapidly, with large language model-based chatbots emerging as potential tools for education and clinical support. However, their performance in complex medical domains such as sedation and general anesthesia remains underexplored. This study aimed to evaluate the accuracy of AI-supported chatbot models (ChatGPT 4.0 Mini, Gemini 1.5 Pro, and Claude 3 Sonnet) in comparison to human experts (Anesthesiologists, Pediatric Dentists, Oral and Maxillofacial Surgeons) within the context of sedation and general anesthesia.

Methods

This descriptive and comparative cross-sectional study utilized a 26-item true/false questionnaire based on ASA, AAP, ADA, and AAPD guidelines. The questionnaire was administered to three chatbots and 72 human participants (24 per specialty). For chatbot evaluation, a zero-shot prompting technique with a uniform command (“Is this statement true or false?”) was applied in separate sessions after clearing the cache to ensure standardization. Responses were coded as correct/incorrect by two independent pediatric dentists. Human participants completed the same questionnaire via Google Forms. Descriptive statistics were calculated. Normality of data was assessed using the Shapiro-Wilk test. Non-parametric data were compared using the Kruskal-Wallis test followed by Bonferroni-adjusted post hoc tests. Categorical variables were analyzed using the Pearson Chi-square test. Statistical analyses were performed using IBM SPSS v27.

Results

Significant differences were found among all groups ($p < 0.001$). In terms of text-based guideline retrieval, Claude 3 Sonnet (92.8%) and Gemini 1.5 Pro (91.7%) demonstrated high factual accuracy, while anesthesiologists achieved the highest performance among human clinicians (87.0%). While the chatbot models showed high proficiency in data retrieval for “General Information” and “Postoperative Monitoring,” anesthesiologists significantly outperformed the AI models in “Patient Evaluation and Preparation” (91.7%), highlighting a critical gap in AI’s clinical reasoning compared to human expertise. Notable knowledge gaps were identified among oral surgeons regarding pharmacological reversal agents.

Conclusions

Claude and Gemini chatbots exhibit high factual accuracy in retrieving guideline-based pediatric sedation theory, suggesting potential exclusively as supplementary informational resources. However, anesthesiologists’ superior performance in pre-operative assessment underscores the irreplaceable role of clinical judgment. AI integration should therefore follow a synergistic model, positioning language models strictly as adjunct tools under mandatory human oversight.