Background <p>Artificial intelligence (AI) chatbots, which use large language models (LLMs) to comprehend user questions and respond to human conversations, are increasingly used in dental education, but their accuracy in core subjects like occlusion remains unexplored. This study evaluated the performance of four AI chatbots, namely, Claude 3.7 Sonnet, GPT-4o, DeepSeek, and Meta LLaMA 3.2, in answering multiple-choice questions (MCQs) related to fundamentals of dental occlusion.</p> Methods <p>Each chatbot was tested across two rounds on a standardized set of MCQs taken from a textbook on dental occlusion. Accuracy, intra-model consistency, and inter-model agreement were assessed using descriptive statistics, Cohen’s Kappa, and McNemar’s test.</p> Results <p>Claude 3.7 sonnet and GPT-4o achieved the highest accuracy in round one at 90%, while in round two, Claude 3.7 sonnet had an accuracy of 91.4%, and GPT-4o scored 84.4%. Meta Llama3.2 recorded the lowest accuracy across both rounds, with 71.4% in the first and 64.3% in the second. The Kappa test showed the intra-model consistency between the two rounds to be highest for Claude 3.7 sonnet (0.746), with a McNemar P-value of 1. For inter-model consistency when comparing the chatbots in round 1, Claude 3.7 sonnet and GPT-4o displayed the highest consistency with a kappa value (<i>p</i> &lt; 0.005) and McNemar P-value of 1. In round 2, Claude 3.7 sonnet with GPT-4o and with DeepSeek had the highest consistency, with kappa values of 0.272 and a McNemar P-value of 0.227.</p> Conclusions <p>Claude 3.7 Sonnet and GPT-4o show strong potential as reliable tools for dental education. DeepSeek offers moderate reliability, while Meta LLaMA 3.2 displayed inconsistent performance. Human oversight remains critical, and further refinement is essential before broader adoption in academic or clinical settings.</p> Trial registration <p>Not applicable as it is not a clinical trial.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessment of artificial intelligence chatbots in responding to dental occlusion questions: a comparative study

  • Hamod Alqahtani

摘要

Background

Artificial intelligence (AI) chatbots, which use large language models (LLMs) to comprehend user questions and respond to human conversations, are increasingly used in dental education, but their accuracy in core subjects like occlusion remains unexplored. This study evaluated the performance of four AI chatbots, namely, Claude 3.7 Sonnet, GPT-4o, DeepSeek, and Meta LLaMA 3.2, in answering multiple-choice questions (MCQs) related to fundamentals of dental occlusion.

Methods

Each chatbot was tested across two rounds on a standardized set of MCQs taken from a textbook on dental occlusion. Accuracy, intra-model consistency, and inter-model agreement were assessed using descriptive statistics, Cohen’s Kappa, and McNemar’s test.

Results

Claude 3.7 sonnet and GPT-4o achieved the highest accuracy in round one at 90%, while in round two, Claude 3.7 sonnet had an accuracy of 91.4%, and GPT-4o scored 84.4%. Meta Llama3.2 recorded the lowest accuracy across both rounds, with 71.4% in the first and 64.3% in the second. The Kappa test showed the intra-model consistency between the two rounds to be highest for Claude 3.7 sonnet (0.746), with a McNemar P-value of 1. For inter-model consistency when comparing the chatbots in round 1, Claude 3.7 sonnet and GPT-4o displayed the highest consistency with a kappa value (p < 0.005) and McNemar P-value of 1. In round 2, Claude 3.7 sonnet with GPT-4o and with DeepSeek had the highest consistency, with kappa values of 0.272 and a McNemar P-value of 0.227.

Conclusions

Claude 3.7 Sonnet and GPT-4o show strong potential as reliable tools for dental education. DeepSeek offers moderate reliability, while Meta LLaMA 3.2 displayed inconsistent performance. Human oversight remains critical, and further refinement is essential before broader adoption in academic or clinical settings.

Trial registration

Not applicable as it is not a clinical trial.