Can chatbots replace experts? Diagnostic accuracy of AI models in classifying impacted mandibular third molars
摘要
Artificial intelligence (AI)-based chatbots are increasingly employed in various fields of dentistry. However, their capability to interpret panoramic radiographs and classify impacted mandibular third molars has not yet been systematically evaluated. This study aimed to assess the diagnostic performance of four widely used large language model-based chatbots in this context. A total of 93 impacted mandibular third molars were assessed using panoramic radiographs. Four chatbots—ChatGPT-4o, Gemini 2.5 Pro, Claude Sonnet 4.0, and Copilot (GPT-4)—were asked to classify each case according to the Pell and Gregory, Winter, and Rood and Shehab systems without any additional guidance. Three experts then rated each chatbot response using a Global Quality Score. Inter-rater reliability and diagnostic agreement were analyzed. Inter-rater agreement was high among human evaluators (ICC: 0.886–0.935, p < 0.001). No significant difference was found in GQS ratings across the chatbots (p > 0.05), although ChatGPT-4o received the highest mean score (2.41 ± 1.03). ChatGPT-4o also performed significantly better in the Winter classification (κ = 0.171; p = 0.005). Gemini 2.5 Pro showed moderate agreement in root-related findings, whereas Copilot (GPT-4) showed more consistency in canal-related parameters (p < 0.05). No chatbot demonstrated acceptable agreement in the Pell and Gregory classification (p > 0.05). While AI-based chatbots showed potential for interpreting panoramic images, their current performance in third molar classification remains suboptimal at present. ChatGPT-4o outperformed other models in certain tasks, but none achieved expert-level accuracy. Further improvements, particularly with multimodal AI models and labeled datasets, are essential for future clinical integration.