Performance of artificial intelligence chatbots in the diagnosis and management of simulated dental trauma cases: an evaluation based on IADT guidelines
摘要
This study aims to comparatively evaluate the performance of four different artificial intelligence-based chatbots (ChatGPT-4o (Free), ChatGPT-5 (Plus), DeepSeek, and Google Gemini) in the diagnosis and treatment processes of dental trauma cases.
Material and methodsBased on the International Association of Dental Traumatology (IADT) guidelines, eleven fictional cases were developed, each representing different diagnostic possibilities of dental trauma. The cases cenarios were based on a standardized dataset including clinical examination, radiographic findings, and pulp sensitivity tests. All AI models were tested for three days for each case. Two researchers analyzed the responses using ablinding method according to five evaluation criteria (diagnostic accuracy, treatment plan appropriateness, splinting duration accuracy, antibiotic indication, and reference accuracy).
ResultsAccording to the results, while no significant difference was found in terms of diagnostic accuracy (p > 0.05), Google Gemini showed the highest performance with 100% accuracy. ChatGPT-4o (Free) stood out with a 97% accuracy rate in antibiotic indication, while in splinting duration prediction, ChatGPT-5 (Plus) was the most successful model with 75.8%. DeepSeek exhibited the highest variability (p < 0.05).
ConclusionsThe findings show that AI chatbots have promising potential as complementary tools in dental trauma diagnosis, but evidence-based validation, expert supervision, and methodological standards are needed for safe use in clinical practice.