Comparative evaluation of ChatGPT and Gemini in detecting external apical root resorption on panoramic radiographs of orthodontic patients
摘要
This study aimed to evaluate the performance of large language models (LLMs), specifically GPT-4o and Gemini 2 Flash, in identifying external apical root resorption (EARR) on panoramic radiographs of orthodontic patients using a standardized prompt.
MethodsThis comparative observational diagnostic study included 52 cropped tooth images obtained from panoramic radiographs of healthy individuals after orthodontic treatment. From each image, the regions corresponding to the permanent maxillary and mandibular incisors were manually cropped to include the apex, surrounding alveolar bone, and crown. An expert in endodontics evaluated each cropped image for the presence and severity of EARR using the Malmgren scale. The same images were submitted to both LLMs (GPT-4o and Gemini 2 Flash) using an identical multimodal prompt (image and text instructions). The models’ responses were compared to the expert ratings using Cohen’s kappa (κ), accuracy, F1-score, mean absolute error (MAE), and confusion matrices. The Wilcoxon signed-rank test was used to compare the MAE between the models. Confidence intervals (95%) were calculated via bootstrapping.
ResultsAccording to the expert evaluation, EARR was identified across all Malmgren grades: 11 teeth (21.2%) showed no resorption, 11 (21.2%) irregular apical contour, 11 (21.2%) small apical resorption, 11 (21.2%) resorption up to one-third of the root length, and 8 (15.4%) exceeding one-third resorption. GPT-4o showed fair agreement with the expert for the binary classification (κ= 0.371), whereas Gemini exhibited a negative κ (−0.152), indicating performance below chance. GPT-4o achieved 36.5% accuracy and a MAE of 1.269 in severity classification, compared to 13.5% accuracy and a MAE of 1.750 for Gemini. Both models performed poorly in detecting moderate to severe EARR. The Wilcoxon test showed no significant difference between the models (p > 0.05).
ConclusionGPT-4o achieved numerically better results than Gemini, with lower error rates and slightly higher agreement with the expert. Nevertheless, both models showed limited accuracy and agreement, particularly in detecting moderate to severe resorption, and neither can be considered suitable for clinical application at this stage.