Purpose <p>This study aimed to evaluate the performance of large language models (LLMs), specifically GPT-4o and Gemini&#xa0;2 Flash, in identifying external apical root resorption (EARR) on panoramic radiographs of orthodontic patients using a&#xa0;standardized prompt.</p> Methods <p>This comparative observational diagnostic study included 52&#xa0;cropped tooth images obtained from panoramic radiographs of healthy individuals after orthodontic treatment. From each image, the regions corresponding to the permanent maxillary and mandibular incisors were manually cropped to include the apex, surrounding alveolar bone, and crown. An expert in endodontics evaluated each cropped image for the presence and severity of EARR using the Malmgren scale. The same images were submitted to both LLMs (GPT-4o and Gemini 2&#xa0;Flash) using an identical multimodal prompt (image and text instructions). The models’ responses were compared to the expert ratings using Cohen’s kappa (κ), accuracy, F1-score, mean absolute error (MAE), and confusion matrices. The Wilcoxon signed-rank test was used to compare the MAE between the models. Confidence intervals (95%) were calculated via bootstrapping.</p> Results <p>According to the expert evaluation, EARR was identified across all Malmgren grades: 11&#xa0;teeth (21.2%) showed no resorption, 11&#xa0;(21.2%) irregular apical contour, 11&#xa0;(21.2%) small apical resorption, 11&#xa0;(21.2%) resorption up to one-third of the root length, and 8&#xa0;(15.4%) exceeding one-third resorption. GPT-4o showed fair agreement with the expert for the binary classification (κ= 0.371), whereas Gemini exhibited a&#xa0;negative κ (−0.152), indicating performance below chance. GPT-4o achieved 36.5% accuracy and a&#xa0;MAE of 1.269 in severity classification, compared to 13.5% accuracy and a&#xa0;MAE of 1.750 for Gemini. Both models performed poorly in detecting moderate to severe EARR. The Wilcoxon test showed no significant difference between the models (<i>p</i> &gt; 0.05).</p> Conclusion <p>GPT-4o achieved numerically better results than Gemini, with lower error rates and slightly higher agreement with the expert. Nevertheless, both models showed limited accuracy and agreement, particularly in detecting moderate to severe resorption, and neither can be considered suitable for clinical application at this stage.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative evaluation of ChatGPT and Gemini in detecting external apical root resorption on panoramic radiographs of orthodontic patients

  • Allan Abuabara,
  • Flares Baratto-Filho,
  • Giancarlo Roos Gallego,
  • Luana Beatriz das Portas Luiz,
  • Michelle Nascimento Meger,
  • Rafaela Scariot,
  • Svenja Beisel-Memmert,
  • Cristiano Miranda de Araujo,
  • Erika Calvano Küchler,
  • Bianca Marques de Mattos de Araujo

摘要

Purpose

This study aimed to evaluate the performance of large language models (LLMs), specifically GPT-4o and Gemini 2 Flash, in identifying external apical root resorption (EARR) on panoramic radiographs of orthodontic patients using a standardized prompt.

Methods

This comparative observational diagnostic study included 52 cropped tooth images obtained from panoramic radiographs of healthy individuals after orthodontic treatment. From each image, the regions corresponding to the permanent maxillary and mandibular incisors were manually cropped to include the apex, surrounding alveolar bone, and crown. An expert in endodontics evaluated each cropped image for the presence and severity of EARR using the Malmgren scale. The same images were submitted to both LLMs (GPT-4o and Gemini 2 Flash) using an identical multimodal prompt (image and text instructions). The models’ responses were compared to the expert ratings using Cohen’s kappa (κ), accuracy, F1-score, mean absolute error (MAE), and confusion matrices. The Wilcoxon signed-rank test was used to compare the MAE between the models. Confidence intervals (95%) were calculated via bootstrapping.

Results

According to the expert evaluation, EARR was identified across all Malmgren grades: 11 teeth (21.2%) showed no resorption, 11 (21.2%) irregular apical contour, 11 (21.2%) small apical resorption, 11 (21.2%) resorption up to one-third of the root length, and 8 (15.4%) exceeding one-third resorption. GPT-4o showed fair agreement with the expert for the binary classification (κ= 0.371), whereas Gemini exhibited a negative κ (−0.152), indicating performance below chance. GPT-4o achieved 36.5% accuracy and a MAE of 1.269 in severity classification, compared to 13.5% accuracy and a MAE of 1.750 for Gemini. Both models performed poorly in detecting moderate to severe EARR. The Wilcoxon test showed no significant difference between the models (p > 0.05).

Conclusion

GPT-4o achieved numerically better results than Gemini, with lower error rates and slightly higher agreement with the expert. Nevertheless, both models showed limited accuracy and agreement, particularly in detecting moderate to severe resorption, and neither can be considered suitable for clinical application at this stage.