Background <p>This study evaluated the diagnostic accuracy and consistency of ChatGPT-4o in salivary gland disorders compared to experienced clinicians.</p> Methods <p>Eighty anonymized salivary gland cases from peer-reviewed reports were evaluated by ChatGPT-4o using standardized multimodal prompts and by three oral medicine specialists who provided Top-5 differentials. The primary outcome was diagnostic accuracy at the most likely diagnosis (Top-1), within the top three (Top-3), and within the top five (Top-5) differential diagnoses, with agreement measured by Cohen’s kappa and subgroup analyses by gland type, imaging, and case difficulty.</p> Results <p>At Top-3 and Top-5, ChatGPT showed perfect sensitivity (100%) and Top-1 86.67%. Experts surpassed ChatGPT at Top-5 (77.5% vs. 67.5%, <i>p</i> &lt; 0.0001), but ChatGPT outperformed experts at Top-1 (50.0% vs. 37.5%, <i>p</i> = 0.0309) and Top-3 (62.5% vs. 62.5%, <i>p</i> = 1.000). At Top-1, Cohen’s Kappa indicated moderate agreement (0.55). Experts showed notable variation by modality (<i>p</i> = 0.0174) and gland (<i>p</i> = 0.053). Although initial subgroup analyses found no notable heterogeneity of ChatGPT’s performance across imaging modalities, multivariate regression identified gland type to be an independent predictor of its Top-1 accuracy.</p> Conclusions <p>This first study shows ChatGPT can provide expert-level differential diagnoses for salivary gland disorders, suggesting promise as a supportive tool, though further research is needed to confirm its clinical role.</p> Clinical significance <p>ChatGPT-4o shows promise as a reliable supportive tool for differential diagnosis in oral medicine. Compared to experts, it performed more consistently across imaging modalities, although the particular salivary gland involved had a significant impact on its accuracy. Further validation through larger studies is needed for its integration into routine clinical practice.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative diagnostic accuracy of ChatGPT models in salivary gland disease: a multimodal vignette-based evaluation

  • Asmaa Abou-Bakr,
  • Ahmed Adel Eissa,
  • Basma Alshikh,
  • Yousra Ahmed,
  • Eslam Farid AbuShady,
  • Melek Tassoker,
  • Fatma E. A. Hassanein

摘要

Background

This study evaluated the diagnostic accuracy and consistency of ChatGPT-4o in salivary gland disorders compared to experienced clinicians.

Methods

Eighty anonymized salivary gland cases from peer-reviewed reports were evaluated by ChatGPT-4o using standardized multimodal prompts and by three oral medicine specialists who provided Top-5 differentials. The primary outcome was diagnostic accuracy at the most likely diagnosis (Top-1), within the top three (Top-3), and within the top five (Top-5) differential diagnoses, with agreement measured by Cohen’s kappa and subgroup analyses by gland type, imaging, and case difficulty.

Results

At Top-3 and Top-5, ChatGPT showed perfect sensitivity (100%) and Top-1 86.67%. Experts surpassed ChatGPT at Top-5 (77.5% vs. 67.5%, p < 0.0001), but ChatGPT outperformed experts at Top-1 (50.0% vs. 37.5%, p = 0.0309) and Top-3 (62.5% vs. 62.5%, p = 1.000). At Top-1, Cohen’s Kappa indicated moderate agreement (0.55). Experts showed notable variation by modality (p = 0.0174) and gland (p = 0.053). Although initial subgroup analyses found no notable heterogeneity of ChatGPT’s performance across imaging modalities, multivariate regression identified gland type to be an independent predictor of its Top-1 accuracy.

Conclusions

This first study shows ChatGPT can provide expert-level differential diagnoses for salivary gland disorders, suggesting promise as a supportive tool, though further research is needed to confirm its clinical role.

Clinical significance

ChatGPT-4o shows promise as a reliable supportive tool for differential diagnosis in oral medicine. Compared to experts, it performed more consistently across imaging modalities, although the particular salivary gland involved had a significant impact on its accuracy. Further validation through larger studies is needed for its integration into routine clinical practice.