Background <p>Artificial intelligence (AI)-driven large language models (LLMs) are increasingly used to provide patient-oriented oral health information; however, their performance in answering periodontal–orthodontic patient questions remains unclear. This study aimed to compare the performance of contemporary AI-driven LLMs in answering periodontal–orthodontic patient questions by assessing response accuracy, comprehensiveness, readability, and temporal consistency.</p> Methods <p>Thirty patient-oriented periodontal–orthodontic questions were submitted to six contemporary LLMs (ChatGPT 5.5, Claude Sonnet 4.6, DeepSeek 3.2, Gemini 3.1 Pro, Grok 4.20, and PerioGPT) under standardized conditions, and the generated responses were analyzed. Accuracy and comprehensiveness were evaluated using a modified five-point Likert scale, whereas readability was assessed using the Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) indices. Temporal consistency was assessed by repeating the response-generation process one week later. Inter-model comparisons and temporal consistency analyses were performed using repeated-measures ANOVA or Friedman tests with Bonferroni-adjusted post hoc comparisons, as appropriate.</p> Results <p>Significant differences were observed among the evaluated LLMs in accuracy (<i>p</i> &lt; 0.001, Kendall’s W = 0.401), comprehensiveness (<i>p</i> &lt; 0.001, Kendall’s W = 0.575), FRE (<i>p</i> &lt; 0.001, partial η<sup>2</sup> = 0.807), and FKGL (<i>p</i> &lt; 0.001, partial η<sup>2</sup> = 0.688). PerioGPT and Gemini achieved the highest accuracy (4.93 ± 0.13 and 4.90 ± 0.19, respectively) and comprehensiveness scores (4.85 ± 0.21 and 4.90 ± 0.22, respectively). PerioGPT generated highly accurate and comprehensive responses but exhibited the lowest FRE (12.7 ± 10.2) and highest FKGL (15.1 ± 2.1) scores, whereas Claude, DeepSeek, and Grok produced more readable outputs. Although some significant differences in temporal consistency were observed among the evaluated models, the associated effect sizes were generally small (0.041–0.143), indicating broadly similar performance over time.</p> Conclusions <p>While PerioGPT and Gemini achieved the highest accuracy and comprehensiveness scores, and Claude, DeepSeek, and Grok produced more readable responses, the evaluated models generally demonstrated similar temporal consistency over time; however, no model consistently outperformed the others across all evaluated criteria. Therefore, AI-driven LLMs may support patient education and oral health literacy in periodontal–orthodontic care; however, their use as complementary tools under expert supervision is recommended.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative evaluation of AI-driven large language models for periodontal–orthodontic patient questions: accuracy, comprehensiveness, readability, and temporal consistency

  • Resül Çolak,
  • Merve Küçükoğlu Çolak,
  • Orhan Cicek

摘要

Background

Artificial intelligence (AI)-driven large language models (LLMs) are increasingly used to provide patient-oriented oral health information; however, their performance in answering periodontal–orthodontic patient questions remains unclear. This study aimed to compare the performance of contemporary AI-driven LLMs in answering periodontal–orthodontic patient questions by assessing response accuracy, comprehensiveness, readability, and temporal consistency.

Methods

Thirty patient-oriented periodontal–orthodontic questions were submitted to six contemporary LLMs (ChatGPT 5.5, Claude Sonnet 4.6, DeepSeek 3.2, Gemini 3.1 Pro, Grok 4.20, and PerioGPT) under standardized conditions, and the generated responses were analyzed. Accuracy and comprehensiveness were evaluated using a modified five-point Likert scale, whereas readability was assessed using the Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) indices. Temporal consistency was assessed by repeating the response-generation process one week later. Inter-model comparisons and temporal consistency analyses were performed using repeated-measures ANOVA or Friedman tests with Bonferroni-adjusted post hoc comparisons, as appropriate.

Results

Significant differences were observed among the evaluated LLMs in accuracy (p < 0.001, Kendall’s W = 0.401), comprehensiveness (p < 0.001, Kendall’s W = 0.575), FRE (p < 0.001, partial η2 = 0.807), and FKGL (p < 0.001, partial η2 = 0.688). PerioGPT and Gemini achieved the highest accuracy (4.93 ± 0.13 and 4.90 ± 0.19, respectively) and comprehensiveness scores (4.85 ± 0.21 and 4.90 ± 0.22, respectively). PerioGPT generated highly accurate and comprehensive responses but exhibited the lowest FRE (12.7 ± 10.2) and highest FKGL (15.1 ± 2.1) scores, whereas Claude, DeepSeek, and Grok produced more readable outputs. Although some significant differences in temporal consistency were observed among the evaluated models, the associated effect sizes were generally small (0.041–0.143), indicating broadly similar performance over time.

Conclusions

While PerioGPT and Gemini achieved the highest accuracy and comprehensiveness scores, and Claude, DeepSeek, and Grok produced more readable responses, the evaluated models generally demonstrated similar temporal consistency over time; however, no model consistently outperformed the others across all evaluated criteria. Therefore, AI-driven LLMs may support patient education and oral health literacy in periodontal–orthodontic care; however, their use as complementary tools under expert supervision is recommended.