Quality, readability, originality, and guideline concordance of ChatGPT 5.5 Plus and Gemini 3.1 Pro responses to parent-oriented questions in pediatric dentistry: a comparative cross-sectional study
摘要
Artificial intelligence-based chatbots are increasingly used by caregivers to obtain pediatric oral health information. This study compared the quality, readability, textual similarity, and guideline-based clinical concordance of ChatGPT 5.5 Plus and Gemini 3.1 Pro responses to parent-oriented pediatric dentistry questions.
MethodsFifty parent-oriented questions were developed through expert consensus by five specialist pediatric dentists and submitted to ChatGPT 5.5 Plus and Gemini 3.1 Pro on 28 April 2026. Each question was asked once in a separate chat session, and the first response was recorded. Responses were anonymized and evaluated by three blinded pediatric dentists using the Global Quality Scale (GQS). Readability was assessed using the Flesch Reading Ease Score (FRES) and Flesch-Kincaid Reading Grade Level (FKRGL), and textual similarity using the iThenticate Similarity Index. Guideline-based expert reference answers were developed from AAPD, EAPD, and IADT guidance, and each response was scored across six clinical domains. Paired statistical analyses were performed as appropriate.
ResultsGuideline-based clinical concordance was high for both models, with total scores of 11.82 ± 0.52 for ChatGPT and 11.80 ± 0.53 for Gemini out of 12, with no significant difference (p = 0.705). No potentially unsafe response was identified in the clinically sensitive subset. Mean GQS scores were similar for ChatGPT and Gemini (4.46 ± 0.45 vs. 4.48 ± 0.54; p = 0.704), and high-quality responses were observed in 46/50 (92%) and 47/50 (94%) responses, respectively. ChatGPT had higher FRES values than Gemini (61.45 ± 8.43 vs. 57.83 ± 8.19; p < 0.001), whereas FKRGL did not differ significantly (8.89 ± 1.49 vs. 9.08 ± 1.39; p = 0.334). Similarity values were low in both groups but higher for Gemini (0.00 ± 0.00% vs. 2.00 ± 2.47%; p < 0.001).
ConclusionsBoth chatbots generated responses with high perceived quality, low textual similarity, and high guideline-based clinical concordance. ChatGPT produced more readable responses, whereas Gemini showed slightly higher but still low similarity values. AI chatbots may support caregiver education but should not replace professional dental evaluation or individualized clinical advice.