Objective <p>Open-source AI chatbots are increasingly used for patient education, but their performance must be evaluated to avoid misinformation. This study aimed to compare the performance of five widely used AI chatbots—ChatGPT, Google Gemini, Copilot, Meta AI and Claude in answering frequently asked questions (FAQs) about oral submucous fibrosis (OSMF). The primary objective was to compare the accuracy of responses generated by these chatbots. The secondary objective was to compare the readability of the responses using the Flesch–Kincaid Reading Ease (FKRE) and Grade Level (FKGL) scores.</p> Materials and Methods <p>This cross-sectional in silico study assessed computer-generated answers to a&#xa0;set of 26 FAQs compiled through web searches and reviewed by oral and maxillofacial surgeons. Each chatbot was prompted to respond to the individual queries. Accuracy was assessed using a five-point modified Likert scale. Readability was measured using the Flesch–Kincaid Reading Ease (FKRE) and Grade Level (FKGL) scores. One-way ANOVA with post hoc analyses was conducted to compare chatbot performance.</p> Results <p>Mean Likert scores ranged from 4.50 (SD = 0.52; 95% CI [4.36, 4.78]) to 4.80 (SD = 0.37; 95% CI [4.65, 4.96]), with no statistically significant differences in accuracy among the five AI chatbots (<i>P</i> = 0.10). FKRE scores were highest for Gemini (mean = 32.34; SD = 13.81; 95% CI [26.75, 37.91]) and lowest for Claude (mean = 18.16; SD = 17.78; 95% CI [10.97, 25.33]), showing a significant difference (<i>P</i> &lt; 0.005). FKGL scores also differed significantly (<i>P</i> &lt; 0.001), with Claude producing the most complex text (mean = 17.57; SD = 4.29; 95% CI [15.83, 19.30]), while Gemini (mean = 13.07; SD = 2.83; 95% CI [11.93, 14.21]) and Copilot (mean = 13.45; SD = 2.69; 95% CI [12.36, 14.54]) had the lowest scores, indicating better readability.</p> Conclusion <p>All five AI chatbots generated accurate responses; however, none reached patient-friendly readability scores. Google Gemini demonstrated the best balance between accuracy and readability, making it more suitable for patient education. While AI tools can enhance healthcare communication, clinicians should carefully validate their responses before clinical use.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Accuracy and Readability of Large Language Model-Based AI Chatbot’s Response to Patient Queries on Oral Submucous Fibrosis

  • Aafiya Ambereen,
  • Jitendra Chawla,
  • Babu Lal,
  • Ragavi Alagarsamy,
  • Sukanth Tatapudi,
  • Sujata Mohanty

摘要

Objective

Open-source AI chatbots are increasingly used for patient education, but their performance must be evaluated to avoid misinformation. This study aimed to compare the performance of five widely used AI chatbots—ChatGPT, Google Gemini, Copilot, Meta AI and Claude in answering frequently asked questions (FAQs) about oral submucous fibrosis (OSMF). The primary objective was to compare the accuracy of responses generated by these chatbots. The secondary objective was to compare the readability of the responses using the Flesch–Kincaid Reading Ease (FKRE) and Grade Level (FKGL) scores.

Materials and Methods

This cross-sectional in silico study assessed computer-generated answers to a set of 26 FAQs compiled through web searches and reviewed by oral and maxillofacial surgeons. Each chatbot was prompted to respond to the individual queries. Accuracy was assessed using a five-point modified Likert scale. Readability was measured using the Flesch–Kincaid Reading Ease (FKRE) and Grade Level (FKGL) scores. One-way ANOVA with post hoc analyses was conducted to compare chatbot performance.

Results

Mean Likert scores ranged from 4.50 (SD = 0.52; 95% CI [4.36, 4.78]) to 4.80 (SD = 0.37; 95% CI [4.65, 4.96]), with no statistically significant differences in accuracy among the five AI chatbots (P = 0.10). FKRE scores were highest for Gemini (mean = 32.34; SD = 13.81; 95% CI [26.75, 37.91]) and lowest for Claude (mean = 18.16; SD = 17.78; 95% CI [10.97, 25.33]), showing a significant difference (P < 0.005). FKGL scores also differed significantly (P < 0.001), with Claude producing the most complex text (mean = 17.57; SD = 4.29; 95% CI [15.83, 19.30]), while Gemini (mean = 13.07; SD = 2.83; 95% CI [11.93, 14.21]) and Copilot (mean = 13.45; SD = 2.69; 95% CI [12.36, 14.54]) had the lowest scores, indicating better readability.

Conclusion

All five AI chatbots generated accurate responses; however, none reached patient-friendly readability scores. Google Gemini demonstrated the best balance between accuracy and readability, making it more suitable for patient education. While AI tools can enhance healthcare communication, clinicians should carefully validate their responses before clinical use.