Introduction <p>Artificial intelligence (AI) is quickly transforming healthcare by improving patient and clinician access to and understanding of medical information. Generative AI models answer healthcare queries and provide tailored and quick responses. This research evaluates the readability and quality of bladder cancer (BC) patient information in 10 popular AI-enabled chatbots.</p> Materials and methods <p>We used the latest versions of ten popular chatbots: OpenAI’s GPT-4o, Microsoft’s Copilot Pro, Claude-3.5 Haiku, Sonar Large, Grok 2, Gemini Advanced 1.5 Pro, Mistral Large, Google Palm 2 (Google Bard), Meta’s Llama 3.3, and Meta AI v2. Prompts were developed to provide texts about BC, non-muscle-invasive BC, muscle-invasive BC, and metastatic BC. The modified Ensuring Quality Information for Patients (mEQIP), the Quality Evaluating Scoring Tool (QUEST), and DISCERN were used to assess quality. The Average Reading Level Consensus (ARLC), Flesch Reading Ease (FKRE), and Flesch-Kincaid Grade Level (FKGL) were used to evaluate readability.</p> Results <p>Ten chatbots exhibited statistically significant differences in mean mEQIP, DISCERN, and QUEST scores (<i>p</i> = <b>0.048</b>, <i>p</i> = <b>0.025</b>, and <i>p</i> = <b>0.021</b>, respectively). Meta scored lowest on the average mEQIP, DISCERN, and QUEST, while Llama attained the highest. Statistically significant differences were also seen in the chatbots’ average ARLC, FKGL, and FKRE scores (<i>p</i> = <b>0.002</b>, <i>p</i> = <b>0.001</b>, and <i>p</i> = <b>0.002</b>, respectively), in which Google Palm produced texts that are easiest to read, and Llama is the most difficult chatbot&#xa0;to understand.</p> Conclusion <p>AI chatbots can produce information on BC that is of moderate quality and readability, while there is significant variability among platforms. Results should be evaluated with caution due to the single-query approach and the continuously advancing AI models. Clinicians can support safety in implementation by delivering structured feedback and incorporating content review stages into patient education processes. Continuous collaboration between healthcare practitioners and AI developers is crucial to maintain the accuracy, currency, and clarity of AI-generated content.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

What do the current popular artificial intelligence chatbots offer us regarding patient information? Comparison of responses from the ten most popular chatbots about bladder cancer

  • Mehmet Fatih Sahin,
  • Murat Akgül,
  • Çağrı Akpınar,
  • Deniz Bolat,
  • Ozan Bozkurt,
  • Volkan İzol,
  • Nihat Karakoyunlu,
  • Evren Süer,
  • İlker Tınay

摘要

Introduction

Artificial intelligence (AI) is quickly transforming healthcare by improving patient and clinician access to and understanding of medical information. Generative AI models answer healthcare queries and provide tailored and quick responses. This research evaluates the readability and quality of bladder cancer (BC) patient information in 10 popular AI-enabled chatbots.

Materials and methods

We used the latest versions of ten popular chatbots: OpenAI’s GPT-4o, Microsoft’s Copilot Pro, Claude-3.5 Haiku, Sonar Large, Grok 2, Gemini Advanced 1.5 Pro, Mistral Large, Google Palm 2 (Google Bard), Meta’s Llama 3.3, and Meta AI v2. Prompts were developed to provide texts about BC, non-muscle-invasive BC, muscle-invasive BC, and metastatic BC. The modified Ensuring Quality Information for Patients (mEQIP), the Quality Evaluating Scoring Tool (QUEST), and DISCERN were used to assess quality. The Average Reading Level Consensus (ARLC), Flesch Reading Ease (FKRE), and Flesch-Kincaid Grade Level (FKGL) were used to evaluate readability.

Results

Ten chatbots exhibited statistically significant differences in mean mEQIP, DISCERN, and QUEST scores (p = 0.048, p = 0.025, and p = 0.021, respectively). Meta scored lowest on the average mEQIP, DISCERN, and QUEST, while Llama attained the highest. Statistically significant differences were also seen in the chatbots’ average ARLC, FKGL, and FKRE scores (p = 0.002, p = 0.001, and p = 0.002, respectively), in which Google Palm produced texts that are easiest to read, and Llama is the most difficult chatbot to understand.

Conclusion

AI chatbots can produce information on BC that is of moderate quality and readability, while there is significant variability among platforms. Results should be evaluated with caution due to the single-query approach and the continuously advancing AI models. Clinicians can support safety in implementation by delivering structured feedback and incorporating content review stages into patient education processes. Continuous collaboration between healthcare practitioners and AI developers is crucial to maintain the accuracy, currency, and clarity of AI-generated content.