Large language model chatbots for patient education in kidney stones: a scoping review
摘要
In 2024, 17% of adults reported using an artificial intelligence (AI) chatbot at least once a month as a source of health information, rising to 25% among those under 30. We aim to conduct a scoping review of the existing literature assessing the performance of large language model (LLM) chatbots for patient education in kidney stone disease (KSD).
MethodsThe Joanna Briggs Institute methodology was followed. Ovid MEDLINE, Embase, CENTRAL, Web of Science, CINAHL, and Google Scholar were searched for studies in all languages, from 2015 up to February 16th 2025, evaluating LLM chatbots to create educational content on KSD. Two independent reviewers completed screening, full-text review, and data extraction, with conflicts resolved by a third reviewer.
ResultsOf the 281 search results, 17 were included. Five of six studies assessing readability found that LLM responses exceeded the recommended 6th-8th grade reading level, though effective prompting can help meet this target. Six out of eight studies reported adequate to very good accuracy performance, with two showing comparable performance to traditional information sources. Understandability and actionability performance was poor. Quality performance was variable across seven studies. Three studies assessed patients’ perception, revealing a mixed but generally favorable experience. Two studies noted notable deviations from established clinical guidelines.
ConclusionLLM chatbots show potential for KSD patient education and physician workload reduction, but currently have limitations in readability, guideline adherence, understandability, and actionability. This could be mitigated by prompting and the development of urology-specific tools trained on validated content and evaluated with patient involvement.