AI in patient education: a comparative analysis of the quality and readability of AI generated patient education material on spinal fusion
摘要
Generative artificial intelligence (AI) models are increasingly used to create patient education materials (PEM), offering on-demand health information. These tools hold the potential to democratise access to medical information, but concerns remain regarding the quality, readability, and reliability of AI-generated PEM. Spinal fusion surgery, a complex procedure, necessitates clear, accurate educational materials to support informed decision-making. Despite their promise, the capacity of AI models to meet health literacy needs remains underexplored. This study aimed to evaluate and compare the readability and quality of PEM on spinal fusion surgery sourced from institution and society websites and those generated by three AI models (ChatGPT, Gemini, and Co-Pilot).
MethodsPatient information on spinal fusion surgery was sourced from the British Association of Spinal Surgeons (BASS), American Academy of Orthopaedic Surgeons (AAOS), Cleveland Clinic, Mayo Clinic, and John Hopkins websites, and generated by AI models (Chat GPT, Co-Pilot & Gemini) using a standard prompt on 15/12/24. Patient Education Materials Assessment Tool (PEMAT), JAMA benchmark criteria, and the DISCERN tool assessed quality. Readability was evaluated with the Flesch-Kincaid Grade Level (FKGL), Reading Ease (FKRE), and Gunning Fog Index (GFI). Mean quality and readability outcome measures of AI generated and institution or society sourced PEM were compared. Post Bonferroni correction, statistical significance was set at P < .0125 for quality and follow-up prompting assessments, and P < .0167 for readability assessments.
ResultsSociety and institution sourced PEM outperformed AI-generated content in readability and quality. Website-sourced materials scored significantly higher PEMAT (77.1% ± 9.8%, vs. 57.6% ± 6.3% for AI-generated PEM (P < .001)) and DISCERN scores (59.667 ± 9.686 vs. 49 ± 4.884, P = .006). Website-sourced PEM exhibited superior readability with lower FKGL and GFI scores and higher FKRE scores. AI-generated PEM quality varied by model, Co-Pilot scored the highest overall PEMAT (62.2% ± 3.3%) and highest DISCERN scores (50.8 +/- 4.087). Follow-up prompting resulted in an overall improvement in the AI content quality as assessed by the DISCERN, however this did not reach statistical significance (P = .024).
ConclusionProfessional website sourced PEM on spinal fusion surgery remains superior in terms of readability and reliability compared to AI-generated content. AI models produced PEM of variable quality. These findings highlight the need for further improvements in AI tools to ensure reliable and accessible health information for patients.