Quality of AI-generated exercise awareness messages for older adults aligned with the ICFSR consensus: a comparative study of three LLMs
摘要
Large Language Models (LLMs) are emerging as potential tools for health communication and patient education. However, their ability to translate complex medical guidelines into accessible, safe, and accurate messages for older adults remains insufficiently evaluated. To assess the capacity of ChatGPT-4, Claude AI, and Deepseek to generate exercise awareness messages for older adults, aligned with international expert consensus.
Materials and methodsThis was a cross-sectional observational study conducted between April and August 2025. Using the ICFSR 2021 Global Consensus on optimal exercise recommendations as the gold standard, we evaluated messages generated by three LLMs for 13 common chronic conditions in older adults. A standardized prompt was used to generate messages addressing exercise prescription considerations, disease progression, and recommended modalities. Two independent expert evaluators; a geriatrician and a sports medicine physician assessed five dimensions using a 5-point Likert scale; accuracy, clarity, safety, behavioral relevance, and absence of fabrication. Readability was measured using Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG). Inter-rater agreement was assessed using intraclass correlation coefficient (ICC).
ResultsClaude achieved the overall scores from both evaluators at 4.63 ± 0.55 and 4.60 ± 0.55, followed by ChatGPT-4 at 4.55 ± 0.50 and 4.48 ± 0.53 and Deepseek at 4.15 ± 0.87 and 4.17 ± 0.86. Inter-rater agreement was moderate was 0.668. Claude demonstrated accuracy scores at 4.58 ± 0.64, while ChatGPT-4 excelled in clarity with 5 ± 0. All models achieved perfect scores for absence of fabrication (5 ± 0). Readability indices revealed high complexity across all LLMs, with median FKGL values ranging from 10.84 to 10.93, corresponding to 10th-11th grade reading level, exceeding recommended levels for older populations. A significant correlation was found between Claude’s accuracy scores and FKGL (r = 0.599, p = 0.030).
ConclusionLLMs, particularly Claude, ChatGPT-4 and Deepseek demonstrate strong potential for generating accurate, safe, and hallucination-free exercise awareness messages for older adults. However, readability remains above recommended levels, requiring optimization.