Analyses of different prescriptions for health using artificial intelligence: a critical approach based on the international guidelines of health institutions
摘要
Large-language models (LLMs) are increasingly used for health advice, but their alignment with evidence-based guidelines and sensitivity to question phrasing remain unclear.
MethodsIn May 2025, we evaluated ChatGPT 4.0, ChatGPT 4.5, and DeepSeek V3 using four clinical vignettes: major depression with polysubstance use, irritable bowel syndrome flare, new-onset hypertension requiring exercise counseling, and chronic low back pain. Each scenario was tested with clinician- and patient-style prompts, generating 24 responses. Outputs were benchmarked against 89 guideline-derived recommendations from three authoritative sources per domain. Two blinded reviewers scored concordance (1 = actionable detail, 0.5 = generic mention, 0 = absent), with adjudication by a third reviewer. Inter-rater reliability was measured using Cronbach’s α.
ResultsChatGPT 4.5 achieved the highest guideline concordance (61.9%), followed by DeepSeek V3 (60.7%) and ChatGPT 4.0 (53.7%). Performance varied by domain, exceeding 67% in mental health but dropping below 45% in nutrition. Prompt phrasing influenced capture rates, with clinician-style prompts improving scores in exercise and pain domains, while patient-style prompts outperformed in nutrition. Reviewer agreement was high (α = 0.97 for chatbot scoring; 0.80 for matrix coding).
ConclusionLLMs can rapidly generate draft care plans that reflect clinical guidelines, though they favor generic over individualized advice. By introducing a unique, domain-agnostic scoring rubric that aligns AI-generated 30-day care plans with gold-standard guidelines, and by applying it in parallel to mental health, nutrition, exercise, and physical therapy scenarios, our study delivers the first prompt-sensitive audit showing where current LLMs exceed, match, or fall short of multidisciplinary best practices.