Purpose <p>Large-language models (LLMs) are increasingly used for health advice, but their alignment with evidence-based guidelines and sensitivity to question phrasing remain unclear.</p> Methods <p>In May 2025, we evaluated ChatGPT 4.0, ChatGPT 4.5, and DeepSeek V3 using four clinical vignettes: major depression with polysubstance use, irritable bowel syndrome flare, new-onset hypertension requiring exercise counseling, and chronic low back pain. Each scenario was tested with clinician- and patient-style prompts, generating 24 responses. Outputs were benchmarked against 89 guideline-derived recommendations from three authoritative sources per domain. Two blinded reviewers scored concordance (1 = actionable detail, 0.5 = generic mention, 0 = absent), with adjudication by a third reviewer. Inter-rater reliability was measured using Cronbach’s α.</p> Results <p>ChatGPT 4.5 achieved the highest guideline concordance (61.9%), followed by DeepSeek V3 (60.7%) and ChatGPT 4.0 (53.7%). Performance varied by domain, exceeding 67% in mental health but dropping below 45% in nutrition. Prompt phrasing influenced capture rates, with clinician-style prompts improving scores in exercise and pain domains, while patient-style prompts outperformed in nutrition. Reviewer agreement was high (α = 0.97 for chatbot scoring; 0.80 for matrix coding).</p> Conclusion <p>LLMs can rapidly generate draft care plans that reflect clinical guidelines, though they favor generic over individualized advice. By introducing a unique, domain-agnostic scoring rubric that aligns AI-generated 30-day care plans with gold-standard guidelines, and by applying it in parallel to mental health, nutrition, exercise, and physical therapy scenarios, our study delivers the first prompt-sensitive audit showing where current LLMs exceed, match, or fall short of multidisciplinary best practices.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyses of different prescriptions for health using artificial intelligence: a critical approach based on the international guidelines of health institutions

  • Vítor Marcelo Soares Campos,
  • Tiago Paiva Prudente,
  • Luana Lemos Leão,
  • Maurício Silva da Costa,
  • Henrique Nunes Pereira Oliva,
  • Renato Sobral Monteiro-Junior

摘要

Purpose

Large-language models (LLMs) are increasingly used for health advice, but their alignment with evidence-based guidelines and sensitivity to question phrasing remain unclear.

Methods

In May 2025, we evaluated ChatGPT 4.0, ChatGPT 4.5, and DeepSeek V3 using four clinical vignettes: major depression with polysubstance use, irritable bowel syndrome flare, new-onset hypertension requiring exercise counseling, and chronic low back pain. Each scenario was tested with clinician- and patient-style prompts, generating 24 responses. Outputs were benchmarked against 89 guideline-derived recommendations from three authoritative sources per domain. Two blinded reviewers scored concordance (1 = actionable detail, 0.5 = generic mention, 0 = absent), with adjudication by a third reviewer. Inter-rater reliability was measured using Cronbach’s α.

Results

ChatGPT 4.5 achieved the highest guideline concordance (61.9%), followed by DeepSeek V3 (60.7%) and ChatGPT 4.0 (53.7%). Performance varied by domain, exceeding 67% in mental health but dropping below 45% in nutrition. Prompt phrasing influenced capture rates, with clinician-style prompts improving scores in exercise and pain domains, while patient-style prompts outperformed in nutrition. Reviewer agreement was high (α = 0.97 for chatbot scoring; 0.80 for matrix coding).

Conclusion

LLMs can rapidly generate draft care plans that reflect clinical guidelines, though they favor generic over individualized advice. By introducing a unique, domain-agnostic scoring rubric that aligns AI-generated 30-day care plans with gold-standard guidelines, and by applying it in parallel to mental health, nutrition, exercise, and physical therapy scenarios, our study delivers the first prompt-sensitive audit showing where current LLMs exceed, match, or fall short of multidisciplinary best practices.