Introduction <p>The study assesses the quality, readability, reliability, and usefulness of exercise-related information generated by two large language models (LLMs), ChatGPT-4 and DeepSeek-V3, in response to frequently asked questions by patients with ankylosing spondylitis (AS).</p> Method <p>This cross-sectional comparative study developed a structured assessment framework using a set of exercise and rehabilitation-related questions, distributed across four key domains: exercise and physical activity (C1; 33 items), posture and mobility (C2; 6 items), breathing and pulmonary health (C3; 6 items), and general topics (C4; 5 items). Information quality was assessed using the modified DISCERN (mDISCERN) tool, while content reliability was evaluated with the Reliability Score and perceived usefulness was measured using the Usefulness Score. Readability was assessed using the Flesch Reading Ease (FRE) scale. Three independent physiotherapists with expertise in rheumatologic rehabilitation independently evaluated the responses.</p> Results <p>In total score comparisons, DeepSeek-V3 achieved significantly higher scores than ChatGPT-4 on the mDISCERN (4(3–4) vs. 3(3–3); <i>p</i> &lt; 0.001), reliability (5(5–6) vs. 5(4–5); <i>p</i> &lt; 0.001), and usefulness (6(5–6) vs. 5(5–6); <i>p</i> &lt; 0.001). Domain-specific analysis showed higher usefulness scores for DeepSeek-V3 in C1 (<i>p</i> = 0.004), C2 (<i>p</i> = 0.019), and C4 (<i>p</i> = 0.005). Mean FRE scores were 30.4 ± 14.37 for ChatGPT-4 and 28.77 ± 17.77 for DeepSeek-V3, both classified as very difficult (<i>p</i> &gt; 0.05).</p> Conclusion <p>This study highlighted that responses generated by DeepSeek-V3 related to AS were generally more accurate and demonstrated greater reliability compared to those produced by ChatGPT-4. However, the complex language used by both LLMs may reduce accessibility for patients with limited health literacy. These limitations highlight the importance of healthcare professional oversight in exercise planning.</p> <p><Table Float="No" ID="Taba"> <tgroup cols="2"> <colspec align="left" colname="c1" colnum="1" /> <colspec align="left" colname="c2" colnum="2" /> <tbody> <row> <entry align="left" nameend="c2" namest="c1"> <p>Key Points</p> <p><i>• DeepSeek-V3 provided more accurate and reliable responses than ChatGPT-4 regarding exercise in AS.</i></p> <p><i>• Domain-specific analysis showed DeepSeek-V3 was particularly more useful in exercise, posture, and general topics.</i></p> <p><i>• Both LLMs generated content with very difficult readability, requiring college-level comprehension.</i></p> <p><i>• Healthcare professional supervision is essential when using LLMs in patient education.</i></p> </entry> </row> </tbody> </tgroup> </Table></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ChatGPT-4 vs. DeepSeek-V3: a comparative study of response quality, reliability, usefulness, and readability for exercise and rehabilitation strategies in patients with ankylosing spondylitis

  • Fulden Sari,
  • Zeliha Çelik,
  • Yasemin Mirza

摘要

Introduction

The study assesses the quality, readability, reliability, and usefulness of exercise-related information generated by two large language models (LLMs), ChatGPT-4 and DeepSeek-V3, in response to frequently asked questions by patients with ankylosing spondylitis (AS).

Method

This cross-sectional comparative study developed a structured assessment framework using a set of exercise and rehabilitation-related questions, distributed across four key domains: exercise and physical activity (C1; 33 items), posture and mobility (C2; 6 items), breathing and pulmonary health (C3; 6 items), and general topics (C4; 5 items). Information quality was assessed using the modified DISCERN (mDISCERN) tool, while content reliability was evaluated with the Reliability Score and perceived usefulness was measured using the Usefulness Score. Readability was assessed using the Flesch Reading Ease (FRE) scale. Three independent physiotherapists with expertise in rheumatologic rehabilitation independently evaluated the responses.

Results

In total score comparisons, DeepSeek-V3 achieved significantly higher scores than ChatGPT-4 on the mDISCERN (4(3–4) vs. 3(3–3); p < 0.001), reliability (5(5–6) vs. 5(4–5); p < 0.001), and usefulness (6(5–6) vs. 5(5–6); p < 0.001). Domain-specific analysis showed higher usefulness scores for DeepSeek-V3 in C1 (p = 0.004), C2 (p = 0.019), and C4 (p = 0.005). Mean FRE scores were 30.4 ± 14.37 for ChatGPT-4 and 28.77 ± 17.77 for DeepSeek-V3, both classified as very difficult (p > 0.05).

Conclusion

This study highlighted that responses generated by DeepSeek-V3 related to AS were generally more accurate and demonstrated greater reliability compared to those produced by ChatGPT-4. However, the complex language used by both LLMs may reduce accessibility for patients with limited health literacy. These limitations highlight the importance of healthcare professional oversight in exercise planning.

Key Points

• DeepSeek-V3 provided more accurate and reliable responses than ChatGPT-4 regarding exercise in AS.

• Domain-specific analysis showed DeepSeek-V3 was particularly more useful in exercise, posture, and general topics.

• Both LLMs generated content with very difficult readability, requiring college-level comprehension.

• Healthcare professional supervision is essential when using LLMs in patient education.