Purpose <p>Hyperthyroidism is a common endocrine disorder that requires long-term management, heavily relying on effective patient education. This study aims to evaluate the application value of three mainstream large language models (LLMs) ChatGPT, Gemini, and DeepSeek in the education of hyperthyroidism patients.</p> Methods <p>We developed a standardized question bank containing 20 issues related to hyperthyroidism. The three LLMs were prompted to generate responses for each question. Five endocrinology experts performed a double-blind evaluation using a Likert 5-point scale across five dimensions: relevance, accuracy, comprehensibility, comprehensiveness, and humanistic care. Additionally, the Flesch-Kincaid readability formula was used to analyze text complexity.</p> Results <p>The scores of the three large language models demonstrated statistically significant differences (p &lt; 0.05). DeepSeek achieved the highest scores in relevance (4.80 ± 0.45), accuracy (4.59 ± 0.57), and comprehensiveness (4.70 ± 0.46). Gemini excelled in comprehensibility (4.49 ± 0.58) and humanistic care (4.27 ± 0.33). Notably, The Flesch Reading Ease Index classified the text generated by all three LLMs as ‘Difficult’ to read (DeepSeek: 38, Gemini: 38, ChatGPT: 37).</p> Conclusions <p>LLMs show potential in hyperthyroidism patient education, but there is still room for improvement across various dimensions. Patients and healthcare professionals should consider these models as supplementary tools rather than replacements for professional medical personnel. Further research is needed to explore the clinical application of artificial intelligence in healthcare.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of large language models in patient education for hyperthyroidism: A comparative study of chatgpt, gemini, and deepseek

  • Xin Hu,
  • Ruhai Lin,
  • Yinqiong Huang,
  • Xiaohong Wu,
  • Yiran Gong,
  • Huibin Huang

摘要

Purpose

Hyperthyroidism is a common endocrine disorder that requires long-term management, heavily relying on effective patient education. This study aims to evaluate the application value of three mainstream large language models (LLMs) ChatGPT, Gemini, and DeepSeek in the education of hyperthyroidism patients.

Methods

We developed a standardized question bank containing 20 issues related to hyperthyroidism. The three LLMs were prompted to generate responses for each question. Five endocrinology experts performed a double-blind evaluation using a Likert 5-point scale across five dimensions: relevance, accuracy, comprehensibility, comprehensiveness, and humanistic care. Additionally, the Flesch-Kincaid readability formula was used to analyze text complexity.

Results

The scores of the three large language models demonstrated statistically significant differences (p < 0.05). DeepSeek achieved the highest scores in relevance (4.80 ± 0.45), accuracy (4.59 ± 0.57), and comprehensiveness (4.70 ± 0.46). Gemini excelled in comprehensibility (4.49 ± 0.58) and humanistic care (4.27 ± 0.33). Notably, The Flesch Reading Ease Index classified the text generated by all three LLMs as ‘Difficult’ to read (DeepSeek: 38, Gemini: 38, ChatGPT: 37).

Conclusions

LLMs show potential in hyperthyroidism patient education, but there is still room for improvement across various dimensions. Patients and healthcare professionals should consider these models as supplementary tools rather than replacements for professional medical personnel. Further research is needed to explore the clinical application of artificial intelligence in healthcare.