<p>Adult type 1 diabetes mellitus (T1DM) involves complex diagnosis, treatment, and long-term self-management, creating a need for accurate and accessible health education. Large language models (LLMs) are increasingly used for medical information seeking, yet their accuracy and consistency in adult T1DM-related queries remain insufficiently evaluated. A guideline-based comparative evaluation assessed DeepSeek-V3.2 and ChatGPT-5.0 using 22 English-language prompts derived from the 2021 ADA/EASD consensus report, covering basic knowledge, diagnosis and differential diagnosis, treatment, and complications. The prompts were submitted to both models twice, two weeks apart. Responses were independently evaluated by two blinded endocrinology specialists using a predefined four-point scoring rubric, with disagreements adjudicated by a third senior endocrinologist. Short-term consistency was assessed by expert judgment and TF-IDF cosine similarity. Inter-rater agreement was good (Cohen’s κ = 0.71). Expert-judged consistency was 95.45% (21/22) for both models; TF-IDF cosine similarity was 0.52 ± 0.09 for DeepSeek and 0.54 ± 0.09 for ChatGPT. Overall accuracy scores were 3.59 ± 0.59 and 3.77 ± 0.43, respectively, with no statistically significant difference (p = 0.102). Comprehensive ratings accounted for 63.64% and 77.27%, respectively, and mixed correct and incorrect or outdated information accounted for 4.55% and 0.00%. Both models may support adult T1DM-related health education, but outputs should be interpreted as supplementary educational material under professional guidance rather than as diagnostic or therapeutic advice.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of the accuracy and consistency of DeepSeek and ChatGPT in addressing type 1 diabetes mellitus-related queries in adults

  • Junyun Feng,
  • Xiao Fei,
  • Tingyi Qian,
  • Huilan Luo,
  • Suying Wang,
  • Yihuan Cai,
  • Yaobin Ouyang,
  • Foqiang Liao,
  • Lingyan Zhu

摘要

Adult type 1 diabetes mellitus (T1DM) involves complex diagnosis, treatment, and long-term self-management, creating a need for accurate and accessible health education. Large language models (LLMs) are increasingly used for medical information seeking, yet their accuracy and consistency in adult T1DM-related queries remain insufficiently evaluated. A guideline-based comparative evaluation assessed DeepSeek-V3.2 and ChatGPT-5.0 using 22 English-language prompts derived from the 2021 ADA/EASD consensus report, covering basic knowledge, diagnosis and differential diagnosis, treatment, and complications. The prompts were submitted to both models twice, two weeks apart. Responses were independently evaluated by two blinded endocrinology specialists using a predefined four-point scoring rubric, with disagreements adjudicated by a third senior endocrinologist. Short-term consistency was assessed by expert judgment and TF-IDF cosine similarity. Inter-rater agreement was good (Cohen’s κ = 0.71). Expert-judged consistency was 95.45% (21/22) for both models; TF-IDF cosine similarity was 0.52 ± 0.09 for DeepSeek and 0.54 ± 0.09 for ChatGPT. Overall accuracy scores were 3.59 ± 0.59 and 3.77 ± 0.43, respectively, with no statistically significant difference (p = 0.102). Comprehensive ratings accounted for 63.64% and 77.27%, respectively, and mixed correct and incorrect or outdated information accounted for 4.55% and 0.00%. Both models may support adult T1DM-related health education, but outputs should be interpreted as supplementary educational material under professional guidance rather than as diagnostic or therapeutic advice.