Background <p> Artificial intelligence (AI) models such as ChatGPT and DeepSeek have gained increasing attention for their potential to enhance patient education by delivering accessible and evidence-based health information. We designed the following study to evaluate the AI models—ChatGPT and DeepSeek—in generating patient education materials for bariatric surgery.</p> Methods <p>Thirty commonly asked patient questions related to bariatric surgery were classified into four thematic domains: (1) surgical planning and technical considerations, (2) preoperative assessment and optimization, (3) postoperative care and complication management, and (4) long-term follow-up and disease management. Responses generated by ChatGPT and DeepSeek were evaluated using three key metrics: (1) response quality, assessed by the Global Quality Score, rated on a 5-point scale from 1 (poor) to 5 (excellent); (2) reliability, measured using modified DISCERN criteria, which assess adherence to clinical guidelines and evidence-based standards, with scores ranging from 5 (low) to 25 (high); and (3) readability, evaluated using two validated formulas: the Flesch-Kincaid Grade Level and the Flesch Reading Ease Score.</p> Results <p>ChatGPT significantly outperformed DeepSeek in response quality, with a median (IQR) Global Quality Score of 5.00 (4.00, 5.00) vs. 4.00 (4.00, 5.00) (<i>P</i> = 0.002). Higher reliability was also observed in ChatGPT, as reflected by mDISCERN scores across all four domains (median [IQR], 22.0 [21.0, 23.25] vs. 19.7 [19.0, 20.75]; <i>P</i> &lt; 0.001). While no significant difference was found in the Flesch Reading Ease Score (mean [SD], 26.11 [12.84] vs. 20.87 [12.20]; <i>P</i> = 0.110), ChatGPT yielded significantly higher Flesch-Kincaid Grade Level Scores (meaning its text was more complex) (mean [SD], 16.40 [2.43] vs. 13.48 [2.35]; <i>P</i> &lt; 0.001). Both models produced responses at a readability level corresponding to college education.</p> Conclusions <p>ChatGPT provided higher-quality and more reliable responses, while DeepSeek’s answers were slightly easier to read. However, both models’ answers lacked attention to psychosocial and cultural aspects of patient care, highlighting the need for more empathetic, adaptive AI to support inclusive patient education.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Artificial Intelligence-Generated Patient Education Materials for Bariatric Surgery: Comparative Analysis of Response Quality, Reliability, and Readability Across ChatGPT and DeepSeek Models

  • Shuai Guo,
  • Cheng-Li Yang,
  • Xiang-Ping Lin,
  • Man Jiang,
  • Ji Chen,
  • Kang-Xiu Tuo,
  • Wei-Wei Yang,
  • Qian Wang,
  • Xiang-Ren Jin,
  • Pei Li

摘要

Background

Artificial intelligence (AI) models such as ChatGPT and DeepSeek have gained increasing attention for their potential to enhance patient education by delivering accessible and evidence-based health information. We designed the following study to evaluate the AI models—ChatGPT and DeepSeek—in generating patient education materials for bariatric surgery.

Methods

Thirty commonly asked patient questions related to bariatric surgery were classified into four thematic domains: (1) surgical planning and technical considerations, (2) preoperative assessment and optimization, (3) postoperative care and complication management, and (4) long-term follow-up and disease management. Responses generated by ChatGPT and DeepSeek were evaluated using three key metrics: (1) response quality, assessed by the Global Quality Score, rated on a 5-point scale from 1 (poor) to 5 (excellent); (2) reliability, measured using modified DISCERN criteria, which assess adherence to clinical guidelines and evidence-based standards, with scores ranging from 5 (low) to 25 (high); and (3) readability, evaluated using two validated formulas: the Flesch-Kincaid Grade Level and the Flesch Reading Ease Score.

Results

ChatGPT significantly outperformed DeepSeek in response quality, with a median (IQR) Global Quality Score of 5.00 (4.00, 5.00) vs. 4.00 (4.00, 5.00) (P = 0.002). Higher reliability was also observed in ChatGPT, as reflected by mDISCERN scores across all four domains (median [IQR], 22.0 [21.0, 23.25] vs. 19.7 [19.0, 20.75]; P < 0.001). While no significant difference was found in the Flesch Reading Ease Score (mean [SD], 26.11 [12.84] vs. 20.87 [12.20]; P = 0.110), ChatGPT yielded significantly higher Flesch-Kincaid Grade Level Scores (meaning its text was more complex) (mean [SD], 16.40 [2.43] vs. 13.48 [2.35]; P < 0.001). Both models produced responses at a readability level corresponding to college education.

Conclusions

ChatGPT provided higher-quality and more reliable responses, while DeepSeek’s answers were slightly easier to read. However, both models’ answers lacked attention to psychosocial and cultural aspects of patient care, highlighting the need for more empathetic, adaptive AI to support inclusive patient education.