Objective <p>This study evaluated the accuracy and comprehensiveness of responses generated by ChatGPT-4o and DeepSeek regarding commonly asked questions about myopia.</p> Methods <p>Thirty myopia-related questions spanning six clinical domains were submitted to both chatbots. Three medical professionals independently rated each response for accuracy and comprehensiveness. Inter-rater reliability was assessed using Fleiss’ Kappa, and Shapiro-Wilk tests were conducted to examine normality in rating distributions. Statistical comparisons were performed using the Chi-square test, with significance set at <i>p</i> &lt; 0.05.</p> Results <p>DeepSeek outperformed ChatGPT-4o in overall accuracy, with significantly more responses rated as “Good” (<i>p</i> &lt; 0.0001). Both models demonstrated high comprehensiveness scores when accuracy was rated “Good,” though performance declined in treatment-related queries, particularly regarding commercial products like DIMS lenses. Fleiss’ Kappa values indicated poor inter-rater agreement (DeepSeek: <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12886_2025_4328_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="11" /> </InlineMediaObject> <EquationSource Format="TEX">\(\:{\upkappa\:}\)</EquationSource> </InlineEquation> = 0.106; ChatGPT-4o: <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12886_2025_4328_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="11" /> </InlineMediaObject> <EquationSource Format="TEX">\(\:{\upkappa\:}\)</EquationSource> </InlineEquation> = − 0.0221), and normality tests showed non-normal score distributions (<i>p</i> &lt; 0.0001 across domains).</p> Conclusion <p>Both ChatGPT-4o and DeepSeek can deliver useful responses to myopia-related questions, though limitations remain in areas requiring up-to-date, region-specific treatment information. DeepSeek’s stronger performance suggests that localized LLMs may offer competitive advantages. Ongoing refinement, regular data updates, and domain-specific fine-tuning are essential for improving the reliability of AI chatbots in clinical communication.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmark analysis of myopia-related issues using large language models: a comparison of ChatGPT-4o and deepseek

  • Jinglei Yao,
  • Sun Chen Hsin,
  • Luxi Li,
  • Xiaofang Ren,
  • Wen Liu

摘要

Objective

This study evaluated the accuracy and comprehensiveness of responses generated by ChatGPT-4o and DeepSeek regarding commonly asked questions about myopia.

Methods

Thirty myopia-related questions spanning six clinical domains were submitted to both chatbots. Three medical professionals independently rated each response for accuracy and comprehensiveness. Inter-rater reliability was assessed using Fleiss’ Kappa, and Shapiro-Wilk tests were conducted to examine normality in rating distributions. Statistical comparisons were performed using the Chi-square test, with significance set at p < 0.05.

Results

DeepSeek outperformed ChatGPT-4o in overall accuracy, with significantly more responses rated as “Good” (p < 0.0001). Both models demonstrated high comprehensiveness scores when accuracy was rated “Good,” though performance declined in treatment-related queries, particularly regarding commercial products like DIMS lenses. Fleiss’ Kappa values indicated poor inter-rater agreement (DeepSeek: \(\:{\upkappa\:}\) = 0.106; ChatGPT-4o: \(\:{\upkappa\:}\) = − 0.0221), and normality tests showed non-normal score distributions (p < 0.0001 across domains).

Conclusion

Both ChatGPT-4o and DeepSeek can deliver useful responses to myopia-related questions, though limitations remain in areas requiring up-to-date, region-specific treatment information. DeepSeek’s stronger performance suggests that localized LLMs may offer competitive advantages. Ongoing refinement, regular data updates, and domain-specific fine-tuning are essential for improving the reliability of AI chatbots in clinical communication.