Background <p>This study aimed to evaluate the performance of four large language models (LLMs)—ChatGPT, Gemini, Copilot, and Claude—in responding to upper eyelid blepharoplasty-related questions, focusing on medical accuracy, clinical relevance, response length, and readability<b>.</b></p> Methods <p>A set of queries regarding upper eyelid blepharoplasty, covering six categories (anatomy, surgical procedure, additional intraoperative procedures, postoperative monitoring, follow-up, and postoperative complications) were posed to each LLM. An identical prompt establishing clinical context was provided before each question. Responses were evaluated by three ophthalmologists using a 5-point Likert scale for medical accuracy and a 3-point Likert scale for clinical relevance. The length of the responses was assessed. Readability was also evaluated using the Flesch Reading Ease Score, Flesch-Kincaid Grade Level, Coleman-Liau Index, Gunning Fog Index, and Simple Measure of Gobbledygook grade.</p> Results <p>A total of 30 standardized questions were presented to each LLM. None of the responses from any LLM received a score of 1 regarding medical accuracy for any question. ChatGPT achieved an 80% ‘highly accurate’ response rate, followed by Claude (60%), Gemini (40%), and Copilot (20%). None of the responses from ChatGPT and Claude received a score of&#xa0;1 regarding clinical relevance, whereas 10% of Gemini’s responses and 26.7% of Copilot’s responses received a score of 1. ChatGPT also provided the most clinically ‘relevant’ responses (86.7%), outperforming the other LLMs. Copilot generated the shortest responses, while ChatGPT generated the longest. Readability analyses revealed that all responses required advanced reading skills at a ‘college graduate’ level or higher, with Copilot’s responses being the most complex.</p> Conclusion <p>ChatGPT demonstrated superior performance in both medical accuracy and clinical relevance among evaluated LLMs regarding upper eyelid blepharoplasty, particularly excelling in postoperative monitoring and follow-up categories. While all models generated complex texts requiring advanced literacy, ChatGPT’s detailed responses offer valuable guidance for ophthalmologists managing upper eyelid blepharoplasty cases.</p> Level of Evidence V <p>This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors <a href="http://www.springer.com/00266">www.springer.com/00266</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Accuracy of ChatGPT, Gemini, Copilot, and Claude to Blepharoplasty-Related Questions

  • Seher Köksaldı,
  • Mustafa Kayabaşı,
  • Ceren Durmaz Engin,
  • Andrzej Grzybowski

摘要

Background

This study aimed to evaluate the performance of four large language models (LLMs)—ChatGPT, Gemini, Copilot, and Claude—in responding to upper eyelid blepharoplasty-related questions, focusing on medical accuracy, clinical relevance, response length, and readability.

Methods

A set of queries regarding upper eyelid blepharoplasty, covering six categories (anatomy, surgical procedure, additional intraoperative procedures, postoperative monitoring, follow-up, and postoperative complications) were posed to each LLM. An identical prompt establishing clinical context was provided before each question. Responses were evaluated by three ophthalmologists using a 5-point Likert scale for medical accuracy and a 3-point Likert scale for clinical relevance. The length of the responses was assessed. Readability was also evaluated using the Flesch Reading Ease Score, Flesch-Kincaid Grade Level, Coleman-Liau Index, Gunning Fog Index, and Simple Measure of Gobbledygook grade.

Results

A total of 30 standardized questions were presented to each LLM. None of the responses from any LLM received a score of 1 regarding medical accuracy for any question. ChatGPT achieved an 80% ‘highly accurate’ response rate, followed by Claude (60%), Gemini (40%), and Copilot (20%). None of the responses from ChatGPT and Claude received a score of 1 regarding clinical relevance, whereas 10% of Gemini’s responses and 26.7% of Copilot’s responses received a score of 1. ChatGPT also provided the most clinically ‘relevant’ responses (86.7%), outperforming the other LLMs. Copilot generated the shortest responses, while ChatGPT generated the longest. Readability analyses revealed that all responses required advanced reading skills at a ‘college graduate’ level or higher, with Copilot’s responses being the most complex.

Conclusion

ChatGPT demonstrated superior performance in both medical accuracy and clinical relevance among evaluated LLMs regarding upper eyelid blepharoplasty, particularly excelling in postoperative monitoring and follow-up categories. While all models generated complex texts requiring advanced literacy, ChatGPT’s detailed responses offer valuable guidance for ophthalmologists managing upper eyelid blepharoplasty cases.

Level of Evidence V

This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266.