Objectives <p>This study aims to evaluate and compare the responses of two large language model (LLM) AI chatbots, ChatGPT and Gemini, against those provided by expert surgeons during consultations for revision rhinoplasty. Given the emotional complexities and relatively low satisfaction rates in revision cases, assessing AI’s effectiveness in providing empathetic and accurate information is essential.</p> Materials and Methods <p>A set of fifteen hypothetical questions reflecting patient concerns were presented to ChatGPT, Gemini, and two expert surgeons. Four academic otolaryngologists rated the responses based on empathy, precision, perfectness, and communication skills using a 5-point Likert scale. The ratings were analyzed using one-way ANOVA and Bonferroni tests to determine statistical significance.</p> Results <p>ChatGPT achieved the highest mean scores across all categories, outperforming both expert surgeons significantly in empathy, precision, perfectness, and communication skills (<i>p</i>&#xa0;&lt;&#xa0;0.01). Gemini also outperformed the expert surgeons in these categories. Notably, ChatGPT excelled in perfectness compared to Gemini, while expert surgeon1 demonstrated superior precision. Evaluators showed consistent ratings in precision, perfectness, and communication skills, but significant differences were found in empathy (<i>p</i>&#xa0;&lt;&#xa0;0.01).</p> Conclusion <p>ChatGPT and Gemini showed remarkable performance in consultation for revision rhinoplasty. However, there are known weak points in LLM chatbots; they can play an under-controlled role in facial plastic surgery and the healthcare system.</p> Level of Evidence IV <p>This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors <a href="http://www.springer.com/00266">www.springer.com/00266</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Insights of ChatGPT, Gemini and Expert Surgeons in Revision Rhinoplasty Consultation

  • Shahin Bastaninejad,
  • Samira Alipour,
  • Luiz Carlos Ishida,
  • Mahdieh Mohebbi,
  • Benyamin Mousavi-asl,
  • Farrokh Heidari,
  • Habib Azimi

摘要

Objectives

This study aims to evaluate and compare the responses of two large language model (LLM) AI chatbots, ChatGPT and Gemini, against those provided by expert surgeons during consultations for revision rhinoplasty. Given the emotional complexities and relatively low satisfaction rates in revision cases, assessing AI’s effectiveness in providing empathetic and accurate information is essential.

Materials and Methods

A set of fifteen hypothetical questions reflecting patient concerns were presented to ChatGPT, Gemini, and two expert surgeons. Four academic otolaryngologists rated the responses based on empathy, precision, perfectness, and communication skills using a 5-point Likert scale. The ratings were analyzed using one-way ANOVA and Bonferroni tests to determine statistical significance.

Results

ChatGPT achieved the highest mean scores across all categories, outperforming both expert surgeons significantly in empathy, precision, perfectness, and communication skills (p < 0.01). Gemini also outperformed the expert surgeons in these categories. Notably, ChatGPT excelled in perfectness compared to Gemini, while expert surgeon1 demonstrated superior precision. Evaluators showed consistent ratings in precision, perfectness, and communication skills, but significant differences were found in empathy (p < 0.01).

Conclusion

ChatGPT and Gemini showed remarkable performance in consultation for revision rhinoplasty. However, there are known weak points in LLM chatbots; they can play an under-controlled role in facial plastic surgery and the healthcare system.

Level of Evidence IV

This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266.