Background <p>Children with microtia and their parents require comprehensive information to make informed decisions about treatment options.</p> Objective <p>We evaluated the effectiveness of various large language models (LLMs) in providing preoperative education for congenital microtia reconstruction (CMR) by analyzing their responses to related inquiries.</p> Methods <p>Ten plastic surgeons developed 13 CMR-related preoperative education strategies and input 14 text commands into Claude-3-Opus, GPT-4-Turbo, and Gemini-1.5-Pro during an online session. Five experts evaluated these language model’s responses for correctness, completeness, logic, and potential harm, while five postoperative patients’ parent reviewed the education materials for readability and value. All responses were also analyzed for readability using the context package.</p> Results <p>The results showed no statistically significant differences among Gemini, Claude, and GPT in the evaluation metrics of accuracy, completeness, and potential risk. In terms of logicality and overall rating, Gemini’s responses were significantly superior to GPT. Preoperative patient education materials generated by GPT received the highest DISCERN scores, significantly outperforming those from Claude and Gemini. From the perspective of patient’s parent’s, there are no statistically significant differences among Gemini, Claude, and GPT. Objective assessments of readability confirmed that Claude’s materials were easier to understand compared to those from the other models.</p> Conclusion <p>Claude-3-Opus, GPT-4-Turbo, and Gemini-1.5-Pro effectively addressed patient inquiries and produced clear pre-surgical education materials. However, these LLMs should not be used independently for patient education without expert supervision to ensure accuracy and completeness.</p> Level of Evidence IV <p>This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors <a href="http://www.springer.com/00266">www.springer.com/00266</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Utility of Large Language Models for Congenital Microtia Reconstruction Education: Comparison of the Performance of Claude, GPT, and Gemini

  • Zhifeng Liao,
  • Jiadong Huang,
  • Yukang Liu,
  • Fangwei Li,
  • Li Tang,
  • Liyao Cong,
  • Haibin Wang,
  • Sheng-kang Luo

摘要

Background

Children with microtia and their parents require comprehensive information to make informed decisions about treatment options.

Objective

We evaluated the effectiveness of various large language models (LLMs) in providing preoperative education for congenital microtia reconstruction (CMR) by analyzing their responses to related inquiries.

Methods

Ten plastic surgeons developed 13 CMR-related preoperative education strategies and input 14 text commands into Claude-3-Opus, GPT-4-Turbo, and Gemini-1.5-Pro during an online session. Five experts evaluated these language model’s responses for correctness, completeness, logic, and potential harm, while five postoperative patients’ parent reviewed the education materials for readability and value. All responses were also analyzed for readability using the context package.

Results

The results showed no statistically significant differences among Gemini, Claude, and GPT in the evaluation metrics of accuracy, completeness, and potential risk. In terms of logicality and overall rating, Gemini’s responses were significantly superior to GPT. Preoperative patient education materials generated by GPT received the highest DISCERN scores, significantly outperforming those from Claude and Gemini. From the perspective of patient’s parent’s, there are no statistically significant differences among Gemini, Claude, and GPT. Objective assessments of readability confirmed that Claude’s materials were easier to understand compared to those from the other models.

Conclusion

Claude-3-Opus, GPT-4-Turbo, and Gemini-1.5-Pro effectively addressed patient inquiries and produced clear pre-surgical education materials. However, these LLMs should not be used independently for patient education without expert supervision to ensure accuracy and completeness.

Level of Evidence IV

This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266.