The Utility of ChatGPT for Medical Students on Plastic and Reconstructive Surgery Rotations
摘要
Subinternships are a key component of the competitive plastic and reconstructive surgery (PRS) match. While preparedness is highly valued, there is no standardized method for students to prepare for surgical cases. ChatGPT, a widely used large language model (LLM), has shown promise in medical education, but its reliability in surgical contexts remains unclear. We aim to evaluate ChatGPT’s potential to help students prepare for PRS rotations and subinternships by analyzing its responses to common PRS questions.
MethodsA dataset of 267 questions was generated across PRS subtopics. Three procedures were selected per subtopic, each with 12 identical questions categorized by type and format. Each prompt began with, "I am a medical student preparing for my PRS rotation," to contextualize responses. GPT-4o responses were scored for accuracy, completeness, usefulness, relevance, and overall quality with a five-point Likert scale. Low-scoring responses were re-queried using OpenA1 o1.
ResultsTwenty-one responses about pre-case reading were excluded due to fabricated articles (average score 1.17). Among the included responses, scores were: accuracy 4.12, completeness 3.88, usefulness 3.96, relevance 4.19, and quality 4.00. Lymphatics scored highest, while head and neck were lowest (p=0.03). Educational questions outperformed surgical questions (p<0.0001). No significant difference was found between question formats. Re-queried responses showed improved scores (p = 0.0001).
ConclusionChatGPT performed best in lymphatics and providing educational tools, but had limitations in generating literature and procedural questions. The newer o1 version demonstrated significantly improved performance, suggesting that with continued model refinement, LLMs are poised to become increasingly valuable in surgical education.
Level of Evidence VThis journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266.