Background <p>Manual chart review is labor-intensive and error-prone. Integrating Large Language Models (LLMs), like GPT, may reduce errors and accelerate the scalability of plastic surgery studies. This study compares the performance of GPT-3.5 Turbo and GPT-4 Turbo to manual review for second stage implant-based breast reconstructions requiring capsule revision.</p> Methods <p>This retrospective cohort study identified 101&#xa0;s stage breast reconstruction surgeries at Stanford Hospital (2018–2021) using CPT codes. Chart reviews collected patient demographics and operative details. Capsule procedures included capsulectomy, capsulorrhaphy, and ‘extensive’ capsulotomy, defined as any capsulotomy that changed the breast pocket position beyond entry into the capsule. GPT-3.5 Turbo and GPT-4 Turbo filtered operative reports and identified revisions. Performance metrics were averaged over 10 runs for each LLM.</p> Results <p>The GPT-3.5 Turbo model performed well in determining whether a second stage breast reconstruction occurred, incorrectly categorizing just one operative report. The model correctly predicted 85.2% of capsulectomies, 76.9% of capsulorrhaphies and 78.1% of extensive capsulotomies. The overall success rate was 80.1%, with a recall of 0.76, precision of 0.71, and F-score of 0.72. Re-running with GPT-4 Turbo considerably improved performance, correctly identifying 86.9% of capsulectomies, 88.3% of capsulorrhaphies, and 85.3% of extensive capsulotomies, with an overall success rate of 86.8%, recall of 0.94, precision of 0.69, and F-score of 0.78.</p> Conclusions <p>While GPT-4 Turbo outperformed GPT-3.5 Turbo, both models exhibit limitations in precisely matching manual chart review. These findings emphasize the potential as first pass tools, but highlight the need for cautious interpretation and complementary manual review in ensuring accurate data collection.</p> Level of Evidence <p>Not gradable.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating large language models for surgical chart review of second stage implant-based breast reconstruction: a comparative analysis of manual review, GPT-3.5 Turbo, and GPT-4 Turbo

  • Devi Lakhlani,
  • Dhruv Dadhania,
  • Rahim Nazerali

摘要

Background

Manual chart review is labor-intensive and error-prone. Integrating Large Language Models (LLMs), like GPT, may reduce errors and accelerate the scalability of plastic surgery studies. This study compares the performance of GPT-3.5 Turbo and GPT-4 Turbo to manual review for second stage implant-based breast reconstructions requiring capsule revision.

Methods

This retrospective cohort study identified 101 s stage breast reconstruction surgeries at Stanford Hospital (2018–2021) using CPT codes. Chart reviews collected patient demographics and operative details. Capsule procedures included capsulectomy, capsulorrhaphy, and ‘extensive’ capsulotomy, defined as any capsulotomy that changed the breast pocket position beyond entry into the capsule. GPT-3.5 Turbo and GPT-4 Turbo filtered operative reports and identified revisions. Performance metrics were averaged over 10 runs for each LLM.

Results

The GPT-3.5 Turbo model performed well in determining whether a second stage breast reconstruction occurred, incorrectly categorizing just one operative report. The model correctly predicted 85.2% of capsulectomies, 76.9% of capsulorrhaphies and 78.1% of extensive capsulotomies. The overall success rate was 80.1%, with a recall of 0.76, precision of 0.71, and F-score of 0.72. Re-running with GPT-4 Turbo considerably improved performance, correctly identifying 86.9% of capsulectomies, 88.3% of capsulorrhaphies, and 85.3% of extensive capsulotomies, with an overall success rate of 86.8%, recall of 0.94, precision of 0.69, and F-score of 0.78.

Conclusions

While GPT-4 Turbo outperformed GPT-3.5 Turbo, both models exhibit limitations in precisely matching manual chart review. These findings emphasize the potential as first pass tools, but highlight the need for cautious interpretation and complementary manual review in ensuring accurate data collection.

Level of Evidence

Not gradable.