From OR dialogue to actionable feedback: a scalable method for analyzing intraoperative teaching
摘要
Intraoperative feedback is essential to surgical training, yet it is transient, inconsistently documented, and difficult to analyze at scale. Traditional qualitative methods for studying operative dialogue are labor-intensive, limiting their use for educational research and surgical trainee feedback. This study describes the development and evaluation of an AI-assisted qualitative coding pipeline for analyzing intraoperative attending-trainee communication. We frame this work as a proof-of-concept feasibility study.
MethodsAudio-recorded dialogue from 25 general and colorectal surgical operations was transcribed and formatted for line-by-line coding using a 38-category thematic codebook. A custom GPT built on GPT-4o with retrieval-augmented generation was developed to assign codes to each utterance within a human-in-the-loop workflow. AI reliability was assessed through run-to-run agreement on a purposive sample of five transcripts. Coding efficiency and error rates were tracked across AI-assisted and manual workflows.
ResultsThe AI-assisted workflow reduced coding time to less than 20% of manual effort, with an overall error rate of 2.0%. Run-to-run AI agreement across 2,245 lines demonstrated 79% agreement with substantial concordance (κ = 0.75). Agreement was highest for explicit utterance types and lowest for codes requiring contextual interpretation.
ConclusionsIn this proof-of-concept evaluation, an AI-supported, human-verified coding workflow achieved the efficiency and reliability needed to make scalable analysis of intraoperative teaching dialogue feasible. This approach has the potential to convert routine operative communication into structured, actionable feedback for trainees and faculty.