Efficacy of large language models in detecting postoperative delirium from unstructured clinical notes: A retrospective cohort study
摘要
Early identification of postoperative delirium (POD) remains challenging. This retrospective observational study compared the performance of large language models (LLMs), Llama-3-70B and GPT-4o, and physicians in predicting clinically significant POD, defined as either requiring antipsychotics or diagnosis of delirium by neurologists following consultation for delirium-related symptoms. The c-statistics of Llama-3-70B and GPT-4o were 0.74 and 0.76, respectively. LLMs showed higher sensitivity (Llama-3-70B, 0.900; GPT-4o, 0.868; physicians, 0.723) and lower specificity (0.463, 0.547, and 0.814, respectively) than physicians. Inter-rater agreement was almost perfect for both Llama-3-70B and GPT-4o (Fleiss’ kappa = 0.852 and 0.854, respectively) but fair for physicians (0.219). Both LLMs detected clinically significant POD approximately one day earlier than physicians (Kaplan-Meier analysis, median time to diagnosis: Llama-3-70B, 34.5 h; GPT-4o, 37.5 h; physicians, 62.9 h; log-rank P < 0.001). The integration of LLMs as a complementary screening tool under physician supervision may improve the early, reproducible diagnosis of clinically significant POD.