Introduction <p>Operative notes play a critical role in documenting surgical procedures and supporting medical communication. However, due to their technical language, these documents are often complex and difficult to understand for patients, non-medical individuals, and even some healthcare professionals. Large Language Models (LLMs) offer a novel opportunity to simplify such documents and make them more accessible. This study aims to quantify how six LLMs simplify otolaryngology operative notes and to compare readability, clinical accuracy and clarity.</p> Materials and methods <p>In this study, 39 fictional operative notes specific to otolaryngologic surgery were simplified using six LLMs (GPT-4, GPT-4o, Claude 3.7, Gemini 2.0, DeepSeek, and Microsoft Copilot). The outputs were analyzed using eight different readability metrics and evaluated by two expert physicians in terms of medical accuracy and comprehensibility. Correlation analyses were also conducted across clinical subgroups (rhinology, otology, head and neck surgery).</p> Results <p>Claude 3.7 produced the most complex outputs, whereas GPT-4o, Gemini, and DeepSeek generated the most readable texts. According to expert evaluations, GPT-4 achieved the highest scores for medical accuracy, while GPT-4o received the highest ratings for clarity. Model performance varied across clinical subgroups.</p> Conclusion <p>LLMs are effective tools for simplifying medical texts; however, model selection should consider the target audience and clinical context, and all outputs must be verified by medical experts. When used in a controlled and validated manner, LLMs may contribute significantly to a new era of health communication.</p> Level of evidence <p>N/A.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Can LLMs simplify operative notes? A comparative analysis in otorhinolaryngology

  • Ahmet Ufuk Kılıçtaş,
  • Oğuz Gül,
  • Bilgeşah Kılıçtaş,
  • Esat Kaba,
  • Basar Erdivanli

摘要

Introduction

Operative notes play a critical role in documenting surgical procedures and supporting medical communication. However, due to their technical language, these documents are often complex and difficult to understand for patients, non-medical individuals, and even some healthcare professionals. Large Language Models (LLMs) offer a novel opportunity to simplify such documents and make them more accessible. This study aims to quantify how six LLMs simplify otolaryngology operative notes and to compare readability, clinical accuracy and clarity.

Materials and methods

In this study, 39 fictional operative notes specific to otolaryngologic surgery were simplified using six LLMs (GPT-4, GPT-4o, Claude 3.7, Gemini 2.0, DeepSeek, and Microsoft Copilot). The outputs were analyzed using eight different readability metrics and evaluated by two expert physicians in terms of medical accuracy and comprehensibility. Correlation analyses were also conducted across clinical subgroups (rhinology, otology, head and neck surgery).

Results

Claude 3.7 produced the most complex outputs, whereas GPT-4o, Gemini, and DeepSeek generated the most readable texts. According to expert evaluations, GPT-4 achieved the highest scores for medical accuracy, while GPT-4o received the highest ratings for clarity. Model performance varied across clinical subgroups.

Conclusion

LLMs are effective tools for simplifying medical texts; however, model selection should consider the target audience and clinical context, and all outputs must be verified by medical experts. When used in a controlled and validated manner, LLMs may contribute significantly to a new era of health communication.

Level of evidence

N/A.