Background <p>Large language models (LLMs) have shown promising performance in medical knowledge assessment; however, their capacity to assist real-world clinical decision-making in rhinology, particularly when integrating clinical, endoscopic, and radiologic data remains insufficiently investigated.</p> Objective <p>To evaluate the diagnostic classification and management recommendations of a large language model using real-world rhinologic cases that incorporate comprehensive clinical and written CT reports findings.</p> Methods <p>This retrospective, single-center diagnostic accuracy study included 301 adults with sinonasal disease evaluated at a tertiary care center. Structured clinical cases encompassing symptoms, endoscopic findings, and written CT reports findings were entered into ChatGPT-4 and compared with physician diagnoses and management plans serving as the reference standard. Diagnostic performance was assessed using sensitivity, specificity, accuracy, area under the curve, and Cohen’s kappa coefficient.</p> Results <p>Diagnostic performance varied by disease phenotype, with a marked discrepancy between CRSwNP and CRSsNP.</p> <p>The model demonstrated high sensitivity and accuracy for chronic rhinosinusitis with nasal polyps (sensitivity, 95.3%; accuracy, 89.0%), whereas performance for chronic rhinosinusitis without nasal polyps showed lower sensitivity (58.3%) but high specificity (97.7%). This clinically important discrepancy indicates that the model performs more reliably in phenotypes with distinct endoscopic and radiologic features than in those requiring more nuanced clinical interpretation. In contrast, management recommendations demonstrated consistently high accuracy for both medical and surgical decisions (overall accuracy, 92.7%), with substantial agreement between AI-generated and physician recommendations.</p> Conclusion <p>When evaluated using real-world rhinologic cases integrating clinical, endoscopic, and written CT reports findings, a large language model demonstrated strong diagnostic agreement for CRSwNP and consistently reliable management recommendations. However, further validation is required before routine clinical implementation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Real-world diagnostic and management decisions in rhinology by a large language model: a retrospective study

  • Noura Farhan Alanazi,
  • Tariq Abdulaziz Aldawood,
  • Abdulrahman Alfayez,
  • Naif AlOsaimi,
  • Riyadh Alhedaithy

摘要

Background

Large language models (LLMs) have shown promising performance in medical knowledge assessment; however, their capacity to assist real-world clinical decision-making in rhinology, particularly when integrating clinical, endoscopic, and radiologic data remains insufficiently investigated.

Objective

To evaluate the diagnostic classification and management recommendations of a large language model using real-world rhinologic cases that incorporate comprehensive clinical and written CT reports findings.

Methods

This retrospective, single-center diagnostic accuracy study included 301 adults with sinonasal disease evaluated at a tertiary care center. Structured clinical cases encompassing symptoms, endoscopic findings, and written CT reports findings were entered into ChatGPT-4 and compared with physician diagnoses and management plans serving as the reference standard. Diagnostic performance was assessed using sensitivity, specificity, accuracy, area under the curve, and Cohen’s kappa coefficient.

Results

Diagnostic performance varied by disease phenotype, with a marked discrepancy between CRSwNP and CRSsNP.

The model demonstrated high sensitivity and accuracy for chronic rhinosinusitis with nasal polyps (sensitivity, 95.3%; accuracy, 89.0%), whereas performance for chronic rhinosinusitis without nasal polyps showed lower sensitivity (58.3%) but high specificity (97.7%). This clinically important discrepancy indicates that the model performs more reliably in phenotypes with distinct endoscopic and radiologic features than in those requiring more nuanced clinical interpretation. In contrast, management recommendations demonstrated consistently high accuracy for both medical and surgical decisions (overall accuracy, 92.7%), with substantial agreement between AI-generated and physician recommendations.

Conclusion

When evaluated using real-world rhinologic cases integrating clinical, endoscopic, and written CT reports findings, a large language model demonstrated strong diagnostic agreement for CRSwNP and consistently reliable management recommendations. However, further validation is required before routine clinical implementation.