Objective <p>To compare the performance of clinicians and two generations of multimodal large language models (LLMs) in Kellgren–Lawrence (KL) grading of knee osteoarthritis (KOA), including feature-level assessment and intraobserver repeatability.</p> Materials and Methods <p>In this retrospective single-center study, 348 knee radiographs were graded by a senior musculoskeletal radiologist (reference standard), a radiologist, an orthopedic surgeon, a radiology resident, and LLMs (ChatGPT-4o and ChatGPT-5.0). Binary KOA detection (KL 0–1 vs ≥ 2), feature-level interpretation (joint space narrowing, osteophytes, subchondral sclerosis), and intraobserver repeatability were evaluated. Agreement metrics included weighted κ, accuracy, and standard diagnostic measures.</p> Results <p>Agreement with the reference standard was highest for the radiologist (κ = 0.87), followed by the orthopedic surgeon and radiology resident. Both LLMs demonstrated moderate agreement, with ChatGPT-5.0 outperforming ChatGPT-4o. For binary KOA detection, ChatGPT-5.0 showed very high sensitivity (0.96) but reduced specificity. Per-grade classification was most accurate for KL 0 and KL 4, but remained limited for KL 1–2. Feature-level concordance was modest across all radiographic findings. Intraobserver repeatability was highest for the reference reader (κ = 0.881), followed by the orthopedic surgeon (κ = 0.634) and radiologist (κ = 0.626), while lower agreement was observed for the resident (κ = 0.478) and LLMs, with ChatGPT-5.0 showing higher consistency than ChatGPT-4o (κ = 0.591 vs. 0.485).</p> Conclusion <p>Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility. Current multimodal LLMs show high sensitivity but limited specificity and are not suitable for standalone radiographic KOA assessment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance of multimodal large language models versus clinicians for radiographic knee osteoarthritis grading: A multiobserver study

  • Asli Irmak Akdogan,
  • Efe Kemal Akdogan,
  • Mehmet Fatih Tumer,
  • Sılanaz Kutlu,
  • Mustafa Agah Tekindal,
  • Ozgur Tosun

摘要

Objective

To compare the performance of clinicians and two generations of multimodal large language models (LLMs) in Kellgren–Lawrence (KL) grading of knee osteoarthritis (KOA), including feature-level assessment and intraobserver repeatability.

Materials and Methods

In this retrospective single-center study, 348 knee radiographs were graded by a senior musculoskeletal radiologist (reference standard), a radiologist, an orthopedic surgeon, a radiology resident, and LLMs (ChatGPT-4o and ChatGPT-5.0). Binary KOA detection (KL 0–1 vs ≥ 2), feature-level interpretation (joint space narrowing, osteophytes, subchondral sclerosis), and intraobserver repeatability were evaluated. Agreement metrics included weighted κ, accuracy, and standard diagnostic measures.

Results

Agreement with the reference standard was highest for the radiologist (κ = 0.87), followed by the orthopedic surgeon and radiology resident. Both LLMs demonstrated moderate agreement, with ChatGPT-5.0 outperforming ChatGPT-4o. For binary KOA detection, ChatGPT-5.0 showed very high sensitivity (0.96) but reduced specificity. Per-grade classification was most accurate for KL 0 and KL 4, but remained limited for KL 1–2. Feature-level concordance was modest across all radiographic findings. Intraobserver repeatability was highest for the reference reader (κ = 0.881), followed by the orthopedic surgeon (κ = 0.634) and radiologist (κ = 0.626), while lower agreement was observed for the resident (κ = 0.478) and LLMs, with ChatGPT-5.0 showing higher consistency than ChatGPT-4o (κ = 0.591 vs. 0.485).

Conclusion

Although ChatGPT-5.0 outperformed ChatGPT-4o, both models remained inferior to clinicians in detailed KL grading, feature-level interpretation, and reproducibility. Current multimodal LLMs show high sensitivity but limited specificity and are not suitable for standalone radiographic KOA assessment.