Background <p>The aim of this study was to examine whether ChatGPT’s performance in a morphological classification task based solely on a single CT slice and without access to any additional health record information varies depending on the inclusion of metric datasets.</p> Methods <p>This study is based on a CT-based classification system in which categorized the crista galli morphometrically and morphologically. Paranasal CT (PNCT) images of 101 randomly selected patients who underwent PNCT were retrieved from the hospital’s database. The PNCT slices used by the radiologists for classification were uploaded individually to GPT-4o in a randomized patient order, and it was asked to perform morphological classification based solely on the CT images. The accuracy of the morphological classifications performed by GPT-4o based on the morphometric data and PNCT slices was statistically compared to the radiologists’ gold standard classifications.</p> Results <p>GPT-4o correctly classified CG types in 27.72% of cases using PNCT slices, whereas its classification accuracy reached 100% when using the morphometric dataset. The accuracy of GPT-4o’s classification based on PNCT slices was significantly lower compared to its performance with morphometric data (<i>p</i> &lt; 0.001). When examining the classification accuracy of GPT-4o based on PNCT slices by CG type, the highest accuracy was observed for the ossified type at 60%, while the lowest was for the tubular type at 6.06%.</p> Conclusions <p>Although ChatGPT has demonstrated promising performance in diagnostic radiology when provided with textual or numerical datasets, in this controlled CT-based crista galli morphologic classification task, GPT-4o demonstrated substantially lower performance when relying solely on image input compared with structured morphometric data. These findings suggest that the performance of multimodal LLMs may strongly depend on structured information. Further studies involving broader diagnostic settings are required before conclusions regarding clinical applicability can be drawn.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance of GPT-4o in CT-based crista galli morphological classification: image-only versus morphometric dataset-driven approaches: a controlled comparative study

  • Erdal Komut,
  • Serkan Günay,
  • Ahmet Öztürk,
  • Gurbet Yanarateş,
  • Onur Karacif,
  • Seval Komut

摘要

Background

The aim of this study was to examine whether ChatGPT’s performance in a morphological classification task based solely on a single CT slice and without access to any additional health record information varies depending on the inclusion of metric datasets.

Methods

This study is based on a CT-based classification system in which categorized the crista galli morphometrically and morphologically. Paranasal CT (PNCT) images of 101 randomly selected patients who underwent PNCT were retrieved from the hospital’s database. The PNCT slices used by the radiologists for classification were uploaded individually to GPT-4o in a randomized patient order, and it was asked to perform morphological classification based solely on the CT images. The accuracy of the morphological classifications performed by GPT-4o based on the morphometric data and PNCT slices was statistically compared to the radiologists’ gold standard classifications.

Results

GPT-4o correctly classified CG types in 27.72% of cases using PNCT slices, whereas its classification accuracy reached 100% when using the morphometric dataset. The accuracy of GPT-4o’s classification based on PNCT slices was significantly lower compared to its performance with morphometric data (p < 0.001). When examining the classification accuracy of GPT-4o based on PNCT slices by CG type, the highest accuracy was observed for the ossified type at 60%, while the lowest was for the tubular type at 6.06%.

Conclusions

Although ChatGPT has demonstrated promising performance in diagnostic radiology when provided with textual or numerical datasets, in this controlled CT-based crista galli morphologic classification task, GPT-4o demonstrated substantially lower performance when relying solely on image input compared with structured morphometric data. These findings suggest that the performance of multimodal LLMs may strongly depend on structured information. Further studies involving broader diagnostic settings are required before conclusions regarding clinical applicability can be drawn.