Performance of GPT-4o in CT-based crista galli morphological classification: image-only versus morphometric dataset-driven approaches: a controlled comparative study
摘要
The aim of this study was to examine whether ChatGPT’s performance in a morphological classification task based solely on a single CT slice and without access to any additional health record information varies depending on the inclusion of metric datasets.
MethodsThis study is based on a CT-based classification system in which categorized the crista galli morphometrically and morphologically. Paranasal CT (PNCT) images of 101 randomly selected patients who underwent PNCT were retrieved from the hospital’s database. The PNCT slices used by the radiologists for classification were uploaded individually to GPT-4o in a randomized patient order, and it was asked to perform morphological classification based solely on the CT images. The accuracy of the morphological classifications performed by GPT-4o based on the morphometric data and PNCT slices was statistically compared to the radiologists’ gold standard classifications.
ResultsGPT-4o correctly classified CG types in 27.72% of cases using PNCT slices, whereas its classification accuracy reached 100% when using the morphometric dataset. The accuracy of GPT-4o’s classification based on PNCT slices was significantly lower compared to its performance with morphometric data (p < 0.001). When examining the classification accuracy of GPT-4o based on PNCT slices by CG type, the highest accuracy was observed for the ossified type at 60%, while the lowest was for the tubular type at 6.06%.
ConclusionsAlthough ChatGPT has demonstrated promising performance in diagnostic radiology when provided with textual or numerical datasets, in this controlled CT-based crista galli morphologic classification task, GPT-4o demonstrated substantially lower performance when relying solely on image input compared with structured morphometric data. These findings suggest that the performance of multimodal LLMs may strongly depend on structured information. Further studies involving broader diagnostic settings are required before conclusions regarding clinical applicability can be drawn.