Background <p>Cardiotocography (CTG) remains a cornerstone in fetal monitoring, but its interpretation is subject to considerable inter- and intra-observer variability. Artificial intelligence (AI) tools, particularly large language models (LLMs), offer potential to improve diagnostic consistency and reduce clinician workload.</p> Objectives <p>This study aims to evaluate and compare the accuracy of various LLMs in CTG interpretation based on Federation of Gynecology and Obstetrics (FIGO 2015) criteria.</p> Study design <p>An analysis of sixty CTG traces previously classified by clinicians at the University Hospital Basel according to FIGO guidelines was conducted. In a two-run protocol, 30 normal CTG traces were initially presented as screenshots to Chat-GPT-4.0, Google Gemini, Bing Copilot, and DeepSeek. Subsequently, the LLMs that demonstrated adequate interpretation of normal CTGs were tasked to classify another 30 suspicious or pathological CTG traces. Each LLM was asked to classify each CTG trace as normal or abnormal.</p> Results <p>DeepSeek was unable to interpret CTGs and was excluded. Google Gemini showed poor performance (6.7%) on normal CTGs. Chat-GPT-4.0 partially succeeded in correctly classifying the provided CTG traces as normal (46.7%) or abnormal (50%). Bing Copilot accurately interpreted normal CTGs (96.6%) but failed on abnormal ones (0%).</p> Conclusions <p>LLMs show major limitations in the interpretation of CTG traces according to the FIGO criteria.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A comparative evaluation of publicly available large language models in the assessment of CTG traces according to the FIGO criteria

  • Iason Psilopatis,
  • Cécile Monod,
  • Valeria Filippi,
  • Rebecca Tschudin,
  • Olaf Lapaire,
  • Julius Emons,
  • Beatrice Mosimann,
  • Tibor A. Zwimpfer

摘要

Background

Cardiotocography (CTG) remains a cornerstone in fetal monitoring, but its interpretation is subject to considerable inter- and intra-observer variability. Artificial intelligence (AI) tools, particularly large language models (LLMs), offer potential to improve diagnostic consistency and reduce clinician workload.

Objectives

This study aims to evaluate and compare the accuracy of various LLMs in CTG interpretation based on Federation of Gynecology and Obstetrics (FIGO 2015) criteria.

Study design

An analysis of sixty CTG traces previously classified by clinicians at the University Hospital Basel according to FIGO guidelines was conducted. In a two-run protocol, 30 normal CTG traces were initially presented as screenshots to Chat-GPT-4.0, Google Gemini, Bing Copilot, and DeepSeek. Subsequently, the LLMs that demonstrated adequate interpretation of normal CTGs were tasked to classify another 30 suspicious or pathological CTG traces. Each LLM was asked to classify each CTG trace as normal or abnormal.

Results

DeepSeek was unable to interpret CTGs and was excluded. Google Gemini showed poor performance (6.7%) on normal CTGs. Chat-GPT-4.0 partially succeeded in correctly classifying the provided CTG traces as normal (46.7%) or abnormal (50%). Bing Copilot accurately interpreted normal CTGs (96.6%) but failed on abnormal ones (0%).

Conclusions

LLMs show major limitations in the interpretation of CTG traces according to the FIGO criteria.