<p>This retrospective bi-national multicenter diagnostic accuracy study evaluated the locked default-threshold performance of a commercial chest radiograph AI system for selected thoracic findings using a CT-anchored, radiographic-detectability reference framework. Consecutive eligible adult patients who underwent frontal chest radiography and chest CT at four tertiary-care centers in Turkey and Bulgaria between June and December 2024 were included. Reference labels were assigned using temporally paired CT, with adjudication of whether CT-confirmed abnormalities had a corresponding radiographic manifestation on the paired chest radiograph. For potentially dynamic findings, the allowable CT–CXR interval was restricted to ≤ 2 days. The AI system was evaluated at the manufacturer’s default threshold of 0.50. Diagnostic performance was summarized using prevalence, sensitivity, specificity, positive predictive value, negative predictive value, and false-positive burden with 95% confidence intervals. The study included 940 patients with paired CXR-CT examinations. Reference-positive prevalence ranged from 2.2% for pneumothorax to 21.0% for pleural effusion. Sensitivity was highest for fracture (91%) and pneumothorax (90%) and lowest for atelectasis (64%); specificity ranged from 82% for consolidation/opacity to 98% for fracture and pneumothorax. PPV ranged from 43% to 63%, indicating a non-trivial false-positive burden at the evaluated operating point. Because continuous probability scores, human-reader comparison, and workflow outcomes were unavailable, these findings should be interpreted as default-threshold technical validation and support further evaluation of the system as radiologist-supervised decision support with local performance monitoring, rather than standalone diagnosis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Default-threshold operating-point validation of a commercial chest radiograph AI system for selected CT-anchored thoracic findings: a bi-national multicenter retrospective study

  • Yeliz Basar,
  • Mustafa Ege Seker,
  • Galina Ivanova Kirova-Nedyalkova,
  • Elina Plamenova Milkovska,
  • Boyan Emilov Manov,
  • Atahan Tamturk,
  • Ilke Tasci,
  • Melih Karadag,
  • Nuri Sarac,
  • Selin Ardali Duzgun,
  • Erencan Karakoc,
  • Ahmet Karabulut,
  • Cemre Kavi,
  • Kaan Buyukkirli,
  • Mehmet Onur Onal,
  • Deniz Alis,
  • Recep Savas,
  • Ercan Karaarslan,
  • Gamze Durhan

摘要

This retrospective bi-national multicenter diagnostic accuracy study evaluated the locked default-threshold performance of a commercial chest radiograph AI system for selected thoracic findings using a CT-anchored, radiographic-detectability reference framework. Consecutive eligible adult patients who underwent frontal chest radiography and chest CT at four tertiary-care centers in Turkey and Bulgaria between June and December 2024 were included. Reference labels were assigned using temporally paired CT, with adjudication of whether CT-confirmed abnormalities had a corresponding radiographic manifestation on the paired chest radiograph. For potentially dynamic findings, the allowable CT–CXR interval was restricted to ≤ 2 days. The AI system was evaluated at the manufacturer’s default threshold of 0.50. Diagnostic performance was summarized using prevalence, sensitivity, specificity, positive predictive value, negative predictive value, and false-positive burden with 95% confidence intervals. The study included 940 patients with paired CXR-CT examinations. Reference-positive prevalence ranged from 2.2% for pneumothorax to 21.0% for pleural effusion. Sensitivity was highest for fracture (91%) and pneumothorax (90%) and lowest for atelectasis (64%); specificity ranged from 82% for consolidation/opacity to 98% for fracture and pneumothorax. PPV ranged from 43% to 63%, indicating a non-trivial false-positive burden at the evaluated operating point. Because continuous probability scores, human-reader comparison, and workflow outcomes were unavailable, these findings should be interpreted as default-threshold technical validation and support further evaluation of the system as radiologist-supervised decision support with local performance monitoring, rather than standalone diagnosis.