Evaluating the efficacy of a federated AI algorithm in assisting in the detection and segmentation of brain metastases: limitations and opportunities
摘要
Accurate detection and segmentation of brain metastases is critical for stereotactic radiosurgery (SRS) planning, but manual approaches remain time-consuming and subject to interobserver variability. Artificial intelligence (AI)-based auto-segmentation tools have shown promise but independent external validation on disparate datasets is limited.
MethodsThis retrospective study included 50 patients with 236 brain metastases who underwent SRS using post-contrast 3 Tesla MPRAGE and 3D-TSE sequences. A cloud-based, federated AI algorithm was evaluated for detection sensitivity, false positivity (FP), dice similarity coefficient (DSC), and qualitative physician review. Ground truth contours from clinical treatment plans developed after multi-physician review.
ResultsThe AI algorithm successfully identified 123 out of 236 lesions treated, corresponding to a lesion-wise sensitivity of 52.1% and a mean patient-wise sensitivity of 69.2%, with a positive predictive value (PPV) of 78.0%. While the median volume of all lesions in the cohort was 0.045 cc, the lesions detected by the algorithm were significantly larger with a median volume of 0.250 cc (p < 0.001). The mean Dice similarity coefficient (DSC) was 0.60, with consistently higher scores observed for lesions exceeding 1.0 cm in diameter. Qualitative evaluation revealed that 46.3% of AI-generated segmentations required major revisions. Notably, most AI-generated contours were entirely enclosed within the ground truth annotations, often leading to underestimation of lesion margins.
ConclusionThe AI algorithm demonstrated moderate sensitivity with excellent false positivity. However, performance was limited for sub-centimeter lesions, likely due to the composition of the initial training dataset, which underscores the need/value of independent validation on diverse external datasets.