Predicting rupture risk in unruptured intracranial aneurysms: a comparative evaluation of machine learning models and conventional risk scores
摘要
Unruptured intracranial aneurysms (UIAs) are increasingly detected, yet individualized rupture risk prediction remains imprecise using conventional clinical scores. Machine learning (ML) models integrating high-dimensional imaging and clinical data have been proposed, but their incremental prognostic value remains uncertain. The objectives were to systematically evaluate ML-based models for predicting rupture or instability of UIAs and compare their performance with established clinical risk tools.
MethodsA systematic review and meta-analysis were conducted according to PRISMA 2020. PubMed, Embase, Scopus, Web of Science, and the Cochrane Library were searched (2010–2025). Studies developing or validating ML or deep learning models for rupture or growth prediction in UIAs were included. Discrimination was assessed using area under the receiver operating characteristic curve (AUC), and risk of bias using PROBAST-AI.
ResultsTwenty-five studies comprising 18,569 aneurysms were included. ML models demonstrated good to excellent discrimination, with AUCs approximately 0.70–0.93. Performance increased with predictor complexity, with multimodal and imaging-rich models achieving the highest discrimination. In head-to-head comparisons, ML models consistently outperformed the PHASES score, whereas comparisons with logistic regression showed largely comparable performance without consistent ML superiority, particularly in externally validated cohorts. External validation was associated with modest attenuation of AUCs, and calibration reporting was infrequent.
ConclusionsML-based models show discriminative capability for predicting rupture or instability in UIAs, driven primarily by enriched predictor domains rather than algorithmic complexity. However, limitations in external validation, calibration, and clinical utility reporting constrain translation into practice. Future research should prioritize prospective validation, calibrated absolute risk estimation, and integration into clinical decision-support frameworks.