Diagnostic accuracy of image-based deep learning for glaucomatous optic neuropathy detection: a systematic review and meta-analysis
摘要
Several studies have explored the use of deep learning (DL) algorithms based on fundus photography (FP) and optical coherence tomography (OCT) for the detection of glaucomatous optic neuropathy (GON). However, high-quality evidence regarding their diagnostic accuracy remains insufficient. Therefore, this study systematically evaluated the diagnostic performance of DL algorithms for GON detection and compared their accuracy with that of clinical experts to inform the development and optimization of intelligent diagnostic systems.
MethodsPubMed, the Cochrane Library, Embase, and Web of Science were systematically searched until May 28, 2026. The methodological quality of included studies was assessed using the QUADAS-2 tool. Meta-analyses were conducted based on validation datasets, and subgroup analyses were performed according to imaging modality.
ResultsA total of 30 eligible studies were included, of which 24 were incorporated into the meta-analysis. For DL algorithms based on FP, the pooled sensitivity (SEN) and specificity (SPC) for GON detection were 0.92 (95% CI: 0.89–0.94) and 0.92 (95% CI: 0.89–0.95), respectively. The pooled positive likelihood ratio (PLR) and negative likelihood ratio (NLR) were 11.9 (95% CI: 8.3–17.0) and 0.09 (95% CI: 0.06–0.12), respectively. For clinical experts diagnosing GON using FP, the pooled SEN, SPC, PLR, and NLR were 0.89 (95% CI: 0.78–0.95), 0.92 (95% CI: 0.85–0.96), 10.7 (95% CI: 5.3–21.5), and 0.12 (95% CI: 0.05–0.26), respectively. For DL models based on OCT, the pooled SEN, SPC, PLR, and NLR were 0.87 (95% CI: 0.80–0.92), 0.91 (95% CI: 0.86–0.94), 9.5 (95% CI: 6.3–14.2), and 0.14 (95% CI: 0.09–0.22), respectively. Likelihood ratios presented in the nomogram were rounded to integers to facilitate rapid clinical interpretation.
ConclusionsDL models based on FP and OCT demonstrated diagnostic accuracy comparable to that of clinical experts in identifying GON, with FP-based models showing a tendency toward higher SEN. However, in the current application of DL for the detection of glaucomatous optic neuropathy, the generation of validation datasets relies predominantly on internal validation. Although the diagnostic performance of these models is not inferior to that of clinical experts, the results should not be interpreted with undue optimism, and a cautious attitude should be maintained.