Background <p>Translating the intricate anatomical signatures of retinal disease from optical coherence tomography (OCT) B-scans into clear, accurate clinical narratives demands algorithms that seamlessly fuse visual features with domain expertise.</p> Methods <p>We curated a multimodal dataset of 40,000 OCT B-scans from public repositories and private clinical cohorts, each paired with expert-validated summaries spanning six conditions: diabetic macular edema, diabetic retinopathy, geographic atrophy, drusen, choroidal neovascularization, and healthy retina. We introduce LO-VLM, a compact (247M parameter) vision-language model (VLM) that infuses anatomical guidance into both encoder and decoder for free-form summary generation and multiclass disease classification. Benchmarking against state-of-the-art RetinaVLM, LLaVA-Med, and a ViT vision only model demonstrates superior performance.</p> Results <p>In a blinded evaluation by three board certified retina specialists, LO-VLM narratives achieves a mean = 8.5 (standard deviation = 1.15) out of 10, compared to a mean = 5.5 (standard 32 deviation = 1.13) for RetinaVLM (p &lt; 0.0001). In quantitative evaluations, LO-VLM achieves an SBERT similarity of 80.3% and a BERTScore F1 of 71.5%, representing improvements of 8.2% and 28.8% over specialized VLM baselines. For disease classification, LO-VLM reaches 96% accuracy (F1 = 96%), outperforming ViT by 13% and exceeding medical VLM benchmarks by over 62% (p &lt; 0.05).</p> Conclusions <p>By reconciling interpretability with computational efficiency, LO-VLM establishes a paradigm for efficient AI models in OCT interpretation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Compact vision language models enable efficient and interpretable optical coherence tomography through layer-specific multimodal learning

  • Tania Haghighi,
  • Sina Gholami,
  • Jared Todd Sokol,
  • Aayush Biswas,
  • Jennifer I. Lim,
  • Theodore Leng,
  • Atalie C. Thompson,
  • Hamed Tabkhi,
  • Minhaj Nur Alam

摘要

Background

Translating the intricate anatomical signatures of retinal disease from optical coherence tomography (OCT) B-scans into clear, accurate clinical narratives demands algorithms that seamlessly fuse visual features with domain expertise.

Methods

We curated a multimodal dataset of 40,000 OCT B-scans from public repositories and private clinical cohorts, each paired with expert-validated summaries spanning six conditions: diabetic macular edema, diabetic retinopathy, geographic atrophy, drusen, choroidal neovascularization, and healthy retina. We introduce LO-VLM, a compact (247M parameter) vision-language model (VLM) that infuses anatomical guidance into both encoder and decoder for free-form summary generation and multiclass disease classification. Benchmarking against state-of-the-art RetinaVLM, LLaVA-Med, and a ViT vision only model demonstrates superior performance.

Results

In a blinded evaluation by three board certified retina specialists, LO-VLM narratives achieves a mean = 8.5 (standard deviation = 1.15) out of 10, compared to a mean = 5.5 (standard 32 deviation = 1.13) for RetinaVLM (p < 0.0001). In quantitative evaluations, LO-VLM achieves an SBERT similarity of 80.3% and a BERTScore F1 of 71.5%, representing improvements of 8.2% and 28.8% over specialized VLM baselines. For disease classification, LO-VLM reaches 96% accuracy (F1 = 96%), outperforming ViT by 13% and exceeding medical VLM benchmarks by over 62% (p < 0.05).

Conclusions

By reconciling interpretability with computational efficiency, LO-VLM establishes a paradigm for efficient AI models in OCT interpretation.