Vision and convolutional transformers for Alzheimer’s disease diagnosis: a systematic review of architectures, multimodal fusion and critical gaps
摘要
Alzheimer’s disease (AD), a significant public health challenge, requires accurate early diagnosis to improve patient outcomes. Vision Transformers (ViTs) and Convolutional Vision Transformers (CViTs) have emerged as powerful Deep Learning architectures for this task. Following PRISMA guidelines, this systematic review analyzes 68 studies selected from 564 publications (2021–2025) across five major databases: Scopus, Web of Science, ScienceDirect, IEEE Xplore, and PubMed. We introduce novel taxonomies to systematically categorize these works by model architecture, data modality, fusion strategy, and diagnostic objective. Our analysis reveals key trends, such as the rise of hybrid CViT frameworks, and critical gaps, including a limited focus on Mild Cognitive Impairment-to-AD progression. Critically, we also assess practical implementation details, revealing widespread challenges in algorithmic reproducibility. The discussion culminates in a forward-looking analysis of Large Vision Models and proposes future directions emphasizing the need for robust multimodal integration, lightweight transformer designs, and Explainable AI to advance AD research and bridge the critical gap between high-performance modeling and clinical applicability.