Advancing Alzheimer’s disease diagnosis: a hierarchical vision transformer classification framework with hybrid image processing and generative AI augmentation
摘要
Alzheimer’s disease (AD) poses a growing global health challenge, with its progressive cognitive decline affecting millions, particularly as populations age. Early detection is crucial for prompt interventions that can slow disease progression and enhance patient quality of life. However, challenges like subtle brain biomarkers, noisy Magnetic Resonance Imaging (MRI) scans, limited datasets, and class imbalances hinder accurate diagnosis. This study proposes an innovative framework, ADVision (Alzheimer’s Disease Vision Transformer), utilizing the Open Access Series of Imaging Studies (OASIS) MRI dataset to overcome these barriers. The approach integrates advanced image processing and artificial intelligence to enhance diagnostic precision. A Hybrid Wavelet-Autoencoder (HWA) denoises MRI scans, preserving critical features like cortical thinning. StyleGAN2 with Adaptive Discriminator Augmentation (ADA) generates high-quality synthetic images to address data scarcity and mitigate the severe class imbalance, particularly for underrepresented Moderate Demented cases. A perceptually guided super-resolution network reconstructs high-resolution images from low-quality inputs, ensuring sharp details for accurate biomarker detection. The core classification utilizes a hierarchical Vision Transformer (ViT) model to categorize AD stages—Non-Demented, Very Mild Demented, Mild Demented, and Moderate Demented—with a two-stage approach that mimics clinical workflows. This achieves a classification accuracy of 92.4%, surpassing traditional convolutional neural networks. Gradient-weighted Class Activation Mapping (Grad-CAM) enhances interpretability by visualizing key brain regions, such as the hippocampus and ventricles, critical for AD diagnosis. Key contributions include robust denoising with a Peak Signal-to-Noise Ratio (PSNR) of 49.1 dB, effective data augmentation to improve model generalization, and a scalable, interpretable classification system. The framework demonstrates superior performance in denoising (with a Structural Similarity Index of 0.999), classification metrics (including precision, recall, and F1-score), and computational efficiency compared to existing methods. It offers a clinically viable, automated tool for early AD detection, supporting timely interventions and better patient outcomes.