Glaucoma, a progressive eye condition leading to potential blindness, often lacks early symptoms, necessitating timely identification. However, manual prediction systems are significantly limited by subjectivity and the inherent risk of human errors. This study addresses these challenges by introducing the ViT-CNN approach, a hybrid model that fuses features from a custom Convolutional Neural Network (CNN) and a Vision Transformer (ViT) for the analysis of Optical Coherence Tomography (OCT) images. The final classification task is executed using several Machine Learning (ML) classifiers to assess the effectiveness of the proposed ViT-CNN model. Experimental results on four datasets demonstrate that the proposed hybrid ViT-CNN model surpasses standalone CNN or ViT models. The synergy between localized feature extraction by CNN and global contextual comprehension by ViT models enhances feature complexity, as illustrated through Gradient Class Activation Map (GRAD-CAM) visualization. This research advances automated glaucoma detection with a robust and versatile model showing competitive performance across diverse datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Breaking the Mold: ViT-CNN Fusion for Enhanced Glaucoma Prediction in OCT Images

  • Mahjabin Rahman Oishe,
  • S. M. Mahedy Hasan,
  • Minhaz F. Zibran

摘要

Glaucoma, a progressive eye condition leading to potential blindness, often lacks early symptoms, necessitating timely identification. However, manual prediction systems are significantly limited by subjectivity and the inherent risk of human errors. This study addresses these challenges by introducing the ViT-CNN approach, a hybrid model that fuses features from a custom Convolutional Neural Network (CNN) and a Vision Transformer (ViT) for the analysis of Optical Coherence Tomography (OCT) images. The final classification task is executed using several Machine Learning (ML) classifiers to assess the effectiveness of the proposed ViT-CNN model. Experimental results on four datasets demonstrate that the proposed hybrid ViT-CNN model surpasses standalone CNN or ViT models. The synergy between localized feature extraction by CNN and global contextual comprehension by ViT models enhances feature complexity, as illustrated through Gradient Class Activation Map (GRAD-CAM) visualization. This research advances automated glaucoma detection with a robust and versatile model showing competitive performance across diverse datasets.