From Pixels to Predictions: Medical Image Analysis and Prediction of Brain Tumor and Chest Cancer with Vision Transformers
摘要
This study provides a thorough exploration of Vision Transformer (ViT) models in the realm of medical image analysis, specifically focusing on datasets with smaller size and greater complexity than traditional ViT evaluations. Demonstrating the transformative potential of ViT, our research highlights its capacity to enhance diagnostic accuracy and predictive capabilities in healthcare. These analyses were performed on Google Colaboratory, leveraging its cloud-based infrastructure for computational efficiency and accessibility. In the Brain Tumor model, ViT excelled with an accuracy of 85.71%, surpassing the 82.05% achieved by its CNN counterpart. ViT exhibited superior precision, recall, and F1-score, registering 0.88, 0.83, and 0.85, respectively, compared to the CNN model's 0.84, 0.79, and 0.81. The Lung Cancer Model similarly achieved an accuracy of 89.02% as compared to 85.05% and 83.01% compared to CNN models which showcases ViT's effectiveness in diverse medical imaging applications. Our work signifies a significant step towards leveraging ViT models in real-world medical scenarios, where complex data and nuanced patterns demand advanced analytical tools for improved patient outcomes.