<p>Brain tumor classification is a challenging task in medical image analysis, with significant implications for patient diagnosis and treatment. The objective of this paper is to propose a novel approach to brain tumor classification using a Vision Transformer (ViT) with a novel cross-attention mechanism. Our approach leverages the strengths of transformers in modeling long-range dependencies and multi-scale feature fusion. We introduce two new mechanisms to improve the performance of the cross-attention fusion module: Feature Calibration Mechanism (FCM) and Selective Cross-Attention (SCA). The FCM calibrates the features from different branches to make them more compatible, while The SCA selectively attends to the most informative features. Our experimental results demonstrate that the proposed approach outperforms other state-of-the-art methods, including Convolutional Neural Networks (CNN), ViT, and Cross-attention ViT, achieving an accuracy of 98.93% and an F1-score of 98.83%. Furthermore, the addition of Stochastic Depth mechanism improves the accuracy to 99.24% and the F1-score to 99.23%. The proposed FCM and SCA mechanisms can be easily integrated into other ViT architectures, making them a promising direction for future research in medical image analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vision transformer with feature calibration and selective cross-attention for brain tumor classification

  • Mohammad Ali Labbaf Khaniki,
  • Marzieh Mirzaeibonehkhater,
  • Mohammad Manthouri,
  • Elham Hasani

摘要

Brain tumor classification is a challenging task in medical image analysis, with significant implications for patient diagnosis and treatment. The objective of this paper is to propose a novel approach to brain tumor classification using a Vision Transformer (ViT) with a novel cross-attention mechanism. Our approach leverages the strengths of transformers in modeling long-range dependencies and multi-scale feature fusion. We introduce two new mechanisms to improve the performance of the cross-attention fusion module: Feature Calibration Mechanism (FCM) and Selective Cross-Attention (SCA). The FCM calibrates the features from different branches to make them more compatible, while The SCA selectively attends to the most informative features. Our experimental results demonstrate that the proposed approach outperforms other state-of-the-art methods, including Convolutional Neural Networks (CNN), ViT, and Cross-attention ViT, achieving an accuracy of 98.93% and an F1-score of 98.83%. Furthermore, the addition of Stochastic Depth mechanism improves the accuracy to 99.24% and the F1-score to 99.23%. The proposed FCM and SCA mechanisms can be easily integrated into other ViT architectures, making them a promising direction for future research in medical image analysis.