<p>Brain tumor segmentation is critical for accurate identification and treatment planning. However, delineating intricate tumor areas like Edema, Tumor Core, and Enhancing Tumor remains challenging owing to diversity in dimensions, morphology, and delineation. This introduces a novel deep learning model integrating a pretrained Vision Transformer (ViT) for global context representation, Atrous Spatial Pyramid Pooling (ASPP) in both encoder and merging stages, and ReLU6. Additionally, attention mechanisms and Conditional Random Fields (CRF) were incorporated for refining feature emphasis and segmentation boundaries. The model was trained and tested on the BraTS 2020 dataset and compared with state-of-the-art methods. The suggested model obtained the highest Dice coefficients across all tumor regions, with improvements of up to 4% over nnU-Net and SwinBTS in the Enhancing Tumor and Edema region. Grad-CAM heatmaps illustrated the model’s proficiency in precisely identifying tumor areas, hence enhancing the interpretability of predictions. The incorporation of multi-scale feature extraction, attention mechanism, and global context representation markedly improved segmentation accuracy and interpretability. The findings underscore the proposed model’s capacity to enhance clinical processes and establish a standard for brain tumor segmentation tasks. Future endeavors involve testing on varied datasets and investigating lightweight structures for real-time applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable deep learning framework for brain tumor segmentation using vision transformer and conditional random fields

  • Homayoun Safarpour,
  • Soroush Sadeghi,
  • Payam Zarbakhsh,
  • Mohammadreza Kamsari,
  • Marjan Kia,
  • Ramin Ranjbarzadeh

摘要

Brain tumor segmentation is critical for accurate identification and treatment planning. However, delineating intricate tumor areas like Edema, Tumor Core, and Enhancing Tumor remains challenging owing to diversity in dimensions, morphology, and delineation. This introduces a novel deep learning model integrating a pretrained Vision Transformer (ViT) for global context representation, Atrous Spatial Pyramid Pooling (ASPP) in both encoder and merging stages, and ReLU6. Additionally, attention mechanisms and Conditional Random Fields (CRF) were incorporated for refining feature emphasis and segmentation boundaries. The model was trained and tested on the BraTS 2020 dataset and compared with state-of-the-art methods. The suggested model obtained the highest Dice coefficients across all tumor regions, with improvements of up to 4% over nnU-Net and SwinBTS in the Enhancing Tumor and Edema region. Grad-CAM heatmaps illustrated the model’s proficiency in precisely identifying tumor areas, hence enhancing the interpretability of predictions. The incorporation of multi-scale feature extraction, attention mechanism, and global context representation markedly improved segmentation accuracy and interpretability. The findings underscore the proposed model’s capacity to enhance clinical processes and establish a standard for brain tumor segmentation tasks. Future endeavors involve testing on varied datasets and investigating lightweight structures for real-time applications.