Background <p>Given that breast cancer remains the most lethal malignancy among women globally, there is an urgent need for reliable and comprehensible diagnostic technology. Deep learning models have demonstrated encouraging outcomes in breast imaging; nevertheless, existing methodologies struggle to encompass global context and deliver clinically relevant justifications for their conclusions.</p> Method <p>The ExViT-Breast framework was trained and internally validated on CBIS-DDSM (<i>n</i> = 1,566 subjects; masses and calcifications), externally validated on INbreast, and evaluated across modalities on the BUSI ultrasound dataset. A Vision Transformer backbone with a cross-view attention fusion module integrated paired craniocaudal (CC) and mediolateral oblique (MLO) mammographic views. Four explainability methods were compared: Grad-CAM, Grad-CAM++, LayerCAM, and Transformer-LRP. Classification performance was assessed by AUC and sensitivity at one false positive per examination. Explanation quality was evaluated using faithfulness (deletion/insertion curves), clinical plausibility (pointing game against expert-defined lesion regions), and ensemble-based uncertainty estimation. Model calibration was quantified using Expected Calibration Error (ECE).</p> Results <p>ExViT-Breast achieved superior diagnostic performance with an AUC of 0.94 ± 0.02 on CBIS-DDSM, significantly outperforming conventional CNN architectures (ResNet-50: AUC = 0.89, <i>p</i> &lt; 0.001) and single-view Vision Transformers (AUC = 0.91, <i>p</i> = 0.003). External validation on INbreast demonstrated robust generalization (AUC = 0.91 ± 0.03), while cross-modality evaluation on BUSI yielded AUC = 0.88 ± 0.04. Transformer-LRP consistently provided the highest explanation quality across all datasets, achieving faithfulness scores of 0.71 ± 0.05 and plausibility rates of 79.5% on CBIS-DDSM, significantly superior to Grad-CAM variants (<i>p</i> &lt; 0.001).</p> Conclusion <p>The ExViT-Breast framework demonstrates the potential of explainable Vision Transformers in breast imaging, combining high classification performance with clinically interpretable explanations. The integration of multi-view analysis and quantitative explainability assessment positions this approach as a promising tool for computer-aided diagnosis in breast cancer screening and detection.</p> Clinical trial <p>Clinical trial number: not applicable.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable vision transformer framework for breast lesion classification with Grad-CAM support

  • Ajay Kumar,
  • Mudassir Khan,
  • Gunjan Mittal,
  • Izhar Husain,
  • Sambhavi Shukla,
  • Syed Arshad Ali,
  • Anu Sayal,
  • Janhvi Jha,
  • Riaz Ahmad Ziar

摘要

Background

Given that breast cancer remains the most lethal malignancy among women globally, there is an urgent need for reliable and comprehensible diagnostic technology. Deep learning models have demonstrated encouraging outcomes in breast imaging; nevertheless, existing methodologies struggle to encompass global context and deliver clinically relevant justifications for their conclusions.

Method

The ExViT-Breast framework was trained and internally validated on CBIS-DDSM (n = 1,566 subjects; masses and calcifications), externally validated on INbreast, and evaluated across modalities on the BUSI ultrasound dataset. A Vision Transformer backbone with a cross-view attention fusion module integrated paired craniocaudal (CC) and mediolateral oblique (MLO) mammographic views. Four explainability methods were compared: Grad-CAM, Grad-CAM++, LayerCAM, and Transformer-LRP. Classification performance was assessed by AUC and sensitivity at one false positive per examination. Explanation quality was evaluated using faithfulness (deletion/insertion curves), clinical plausibility (pointing game against expert-defined lesion regions), and ensemble-based uncertainty estimation. Model calibration was quantified using Expected Calibration Error (ECE).

Results

ExViT-Breast achieved superior diagnostic performance with an AUC of 0.94 ± 0.02 on CBIS-DDSM, significantly outperforming conventional CNN architectures (ResNet-50: AUC = 0.89, p < 0.001) and single-view Vision Transformers (AUC = 0.91, p = 0.003). External validation on INbreast demonstrated robust generalization (AUC = 0.91 ± 0.03), while cross-modality evaluation on BUSI yielded AUC = 0.88 ± 0.04. Transformer-LRP consistently provided the highest explanation quality across all datasets, achieving faithfulness scores of 0.71 ± 0.05 and plausibility rates of 79.5% on CBIS-DDSM, significantly superior to Grad-CAM variants (p < 0.001).

Conclusion

The ExViT-Breast framework demonstrates the potential of explainable Vision Transformers in breast imaging, combining high classification performance with clinically interpretable explanations. The integration of multi-view analysis and quantitative explainability assessment positions this approach as a promising tool for computer-aided diagnosis in breast cancer screening and detection.

Clinical trial

Clinical trial number: not applicable.