<p>Breast cancer (BC) remains one of the primary causes of death among women worldwide, underscoring the importance of early detection, as it can significantly reduce the risk of cancer progression. While histopathological examination is considered the highest standard, it is complex, time-consuming, and prone to human error. Deep learning (DL) models, especially&#xa0;Convolutional Neural Networks (CNNs), have shown strong potential for BC detection using histopathological images but often struggle to capture long-range dependencies and fine-grained structural details. To address these limitations, Vision Transformers (ViTs) have been introduced, enabling the modeling of global context. However, they require large datasets and substantial computational resources. In this study, a DL-based framework is suggested to evaluate several Transformer networks and pre-trained CNN models — Swin Transformer V2, Vision Transformer (ViT_b32), ResNet-50, AlexNet, Inception V3, and VGG16 — for BC histopathological image classification using the BreaKHis and BACH 2018 datasets. Experimental results show that Swin Transformer V2 achieved superior performance, with an accuracy of 99.0%, a Matthews correlation coefficient of 97.6%, a balanced accuracy of 98.8%, an F1-score of 99.0%, a precision of 99.0%, and a recall of 99.0% on the BreaKHis dataset (40 × magnification). On the BACH 2018 dataset, it reached an accuracy of 93.7%, a Matthews correlation coefficient of 87.4%, a balanced accuracy of 93.6%, an F1-score of 93.7%, a precision of 93.7%, and a recall of 93.7%. These results demonstrate that Swin Transformer V2 effectively overcomes the limitations of CNNs and traditional ViTs by leveraging a hierarchical architecture with shifted window mechanisms, enabling it to efficiently capture both local and global tissue features while maintaining computational efficiency. To enhance interpretability, Local Interpretable Model-Agnostic Explanations (LIME) were employed to visualize and explain the model’s decision-making process, providing greater transparency and reliability in diagnostic predictions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An accurate framework based on transformer networks for breast cancer classification using histopathological image

  • Marwa Naas,
  • Hiba Mzoughi,
  • Ines Njeh,
  • Mohamed Ben Slima

摘要

Breast cancer (BC) remains one of the primary causes of death among women worldwide, underscoring the importance of early detection, as it can significantly reduce the risk of cancer progression. While histopathological examination is considered the highest standard, it is complex, time-consuming, and prone to human error. Deep learning (DL) models, especially Convolutional Neural Networks (CNNs), have shown strong potential for BC detection using histopathological images but often struggle to capture long-range dependencies and fine-grained structural details. To address these limitations, Vision Transformers (ViTs) have been introduced, enabling the modeling of global context. However, they require large datasets and substantial computational resources. In this study, a DL-based framework is suggested to evaluate several Transformer networks and pre-trained CNN models — Swin Transformer V2, Vision Transformer (ViT_b32), ResNet-50, AlexNet, Inception V3, and VGG16 — for BC histopathological image classification using the BreaKHis and BACH 2018 datasets. Experimental results show that Swin Transformer V2 achieved superior performance, with an accuracy of 99.0%, a Matthews correlation coefficient of 97.6%, a balanced accuracy of 98.8%, an F1-score of 99.0%, a precision of 99.0%, and a recall of 99.0% on the BreaKHis dataset (40 × magnification). On the BACH 2018 dataset, it reached an accuracy of 93.7%, a Matthews correlation coefficient of 87.4%, a balanced accuracy of 93.6%, an F1-score of 93.7%, a precision of 93.7%, and a recall of 93.7%. These results demonstrate that Swin Transformer V2 effectively overcomes the limitations of CNNs and traditional ViTs by leveraging a hierarchical architecture with shifted window mechanisms, enabling it to efficiently capture both local and global tissue features while maintaining computational efficiency. To enhance interpretability, Local Interpretable Model-Agnostic Explanations (LIME) were employed to visualize and explain the model’s decision-making process, providing greater transparency and reliability in diagnostic predictions.