Breast cancer is a major global health issue that particularly affects women. Many breast ultrasound (BUS) segmentation methods have been developed based on convolution neural networks (CNNs) due to the need for accurate diagnosis of medical images. Nevertheless, CNNs have restrictions in modeling long-range relationships, which can lead to reduced accuracy in BUS segmentation. Transformers, which can obtain global information, also have limitations in extracting local details and require training on huge datasets. In this paper, we propose a Hybrid CNNs-Transformer architecture for segmenting BUS images. Our proposed method is built upon the U-Net architecture and comprises of two encoders. The first encoder utilizes pure CNN blocks to extract local and fine-grained features, while the second encoder leverages the powerful Swin-Transformer architecture to effectively model global dependencies. The decoder part of the architecture utilizes the merged extracted features and leverages skip-connections to reconstruct the segmented BUS images. Through experimentation and evaluation on two publicly available BUS datasets, it has been demonstrated that our proposed method outperforms recent BUS image segmentation methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Combining CNNs and Transformers Networks for Improved Breast Ultrasound Image Segmentation

  • Jaouad Tagnamas,
  • Hiba Ramadan,
  • Ali Yahyaouy,
  • Hamid Tairi

摘要

Breast cancer is a major global health issue that particularly affects women. Many breast ultrasound (BUS) segmentation methods have been developed based on convolution neural networks (CNNs) due to the need for accurate diagnosis of medical images. Nevertheless, CNNs have restrictions in modeling long-range relationships, which can lead to reduced accuracy in BUS segmentation. Transformers, which can obtain global information, also have limitations in extracting local details and require training on huge datasets. In this paper, we propose a Hybrid CNNs-Transformer architecture for segmenting BUS images. Our proposed method is built upon the U-Net architecture and comprises of two encoders. The first encoder utilizes pure CNN blocks to extract local and fine-grained features, while the second encoder leverages the powerful Swin-Transformer architecture to effectively model global dependencies. The decoder part of the architecture utilizes the merged extracted features and leverages skip-connections to reconstruct the segmented BUS images. Through experimentation and evaluation on two publicly available BUS datasets, it has been demonstrated that our proposed method outperforms recent BUS image segmentation methods.