U-Net leverages skip connections to link the encoder and decoder, effectively combining low-level spatial features with high-level semantic features. However, this architecture is constrained by its limited ability to incorporate global semantic information from the entire image. To overcome this constraint, we propose the Symmetric Cross-Attention (SCA) network. The SCA network incorporates a novel decoder branch engineered for semantic feature extraction. These features are subsequently integrated with the spatial characteristics originating from the primary decoder. At each level, Spatial features serve as the keys, whereas semantic features function as the queries. Spatial cross-attention is utilized to refine low-level spatial information, while channel cross-attention is applied to enrich high-level semantic representations. The results from both attention directions are fused, and the decoder is restructured to more effectively capture long-range dependencies. The SCA module establishes mutual correlations between semantic and spatial responses, thereby enhancing the representation of specific semantic features. We integrate the SCA module into four models: U-Net, V-Net, R2-Unet, and ResUnet. Experiments conducted on two benchmark medical image segmentation datasets, Kvasir-Seg and CVC-ClinicDB, demonstrate that the SCA module significantly improves segmentation performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SCAUnet: Symmetric Cross-Attention U-net Model for Semantic Segementation

  • Chunbo Yang,
  • Hailin Liu

摘要

U-Net leverages skip connections to link the encoder and decoder, effectively combining low-level spatial features with high-level semantic features. However, this architecture is constrained by its limited ability to incorporate global semantic information from the entire image. To overcome this constraint, we propose the Symmetric Cross-Attention (SCA) network. The SCA network incorporates a novel decoder branch engineered for semantic feature extraction. These features are subsequently integrated with the spatial characteristics originating from the primary decoder. At each level, Spatial features serve as the keys, whereas semantic features function as the queries. Spatial cross-attention is utilized to refine low-level spatial information, while channel cross-attention is applied to enrich high-level semantic representations. The results from both attention directions are fused, and the decoder is restructured to more effectively capture long-range dependencies. The SCA module establishes mutual correlations between semantic and spatial responses, thereby enhancing the representation of specific semantic features. We integrate the SCA module into four models: U-Net, V-Net, R2-Unet, and ResUnet. Experiments conducted on two benchmark medical image segmentation datasets, Kvasir-Seg and CVC-ClinicDB, demonstrate that the SCA module significantly improves segmentation performance.