<p>Accurate medical image segmentation is crucial for early disease diagnosis; yet, challenges persist due to shape variations and low contrast with the background. In this study, a novel method for medical image segmentation has been proposed, leveraging the strengths of both convolutional neural networks (CNNs) and transformers through a parallel encoder–decoder structure. Our approach, AttTransCFusion, employs a dual encoder architecture consisting of a CNN encoder branch and a transformer encoder branch to capture both local and global contextual information. An attention decoder module refines features across different network levels, while a gated fusion module selectively fuses local and long-range dependency information to produce an accurate prediction map. When evaluated on benchmark datasets for polyp, skin lesion, and cardiac segmentation, AttTransCFusion achieved state-of-the-art performance, with average Dice coefficients reaching 0.927, 0.882, and 0.923, respectively. Ablation studies validated the effectiveness of our key modules in improving segmentation accuracy. The proposed AttTransCFusion method demonstrates superior generalization performance and efficiency, highlighting its potential for practical applications in medical image segmentation. Furthermore, its versatility allows for extension to other challenging tasks within the medical domain. The source code is released on (<a href="https://github.com/syzhou1226/AttTransCFusion">https://github.com/syzhou1226/AttTransCFusion</a>).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-guided fusion of transformers and CNNs for enhanced medical image segmentation

  • Shiyao Zhou,
  • Mengyao Wang,
  • Zhiyong Huang,
  • Yuqin He,
  • Xiao Han,
  • Yunlan Zhao,
  • Zhiyu Zhao

摘要

Accurate medical image segmentation is crucial for early disease diagnosis; yet, challenges persist due to shape variations and low contrast with the background. In this study, a novel method for medical image segmentation has been proposed, leveraging the strengths of both convolutional neural networks (CNNs) and transformers through a parallel encoder–decoder structure. Our approach, AttTransCFusion, employs a dual encoder architecture consisting of a CNN encoder branch and a transformer encoder branch to capture both local and global contextual information. An attention decoder module refines features across different network levels, while a gated fusion module selectively fuses local and long-range dependency information to produce an accurate prediction map. When evaluated on benchmark datasets for polyp, skin lesion, and cardiac segmentation, AttTransCFusion achieved state-of-the-art performance, with average Dice coefficients reaching 0.927, 0.882, and 0.923, respectively. Ablation studies validated the effectiveness of our key modules in improving segmentation accuracy. The proposed AttTransCFusion method demonstrates superior generalization performance and efficiency, highlighting its potential for practical applications in medical image segmentation. Furthermore, its versatility allows for extension to other challenging tasks within the medical domain. The source code is released on (https://github.com/syzhou1226/AttTransCFusion).