<p>Colorectal polyps are the main precursors of colorectal cancer, and early detection of polyps relies on accurate image segmentation techniques. However, existing segmentation methods often face several key challenges. These include semantic information loss, background noise interference, and insufficient multi-scale modeling capability. Such limitations are particularly obvious when dealing with polyps that exhibit variable shapes, blurred borders, or high similarity between foreground and background. To address these issues, this study proposes a high-performance polyp segmentation model, <b>PolypFormer</b>, which introduces three key modules: the <b>Parallel Self-Attention Module (PSM)</b>, designed to enhance the global contextual representation capability of the encoder; the <b>Multi-Attention Mechanism Module (MAMM)</b>, which improves attention to critical regions and suppresses background noise through multi-scale mechanisms; and the <b>Cross-Dimensional Cross-Fusion Module (CCM)</b>, which adapts to lesion scale variations via cross-layer feature fusion. We validated all models on five open-source datasets (Kvasir, CVC-ClinicDB, CVC-ColonDB, CVC-300, and ETIS) and demonstrated that PolypFormer achieves SOTA in terms of mDice, mIoU, Precision, and Recall metrics, with maximum mDice of 0.939 and very high inference efficiency (137 FPS).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic feature scale and multi-scale fusion networks for polyp segmentation

  • Nannan Huang,
  • Abdul Hadi Abd Rahman,
  • Kauthar Mohd Daud,
  • Liantao Shi,
  • Hongqing Wang

摘要

Colorectal polyps are the main precursors of colorectal cancer, and early detection of polyps relies on accurate image segmentation techniques. However, existing segmentation methods often face several key challenges. These include semantic information loss, background noise interference, and insufficient multi-scale modeling capability. Such limitations are particularly obvious when dealing with polyps that exhibit variable shapes, blurred borders, or high similarity between foreground and background. To address these issues, this study proposes a high-performance polyp segmentation model, PolypFormer, which introduces three key modules: the Parallel Self-Attention Module (PSM), designed to enhance the global contextual representation capability of the encoder; the Multi-Attention Mechanism Module (MAMM), which improves attention to critical regions and suppresses background noise through multi-scale mechanisms; and the Cross-Dimensional Cross-Fusion Module (CCM), which adapts to lesion scale variations via cross-layer feature fusion. We validated all models on five open-source datasets (Kvasir, CVC-ClinicDB, CVC-ColonDB, CVC-300, and ETIS) and demonstrated that PolypFormer achieves SOTA in terms of mDice, mIoU, Precision, and Recall metrics, with maximum mDice of 0.939 and very high inference efficiency (137 FPS).