<p>In recent years, deep learning has driven significant advances in computer vision, making its application to surface segmentation a critical research focus in many industrial systems. However, accurate segmentation in industrial scenarios remains challenging due to complex textures, noise, and the difficulty of extracting refined features from coarse-grained representations. To address these issues, we propose a diffusion model with anchor condition hybrid transformer (Diff-ACHT) for industrial image segmentation. Diff-ACHT combines a conditional hybrid transformer (CHT) with an anchored conditional model based on the diffusion framework. The anchoring condition extracts coarse-grained features, narrowing the feature extraction scope and enabling better refinement of features in subsequent stages. CHT includes two key modules: an overlapping cross-attention block that refines features by fusing diffusion and anchor condition information, and a compact spatial attention block that integrates multiple attention mechanisms to activate more relevant pixels. Additionally, we incorporate up-sampling within the decoder of the diffusion model to enhance feature capture at each diffusion stage. Extensive experiments on the KolektorSDD2 and Blowhole &amp; Crack datasets demonstrate that Diff-ACHT achieves state-of-the-art results in both quantitative metrics and visual quality, delivering high-quality segmentation performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Diff-ACHT: a diffusion model with anchor condition hybrid transformer for industrial image segmentation

  • Lisen Wang,
  • Jialiang Shi

摘要

In recent years, deep learning has driven significant advances in computer vision, making its application to surface segmentation a critical research focus in many industrial systems. However, accurate segmentation in industrial scenarios remains challenging due to complex textures, noise, and the difficulty of extracting refined features from coarse-grained representations. To address these issues, we propose a diffusion model with anchor condition hybrid transformer (Diff-ACHT) for industrial image segmentation. Diff-ACHT combines a conditional hybrid transformer (CHT) with an anchored conditional model based on the diffusion framework. The anchoring condition extracts coarse-grained features, narrowing the feature extraction scope and enabling better refinement of features in subsequent stages. CHT includes two key modules: an overlapping cross-attention block that refines features by fusing diffusion and anchor condition information, and a compact spatial attention block that integrates multiple attention mechanisms to activate more relevant pixels. Additionally, we incorporate up-sampling within the decoder of the diffusion model to enhance feature capture at each diffusion stage. Extensive experiments on the KolektorSDD2 and Blowhole & Crack datasets demonstrate that Diff-ACHT achieves state-of-the-art results in both quantitative metrics and visual quality, delivering high-quality segmentation performance.