This paper presents a novel training-free semantic segmentation method that leverages a pre-trained large-scale image generation model incorporating the Multi-modal Diffusion Transformer (MM-DiT) architecture. Inspired by training-free segmentation techniques using the U-Net-based noise removal model in the Stable Diffusion framework, our approach extracts cross-attention maps between textual and visual features during the inference stages of the MM-DiT to generate mask images. Experimental results demonstrate that our method achieves segmentation accuracy comparable to CLIP-based and U-Net-based stable diffusion methods. While the direct segmentation scores are relatively modest, the significance of our work lies in the exploration of cross-attention maps within the DiT. This investigation provides critical insights that could advance training-free segmentation methodologies and enhance the interpretability of diffusion-based models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Cross-Attention Maps in Multi-modal Diffusion Transformers for Training-Free Semantic Segmentation

  • Rento Yamaguchi,
  • Keiji Yanai

摘要

This paper presents a novel training-free semantic segmentation method that leverages a pre-trained large-scale image generation model incorporating the Multi-modal Diffusion Transformer (MM-DiT) architecture. Inspired by training-free segmentation techniques using the U-Net-based noise removal model in the Stable Diffusion framework, our approach extracts cross-attention maps between textual and visual features during the inference stages of the MM-DiT to generate mask images. Experimental results demonstrate that our method achieves segmentation accuracy comparable to CLIP-based and U-Net-based stable diffusion methods. While the direct segmentation scores are relatively modest, the significance of our work lies in the exploration of cross-attention maps within the DiT. This investigation provides critical insights that could advance training-free segmentation methodologies and enhance the interpretability of diffusion-based models.