<p>Zero-shot referring image segmentation (RIS) is a challenging task that involves identifying an instance segmentation mask from referring texts, without being trained on paired image-text data. Current zero-shot RIS methods mainly rely on pre-trained discriminative models (e.g., CLIP). In contrast, this study investigates the potential of generative models (e.g., Stable Diffusion) to understand relationships between various visual elements and text descriptions, an area that remains unexplored in this context. In this work, we introduce the Referring Diffusional segmentor (Ref-Diff), a model that harnesses the fine-grained multimodal information provided by generative models. Our results demonstrate that Ref-Diff, using only a generative model and no external proposal generator, outperforms state-of-the-art weakly supervised models on the RefCOCO+ and RefCOCOg benchmarks. Furthermore, by combining both generative and discriminative models, we present the enhanced version, Ref-Diff+, which significantly surpasses existing methods. This emphasizes the benefits of generative models for discriminative models, thereby improving referring segmentation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ref-Diff: zero-shot referring image segmentation with generative models

  • Minheng Ni,
  • Yabo Zhang,
  • Kailai Feng,
  • Xiaoming Li,
  • Yiwen Guo,
  • Wangmeng Zuo

摘要

Zero-shot referring image segmentation (RIS) is a challenging task that involves identifying an instance segmentation mask from referring texts, without being trained on paired image-text data. Current zero-shot RIS methods mainly rely on pre-trained discriminative models (e.g., CLIP). In contrast, this study investigates the potential of generative models (e.g., Stable Diffusion) to understand relationships between various visual elements and text descriptions, an area that remains unexplored in this context. In this work, we introduce the Referring Diffusional segmentor (Ref-Diff), a model that harnesses the fine-grained multimodal information provided by generative models. Our results demonstrate that Ref-Diff, using only a generative model and no external proposal generator, outperforms state-of-the-art weakly supervised models on the RefCOCO+ and RefCOCOg benchmarks. Furthermore, by combining both generative and discriminative models, we present the enhanced version, Ref-Diff+, which significantly surpasses existing methods. This emphasizes the benefits of generative models for discriminative models, thereby improving referring segmentation.