Visual content creation is increasingly driven by generative models. Most notably, diffusion-based text-to-image (T2I) models have recently seen widespread adoption due to their flexibility and intuitive use. The generative workflow of T2I models often involves extensive iterative refinement of the text prompt, a laborious task that requires detailed knowledge of prompting techniques. This paper explores strategies to incorporate human feedback into the generative process in order to alleviate some of this burden while improving the quality of outputs. We propose FABRIC, a training-free approach to adapt the generative process through attention injection at inference time, incorporating user feedback in the form of positive and negative reference images, without any explicit need of textual guidance. We evaluate the proposed method both quantitatively by using preference and similarity models to emulate human feedback and qualitatively by measuring the subjective experience of real-world FABRIC users through a user study. Our results show that the proposed method improves generation results over multiple rounds of feedback and that users are able to arrive at better results more quickly when using FABRIC.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FABRIC: Personalizing Diffusion Models with Iterative Feedback

  • Dimitri von Rütte,
  • Elisabetta Fedele,
  • Jonathan Thomm,
  • Lukas Wolf

摘要

Visual content creation is increasingly driven by generative models. Most notably, diffusion-based text-to-image (T2I) models have recently seen widespread adoption due to their flexibility and intuitive use. The generative workflow of T2I models often involves extensive iterative refinement of the text prompt, a laborious task that requires detailed knowledge of prompting techniques. This paper explores strategies to incorporate human feedback into the generative process in order to alleviate some of this burden while improving the quality of outputs. We propose FABRIC, a training-free approach to adapt the generative process through attention injection at inference time, incorporating user feedback in the form of positive and negative reference images, without any explicit need of textual guidance. We evaluate the proposed method both quantitatively by using preference and similarity models to emulate human feedback and qualitatively by measuring the subjective experience of real-world FABRIC users through a user study. Our results show that the proposed method improves generation results over multiple rounds of feedback and that users are able to arrive at better results more quickly when using FABRIC.