<p>Denoising diffusion probabilistic models (DDPMs) have been proven capable of synthesizing high-quality images with remarkable diversity when trained on large amounts of data. Unfortunately, they are still vulnerable to overfitting when fine-tuned on limited data. Existing works have explored subject-driven generation with text-to-image (T2I) models using a few samples. However, there is still a lack of effective and stable data-efficient methods to synthesize images in specific domains (e.g. styles or properties), which remains challenging due to ambiguities inherent in natural language and out-of-distribution effects. This paper introduces a few-shot fine-tuning approach named DomainStudio as a domain-driven image generation paradigm, which is designed to retain the subjects from prior knowledge provided by pre-trained models and adapt them to the domain extracted from training data, pursuing high quality and great diversity. We propose to keep the image-level relative distances between adapted samples and enhance the learning of high-frequency details from both pre-trained models and training samples. DomainStudio is compatible with both unconditional and T2I DDPMs. The proposed method achieves better results than current state-of-the-art GAN-based approaches in unconditional few-shot image generation. It also outperforms existing few-shot fine-tuning methods for modern large-scale T2I diffusion models like Textual Inversion and DreamBooth on synthesizing samples in specific domains characterized by few-shot training data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DomainStudio: Fine-Tuning Diffusion Models for Domain-Driven Image Generation Using Limited Data

  • Jingyuan Zhu,
  • Huimin Ma,
  • Jiansheng Chen,
  • Jian Yuan

摘要

Denoising diffusion probabilistic models (DDPMs) have been proven capable of synthesizing high-quality images with remarkable diversity when trained on large amounts of data. Unfortunately, they are still vulnerable to overfitting when fine-tuned on limited data. Existing works have explored subject-driven generation with text-to-image (T2I) models using a few samples. However, there is still a lack of effective and stable data-efficient methods to synthesize images in specific domains (e.g. styles or properties), which remains challenging due to ambiguities inherent in natural language and out-of-distribution effects. This paper introduces a few-shot fine-tuning approach named DomainStudio as a domain-driven image generation paradigm, which is designed to retain the subjects from prior knowledge provided by pre-trained models and adapt them to the domain extracted from training data, pursuing high quality and great diversity. We propose to keep the image-level relative distances between adapted samples and enhance the learning of high-frequency details from both pre-trained models and training samples. DomainStudio is compatible with both unconditional and T2I DDPMs. The proposed method achieves better results than current state-of-the-art GAN-based approaches in unconditional few-shot image generation. It also outperforms existing few-shot fine-tuning methods for modern large-scale T2I diffusion models like Textual Inversion and DreamBooth on synthesizing samples in specific domains characterized by few-shot training data.