SPI2I: Structure-Preserved Image-to-Image Translation with Diffusion Models
摘要
Large-scale text-to-image generative models are already proficient at producing high-quality results that closely match the intended prompts. Nevertheless, the pivotal challenge in image editing tasks lies in the difficulty of confining alterations within the editing region while preserving the structure and details of the source image. In this paper, we propose a zero-shot structure-preserved image-to-image translation approach based on diffusion models. We combine the optimization of the latent code and the injection of the U-Net features to strengthen the structural preservation effect by alleviating the inconsistency between the information contained in the latent code and the injected features. Our method effectively preserves the structural and detailed information of the source image while enhancing the quality of the generated results. We exhibit comprehensive and high-quality experimental results showcasing that our approach surpasses state-of-the-art methods across various image-to-image translation tasks.