Deep generative models, particularly Generative Adversarial Networks (GANs), have made significant strides in the fields of image generation and manipulation. However, current image processing methods face two critical challenges: insufficient control over the outputs of generative models and a lack of interpretability regarding changes in the latent space of these models. To address these challenges, we propose a novel method called Fredit to facilitate controllable image editing in the frequency space. The core idea is to design a learnable Finite-Impulse-Response (FIR) filter to model and manipulate specific frequency components with respect to the semantic changes of the input image data. Since our DSP components are designed for vision tasks and exploit the attention mechanism as in visual transformers, we call the proposed module Visual DSP (VDSP) and propose a contrastive representation loss to supervise the learning of our VDSP filter. We extensively evaluate the Fredit method on two large-scale datasets for image editing tasks. Through comprehensive experiments, we demonstrate that Fredit can significantly enhance the controllability and interpretability of generative image editing, effectively allowing for the manipulation of semantic changes in images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual Differential Signal Processing for Generative Image Editing

  • Peng Wei,
  • Yuanting Zhang,
  • Quanwei Wu,
  • Yi Wang

摘要

Deep generative models, particularly Generative Adversarial Networks (GANs), have made significant strides in the fields of image generation and manipulation. However, current image processing methods face two critical challenges: insufficient control over the outputs of generative models and a lack of interpretability regarding changes in the latent space of these models. To address these challenges, we propose a novel method called Fredit to facilitate controllable image editing in the frequency space. The core idea is to design a learnable Finite-Impulse-Response (FIR) filter to model and manipulate specific frequency components with respect to the semantic changes of the input image data. Since our DSP components are designed for vision tasks and exploit the attention mechanism as in visual transformers, we call the proposed module Visual DSP (VDSP) and propose a contrastive representation loss to supervise the learning of our VDSP filter. We extensively evaluate the Fredit method on two large-scale datasets for image editing tasks. Through comprehensive experiments, we demonstrate that Fredit can significantly enhance the controllability and interpretability of generative image editing, effectively allowing for the manipulation of semantic changes in images.