Portrait drawing generation typically relies on CycleGAN-based methods that utilize cyclic consistency loss to facilitate unpaired image translation. However, these approaches often struggle to preserve accurate facial semantic features when processing highly abstract artistic styles, leading to contour distortions or loss of critical details in the generated portraits. In this paper, we propose an unsupervised asymmetric network for portrait drawing generation. Firstly, we introduce a perceptual cycle consistency loss, which encourages perceptual-level similarity between input photos and reconstructed images, thereby improving the fidelity of style transfer. Secondly, we design a generator architecture that integrates an autoencoder backbone with skip connections and self-attention blocks, which enhances global context modeling while preserving fine-grained facial details. Finally, we incorporate a CLIP-guided semantic perception loss, which aligns high-level semantic features between source images and generated portraits, effectively maintaining structural integrity and key identity features. Extensive experiments on benchmark datasets demonstrate that our method produces portrait drawings with higher visual quality and stronger semantic fidelity compared to state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unpaired Artistic Portrait Drawing Generation with Asymmetric Network

  • Di Sun,
  • Xiaoyu Li,
  • Tingting Yang,
  • Chuanlei Zhang,
  • Chao Ren,
  • Gang Pan

摘要

Portrait drawing generation typically relies on CycleGAN-based methods that utilize cyclic consistency loss to facilitate unpaired image translation. However, these approaches often struggle to preserve accurate facial semantic features when processing highly abstract artistic styles, leading to contour distortions or loss of critical details in the generated portraits. In this paper, we propose an unsupervised asymmetric network for portrait drawing generation. Firstly, we introduce a perceptual cycle consistency loss, which encourages perceptual-level similarity between input photos and reconstructed images, thereby improving the fidelity of style transfer. Secondly, we design a generator architecture that integrates an autoencoder backbone with skip connections and self-attention blocks, which enhances global context modeling while preserving fine-grained facial details. Finally, we incorporate a CLIP-guided semantic perception loss, which aligns high-level semantic features between source images and generated portraits, effectively maintaining structural integrity and key identity features. Extensive experiments on benchmark datasets demonstrate that our method produces portrait drawings with higher visual quality and stronger semantic fidelity compared to state-of-the-art approaches.