<p>In recent years, image style transfer has matured significantly in the field of computer vision. However, current methods for image style transfer still face the challenge of balancing between content edge contours and style texture strokes, as it is difficult to control the degree of stylization. To address this issue, we propose a UNet based on Transformer feature fusion, named Trans-DCUNet. The network adaptively integrates content and style features by taking advantage of the self-learning characteristics of the cross-attention mechanism in the Transformer. The network combines Transformer and CNN, which takes advantage of the global context capture ability of Transformers and the local modeling ability of convolutional neural networks to fully learn image features at multiple levels, thereby generating stylized images. To enhance texture generation, we introduce a deformable convolutional residual module, which allows the convolution kernel to adapt to varying image features, capturing fine texture details more effectively. Additionally, we augment the traditional perception loss with edge detection loss and frequency perception loss, aiming to better preserve the edge contours of the content image and learn the texture strokes of the style image. Our experiments were conducted on the Microsoft COCO and WikiArt datasets. Experimental results show that our method achieves a content retention SSIM of up to 0.8655 and a style similarity LPIPS of 0.5655, outperforming most competing methods, while generating more artistic stylized images with significantly improved visual effects.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Transformer-Based Deformable Convolution UNet for Adaptive Arbitrary Style Transfer

  • Yingjie Zhao,
  • Libo Xu,
  • Chaoyi Pang,
  • Genlang Chen,
  • Huanhuan Wang,
  • Xin Yu

摘要

In recent years, image style transfer has matured significantly in the field of computer vision. However, current methods for image style transfer still face the challenge of balancing between content edge contours and style texture strokes, as it is difficult to control the degree of stylization. To address this issue, we propose a UNet based on Transformer feature fusion, named Trans-DCUNet. The network adaptively integrates content and style features by taking advantage of the self-learning characteristics of the cross-attention mechanism in the Transformer. The network combines Transformer and CNN, which takes advantage of the global context capture ability of Transformers and the local modeling ability of convolutional neural networks to fully learn image features at multiple levels, thereby generating stylized images. To enhance texture generation, we introduce a deformable convolutional residual module, which allows the convolution kernel to adapt to varying image features, capturing fine texture details more effectively. Additionally, we augment the traditional perception loss with edge detection loss and frequency perception loss, aiming to better preserve the edge contours of the content image and learn the texture strokes of the style image. Our experiments were conducted on the Microsoft COCO and WikiArt datasets. Experimental results show that our method achieves a content retention SSIM of up to 0.8655 and a style similarity LPIPS of 0.5655, outperforming most competing methods, while generating more artistic stylized images with significantly improved visual effects.