<p>With the rapid advancement of deep learning, image inpainting has achieved remarkable progress, yet challenges such as blurring persist, especially when restoring large missing regions. In this work, we propose a novel dual-branch architecture that integrates U-Net and Transformer networks: U-Net is leveraged for restoring fine-grained textures, while the Transformer effectively models long-range dependencies to ensure global semantic consistency. To overcome the substantial computational overhead typically associated with Transformer-based models, we further introduce Trans-GAN—a lightweight structure that synergistically combines self-attention mechanisms with efficient convolutional operations, significantly reducing resource consumption without sacrificing contextual modeling capability. Moreover, we incorporate a style loss into the training of the GAN, guiding the restored images to match the original not only in content but also in style, thus delivering inpainted results with enhanced aesthetic and stylistic coherence. Extensive qualitative and quantitative experiments on the CelebA and Paris datasets demonstrate that our approach achieves superior inpainting quality, efficient inference, and exceptional consistency compared to contemporary methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Trans-GAN: lightweight dual-branch transformer for style-preserved image inpainting

  • Huaming Liu,
  • Minglong Zhang,
  • Xiuyou Wang,
  • Xuehui Bi,
  • Yumu Wang

摘要

With the rapid advancement of deep learning, image inpainting has achieved remarkable progress, yet challenges such as blurring persist, especially when restoring large missing regions. In this work, we propose a novel dual-branch architecture that integrates U-Net and Transformer networks: U-Net is leveraged for restoring fine-grained textures, while the Transformer effectively models long-range dependencies to ensure global semantic consistency. To overcome the substantial computational overhead typically associated with Transformer-based models, we further introduce Trans-GAN—a lightweight structure that synergistically combines self-attention mechanisms with efficient convolutional operations, significantly reducing resource consumption without sacrificing contextual modeling capability. Moreover, we incorporate a style loss into the training of the GAN, guiding the restored images to match the original not only in content but also in style, thus delivering inpainted results with enhanced aesthetic and stylistic coherence. Extensive qualitative and quantitative experiments on the CelebA and Paris datasets demonstrate that our approach achieves superior inpainting quality, efficient inference, and exceptional consistency compared to contemporary methods.