Although transformers have achieved remarkable performance across different vision tasks, they have not yet demonstrated the comparable proficiency to CNN (ConvNets) in image generation. This paper presents an innovative Swin Transformer plus VQVAE model for encoding and decoding. Elevating Swin Transformer’s scalability and global dependency modeling, our proposed model addresses these shortcomings. This advancement fills the gap in transformer-based solutions engineered specifically for image generation thereby offering improved quality and efficiency in generating images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vision Transformer Hybrid Model for Enhanced Image Generation: Integrating VQVAE and Swin

  • Saarthak Singh,
  • Satvik Gupta,
  • Shivam Garg,
  • Shailender Kumar

摘要

Although transformers have achieved remarkable performance across different vision tasks, they have not yet demonstrated the comparable proficiency to CNN (ConvNets) in image generation. This paper presents an innovative Swin Transformer plus VQVAE model for encoding and decoding. Elevating Swin Transformer’s scalability and global dependency modeling, our proposed model addresses these shortcomings. This advancement fills the gap in transformer-based solutions engineered specifically for image generation thereby offering improved quality and efficiency in generating images.