<p>Image super-resolution reconstruction is a core and widely studied issue in image processing. To address the challenge of inadequate recovery of texture details in image super-resolution, an algorithm was proposed to enhance the effectiveness of the striped window transformer. First, the design of the stripe window-based self-attention mechanism improves the network’s ability to capture anisotropic features from images. Second, the design of the channel and global transformer module enable the modeling of spatial-channel dimensions and global-local scale features within feature maps. Next, the dilated spatial attention module is designed to capture multiscale spatial feature information and to restore local texture details in images. Finally, during the fine-tuning phase, the combined use of the improved segmentation-based perceptual loss and <i>L</i>1 loss further enhances the network’s performance. Simulation experiment results show that the proposed algorithm outperforms existing methods in terms of both performance and visual fidelity across the Set5, Set14, BSD100, Urban100, and Manga109 benchmark datasets. Notably, when evaluated on Manga109, the proposed algorithm demonstrates substantial impro3"" vements over the leading algorithm SwinIR, achieving a 0.60dB increase in peak signal-to-noise ratio (PSNR) and a 0.0047 improvement in structural similarity (SSIM), while recovering richer image texture details.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Striped Window Transformer Algorithm for Image Super-Resolution

  • Jinyu Han,
  • Mingli Jing,
  • Cheng Zhang,
  • Long Jiao,
  • Lan Li,
  • Yan Zhang

摘要

Image super-resolution reconstruction is a core and widely studied issue in image processing. To address the challenge of inadequate recovery of texture details in image super-resolution, an algorithm was proposed to enhance the effectiveness of the striped window transformer. First, the design of the stripe window-based self-attention mechanism improves the network’s ability to capture anisotropic features from images. Second, the design of the channel and global transformer module enable the modeling of spatial-channel dimensions and global-local scale features within feature maps. Next, the dilated spatial attention module is designed to capture multiscale spatial feature information and to restore local texture details in images. Finally, during the fine-tuning phase, the combined use of the improved segmentation-based perceptual loss and L1 loss further enhances the network’s performance. Simulation experiment results show that the proposed algorithm outperforms existing methods in terms of both performance and visual fidelity across the Set5, Set14, BSD100, Urban100, and Manga109 benchmark datasets. Notably, when evaluated on Manga109, the proposed algorithm demonstrates substantial impro3"" vements over the leading algorithm SwinIR, achieving a 0.60dB increase in peak signal-to-noise ratio (PSNR) and a 0.0047 improvement in structural similarity (SSIM), while recovering richer image texture details.