The Transformer achieves remarkable success in image super-resolution (SR). However, it has limitations in handling complex local details, whereas Convolutional Neural Networks (CNNs) are more advantageous in fine feature extraction. Additionally, we find that the Feedforward Network (FFN) in the Transformer architecture does not effectively integrate spatial and channel information, containing redundant information that hinders feature representation capability. Thus, we propose the Adaptive Multi-scale Fusion Transformer (AMFT), which combines the global advantages of the Transformer with the local fine feature extraction capabilities of CNNs. Through feature shifting, multi-scale feature extraction, and hierarchical feature fusion, the AMFT significantly enhances image detail representation and visual quality. In detail, we incorporate the Multi-Scale Shifting Convolution Module (MSSCM) into the Transformer. MSSCM first shifts the feature positions and then extracts and fuses features at different scales, preserving more details, and through a hybrid attention mechanism achieves global transmission of information. Additionally, we replace FFN with the Spatial Channel Fusion Module (SCFM) to achieve full information integration and reduce computational complexity. Extensive experiments demonstrate the superior performance of AMFT compared to existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Multi-scale Fusion Transformer for Image Super-Resolution

  • Yusheng Sun,
  • Weihai Chen,
  • Xiaogang Zang,
  • Changjiang Wang

摘要

The Transformer achieves remarkable success in image super-resolution (SR). However, it has limitations in handling complex local details, whereas Convolutional Neural Networks (CNNs) are more advantageous in fine feature extraction. Additionally, we find that the Feedforward Network (FFN) in the Transformer architecture does not effectively integrate spatial and channel information, containing redundant information that hinders feature representation capability. Thus, we propose the Adaptive Multi-scale Fusion Transformer (AMFT), which combines the global advantages of the Transformer with the local fine feature extraction capabilities of CNNs. Through feature shifting, multi-scale feature extraction, and hierarchical feature fusion, the AMFT significantly enhances image detail representation and visual quality. In detail, we incorporate the Multi-Scale Shifting Convolution Module (MSSCM) into the Transformer. MSSCM first shifts the feature positions and then extracts and fuses features at different scales, preserving more details, and through a hybrid attention mechanism achieves global transmission of information. Additionally, we replace FFN with the Spatial Channel Fusion Module (SCFM) to achieve full information integration and reduce computational complexity. Extensive experiments demonstrate the superior performance of AMFT compared to existing methods.