<p>Image dehazing is a challenging task in image restoration, aiming to recover clear images from hazy observations. Although many Transformer-based methods improve performance by modifying self-attention mechanisms, the potential of using frequency domain information to enhance Transformer learning has not been fully explored. In this paper, we propose a frequency-domain refined self-attention that refines input features in the frequency domain to guide the generation of attention maps. To preserve high-frequency information, we use this attention only in the decoding stage. Additionally, we find that simple feed-forward networks cannot model the potential local information, and to overcome this problem, we propose a multi-scale enhanced feed-forward network to recover clearer image details. Furthermore, mixed attention fusion block is incorporated in the decoding stage to effectively merge shallow and deep features. By combining these components, we present the Multi-Scale Spectral Refinement Transformer (MSRformer), and experimental results on several dehazing benchmarks show that MSRformer achieves state-of-the-art performance while maintaining a low computational burden.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MSR former: multi-scale spectral refinement transformer for single image dehazing

  • Yufeng Li,
  • Rui Li,
  • Zitian Zhao

摘要

Image dehazing is a challenging task in image restoration, aiming to recover clear images from hazy observations. Although many Transformer-based methods improve performance by modifying self-attention mechanisms, the potential of using frequency domain information to enhance Transformer learning has not been fully explored. In this paper, we propose a frequency-domain refined self-attention that refines input features in the frequency domain to guide the generation of attention maps. To preserve high-frequency information, we use this attention only in the decoding stage. Additionally, we find that simple feed-forward networks cannot model the potential local information, and to overcome this problem, we propose a multi-scale enhanced feed-forward network to recover clearer image details. Furthermore, mixed attention fusion block is incorporated in the decoding stage to effectively merge shallow and deep features. By combining these components, we present the Multi-Scale Spectral Refinement Transformer (MSRformer), and experimental results on several dehazing benchmarks show that MSRformer achieves state-of-the-art performance while maintaining a low computational burden.