MSR former: multi-scale spectral refinement transformer for single image dehazing
摘要
Image dehazing is a challenging task in image restoration, aiming to recover clear images from hazy observations. Although many Transformer-based methods improve performance by modifying self-attention mechanisms, the potential of using frequency domain information to enhance Transformer learning has not been fully explored. In this paper, we propose a frequency-domain refined self-attention that refines input features in the frequency domain to guide the generation of attention maps. To preserve high-frequency information, we use this attention only in the decoding stage. Additionally, we find that simple feed-forward networks cannot model the potential local information, and to overcome this problem, we propose a multi-scale enhanced feed-forward network to recover clearer image details. Furthermore, mixed attention fusion block is incorporated in the decoding stage to effectively merge shallow and deep features. By combining these components, we present the Multi-Scale Spectral Refinement Transformer (MSRformer), and experimental results on several dehazing benchmarks show that MSRformer achieves state-of-the-art performance while maintaining a low computational burden.