<p>Single image dehazing is a fundamental yet challenging task in computer vision. Many studies aim to improve learning-based methods by constructing deep residual networks with multiple small convolutional kernels and attention mechanisms. However, these methods encounter two major limitations. First, small convolutional kernels in redundant layers of deep residual networks may not train effectively, restricting the receptive field’s expansion. Second, haze impacts image information across three dimensions, necessitating consideration of their interrelationships in attention mechanism design for dehazing networks. Most existing methods analyze these dimensional relationships separately, which diminishes their effectiveness in the dehazing task. To address these challenges, this paper introduces the Multi-Scale Large Convolution Triple Attention Network (MLCTA-Net) within the U-Net framework. The proposed architecture consists of two primary modules: the multi-scale feature extraction unit utilizing parallel depth-wise large convolutions, and the triple attention unit with three branches that enhance interactions among single-dimensional information and the other two dimensions. By utilizing large convolutional kernels, MLCTA-Net effectively expands the receptive field, facilitating a more comprehensive understanding of image information. Furthermore, the attention mechanism, designed through interactions among dimensional information and parallel connections, better aligns with the multidimensional effects of haze on image information. Extensive experimental evaluations validate the effectiveness of the proposed method. Compared to the MAXIM-2&#xa0;S and Dehamer networks, MLCTA-Net achieves PSNR improvements of 3.85 dB and 0.71 dB on the SOTS-indoor and SOTS-outdoor test datasets, respectively, while utilizing only 1.65 million parameters.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MLCTA-Net: multi-scale large convolution and triplet attention network for single image dehazing

  • Qingshan Tang,
  • Yongqi Miao,
  • Xinwei Huang,
  • Huang Jiang

摘要

Single image dehazing is a fundamental yet challenging task in computer vision. Many studies aim to improve learning-based methods by constructing deep residual networks with multiple small convolutional kernels and attention mechanisms. However, these methods encounter two major limitations. First, small convolutional kernels in redundant layers of deep residual networks may not train effectively, restricting the receptive field’s expansion. Second, haze impacts image information across three dimensions, necessitating consideration of their interrelationships in attention mechanism design for dehazing networks. Most existing methods analyze these dimensional relationships separately, which diminishes their effectiveness in the dehazing task. To address these challenges, this paper introduces the Multi-Scale Large Convolution Triple Attention Network (MLCTA-Net) within the U-Net framework. The proposed architecture consists of two primary modules: the multi-scale feature extraction unit utilizing parallel depth-wise large convolutions, and the triple attention unit with three branches that enhance interactions among single-dimensional information and the other two dimensions. By utilizing large convolutional kernels, MLCTA-Net effectively expands the receptive field, facilitating a more comprehensive understanding of image information. Furthermore, the attention mechanism, designed through interactions among dimensional information and parallel connections, better aligns with the multidimensional effects of haze on image information. Extensive experimental evaluations validate the effectiveness of the proposed method. Compared to the MAXIM-2 S and Dehamer networks, MLCTA-Net achieves PSNR improvements of 3.85 dB and 0.71 dB on the SOTS-indoor and SOTS-outdoor test datasets, respectively, while utilizing only 1.65 million parameters.