<p>This paper presents CONSRGAN, an enhanced super-resolution (SR) model that addresses SRGAN limitations through architectural innovations and efficiency optimizations. Its generator integrates redesigned residual blocks with depth-wise separable convolutions(DSConv), dynamic layer scaling, and channel-expanded point-wise convolutions. A critical adjustment replaces batch normalization (BatchNorma) with layer normalization (LayerNorm), reducing parameters and enhancing training stability. Multi-scale feature fusion is strengthened through stacked CONSRGANBlocks, channel expansion, and hybrid upsampling (two-stage PixelShuffle with nonlinear activation). Experimental evaluations on 4<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4759_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> SR tasks across COCO 2017, ImageNet-100, Set 5, and Set 14 demonstrate CONSRGAN outperforms SRGAN by an average <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4759_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="80" /> </InlineMediaObject> <EquationSource Format="TEX">\(2.86 \pm 1.08\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>2.86</mn> <mo>±</mo> <mn>1.08</mn> </mrow> </math></EquationSource> </InlineEquation> dB in PSNR and <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4759_Article_IEq3.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="90" /> </InlineMediaObject> <EquationSource Format="TEX">\(7.8\% \pm 3.1\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>7.8</mn> <mo>%</mo> <mo>±</mo> <mn>3.1</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> in SSIM, with statistical significance (t-test, <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4759_Article_IEq4.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="62" /> </InlineMediaObject> <EquationSource Format="TEX">\(p &lt; 0.05\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>p</mi> <mo>&lt;</mo> <mn>0.05</mn> </mrow> </math></EquationSource> </InlineEquation>). Qualitative assessments verify superior edge sharpness, texture reconstruction, and artifact suppression while preserving visual fidelity. Key optimizations include 50% reduced residual block parameters, stable LayerNorm-driven training, and enhanced multi-scale aggregation, effectively mitigating SRGAN’s inefficiencies and batch-dependent variations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Convolutional Optimized Network with DSConv for Image SR via GAN

  • ZeLun Li,
  • ChangJiang Liu

摘要

This paper presents CONSRGAN, an enhanced super-resolution (SR) model that addresses SRGAN limitations through architectural innovations and efficiency optimizations. Its generator integrates redesigned residual blocks with depth-wise separable convolutions(DSConv), dynamic layer scaling, and channel-expanded point-wise convolutions. A critical adjustment replaces batch normalization (BatchNorma) with layer normalization (LayerNorm), reducing parameters and enhancing training stability. Multi-scale feature fusion is strengthened through stacked CONSRGANBlocks, channel expansion, and hybrid upsampling (two-stage PixelShuffle with nonlinear activation). Experimental evaluations on 4 \(\times \) × SR tasks across COCO 2017, ImageNet-100, Set 5, and Set 14 demonstrate CONSRGAN outperforms SRGAN by an average \(2.86 \pm 1.08\) 2.86 ± 1.08 dB in PSNR and \(7.8\% \pm 3.1\%\) 7.8 % ± 3.1 % in SSIM, with statistical significance (t-test, \(p < 0.05\) p < 0.05 ). Qualitative assessments verify superior edge sharpness, texture reconstruction, and artifact suppression while preserving visual fidelity. Key optimizations include 50% reduced residual block parameters, stable LayerNorm-driven training, and enhanced multi-scale aggregation, effectively mitigating SRGAN’s inefficiencies and batch-dependent variations.