<p>Modern CNN architectures can compete with Transformers in various vision tasks. Therefore, we propose a Multi-scale Chunked Feature Mixing Network (MCFM) for efficient image super-resolution. Our approach combines re-parameterization with Incremental Weight Optimization, maintaining multi-branch capability during training while ensuring single-branch efficiency at inference. The core of MCFM is the Multi-Scale Feature Decomposition (MSFD) module, which employs a three-tier strategy: MBRConvS captures fine-grained local features through multi-branch structures, MBRConvM extracts medium-scale contextual information, and decomposed strip convolutions efficiently model global spatial relationships. To enhance feature representation, we introduce the Gated Spatial Attention Unit (GSAU) that integrates spatial attention with gating mechanisms while reducing computational complexity. Additionally, the Hierarchical Dual-Path Attention (HDPA) mechanism realizes adaptive feature selection through collaborative spatial and global attention paths. Experimental results on benchmark datasets show that MCFM outperforms state-of-the-art methods by 0.2 dB on Urban100 <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4728_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>4 super-resolution with 58% parameter reduction, indicating substantial improvements in computational efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCFM: An efficient multi-scale feature mixing network for image super-resolution

  • Xiaoyu Yan,
  • Wei Song,
  • Wei Guo,
  • Wenhao Cui,
  • Keqing Ning

摘要

Modern CNN architectures can compete with Transformers in various vision tasks. Therefore, we propose a Multi-scale Chunked Feature Mixing Network (MCFM) for efficient image super-resolution. Our approach combines re-parameterization with Incremental Weight Optimization, maintaining multi-branch capability during training while ensuring single-branch efficiency at inference. The core of MCFM is the Multi-Scale Feature Decomposition (MSFD) module, which employs a three-tier strategy: MBRConvS captures fine-grained local features through multi-branch structures, MBRConvM extracts medium-scale contextual information, and decomposed strip convolutions efficiently model global spatial relationships. To enhance feature representation, we introduce the Gated Spatial Attention Unit (GSAU) that integrates spatial attention with gating mechanisms while reducing computational complexity. Additionally, the Hierarchical Dual-Path Attention (HDPA) mechanism realizes adaptive feature selection through collaborative spatial and global attention paths. Experimental results on benchmark datasets show that MCFM outperforms state-of-the-art methods by 0.2 dB on Urban100 \(\times \) × 4 super-resolution with 58% parameter reduction, indicating substantial improvements in computational efficiency.