MSFF-RWKV: Single-Structure Multi-stage Feature Fusion Lightweight Super-Resolution Network
摘要
In recent years, residual learning has become the cornerstone of lightweight super-resolution networks. However, these approaches often rely on stacking numerous layers, making it challenging to substantially reduce parameter count and computational complexity. Meanwhile, the Receptance Weighted Key Value (RWKV) architecture—originally developed for natural language processing—has garnered increasing attention for its ability to efficiently handle long sequences and high-resolution data. To address the limitations of conventional residual-based designs and harness the efficiency of RWKV, we introduce a novel multi-stage feature fusion framework built upon a single RWKV block. During training, this block progressively learns diverse semantic representations by recursively integrating its current output with previously processed features, significantly reducing both parameters and FLOPs. Furthermore, we propose a Local Pixel Perception (LPP) layer that promotes high-order information sharing and strengthens interactions between pixels and their surrounding contextual regions. To capture multi-scale features more effectively, we also design the ME-Shift module. Experiments demonstrate the strong potential of our MSFF-RWKV model, achieving a PSNR gain of 0.14 dB and an SSIM improvement of 0.0007, while reducing parameter count by 26.6%, thus exhibiting both strong performance and high efficiency.