Hybrid CNN-Transformer network with multi-scale attention for enhanced image compressive sensing
摘要
Image compressive sensing has emerged as a promising technique for efficient image acquisition and reconstruction. However, existing deep networks for CS reconstruction often suffer from limited contextual modeling capability and inefficient multi-scale feature extraction. Recent hybrid architectures that combine local and global feature representations show promise but relies heavily on static fusion for leveraging their complementary strengths. In this paper, we propose CT-MSA, a novel CNN-Transformer hybrid network that synergistically combines local detail extraction with global contextual modeling for superior CS reconstruction performance. Our approach introduces a Cross-modal Multi-scale Fusion mechanism that adaptively captures multi-scale spatial dependencies, and Multi-Kernel Gated Attention that effectively fuse local and global features through Gated Spatial Attention Units and Multi-scale Large Kernel Attention. Furthermore, we design an enhanced CNN-Transformer fusion strategy that progressively combines features from different scales to achieve optimal reconstruction quality. Extensive experiments on benchmark datasets demonstrate that CT-MSA achieves superior performance compared to state-of-the-art methods, with PSNR improvements ranging from 0.53 to 1.08 dB and SSIM gains up to 0.0091 across various sampling rates (1–25%), particularly excelling in low sampling rate scenarios where reconstruction is most challenging.