Hierarchical Local-Latent Diffusion Model for Efficient Video Deblurring
摘要
Current methods typically employ spatio-temporal transformers to aggregate temporal information, but their practical effectiveness is constrained by temporal window length, fundamentally limiting their ability to capture long-range temporal dependencies. To address these challenges, we present HLLDM: a Hierarchical Local-Latent Video Diffusion Model that introduces a subspace compression attention (SCA) module that projects long-range sequences into a local latent space for efficient shift-window (Swin) transformer, and a subspace expansion attention (SEA) module that reconstructs the original dimensionality. By integrating the SCA-SEA module with the Swin Transformer, we propose a hybrid architecture named Local-Latent Swin Transformer (LLST). This architecture achieves computational complexity growth at half the rate of standard Swin transformer methods for growing sequence length and 4x receptive field, enabling effective processing of extended temporal contexts and enhanced model performance. Extensive experiments demonstrate that HLLDM achieves outstanding perceptual metrics and distortion metrics. Visualization results further validate that balancing perceptual and distortion metrics yields superior reconstruction quality.