Current methods typically employ spatio-temporal transformers to aggregate temporal information, but their practical effectiveness is constrained by temporal window length, fundamentally limiting their ability to capture long-range temporal dependencies. To address these challenges, we present HLLDM: a Hierarchical Local-Latent Video Diffusion Model that introduces a subspace compression attention (SCA) module that projects long-range sequences into a local latent space for efficient shift-window (Swin) transformer, and a subspace expansion attention (SEA) module that reconstructs the original dimensionality. By integrating the SCA-SEA module with the Swin Transformer, we propose a hybrid architecture named Local-Latent Swin Transformer (LLST). This architecture achieves computational complexity growth at half the rate of standard Swin transformer methods for growing sequence length and 4x receptive field, enabling effective processing of extended temporal contexts and enhanced model performance. Extensive experiments demonstrate that HLLDM achieves outstanding perceptual metrics and distortion metrics. Visualization results further validate that balancing perceptual and distortion metrics yields superior reconstruction quality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical Local-Latent Diffusion Model for Efficient Video Deblurring

  • Haoyang Long

摘要

Current methods typically employ spatio-temporal transformers to aggregate temporal information, but their practical effectiveness is constrained by temporal window length, fundamentally limiting their ability to capture long-range temporal dependencies. To address these challenges, we present HLLDM: a Hierarchical Local-Latent Video Diffusion Model that introduces a subspace compression attention (SCA) module that projects long-range sequences into a local latent space for efficient shift-window (Swin) transformer, and a subspace expansion attention (SEA) module that reconstructs the original dimensionality. By integrating the SCA-SEA module with the Swin Transformer, we propose a hybrid architecture named Local-Latent Swin Transformer (LLST). This architecture achieves computational complexity growth at half the rate of standard Swin transformer methods for growing sequence length and 4x receptive field, enabling effective processing of extended temporal contexts and enhanced model performance. Extensive experiments demonstrate that HLLDM achieves outstanding perceptual metrics and distortion metrics. Visualization results further validate that balancing perceptual and distortion metrics yields superior reconstruction quality.