PHANet: Progressive Hybrid Attention Network for Enhanced Video Deraining
摘要
The task of video deraining is crucial and intricate in the field of computer vision due to the detrimental effects rain can have on video recordings, such as blurriness and occlusion. While numerous video deraining methods have shown remarkable accomplishments, two significant issues persistently endure and require attention: (1) How to efficiently and accurately extract spatio-temporal features from consecutive frames for background restoration, leveraging the spatial and temporal correlations within the video sequence, and (2) How to adaptively select motion patterns in the temporal domain to effectively handle complex rain movements. Regarding the challenges above, we propose a Progressive Hybrid Attention Network (PHANet), a novel approach consisting of a two-stage progressive network and multiple attention-based feature processing modules, providing an enhanced solution for video deraining by employing a task division strategy from coarse to fine and hybridizing various attention mechanisms. To be specific, we design a Spatio-Temporal Attention Module (STAM), a Supervised Fusion Attention Module (SFAM) and a Multi-Feature Refine Module (MFRM) to be applied to the network at different stages, which effectively improve the performance of the video deraining task with fewer computing resources. The experimental results on multiple public datasets demonstrate that our proposed PHANet exhibits significant improvements in both performance and speed compared to the previous state-of-the-art (SOTA) methods.