Multi-scale Differential Perception Network for Video Anomaly Detection
摘要
Video anomaly detection(VAD) aims to learn the normal appearance and motion patterns of video data and identify abnormal behaviours that deviate from these expected patterns. Existing methods employ easily fusion strategy to combine appearance and motion features. However, these methods overlook the characteristics of motion features in VAD, resulting in ineffective associations between appearance and motion features. In addition, most approaches employ coarse-grained modeling, which is inadequate for capturing intrinsic representations of normal behaviors. Therefore, we propose a multi-scale differential perception network(MSDPN). Firstly, we propose the differential perception fusion block(DPFB), which employs differential perception attention compute motion-salient regions through differential comparison and attention, enhancing the model’s sensitivity to critical motion cues. Secondly, we design a progressive multi-scale differential perception fusion strategy to establish connections between appearance and motion features across multiple scales, enriching the motion representation of the appearance branch. Finally, we design a memory network with feature refined fusion mechanism(MFRM), which removes redundant information through refined processing of deep features, enabling memory network to get more discriminative prototypes. Extensive experiments demonstrate that our MSDPN outperforms the state-of-the-art methods, achieving AUC improvements of at least 0.2%, 1.6%, and 1.0% on the Ped2, Avenue, and ShanghaiTech datasets, respectively.