Synthetic Video Generation for Weakly Supervised Cross-Domain Video Anomaly Detection
摘要
Video anomaly detection (VAD) plays a pivotal role in crucial applications such as security and surveillance, garnering significant interest from the research community. The utility of cross-domain VAD is critical in practical scenarios, yet most of research remains focused on same-domain VAD. Weakly supervised approaches excel in same-domain contexts but are rarely applied to cross-domain VAD, which typically relies on unsupervised methods. This paper presents a new weakly supervised framework for addressing cross-domain VAD challenges, aiming to improve model generalization across different domains. A key issue is the model’s propensity for overfitting to source domain anomalies, impairing its ability to detect out-of-distribution anomalies. Our approach introduces a video synthesis technique using generative technologies for zero-shot cross-domain VAD. This strategy combats the generative technologies’ limitations, especially their struggle to generate human behavior and object motion accurately—vital aspects of VAD. By merging generative video editing with object synthesizing, we ensure that synthesized videos maintain their original normal or abnormal status. Combining synthesized with original data, our model is trained in a weakly supervised manner. The experimental results demonstrate that our method outperforms existing works for cross-domain scenarios.