Weakly Supervised Video Anomaly Detection with Lightweight Knowledge Token Fusion and Memory Optimization Strategy
摘要
While multimodal fusion has shown promise in video anomaly detection, reliance on audio inputs remains impractical due to hardware limitations and privacy concerns. This paper focuses exclusively on the visual domain, using RGB and optical flow to achieve efficient and accurate detection. We propose a lightweight transformer-based memory mechanism that enhances feature fusion through a knowledge marker weighting scheme. To improve generalization, we introduce a Fusion Feature Enhancement Module (FFEM), which employs random masking, multi-scale convolutions, and attention-based fusion to suppress overfitting and enrich spatial representations. Additionally, we develop Minimum Recently Used Memory (M-LRU), a memory optimization strategy that enables efficient long-duration anomaly detection. Experiments on UCF-Crime and XD-Violence demonstrate the effectiveness and strong cross-dataset generalization of our method, all within a compact model design.