FDC-Net: foreground dynamic capture with deep feature enhancement for video anomaly detection
摘要
This paper proposes an innovative end-to-end trainable framework called Foreground Dynamic Capture with Deep Feature Enhancement (FDC-Net) for video anomaly detection. Existing methods either exhibit distorted predictions for normal objects or overlook the more subtle dynamic differences between anomalous and normal objects in foreground. The framework of FDC-Net has a dual encoder–decoder architecture based on U-Net to extract features with high spatio-temporal dependencies from normal frames. A cross fusion module (CFM) is proposed during encoding. It uses cross self-attention and self-attention mechanisms to learn shared features between appearance and motion to obtain the high regularity of normal events while using spatio-temporal attention to constrain these events. Anomalous events tend to generate larger errors due to their low regularity and lack of adherence to normal spatio-temporal constraints. Furthermore, we propose a hybrid spatio-temporal pyramid module (HSPM), which enables the model to utilize dynamic motion and appearance information from optical flow to help the network dynamically capture subtle differences between normal and anomalous objects in the foreground across multiple scales, focusing on spatial texture, abstract semantics, and motion intensity. We provide multi-supervision and stop-gradient strategy that enable FDC-Net to further exploit the spatio-temporal distribution of normal events, ensuring undistorted predictions of normal objects. Extensive experiments demonstrate that the proposed model achieves a good balance between accurately predicting normal events and the prevention of anomaly generalization.