<p>This paper proposes an innovative end-to-end trainable framework called Foreground Dynamic Capture with Deep Feature Enhancement (FDC-Net) for video anomaly detection. Existing methods either exhibit distorted predictions for normal objects or overlook the more subtle dynamic differences between anomalous and normal objects in foreground. The framework of FDC-Net has a dual encoder–decoder architecture based on U-Net to extract features with high spatio-temporal dependencies from normal frames. A cross fusion module (CFM) is proposed during encoding. It uses cross self-attention and self-attention mechanisms to learn shared features between appearance and motion to obtain the high regularity of normal events while using spatio-temporal attention to constrain these events. Anomalous events tend to generate larger errors due to their low regularity and lack of adherence to normal spatio-temporal constraints. Furthermore, we propose a hybrid spatio-temporal pyramid module (HSPM), which enables the model to utilize dynamic motion and appearance information from optical flow to help the network dynamically capture subtle differences between normal and anomalous objects in the foreground across multiple scales, focusing on spatial texture, abstract semantics, and motion intensity. We provide multi-supervision and stop-gradient strategy that enable FDC-Net to further exploit the spatio-temporal distribution of normal events, ensuring undistorted predictions of normal objects. Extensive experiments demonstrate that the proposed model achieves a good balance between accurately predicting normal events and the prevention of anomaly generalization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FDC-Net: foreground dynamic capture with deep feature enhancement for video anomaly detection

  • Ruinian Shi,
  • Qiang He,
  • Hengyou Wang,
  • Changlun Zhang

摘要

This paper proposes an innovative end-to-end trainable framework called Foreground Dynamic Capture with Deep Feature Enhancement (FDC-Net) for video anomaly detection. Existing methods either exhibit distorted predictions for normal objects or overlook the more subtle dynamic differences between anomalous and normal objects in foreground. The framework of FDC-Net has a dual encoder–decoder architecture based on U-Net to extract features with high spatio-temporal dependencies from normal frames. A cross fusion module (CFM) is proposed during encoding. It uses cross self-attention and self-attention mechanisms to learn shared features between appearance and motion to obtain the high regularity of normal events while using spatio-temporal attention to constrain these events. Anomalous events tend to generate larger errors due to their low regularity and lack of adherence to normal spatio-temporal constraints. Furthermore, we propose a hybrid spatio-temporal pyramid module (HSPM), which enables the model to utilize dynamic motion and appearance information from optical flow to help the network dynamically capture subtle differences between normal and anomalous objects in the foreground across multiple scales, focusing on spatial texture, abstract semantics, and motion intensity. We provide multi-supervision and stop-gradient strategy that enable FDC-Net to further exploit the spatio-temporal distribution of normal events, ensuring undistorted predictions of normal objects. Extensive experiments demonstrate that the proposed model achieves a good balance between accurately predicting normal events and the prevention of anomaly generalization.