<p>Detecting anomalous behavior in surveillance video remains one of the most challenging tasks in current computer vision research. Generative auto-encoders (AE), especially their convolutional counterparts, 2D Convolutional Auto-encoders (2D CAEs), have recently achieved spectacular performance in most recognition tasks. This is due to their ability to learn robust 2D features during the learning process. With the success of deep neural networks, 3D Convolutional Neural Network (3D CNN)-based approaches have been proposed to capture both temporal and contextual patterns in successive video frames. However, this type of network is very costly in terms of computational time and memory. To address this issue, in this paper, we propose a hybrid 2D-CAE-3D-CNN model for anomaly event detection in videos. Specifically, the model leverages the discrimination and reconstruction power of 2D-3D CNNs and auto-encoders to effectively learn complex features from the input video patches, called cuboids. Experiments on two benchmark datasets (UCSD Pedestrian and CUHK Avenue) show that the proposed method achieves promising results compared to recent state-of-the-art approaches.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid 2D CAE-3D CNN-Based Video Anomaly Detection from Spatio-Temporal Cuboids

  • Khaled Bayoudh,
  • Fayçal Hamdaoui,
  • Abdellatif Mtibaa

摘要

Detecting anomalous behavior in surveillance video remains one of the most challenging tasks in current computer vision research. Generative auto-encoders (AE), especially their convolutional counterparts, 2D Convolutional Auto-encoders (2D CAEs), have recently achieved spectacular performance in most recognition tasks. This is due to their ability to learn robust 2D features during the learning process. With the success of deep neural networks, 3D Convolutional Neural Network (3D CNN)-based approaches have been proposed to capture both temporal and contextual patterns in successive video frames. However, this type of network is very costly in terms of computational time and memory. To address this issue, in this paper, we propose a hybrid 2D-CAE-3D-CNN model for anomaly event detection in videos. Specifically, the model leverages the discrimination and reconstruction power of 2D-3D CNNs and auto-encoders to effectively learn complex features from the input video patches, called cuboids. Experiments on two benchmark datasets (UCSD Pedestrian and CUHK Avenue) show that the proposed method achieves promising results compared to recent state-of-the-art approaches.