The identification of anomalies in videos primarily focuses on detecting rare or inappropriate occurrences within specific contexts. The other state-of-the-art models often rely on either future frame prediction or reconstruction approaches. However, these methods exhibit limitations, such as inadequate detection due to the less variation between the generated normal and anomaly video frames. Thus, there is a need that combine the strengths of both future frame prediction and reconstruction models, the Integrated Motion-Appearance Generative Adversarial Networks (iMAppGAN). It is an end-to-end network that sequentially performs the prediction of future frames and further reconstructs the predicted video frame. The prediction of future frames facilitates the detection of abnormal events by emphasizing significant reconstruction errors. This is achieved through the integration of an autoencoder block and a UNET block within the generator to generate the normal video frame and distort the anomaly video frame. The first autoencoder block features a dual-stream encoder designed for extracting both motion and appearance features while the UNET block is responsible for reconstructing the predicted future frames received from the autoencoder. This results in a robust solution for effectively detecting anomalies in video surveillance, overcoming the drawbacks of both the prediction and reconstruction methods, and enhancing overall performance. Experimental results show that our novel iMAppGAN outperforms existing state-of-the-art models, demonstrating remarkable performance with an AUC score of 97.9% on UCSD Ped2, 90.8% on CUHK Avenue, and 75.3% on the ShanghaiTech dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

iMAppGAN: Integrated Motion Appearance Generative Adversarial Networks for Video Anomaly Detection

  • Rituraj Singh,
  • Anikeit Sethi,
  • Mitika Bhadada,
  • Hritika Gautam,
  • Krishanu Saini,
  • Aruna Tiwari,
  • Sumeet Saurav,
  • Sanjay Singh

摘要

The identification of anomalies in videos primarily focuses on detecting rare or inappropriate occurrences within specific contexts. The other state-of-the-art models often rely on either future frame prediction or reconstruction approaches. However, these methods exhibit limitations, such as inadequate detection due to the less variation between the generated normal and anomaly video frames. Thus, there is a need that combine the strengths of both future frame prediction and reconstruction models, the Integrated Motion-Appearance Generative Adversarial Networks (iMAppGAN). It is an end-to-end network that sequentially performs the prediction of future frames and further reconstructs the predicted video frame. The prediction of future frames facilitates the detection of abnormal events by emphasizing significant reconstruction errors. This is achieved through the integration of an autoencoder block and a UNET block within the generator to generate the normal video frame and distort the anomaly video frame. The first autoencoder block features a dual-stream encoder designed for extracting both motion and appearance features while the UNET block is responsible for reconstructing the predicted future frames received from the autoencoder. This results in a robust solution for effectively detecting anomalies in video surveillance, overcoming the drawbacks of both the prediction and reconstruction methods, and enhancing overall performance. Experimental results show that our novel iMAppGAN outperforms existing state-of-the-art models, demonstrating remarkable performance with an AUC score of 97.9% on UCSD Ped2, 90.8% on CUHK Avenue, and 75.3% on the ShanghaiTech dataset.