<p>Anomaly detection in human behavior monitoring is a challenging task which requires capturing contextual information. In this paper, we propose a novel anomaly detection method which leverages multimodal (image and caption) and multiview self-supervised learning objectives. Previous works successfully used deep captioning alongside images. However, their reliance on unimodal pre-trained image and text features revealed deficiencies in capturing contextual information across modalities. Our method learns high-quality multimodal feature representations and captures contextual information across modalities by combining contrastive objectives which exploit complementary and consistent information from different modalities and views. We evaluate our method on four real-world datasets for human monitoring anomaly detection. Our extensive experimental results demonstrate substantial improvements compared to the baseline methods. Specifically, our method achieved higher area under the receiver operating characteristic curve (AUC) scores, increasing from 0.967 to 0.99, 0.973 to 0.987, 0.885 to 0.94, and 0.671 to 0.713. Additionally, the area under the precision-recall curve (AUPRC) scores improved from 0.892 to 0.96, 0.90 to 0.905, 0.512 to 0.661, and 0.89 to 0.907.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting anomalies in human monitoring based on multimodal multiview self-supervised learning

  • Jose Alejandro Avellaneda Gonzalez,
  • Tetsu Matsukawa,
  • Einoshin Suzuki

摘要

Anomaly detection in human behavior monitoring is a challenging task which requires capturing contextual information. In this paper, we propose a novel anomaly detection method which leverages multimodal (image and caption) and multiview self-supervised learning objectives. Previous works successfully used deep captioning alongside images. However, their reliance on unimodal pre-trained image and text features revealed deficiencies in capturing contextual information across modalities. Our method learns high-quality multimodal feature representations and captures contextual information across modalities by combining contrastive objectives which exploit complementary and consistent information from different modalities and views. We evaluate our method on four real-world datasets for human monitoring anomaly detection. Our extensive experimental results demonstrate substantial improvements compared to the baseline methods. Specifically, our method achieved higher area under the receiver operating characteristic curve (AUC) scores, increasing from 0.967 to 0.99, 0.973 to 0.987, 0.885 to 0.94, and 0.671 to 0.713. Additionally, the area under the precision-recall curve (AUPRC) scores improved from 0.892 to 0.96, 0.90 to 0.905, 0.512 to 0.661, and 0.89 to 0.907.