<p>Video anomaly detection under weak supervision is a challenging task due to the overwhelming abundance of normal events compared to the few, often subtle, anomalies present in surveillance videos. Conventional methods formulate the problem as a multiple instance learning (MIL) task, yet they frequently bias the classifier toward easily recognizable, dominant normal instances and fail to capture nuanced abnormal patterns. In this study, a hybrid adaptive contrastive self-paced transformer architecture (HACSPT) is proposed for the video anomaly detection task. The proposed method integrates a local branch that captures fine-grained temporal features with a global branch that employs a transformer encoder to model long-range dependencies, and the resulting features are fused for anomaly classification. A hybrid loss function that combines binary cross-entropy with a contrastive loss based on margin ranking enforces that the average norm of self-paced MIL-selected abnormal features exceeds that of normal features by a fixed margin. Extensive experiments on benchmark datasets such as UCF-Crime and ShanghaiTech demonstrate the efficacy of the proposed architecture that produces ROC-AUC scores of 87.05% and 96.5% respectively. The proposed HACSPT is the first to combine self-paced snippet selection, margin-based contrastive learning, and a dual-branch local and global transformer in one framework.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HACSPT: a hybrid adaptive contrastive self-paced transformer for video anomaly detection

  • Siddharth Shah,
  • Dr. Narendrasinh Chauhan

摘要

Video anomaly detection under weak supervision is a challenging task due to the overwhelming abundance of normal events compared to the few, often subtle, anomalies present in surveillance videos. Conventional methods formulate the problem as a multiple instance learning (MIL) task, yet they frequently bias the classifier toward easily recognizable, dominant normal instances and fail to capture nuanced abnormal patterns. In this study, a hybrid adaptive contrastive self-paced transformer architecture (HACSPT) is proposed for the video anomaly detection task. The proposed method integrates a local branch that captures fine-grained temporal features with a global branch that employs a transformer encoder to model long-range dependencies, and the resulting features are fused for anomaly classification. A hybrid loss function that combines binary cross-entropy with a contrastive loss based on margin ranking enforces that the average norm of self-paced MIL-selected abnormal features exceeds that of normal features by a fixed margin. Extensive experiments on benchmark datasets such as UCF-Crime and ShanghaiTech demonstrate the efficacy of the proposed architecture that produces ROC-AUC scores of 87.05% and 96.5% respectively. The proposed HACSPT is the first to combine self-paced snippet selection, margin-based contrastive learning, and a dual-branch local and global transformer in one framework.