<p>Video Anomaly Detection (VAD) is a pivotal research area in computer vision, focusing on identifying events that deviate from normal behavior, with applications spanning public surveillance, healthcare, transport, and industry. Despite significant advancements in deep learning models such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM), traditional VAD methods face critical limitations, including neglect of temporal information, inadequate capture of long-term dependencies, and poor interpretability. The emergence of language-driven models, encompassing Large Language Models (LLMs), Vision Language Models (VLMs), and Vision-Large Language Models (V-LLMs), has addressed these challenges by leveraging cross-modal alignment, open-vocabulary detection, and semantic understanding. This paper presents a systematic study on language-driven VAD, first providing a comprehensive overview of LLMs, VLMs, and V-LLMs, including their foundational concepts, strengths, and limitations. We then propose a novel taxonomy categorizing 34 existing language-driven VAD methods based on learning paradigm, core architecture, adaptation strategy, and functional output. The taxonomy is evaluated using a structured quality attributes framework. Additionally, we discuss multi-modal datasets and real-time constraints, and conduct a comparative performance analysis of the reviewed methods. Finally, we identify future research prospects, including explainable open-vocabulary detection and tri-modal VAD integration. This work serves as a foundational resource for the VAD community, advancing the understanding and application of language-driven models in addressing complex VAD tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Language-driven models for video anomaly detection: taxonomy, datasets, and future horizons

  • Olfa Saket,
  • Anis Ben Aicha,
  • Habib Fathallah

摘要

Video Anomaly Detection (VAD) is a pivotal research area in computer vision, focusing on identifying events that deviate from normal behavior, with applications spanning public surveillance, healthcare, transport, and industry. Despite significant advancements in deep learning models such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM), traditional VAD methods face critical limitations, including neglect of temporal information, inadequate capture of long-term dependencies, and poor interpretability. The emergence of language-driven models, encompassing Large Language Models (LLMs), Vision Language Models (VLMs), and Vision-Large Language Models (V-LLMs), has addressed these challenges by leveraging cross-modal alignment, open-vocabulary detection, and semantic understanding. This paper presents a systematic study on language-driven VAD, first providing a comprehensive overview of LLMs, VLMs, and V-LLMs, including their foundational concepts, strengths, and limitations. We then propose a novel taxonomy categorizing 34 existing language-driven VAD methods based on learning paradigm, core architecture, adaptation strategy, and functional output. The taxonomy is evaluated using a structured quality attributes framework. Additionally, we discuss multi-modal datasets and real-time constraints, and conduct a comparative performance analysis of the reviewed methods. Finally, we identify future research prospects, including explainable open-vocabulary detection and tri-modal VAD integration. This work serves as a foundational resource for the VAD community, advancing the understanding and application of language-driven models in addressing complex VAD tasks.