Language-driven models for video anomaly detection: taxonomy, datasets, and future horizons
摘要
Video Anomaly Detection (VAD) is a pivotal research area in computer vision, focusing on identifying events that deviate from normal behavior, with applications spanning public surveillance, healthcare, transport, and industry. Despite significant advancements in deep learning models such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM), traditional VAD methods face critical limitations, including neglect of temporal information, inadequate capture of long-term dependencies, and poor interpretability. The emergence of language-driven models, encompassing Large Language Models (LLMs), Vision Language Models (VLMs), and Vision-Large Language Models (V-LLMs), has addressed these challenges by leveraging cross-modal alignment, open-vocabulary detection, and semantic understanding. This paper presents a systematic study on language-driven VAD, first providing a comprehensive overview of LLMs, VLMs, and V-LLMs, including their foundational concepts, strengths, and limitations. We then propose a novel taxonomy categorizing 34 existing language-driven VAD methods based on learning paradigm, core architecture, adaptation strategy, and functional output. The taxonomy is evaluated using a structured quality attributes framework. Additionally, we discuss multi-modal datasets and real-time constraints, and conduct a comparative performance analysis of the reviewed methods. Finally, we identify future research prospects, including explainable open-vocabulary detection and tri-modal VAD integration. This work serves as a foundational resource for the VAD community, advancing the understanding and application of language-driven models in addressing complex VAD tasks.