Hierarchical Attention for Violence Detection
摘要
Recognizing actions through vision has become a prominent topic of research, posing fresh challenges in identifying and analyzing aggressive conduct in surveillance footage. This paper proposes a hierarchical attention-based CNN-ConvLSTM model for detecting fights and aggressive behavior in surveillance videos. The model leverages CNNs for spatial features, ConvLSTMs for temporal dependencies, and attention mechanisms to enhance long-range action recognition. An MLP with fully connected layers performs final classification, preceded by pre-processing steps to standardize data and improve accuracy. The experimental results show the interesting performance through three benchmark databases named Hockey-fight, Movies and Surveillance cameras.