Video Crowd Activity Recognition Based on Stereo Vision
摘要
Extracting distinctive features from crowds has been a key focus in crowd activity recognition. While the use of image and monocular video data has been prominent in crowd activity recognition, the utilization of stereo vision presents significant advantages, particularly in the acquisition of three-dimensional information. However, extracting and representing crowd features from binocular videos present challenges. This study presents a novel approach to crowd activity recognition using deep learning, which captures temporal changes in individual behavior while preserving core variations. The method also accounts for the mutual influence of individual behaviors through information diffusion and accumulation. By integrating long short-term memory networks and graph convolutional neural networks, the proposed method extracts spatiotemporal features of crowds, leveraging binocular videos to capture individual behavior and associations. Experimental results demonstrate that the model achieves superior recognition performance by utilizing crowd features derived from binocular videos.