Collaborative SpatioTemporal Information Networks for Video Action Recognition
摘要
With the advancement of computer vision technology, video action recognition has emerged as a hot topic in the domain of artificial intelligence. This paper proposes a video action recognition method based on spatiotemporal collaboration information, which is used to recognize and classify actions in videos. Inspired by the relational network model [1], we use a combination of spatiotemporal network to extract video features, a combination of collaborative network to process time series data, and a simple and efficient attention mechanism to extract spatial information features to capture the dynamic characteristics of behavior. To demonstrate the model’s excellent recognition ability, we train and validate it using a large set of publicly available data, including Kinetics [2], UCF-101 [3], and HMDB-51 [4]. We improve the recognition accuracy in all three video-based behavior recognition tasks. At the same time, we also compare with other models, which proves that our model has high efficiency for feature extraction.