Enhancing Video Action Recognition with Spatial–Temporal Convolutional Attention Networks
摘要
Effectively harnessing feature correlations is crucial for optimal performance in video action recognition tasks, whether in spatial or temporal dimensions. Convolutional operations excel at capturing local features through correlations among nearby points, while self-attention mechanisms capture global information by enabling interactions among all feature points. However, the limitations of a single convolutional layer in comprehending holistic feature correlations and the tendency of self-attention layers to overlook local positional characteristics necessitate an innovative approach. We introduce the Spatial–Temporal Convolutional Attention Network, comprising the Spatial Convolutional Attention Network and the Temporal Convolutional Attention Network. This approach strategically combines the strengths of self-attention for global context and convolution for local features, mitigating the limitations of individual layers. Key innovations include [explicitly state innovations]. To reinforce the uniqueness of our approach, we provide specific comparisons with existing methods, highlighting its superiority. In rigorous experiments, our proposed network demonstrates enhanced recognition performance, quantified through metrics such as accuracy and F1 score improvements over baseline models. By effectively capturing feature correlations, our approach elevates the neural network's capacity for spatiotemporal modeling, marking a significant advancement in video action recognition. Elevating the neural network's capacity for spatiotemporal modeling.