In this paper, we propose STCA-Net (Spatio-Temporal Convolutional Attention Network), an efficient deep learning model for video action recognition. Accurately recognizing meaningful actions in video data requires effectively capturing both spatial and temporal features. However, traditional 3D CNN-based approaches often suffer from high computational complexity, and existing attention mechanisms have limitations in fully modeling the complex interactions between spatio-temporal features. The proposed STCA-Net builds upon the R(2+1)D network as its backbone while integrating a newly designed STCA module. The STCA module employs a decoupled approach to handle spatial and temporal attention separately. Specifically, it utilizes an averaging operation along both the temporal and spatial dimensions, enabling the use of lightweight convolutional operations. This method effectively captures two crucial aspects of video representation: ‘where’ (spatial focus) and ‘when’ (temporal focus). Experimental results on benchmark datasets UCF-101 and HMDB-51 demonstrate that STCA-Net achieves performance that is comparable to or superior to state-of-the-art models while maintaining computational efficiency.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

STCA-Net: Spatio-Temporal Convolutional Attention Network for Efficient Action Recognition in Videos

  • Jiha Kim,
  • Hyunhee Park

摘要

In this paper, we propose STCA-Net (Spatio-Temporal Convolutional Attention Network), an efficient deep learning model for video action recognition. Accurately recognizing meaningful actions in video data requires effectively capturing both spatial and temporal features. However, traditional 3D CNN-based approaches often suffer from high computational complexity, and existing attention mechanisms have limitations in fully modeling the complex interactions between spatio-temporal features. The proposed STCA-Net builds upon the R(2+1)D network as its backbone while integrating a newly designed STCA module. The STCA module employs a decoupled approach to handle spatial and temporal attention separately. Specifically, it utilizes an averaging operation along both the temporal and spatial dimensions, enabling the use of lightweight convolutional operations. This method effectively captures two crucial aspects of video representation: ‘where’ (spatial focus) and ‘when’ (temporal focus). Experimental results on benchmark datasets UCF-101 and HMDB-51 demonstrate that STCA-Net achieves performance that is comparable to or superior to state-of-the-art models while maintaining computational efficiency.