Spatio-Temporal Domain-Aware Network for Skeleton-Based Action Representation Learning
摘要
Considering its capability to extract implicit patterns from unlabeled data, contrastive learning has been widely employed in unsupervised skeleton-based action recognition. Spatio-temporal modeling is a key component in understanding skeleton sequence. However, existing methods often adopt rudimentary mechanisms or even completely overlook this aspect, leading to suboptimal performance in downstream tasks. In this paper, we propose a Spatio-Temporal Domain-Aware Network (STDA-Net). Firstly, the features extracted from the backbone extractor are further decoded into a triple-stream representation, corresponding to the spatial, temporal and global domains, respectively. Following that, an innovative approach named Triple Attention Transformer Module (TATM) is proposed to achieve customized spatio-temporal reasoning. TATM consists of three independent attention modules and a shared feedforward layer, thus achieving reasoning in different domains in a more efficient manner. Finally, domain-aware projectors are used to obtain richer spatio-temporal representations, providing a basis for the subsequent construction of inter-domain and intra-domain contrasts. Comprehensive experiments on NTU-RGB+D 60 &120 and PKU-MMD datasets demonstrate the superior performance of STDA-Net.