Lista-net: a lightweight spatiotemporal adaptive network for skeleton-based action recognition
摘要
In recent years, skeleton-based action recognition has demonstrated strong robustness in complex scenarios. However, existing methods face limitations in computational efficiency and feature modeling, particularly in capturing fine-grained spatiotemporal dynamics. To address these challenges, this paper proposes a lightweight spatiotemporal adaptive 3D network, LiSTA-Net, designed for efficient and accurate skeleton-based action recognition. LiSTA-Net adopts the X3D network as its backbone, enabling efficient modeling of spatiotemporal features in skeleton sequences, while incorporating three key modules to enhance its modeling capabilities: the Context-Aware Fusion Block, which dynamically integrates contextual information and filters features to improve multi-stage feature representation; the Lightweight Multi-scale 3D Convolution Block (L-M3D), which combines depthwise separable convolutions with multi-scale modeling strategies to efficiently capture local and global feature distributions; and the Lightweight Spatiotemporal Adaptive Module, embedded in the L-M3D block, which uses a decomposed spatiotemporal attention mechanism to dynamically weight key region features, focusing on critical spatiotemporal areas and precisely modeling action dynamics. This study conducted experiments on NTU RGB+D, NTU RGB+D 120, and FineGym datasets. Compared to existing convolutional and graph convolutional methods, LiSTA-Net achieves superior performance with lower complexity, demonstrating its potential for deployment in resource-constrained environments.