Dual-path spatio-temporal Mamba for skeleton-based action recognition
摘要
Mamba emerges as a new paradigm for modeling long sequences, presenting a compelling alternative to the Transformer. Although Mamba has demonstrated its utility in modeling temporal structures within skeleton-based action recognition, its spatio-temporal union context modeling has not been fully explored. To bridge this gap, we propose SkeMamba, a novel approach centered entirely on Mamba. Firstly, we introduce adaptive topology transformation (ATT) to convert skeletal graphs into sequential representations while preserving anatomical semantics. Secondly, we formulate dual-path spatio-temporal mamba (DSTMba) to thoroughly account for the dual spatio-temporal nature of joint features. Lastly, we propose the spatio-temporal gated mamba (ST-GateMba) for the pivotal components of DSTMba, which strategically mediates between transient motion dynamics and sustained behavioral semantics. Experimental results demonstrate that SkeMamba outperforms existing Mamba-based methods on NTU series and NW-UCLA datasets.