MSOAP: multi-scale spatial skeleton representations for online action prediction
摘要
With the growing adoption of skeleton-based action recognition in domains like smart monitoring and interactive human–machine systems, achieving low-latency prediction while maintaining high accuracy has become a critical challenge. This paper introduces a multi-scale online action prediction framework (MSOAP), which simultaneously models local joint dependencies and global motion patterns across spatial scales to continuously capture the dynamic evolution of skeleton sequences. To further enhance predictive performance, the framework incorporates a joint learning mechanism that optimizes action recognition together with future motion forecasting, while employing neural ordinary differential equations (neural ODE) to model the continuous temporal dynamics of skeleton sequences. These computations benefit from GPU-based parallel acceleration for low-latency online inference. This design not only improves the accuracy of predicting unobserved motion segments but also strengthens robustness for long-duration actions. Evaluations performed on standard skeleton datasets, such as NTU RGB+D 60, NTU RGB+D 120, and NW-UCLA, indicate that the presented approach achieves competitive performance compared with representative baseline methods, particularly under low-observation online prediction settings.