A Supervised Variational Auto-encoder for Human Motion Generation Using Convolutional Neural Networks
摘要
Human motion generation is an important research domain addressed by a significant amount of work in the recent years, with the availability of new datasets captured either from sensors or cameras. Most of the existing approaches are based on the Variational Auto-Encoder (VAE) architecture using Recurrent Neural Networks (RNN) or Transformers. In this paper, we propose to handle human motion sequences as Multivariate Time Series (MTS), and construct a VAE based on Convolutional Neural Networks (CNN). Furthermore, the proposed architecture uses an action classification task to add the conditioning aspect to the generative model. Our proposed Supervised VAE (SVAE) achieves competitive results on the HumanAct12 dataset, both in terms of quality and diversity of the generated sequences.