Motion recognition holds significant importance across various domains of application nowadays. Deep architectures represent the gold standard, with astonishing results, at the price, in general, of very high model complexity and poor interpretability. Drawing inspiration from the compositional nature of human motion, in this work we investigate the use of a modular architecture, that includes a VAE for the unsupervised learning of a compositional and minimal action representation, followed by a Transformer to classify the action sentence. We assess our approach on the BABEL dataset, comparing various positional and kinematic features in input. Our results demonstrate that despite the simplicity of the representation, our model provides a good trade-off between effectiveness, efficiency and interpretability. These insights pave the way for employing this methodology in diverse tasks, including motion generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MOSAIC: Skeleton-Based Human Motion Recognition with Compositional Representations

  • Federico Figari Tomenotti,
  • Nicoletta Noceti

摘要

Motion recognition holds significant importance across various domains of application nowadays. Deep architectures represent the gold standard, with astonishing results, at the price, in general, of very high model complexity and poor interpretability. Drawing inspiration from the compositional nature of human motion, in this work we investigate the use of a modular architecture, that includes a VAE for the unsupervised learning of a compositional and minimal action representation, followed by a Transformer to classify the action sentence. We assess our approach on the BABEL dataset, comparing various positional and kinematic features in input. Our results demonstrate that despite the simplicity of the representation, our model provides a good trade-off between effectiveness, efficiency and interpretability. These insights pave the way for employing this methodology in diverse tasks, including motion generation.