<p>Skeleton-based Hand Gesture Production (SHGP) is a challenging task due to the complex patterns in different hand gestures. Directly modeling the temporal transitions of different hand gestures from the original gesture space to the high dimensional space is a complex task. To tackle this challenge, in this paper, we focus on the SHGP, aiming to propose a deep generative model with smooth and diverse transitions on the latent space of hand gesture sequences in a lower dimensionality. To this end, a Bidirectional Generative Adversarial Network (BiGAN) is proposed with the embedded Long Short-Term Model (LSTM) plus a frame-wise decoder in the Generator Network. The Encoder aims to conditionally generating the hand gesture poses using the sequence of independent noise plus a one-hot class vector. Using a Discriminator containing a Bidirectional LSTM, the Generator and Discriminator Networks are trained in the proposed BiGAN framework, leading to accurately generate and also classify the hand gesture sequences. Results on four datasets show the promising results relative to the existing methods in SHGP and Hand Action Production (HAP).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A deep generative Skeleton-based dynamic hand gesture production model

  • Razieh Rastgoo,
  • Kourosh Kiani,
  • Sergio Escalera

摘要

Skeleton-based Hand Gesture Production (SHGP) is a challenging task due to the complex patterns in different hand gestures. Directly modeling the temporal transitions of different hand gestures from the original gesture space to the high dimensional space is a complex task. To tackle this challenge, in this paper, we focus on the SHGP, aiming to propose a deep generative model with smooth and diverse transitions on the latent space of hand gesture sequences in a lower dimensionality. To this end, a Bidirectional Generative Adversarial Network (BiGAN) is proposed with the embedded Long Short-Term Model (LSTM) plus a frame-wise decoder in the Generator Network. The Encoder aims to conditionally generating the hand gesture poses using the sequence of independent noise plus a one-hot class vector. Using a Discriminator containing a Bidirectional LSTM, the Generator and Discriminator Networks are trained in the proposed BiGAN framework, leading to accurately generate and also classify the hand gesture sequences. Results on four datasets show the promising results relative to the existing methods in SHGP and Hand Action Production (HAP).