<p>Generating satisfying dance movements based on music remains an extremely difficult challenge due to the complex relationship between these two modalities. Conventional end-to-end training mechanism relies on accurate dancing decoders with GAN or VAE mechanisms, breaking the implicit correlations between music and dancing. Different from prevailing methods, in this paper, we make attempts to explore two key clues (content and rhythm) with contrastive pretraining to model this multi-modal relationship. In our approach, we first disentangle the style and content features of dual modalities to prevent one-to-one generation mapping on specific styles. To extract the clue of rhythms, we learn from music beats and dance motions and then adopt the content-rhythm correlations as bridges to multi-modal contrastive learning. Although zero-shot predictions achieve superior multi-modal correlations in our method, the predicted dance motions are still fragments. To solve this, we propose a music-aware motion harmonization module to dynamically fuse multiple fragments corresponding to their body motions. With the collaboration of the multi-modal contrastive learning framework and the music-aware motion harmonization module, extensive experiments show our proposed approach outperforms the state-of-the-art dancing generation methods in both subjective and objective studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Content-Rhythm Awareness Contrastive Learning for Music-Driven 3D Dance Generation

  • Xin Guo,
  • Yifan Zhao,
  • Jia Li

摘要

Generating satisfying dance movements based on music remains an extremely difficult challenge due to the complex relationship between these two modalities. Conventional end-to-end training mechanism relies on accurate dancing decoders with GAN or VAE mechanisms, breaking the implicit correlations between music and dancing. Different from prevailing methods, in this paper, we make attempts to explore two key clues (content and rhythm) with contrastive pretraining to model this multi-modal relationship. In our approach, we first disentangle the style and content features of dual modalities to prevent one-to-one generation mapping on specific styles. To extract the clue of rhythms, we learn from music beats and dance motions and then adopt the content-rhythm correlations as bridges to multi-modal contrastive learning. Although zero-shot predictions achieve superior multi-modal correlations in our method, the predicted dance motions are still fragments. To solve this, we propose a music-aware motion harmonization module to dynamically fuse multiple fragments corresponding to their body motions. With the collaboration of the multi-modal contrastive learning framework and the music-aware motion harmonization module, extensive experiments show our proposed approach outperforms the state-of-the-art dancing generation methods in both subjective and objective studies.