PolyMotion-7K: A Multimodal-Driven Polyglot Avatar Motion Dataset
摘要
The potential of multimodal-driven multilingual avatar motion generation for cross-cultural communication and low-resource language processing is becoming increasingly prominent. However, the lack of high-quality datasets covering multiple languages and modalities restricts the adaptability of avatar motion generation to linguistic and cultural diversity. To address this limitation, we present PolyMotion-7K, a comprehensive multimodal-driven multilingual avatar motion dataset designed to enhance motion diversity and detail in avatars within multilingual contexts. It encompasses over 7,000 languages, each represented by more than five hours of upper-body speech video material, collected from diverse sources to ensure cross-cultural applicability. In constructing this dataset, we employ the SHOW model for skeletal parameter initialization, incorporated the MediaPipe module to optimize joint poses, used DeepLabV3 for foreground segmentation, and applied Dreamwaltz-G to enhance visual quality, achieving high fidelity in fine-grained motion and expression rendering. To validate the effectiveness of PolyMotion-7K, we conducted an application experiment in the multilingual dissemination of intangible cultural heritage along the Beijing Central Axis, producing multilingual avatar-based introductory videos that provide narratives about the cultural background of various heritage sites. Experimental results demonstrate that PolyMotion-7K supports high-quality multilingual avatar generation, highlighting its broad potential for cross-cultural communication and digital heritage preservation.