FastTalker: Co-Speech Gesture Generation via Fast-Order Diffusion ODE Solver
摘要
Existing research on co-speech gesture generation is predominantly grounded in high-quality gestures by virtue of diffusion models. Although diffusion-based gesture models with numerous sampling steps can achieve high accuracy, how to design a fast-generation method applied in resource-constrained scenarios still meets huge challenges. To mitigate the sampling burden, in this paper, we propose a novel FastTalker, a high-speed generation method for co-speech gestures that decreases sampling step size while keeping the accuracy. The proposed FastTalker consists of two key components, i.e., the gesture-aware latent diffusion module and fast-order diffusion ODE solver for guided sampling. Motivated by the consistency between speech and gesture, we design a transformer-based gesture-aware latent diffusion module to learn speech and speech text, which can efficiently generate the corresponding gestures based on speech. In order to achieve high-speed generation, we introduce an effective sampling mechanism by taking the sampling quality into account. Specifically, a fast-order diffusion ODE solver is employed for guided sampling enabling the gesture diffusion module with a large sampling step size as well as fewer steps. Experimental results indicate that our FastTalker accelerates the generation speed by a factor of 10 compared to the baseline, without compromising motion fidelity. Additionally, ablation studies also show the significance of fast sampling in high-speed gesture generation.