Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation
Lingting ZhuXian LiuXuanyu LiuRui QianZiwei LiuLequan Yu
Proposes DiffGesture, a diffusion-based framework utilizing a multi-modal Transformer and an annealed noise sampling stabilizer to generate high-fidelity, temporally coherent co-speech gestures aligned with audio.
Generating realistic co-speech gestures for virtual avatars is essential for natural human-machine interaction, digital assistants, and embodied artificial intelligence. Historically, automated gesture generation relied on generative adversarial networks (GANs). However, these conventional frameworks regularly suffer from mode collapse and training instability, resulting in repetitive, rigid, and unnaturally synchronized avatar movements.
The article demonstrates a novel diffusion-based framework, named Diffusion Co-Speech Gesture (DiffGesture), designed to produce high-fidelity, diverse, and temporally coherent upper-body gesture sequences directly from continuous speech audio.
The researchers established an end-to-end conditional diffusion process that converts speech audio and initial reference poses into smooth skeletal movement sequences. To capture long-term temporal dependencies across multiple input modalities, the authors built an attention-based Transformer architecture. To prevent jitter and frame-to-frame inconsistencies typical of standard diffusion sampling, they introduced an annealed noise sampling module called the Diffusion Gesture Stabilizer alongside an implicit guidance mechanism that balances motion diversity against overall sample quality. The framework was evaluated across two benchmark datasets: TED Gesture (focusing on 10 upper-body joints) and TED Expressive (capturing 43 body and detailed finger joints).
The evaluation produced four key findings. First, DiffGesture substantially improved gesture quality over previous state-of-the-art methods, lowering the Fréchet Gesture Distance (FGD)—a key metric measuring similarity to real human movement distribution—from 3.072 to 1.506 on the standard TED dataset and from 5.306 to 2.600 on the expressive dataset, representing an improvement of roughly 50%. Second, DiffGesture achieved higher motion-audio synchrony, scoring 0.718 in beat consistency compared to 0.641 for previous leading baselines on expressive data. Third, diversity scores reached 182.757 on the expressive dataset, outperforming competing models and avoiding the static, repetitive failure modes typical of GANs. Fourth, a user study with 18 participants confirmed that the proposed framework was rated higher in naturalness (4.00 out of 5), smoothness (3.89), and speech-gesture synchrony (3.89) than existing automated methods.
These findings indicate that diffusion architectures can successfully replace adversarial frameworks for complex temporal generation tasks, significantly improving animation fidelity while maintaining computational efficiency. The framework avoids heavy GPU memory overheads found in prior hierarchical models, requiring only 10 to 20 hours of training on a single standard graphical processor.
Organizations developing digital avatars, virtual assistants, or interactive robotics should consider transitioning from GAN-based motion synthesis to diffusion-based pipelines using attention backbones. Further work should focus on piloting the system in real-time interactive environments, expanding beyond upper-body and finger tracking to full-body dynamics, and testing cross-language generalization.
The findings are supported by consistent quantitative metrics, ablation studies, and qualitative user evaluations across standard benchmarks. However, confidence should be tempered by certain boundary conditions: the pipeline relies on pre-extracted 2D and 3D pose estimates derived from monologue video recordings, which may not capture the multi-party turn-taking dynamics and environmental physical constraints present in complex real-world deployments.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). This foundational work establishes transformer-based diffusion models directly on human skeletal motion sequences, providing the core motion-diffusion principles adapted by DiffGesture.
- Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). This paper introduces classifier-free guidance for conditional diffusion models, which forms the basis of the implicit guidance mechanism used in DiffGesture to balance motion quality and diversity.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This paper establishes the foundational mathematics and denoising training formulation of Denoising Diffusion Probabilistic Models (DDPM) upon which subsequent conditional diffusion architectures build.
- Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). This work demonstrates how diffusion architectures surpass generative adversarial networks in quality and mode coverage, motivating the shift from GANs to diffusion models for gesture synthesis.
- Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). This paper provides foundational techniques for cross-modal conditional diffusion and non-autoregressive sequence modeling from audio signals.
- Paper: Diffusion Models: A Comprehensive Survey of Methods and Applications, Ling Yang et al. (2022). This comprehensive survey provides essential background on theoretical formulations, conditioning strategies, and sampling optimizations across the diffusion modeling landscape.
- Paper: Seamless Human Motion Composition with Blended Positional Encodings, Germán Barquero et al. (2024). FlowMDM extends diffusion-based skeletal motion generation to long, multi-prompt action sequences using blended positional encodings to eliminate transition artifacts.
- Paper: Executing your Commands via Motion Diffusion in Latent Space, Xin Chen et al. (2023). This work advances motion diffusion efficiency and fidelity by executing conditional denoising within a compact VAE latent space rather than directly on raw joint positions.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). Real3D-Portrait extends audio-driven digital avatar generation by synthesizing 3D talking portraits that integrate speech-driven facial parameters with torso movement.
