QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation
Sicheng YangZhiyong WuMinglei LiZhensong ZhangLei HaoWeihong BaoHaolin Zhuang
Proposes a speech-driven gesture generation framework that combines a pose VQ-VAE, Levenshtein distance audio alignment, and motion phase guidance to eliminate random jitters and solve speech-gesture asynchrony.
Generating realistic 3D body gestures driven by human speech is a crucial capability for digital avatars, virtual assistants, and human-robot interaction. However, current automated methods struggle with two primary challenges: unnatural random jittering in generated movements and the inherent timing mismatch between speech rhythm and physical gesturing. Most existing end-to-end neural network models map audio directly to continuous motion, which often yields unnatural or frozen animations that perform poorly in standardized benchmarks.
The main objective of the article is to introduce and validate a novel speech-driven gesture generation framework, termed QPGesture, which combines discrete gesture quantization, sequence alignment, and motion phase guidance. The study evaluates how effectively this approach eliminates jitter and synchronizes realistic upper-body gestures with both speech audio and text semantics compared to established baseline methods.
To achieve this, the authors developed a system that first compresses complex human motion into a discrete codebook of distinct gesture units using an unsupervised vector quantized variational autoencoder, effectively filtering out minor random jitters. The framework then matches incoming speech with candidate gestures by using edit distance (Levenshtein distance) on quantized audio features and cosine similarity on text embeddings. Finally, a periodic autoencoder extracts the cyclic rhythm and phase of the motion to guide the smooth, natural selection and transition between audio-matched and text-matched gesture candidates. The model was trained and evaluated using four hours of upper-body motion capture data from two speakers in the large-scale BEAT dataset, using standard objective metrics alongside a blind perceptual user study with 23 participants.
The key findings demonstrate that this structured matching framework significantly outperforms previous approaches. Quantitatively, the method achieved a 39% to 44% reduction in gesture distortion metrics compared to the top baseline model, indicating far higher motion fidelity. In user evaluations, the proposed method scored higher in human-likeness (4.00 out of 5) and speech appropriateness (3.66 out of 5) than all competing neural models, matching or exceeding the perceptual scores of real ground-truth motion. Ablation experiments further confirmed that both audio alignment via edit distance and semantic text integration are essential for delivering coherent, context-appropriate movements.
These results show that combining discrete motion units with phase-guided matching provides a more reliable and natural solution for avatar animation than direct end-to-end continuous generation. Operationally, the framework is efficient, requiring less than one day of training on a single high-performance graphics processor, and offers direct controllability by allowing specific motion codes to be constrained or edited for custom character behavior.
For future development, the source recommends extending the framework to incorporate additional input modalities, such as facial expressions and emotional tone, to enrich communicative expressiveness. However, decision-makers should note that the current evaluation is bounded by upper-body joints excluding hands and fingers, and matching-based systems rely heavily on the diversity of the underlying gesture database. Overall, the evidence provides high confidence that quantization and phase-guided matching represent an effective and controllable architecture for high-fidelity speech-driven animation.
- Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This seminal paper introduces the Vector Quantized Variational AutoEncoder (VQ-VAE), which provides the fundamental discrete representation learning mechanism that QPGesture relies on to compress gestures into codebooks and prevent motion jitter.
- Paper: FaceFormer: Speech-Driven 3D Facial Animation with Transformers, Yingruo Fan et al. (2022). This paper establishes the foundational transformer-based audio-to-motion alignment architecture and periodic encoding strategies that inform cross-modal speech-driven animation systems like QPGesture.
- Paper: Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation, Lingting Zhu et al. (2023). Explores a complementary continuous diffusion-based framework with stabilization modules for audio-driven co-speech gesture generation, directly addressing the jitter and timing challenges tackled by QPGesture's discrete matching.
- Paper: Generating Holistic 3D Human Motion from Speech, Hongwei Yi et al. (2023). Extends speech-driven motion generation beyond upper-body gestures by synthesizing holistic 3D human motion, including separate modeling of face, body, and hand gestures.
- Paper: AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond, Zixiang Zhou et al. (2024). Builds on discrete motion tokenization concepts to create an all-in-one foundation model for multi-task motion understanding, planning, and long-sequence character generation.
