QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation

Sicheng YangZhiyong WuMinglei LiZhensong ZhangLei HaoWeihong BaoHaolin Zhuang

article2023CVPR64 citations

Proposes a speech-driven gesture generation framework that combines a pose VQ-VAE, Levenshtein distance audio alignment, and motion phase guidance to eliminate random jitters and solve speech-gesture asynchrony.

Listen

Generating realistic 3D body gestures driven by human speech is a crucial capability for digital avatars, virtual assistants, and human-robot interaction. However, current automated methods struggle with two primary challenges: unnatural random jittering in generated movements and the inherent timing mismatch between speech rhythm and physical gesturing. Most existing end-to-end neural network models map audio directly to continuous motion, which often yields unnatural or frozen animations that perform poorly in standardized benchmarks.

The main objective of the article is to introduce and validate a novel speech-driven gesture generation framework, termed QPGesture, which combines discrete gesture quantization, sequence alignment, and motion phase guidance. The study evaluates how effectively this approach eliminates jitter and synchronizes realistic upper-body gestures with both speech audio and text semantics compared to established baseline methods.

To achieve this, the authors developed a system that first compresses complex human motion into a discrete codebook of distinct gesture units using an unsupervised vector quantized variational autoencoder, effectively filtering out minor random jitters. The framework then matches incoming speech with candidate gestures by using edit distance (Levenshtein distance) on quantized audio features and cosine similarity on text embeddings. Finally, a periodic autoencoder extracts the cyclic rhythm and phase of the motion to guide the smooth, natural selection and transition between audio-matched and text-matched gesture candidates. The model was trained and evaluated using four hours of upper-body motion capture data from two speakers in the large-scale BEAT dataset, using standard objective metrics alongside a blind perceptual user study with 23 participants.

The key findings demonstrate that this structured matching framework significantly outperforms previous approaches. Quantitatively, the method achieved a 39% to 44% reduction in gesture distortion metrics compared to the top baseline model, indicating far higher motion fidelity. In user evaluations, the proposed method scored higher in human-likeness (4.00 out of 5) and speech appropriateness (3.66 out of 5) than all competing neural models, matching or exceeding the perceptual scores of real ground-truth motion. Ablation experiments further confirmed that both audio alignment via edit distance and semantic text integration are essential for delivering coherent, context-appropriate movements.

These results show that combining discrete motion units with phase-guided matching provides a more reliable and natural solution for avatar animation than direct end-to-end continuous generation. Operationally, the framework is efficient, requiring less than one day of training on a single high-performance graphics processor, and offers direct controllability by allowing specific motion codes to be constrained or edited for custom character behavior.

For future development, the source recommends extending the framework to incorporate additional input modalities, such as facial expressions and emotional tone, to enrich communicative expressiveness. However, decision-makers should note that the current evaluation is bounded by upper-body joints excluding hands and fingers, and matching-based systems rely heavily on the diversity of the underlying gesture database. Overall, the evidence provides high confidence that quantization and phase-guided matching represent an effective and controllable architecture for high-fidelity speech-driven animation.

arXiv: 2305.11094
  • Paper: Neural Discrete Representation Learning, Aäron van den Oord et al. (2017). This seminal paper introduces the Vector Quantized Variational AutoEncoder (VQ-VAE), which provides the fundamental discrete representation learning mechanism that QPGesture relies on to compress gestures into codebooks and prevent motion jitter.
  • Paper: FaceFormer: Speech-Driven 3D Facial Animation with Transformers, Yingruo Fan et al. (2022). This paper establishes the foundational transformer-based audio-to-motion alignment architecture and periodic encoding strategies that inform cross-modal speech-driven animation systems like QPGesture.
Cover for QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation

Abstract

Speech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a novel quantization-based and phase-guided motion matching framework. Specifically, we first present a gesture VQ-VAE module to learn a codebook to summarize meaningful gesture units. With each code representing a unique gesture, random jittering problems are alleviated effectively. We then use Levenshtein distance to align diverse gestures with different speech. Levenshtein distance based on audio quantization as a similarity metric of corresponding speech of gestures helps match more appropriate gestures with speech, and solves the alignment problem of speech and gestures well. Moreover, we introduce phase to guide the optimal gesture matching based on the semantics of context or rhythm of audio. Phase guides when text-based or speech-based gestures should be performed to make the generated gestures more natural. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, database, pre-trained models and demos are available at https://github.com/YoungSeng/QPGesture.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Approach
  • 3.1. Learning a discrete latent space representation
  • 3.2. Motion Matching based on Audio and Text
  • 3.3. Phase-Guided Gesture Generation
  • 4. Experiments
  • 4.1. Comparison to Existing Methods
  • 4.2. Ablation Studies
  • 4.3. Controllability
  • 5. Discussion and Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Quantization- and phase-guided motion-matching framework

    model/method

    QPGesture generates a gesture-motion sequence from speech audio, its corresponding text, an initial seed pose, and optional control masks. The method first converts the input motion, audio, text, and a speech–gesture database into discrete or compact representations. It then retrieves one audio-matched gesture candidate using quantized-audio similarity and one text-matched candidate using semantic text similarity. Finally, a learned motion-phase representation compares the candidates with the seed pose and selects the candidate whose phase transition is most compatible with the ongoing motion. Discrete gesture units reduce small, irregular motion jitter; audio sequence alignment addresses the asynchronous timing of speech and gestures; and phase matching improves continuity at candidate boundaries.

  2. Knowl 2 — Gesture VQ-VAE with velocity- and acceleration-aware reconstruction

    equation

    For a normalized gesture sequence G∈RT×Dg\mathbf{G}\in\mathbb{R}^{T\times D_g}, where TT is the number of pose samples and DgD_g is the pose dimension, a temporal encoder EgE_g produces g=Eg(G)∈RT′×C\mathbf{g}=E_g(\mathbf{G})\in\mathbb{R}^{T'\times C}. Each feature gi\mathbf{g}_i is replaced by its nearest vector in a learned codebook Zg={zj}j=1Cb\mathcal{Z}_g=\{\mathbf{z}_j\}_{j=1}^{C_b}:

    gq,i=q(gi)=arg⁡min⁡zj∈Zg∥gi−zj∥2.\mathbf{g}_{q,i}=\mathbf{q}(\mathbf{g}_i)=\arg\min_{\mathbf{z}_j\in\mathcal{Z}_g}\|\mathbf{g}_i-\mathbf{z}_j\|_2.

    A decoder DgD_g reconstructs the gesture as G^=Dg(gq)\hat{\mathbf{G}}=D_g(\mathbf{g}_q). The encoder, decoder, and codebook are trained with

    Lgesture=Lrec(G^,G)+∥sg⁡[g]−gq∥2+β∥g−sg⁡[gq]∥2,\mathcal{L}_{\mathrm{gesture}}=\mathcal{L}_{\mathrm{rec}}(\hat{\mathbf{G}},\mathbf{G})+\|\operatorname{sg}[\mathbf{g}]-\mathbf{g}_q\|_2+\beta\|\mathbf{g}-\operatorname{sg}[\mathbf{g}_q]\|_2,

    where sg⁡\operatorname{sg} stops gradients, β\beta weights the commitment term, and CbC_b is the number of codebook entries. The reconstruction loss is

    Lrec=∥G^−G∥1+α1∥G^′−G′∥1+α2∥G^′′−G′′∥1,\mathcal{L}_{\mathrm{rec}}=\|\hat{\mathbf{G}}-\mathbf{G}\|_1+\alpha_1\|\hat{\mathbf{G}}'-\mathbf{G}'\|_1+\alpha_2\|\hat{\mathbf{G}}''-\mathbf{G}''\|_1,

    where primes denote temporal velocity and acceleration, respectively. The paper uses root-joint centering, common facing direction, and per-joint zero-mean/unit-variance normalization before encoding. The velocity and acceleration terms explicitly discourage jitter in reconstructed gesture sequences.

  3. Knowl 3 — Discrete audio representation for speech–gesture alignment

    model/method

    Audio is quantized with a pretrained vq-wav2vec Gumbel-Softmax model trained on a clean 100-hour LibriSpeech subset. The model produces two discrete token groups; the two group values are combined into one token for each audio segment, yielding a discrete audio sequence. The audio encoder maps each time step to a representation that is replaced by its nearest entry in an audio codebook Za\mathcal{Z}_a. QPGesture uses these discrete audio sequences rather than continuous acoustic features when measuring correspondence between speech and database gesture clips, enabling insertion/deletion-tolerant Levenshtein matching.

  4. Knowl 4 — Speech- and text-conditioned motion-matching algorithm

    algorithm

    The motion database is constructed by splitting recorded gesture motion at text-transcription word intervals longer than 0.50.5 seconds. Each database item therefore contains a gesture clip, its quantized audio, and its associated text. The matching procedure preserves continuity by considering the previously selected gesture code as well as speech or text similarity.

    Input: audio, text, seed pose, speech–gesture database, optional control masks
    Output: a sequence of selected gesture clips
    Quantize the audio with vq-wav2vec and form the audio feature sequence.
    Encode text contexts with Sentence-BERT and form the text feature sequence.
    Quantize the seed pose with the gesture VQ-VAE and set the previous pose code.
    For each inference speech segment:
        Rank all gesture codes by Euclidean distance from the previous pose code.
        Build an audio context from a centered 0.5-second search window.
        For each database clip, compare its corresponding quantized-audio sequence
            with the current audio context using Levenshtein distance.
        For each gesture code, retain the minimum audio distance found at the
            corresponding temporal positions to form audio pre-candidates.
        Build a text context from the text before and after the current time.
        For each database clip, compare its Sentence-BERT text embedding with the
            current text embedding using cosine similarity.
        For each gesture code, retain the best corresponding text match to form
            text pre-candidates.
        Combine the pose rank with the audio rank to select the audio candidate.
        Combine the pose rank with the text rank to select the text candidate.
        Compare the phase-boundary similarity of the two candidates with the seed.
        Select the candidate with the more compatible phase transition.
        Append the selected gesture clip and update the previous pose code.
    Return the concatenated selected gesture clips.

    The pose, audio, and text terms are combined by adding their ranks rather than by weighting raw distances. In the reported implementation, audio and text contexts cover four pose codes before and after the current position, with a temporal stride of d=32d=32 frames. The initial pose code may be sampled from the gesture codebook or set to the most frequent code; the experiments use random sampling.

  5. Knowl 5 — Learned periodic phase manifold for candidate selection

    model/method

    QPGesture learns a periodic autoencoder that maps motion into a phase-aware latent manifold. A temporal encoder produces L=Ep(G)∈Rn×M\mathbf{L}=E_p(\mathbf{G})\in\mathbb{R}^{n\times M}, where nn is the number of samples in a motion window and MM is the number of phase channels. For channel ii, a differentiable real FFT produces coefficients ci,jc_{i,j} for frequency indices j=0,…,Kj=0,\ldots,K, with K=⌊n/2⌋K=\lfloor n/2\rfloor, and the power spectrum is pi,j=2n∣ci,j∣2p_{i,j}=\frac{2}{n}|c_{i,j}|^2. If the window duration is NN seconds and fj=j/Nf_j=j/N, the channel amplitude, frequency, and offset are

    Ai=2n∑j=1Kpi,j,Fi=∑j=1Kfjpi,j∑j=1Kpi,j,Bi=ci,0n.A_i=\sqrt{\frac{2}{n}\sum_{j=1}^{K}p_{i,j}},\qquad F_i=\frac{\sum_{j=1}^{K}f_jp_{i,j}}{\sum_{j=1}^{K}p_{i,j}},\qquad B_i=\frac{c_{i,0}}{n}.

    A fully connected layer predicts (sx,i,sy,i)(s_{x,i},s_{y,i}) from channel Li\mathbf{L}_i, and its phase shift is Si=atan2⁡(sy,i,sx,i)S_i=\operatorname{atan2}(s_{y,i},s_{x,i}). The periodic decoder reconstructs each latent channel at time τ\tau as

    L^i(τ)=Aisin⁡ ⁣(2π(Fiτ−Si))+Bi,\hat{L}_i(\tau)=A_i\sin\!\left(2\pi(F_i\tau-S_i)\right)+B_i,

    then maps the reconstructed latent sequence back to motion. Training minimizes a phase-reconstruction loss between the original and decoded motion. At inference, each time point is represented by a 2M2M-dimensional phase vector

    P2i−1=Aisin⁡(2πSi),P2i=Aicos⁡(2πSi).\mathcal{P}_{2i-1}=A_i\sin(2\pi S_i),\qquad \mathcal{P}_{2i}=A_i\cos(2\pi S_i).

    The system compares the last NstrideN_{\mathrm{stride}} phase vectors of the seed with the first NphaseN_{\mathrm{phase}} phase vectors of each audio- and text-matched candidate, including the corresponding reversed boundary comparison. The candidate with the higher cosine similarity at the boundary is selected. The reported settings are M=8M=8, Nphase=8N_{\mathrm{phase}}=8, and Nstride=3N_{\mathrm{stride}}=3.

  6. Knowl 6 — BEAT-based training and evaluation protocol

    experimental setup

    Training and evaluation use the BEAT conversational-gesture dataset with an 8:1:1 train/validation/test split across speakers. Because database matching is computationally expensive, the motion-matching experiments use four hours from two speakers, “wayne” and “kieks,” with separate speaker databases. The system represents 15 upper-body joints, excluding hands and fingers, using 3×33\times3 rotation matrices, giving pose dimension Dg=9D_g=9. Gesture VQ-VAE training uses 240-frame clips, temporal downsampling rate 88, a 512-entry codebook with code dimension 512, Adam with learning rate 10−410^{-4}, β1=0.5\beta_1=0.5, β2=0.98\beta_2=0.98, batch size 128, 200 epochs, commitment weight β=0.1\beta=0.1, and velocity/acceleration weights α1=α2=1\alpha_1=\alpha_2=1. The phase network uses rotational velocity as input and is trained with AdamW for 100 epochs, learning rate and weight decay both 10−410^{-4}, and batch size 128.

    Objective evaluation reports average Hellinger distance between generated and natural speed histograms and Fréchet Gesture Distance (FGD) in both learned feature space and raw-data space; lower values are better. For histograms h(1)h^{(1)} and h(2)h^{(2)}, the reported Hellinger metric is

    H ⁣(h(1),h(2))=1−∑ihi(1)hi(2).H\!\left(h^{(1)},h^{(2)}\right)=\sqrt{1-\sum_i\sqrt{h_i^{(1)}h_i^{(2)}}}.

    FGD feature representations are obtained from an autoencoder trained on the Trinity dataset. Subjective evaluation uses 38 speech segments of approximately 8–15 seconds, 23 participants, avatar-rendered videos, and 1–5 mean-opinion scores for human-likeness and speech appropriateness.

  7. Knowl 7 — Performance against existing gesture-generation systems

    data/table

    On the BEAT test set, QPGesture is compared with End2End, Trimodal, StyleGestures, KNN, and CaMN. Lower objective metrics are better, while higher subjective scores are better. QPGesture obtains the best FGD values and matches StyleGestures on the best Hellinger distance. Its feature-space FGD is 15.92115.921 lower than StyleGestures, a 44%44\% improvement, and its raw-data FGD is 3837.0683837.068 lower, a 39%39\% improvement. In the user study, QPGesture has the highest human-likeness score, exceeding the ground-truth avatar-rendered score, while its appropriateness is statistically comparable to ground truth.

    Could not parse LaTeX table
  8. Knowl 8 — Ablation evidence for quantized audio, modalities, phase, and matching

    data/table

    Each ablation removes or replaces one component of QPGesture. Replacing vq-wav2vec plus Levenshtein matching with WavLM plus cosine similarity worsens the Hellinger score from 0.1360.136 to 0.1510.151 and raw-data FGD from 5742.2815742.281 to 6009.8596009.859. Removing text causes a larger feature-space and raw-data FGD degradation than removing audio, supporting the importance of semantic context. Removing motion matching and replacing it with a GRU pose-code generator produces the largest FGD degradation. Removing phase has only a small objective effect, consistent with phase acting primarily as a candidate-selection and continuity guide rather than as the main gesture generator.

    Could not parse LaTeX table
  9. Knowl 9 — Explicit gesture control through code-constrained retrieval

    model/method

    Because QPGesture retrieves discrete gesture codes from a database, generation can be constrained without retraining the model. A user can restrict retrieval to codes satisfying a pose condition, such as requiring the left wrist to remain above a threshold rr, or replace selected codes with another code that represents a preferred gesture. In the reported example, twelve occurrences of code 318 in a 240-frame sequence containing 30 codes are replaced with code 260, changing a habitual right-handed gesture toward a preferred left-handed gesture. The authors note that adding code-frequency scores to the matching objective can help preserve naturalness when applying such constraints.

  10. Knowl 10 — Stated limitation and extension direction

    limitation

    The demonstrated framework conditions gesture retrieval on text and audio, but it does not incorporate other potentially relevant modalities such as emotion or facial expressions. The authors identify multimodal conditioning beyond text and speech as a direction for generating more contextually appropriate gestures.

Coverage note — No substantial contributed material was omitted; background, related work, references, and acknowledgements were excluded.

References

  1. 1.Blender. https://www.blender.org/. 6
  2. 2.Mixamo. https://www.mixamo.com/. 1
  3. 3.Chaitanya Ahuja, Dong Won Lee, Ryo Ishii, and Louis-Philippe Morency. No gestures left behind: Learning relationships between spoken language and freeform gestures. In Findings of the Association for Computational Linguistics: EMNLP, volume EMNLP 2020 of Findings of ACL, pages 1884–1895, 2020. 6
  4. 4.Chaitanya Ahuja, Dong Won Lee, Yukiko I. Nakano, and Louis-Philippe Morency. Style transfer for co-speech gesture animation: A multi-speaker conditional-mixture approach. In Computer Vision - ECCV, volume 12363 of Lecture Notes in Computer Science, pages 248–265, 2020. 2, 3
  5. 5.Simon Alexanderson, Gustav Eje Henter, Taras Kucherenko, and Jonas Beskow. Style-controllable speech-driven gesture synthesis using normalising flows. Comput. Graph. Forum, 39(2):487–496, 2020. 2, 3, 6, 7
  6. 6.Simon Alexanderson, Eva Székely, Gustav Eje Henter, Taras Kucherenko, and Jonas Beskow. Generating coherent spontaneous speech and gesture from text. CoRR, abs/2101.05684, 2021. 2
  7. 7.Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, and Libin Liu. Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings. CoRR, abs/2210.01448, 2022. 3
  8. 8.Alexei Baevski, Steffen Schneider, and Michael Auli. vq-wav2vec: Self-supervised learning of discrete speech representations. In International Conference on Learning Representations, ICLR, 2020. 4
  9. 9.Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, and Dinesh Manocha. Speech2affectivegestures: Synthesizing co-speech gestures with generative adversarial affective expression learning. In MM ’21: ACM Multimedia Conference, pages 2027–2036, 2021. 2
  10. 10.Michael Büttner and Simon Clavet. Motion matching-the road to next gen animation. Proc. of Nucl. ai, 2015(1):2, 2015. 3
  11. 11.Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE J. Sel. Top. Signal Process., 16(6):1505–1518, 2022. 7
  12. 12.Kyunghyun Cho, Bart van Merrienboer, Çağlar Gulçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Processing, EMNLP, A meeting of SIGDAT, a Special Interest Group of the ACL, pages 1724–1734, 2014. 8
  13. 13.Ylva Ferstl and Rachel McDonnell. Iva: Investigating the use of recurrent motion modelling for speech gesture generation. In IVA ’18 Proceedings of the 18th International Conference on Intelligent Virtual Agents, Nov 2018. 6
  14. 14.Ylva Ferstl, Michael Neff, and Rachel McDonnell. Expressgesture: Expressive gesture generation from speech through database matching. Comput. Animat. Virtual Worlds, 32(3-4), 2021. 3
  15. 15.Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik. Learning individual styles of conversational gesture. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, pages 3497–3506, 2019. 2
  16. 16.Chuan Guo, Xinxin Xuo, Sen Wang, and Li Cheng. TM2T: stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts. CoRR, abs/2207.01696, 2022. 3
  17. 17.Ikhsanul Habibie, Mohamed Elgharib, Kripasindhu Sarkar, Ahsan Abdullah, Simbarashe Nyatsanga, Michael Neff, and Christian Theobalt. A motion matching-based framework for controllable gesture synthesis from speech. In SIGGRAPH ’22: Special Interest Group on Computer Graphics and Interactive Techniques Conference, pages 46:1–46:9, 2022. 1, 2, 3, 4, 6, 7
  18. 18.Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, and Christian Theobalt. Learning speech-driven 3d conversational gestures from video. In IVA ’21: ACM International Conference on Intelligent Virtual Agents, pages 101–108, 2021. 1, 2
  19. 19.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, pages 6626–6637, 2017. 6
  20. 20.Daniel Holden, Taku Komura, and Jun Saito. Phase-functioned neural networks for character control. ACM Trans. Graph., 36(4):42:1–42:13, 2017. 5
  21. 21.Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai, Lei Yang, and Ziwei Liu. Avatarclip: zero-shot text-driven generation and animation of 3d avatars. ACM Trans. Graph., 41(4):161:1–161:19, 2022. 3
  22. 22.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR, 2015. 6
  23. 23.Michael Kipp. Gesture generation by imitation: from human behavior to computer character animation. PhD thesis, Saarland University, Saarbrücken, Germany, 2003. 2, 3
  24. 24.Taras Kucherenko, Dai Hasegawa, Naoshi Kaneko, Gustav Eje Henter, and Hedvig Kjellström. Moving fast and slow: Analysis of representations and post-processing in speech-driven automatic gesture generation. Int. J. Hum. Comput. Interact., 37(14):1300–1316, 2021. 1, 2
  25. 25.Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter, Simon Alexandersson, Iolanda Leite, and Hedvig Kjellström. Gesticulator: A framework for semantically-aware speech-driven gesture generation. In ICMI: International Conference on Multimodal Interaction, pages 242–250, 2020. 6
  26. 26.Taras Kucherenko, Patrik Jonell, Youngwoo Yoon, Pieter Wolfert, and Gustav Eje Henter. A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA challenge 2020. In 26th International Conference on Intelligent User Interfaces, pages 11–21, 2021. 1
  27. 27.Taras Kucherenko, Rajmund Nagy, Michael Neff, Hedvig Kjellström, and Gustav Eje Henter. Multimodal analysis of the predictability of hand-gesture properties. In 21st International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2022, pages 770–779, 2022. 1
  28. 28.Vladimir I Levenshtein et al. Binary codes capable of correcting deletions, insertions, and reversals. In Soviet physics doklady, volume 10, pages 707–710. Soviet Union, 1966. 2, 4
  29. 29.Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, and Linchao Bao. Audio2gestures: Generating diverse gestures from speech audio with conditional variational autoencoders. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV, pages 11273–11282, 2021. 2, 3
  30. 30.Yuanzhi Liang, Qianyu Feng, Linchao Zhu, Li Hu, Pan Pan, and Yi Yang. SEEG: semantic energized co-speech gesture generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, pages 10463–10472, 2022. 3
  31. 31.Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng, Zhengqing Li, You Zhou, Elif Bozkurt, and Bo Zheng. BEAT: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis. CoRR, abs/2203.05297, 2022. 1, 2, 3, 5, 6, 7
  32. 32.Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu, Rui Qian, Xinyi Lin, Xiaowei Zhou, Wayne Wu, Bo Dai, and Bolei Zhou. Learning hierarchical cross-modal association for co-speech gesture generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, pages 10452–10462, 2022. 2
  33. 33.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR, 2019. 6
  34. 34.Thomas Lucas, Fabien Baradel, Philippe Weinzaepfel, and Gregory Rogez. Posegpt: Quantization-based 3d human motion generation and forecasting. arXiv preprint arXiv:2210.10542, 2022. 3
  35. 35.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: An ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP, pages 5206–5210, 2015. 4
  36. 36.Shenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu, and Shenghua Gao. Speech drives templates: Co-speech gesture synthesis with learned templates. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV, pages 11057–11066, 2021. 2
  37. 37.Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP, pages 3980–3990, 2019. 4
  38. 38.Li Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. Bailando: 3d dance generation by actor-critic gpt with choreographic memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11050–11059, 2022. 3
  39. 39.Sebastian Starke, Ian Mason, and Taku Komura. Deepphase: Periodic autoencoders for learning motion phase manifolds. ACM Trans. Graph., 41(4), jul 2022. 2, 4
  40. 40.Sebastian Starke, Yiwei Zhao, Taku Komura, and Kazi A. Zaman. Local motion phases for learning multi-contact character movements. ACM Trans. Graph., 39(4):54, 2020. 5
  41. 41.Taras Kucherenko, Youngwoo Yoon. Genea numerical evaluations, 2020. https://github.com/Svitozar/genea numerical evaluations. 7
  42. 42.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017. 2, 3
  43. 43.Siyang Wang, Simon Alexanderson, Joakim Gustafson, Jonas Beskow, Gustav Eje Henter, and Eva Székely. Integrated speech and gesture synthesis. In ICMI: International Conference on Multimodal Interaction, pages 177–185, 2021. 2
  44. 44.Jing Xu, Wei Zhang, Yalong Bai, Qibin Sun, and Tao Mei. Freeform body motion generation from speech. CoRR, abs/2203.02291, 2022. 2
  45. 45.Sicheng Yang, Zhiyong Wu, Minglei Li, Mengchen Zhao, Jiuxin Lin, Liyang Chen, and Weihong Bao. The reprgesture entry to the GENEA challenge 2022. In International Conference on Multimodal Interaction, ICMI, pages 758–763, 2022. 1
  46. 46.Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee. Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Transactions on Graphics, 39(6), 2020. 1, 2, 4, 6, 7
  47. 47.Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee, Jaehong Kim, and Geehyuk Lee. Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots. In International Conference on Robotics and Automation, ICRA, pages 4303–4309, 2019. 1, 2, 6, 7
  48. 48.Youngwoo Yoon, Pieter Wolfert, Taras Kucherenko, Carla Viegas, Teodor Nikolov, Mihail Tsakov, and Gustav Eje Henter. The GENEA challenge 2022: A large evaluation of data-driven co-speech gesture generation. In International Conference on Multimodal Interaction, ICMI, pages 736–747, 2022. 1
  49. 49.He Zhang, Sebastian Starke, Taku Komura, and Jun Saito. Mode-adaptive neural networks for quadruped motion control. ACM Trans. Graph., 37(4):145, 2018. 5
  50. 50.Chi Zhou, Tengyue Bian, and Kang Chen. Gesturemaster: Graph-based speech-driven gesture generation. In International Conference on Multimodal Interaction, ICMI, pages 764–770, 2022. 1, 2, 3, 4
  51. 51.Yang Zhou, Jimei Yang, Dingzeyu Li, Jun Saito, Deepali Aneja, and Evangelos Kalogerakis. Audio-driven neural gesture reenactment with video motion graphs. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, pages 3408–3418, 2022. 3

Citation

MLA
Yang, S., et al. “QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation”. arXiv, 2023, http://arxiv.org/abs/2305.11094v1.
APA
Yang, S., Wu, Z., Li, M., Zhang, Z., Hao, L., Bao, W., & Zhuang, H. (2023). QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation. arXiv. http://arxiv.org/abs/2305.11094v1
Chicago
Yang, S., Z. Wu, M. Li, et al. 2023. “QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation”. arXiv. http://arxiv.org/abs/2305.11094v1.
Harvard
Yang, S. et al. (2023) “QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.11094v1.
Vancouver
1. Yang S, Wu Z, Li M, Zhang Z, Hao L, Bao W, Zhuang H (2023) QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation. arXiv

BibTeX

@article{yang2023qpgesture,
  title = {QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation},
  author = {Yang, Sicheng and Wu, Zhiyong and Li, Minglei and Zhang, Zhensong and Hao, Lei and Bao, Weihong and Zhuang, Haolin},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.11094v1},
  eprint = {2305.11094}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE