Seamless Human Motion Composition with Blended Positional Encodings

Germán BarqueroSergio EscaleraCristina Palmero

article2024CVPR89 citationsBest Poster (AI4CC Workshop @ CVPR 2024)

Presents FlowMDM, a diffusion framework that schedules absolute and relative positional encodings during denoising to generate long, continuous multi-action human animations from sequential text prompts without requiring postprocessing or specialized multi-action training data.

Listen

Generating realistic, continuous three-dimensional human motion from text descriptions is critical for emerging applications in virtual reality, video gaming, and robotics. However, current automated methods primarily synthesize isolated, short movements lasting only a few seconds under a single instruction. Extending these models to produce long sequences driven by a series of distinct actions usually leads to unnatural, abrupt transitions or requires computationally intensive postprocessing and stitching. The core difficulty stems from existing training datasets lacking long-duration sequences with annotations for transitions between consecutive actions.

The article introduces FlowMDM, a generative diffusion framework designed to evaluate and demonstrate the simultaneous synthesis of long, seamless human motion compositions directly from sequential text prompts without requiring manual stitching, postprocessing, or datasets with annotated transitions.

To achieve this, the researchers evaluated their approach on two benchmark datasets, HumanML3D and Babel, comparing it against established sequential and diffusion-based baselines. The methodology introduces two key architectural features: Blended Positional Encodings and Pose-Centric Cross-Attention. Blended Positional Encodings utilize absolute positional information early in the iterative generation process to establish overall motion coherence, gradually shifting to relative positional information in later stages to ensure fluid transitions between actions. The Pose-Centric Cross-Attention mechanism prevents conflicting conditions from tangling at action boundaries, allowing the system to train on datasets with only one text description per sequence while still executing multi-prompt sequences during inference. Additionally, the authors developed two physical metrics based on jerk—the time derivative of acceleration—to quantitatively evaluate transition smoothness.

The findings demonstrate substantial improvements over existing methods across multiple performance dimensions. First, the proposed framework achieved superior motion quality and text alignment, reducing motion distribution error scores by over 60% compared to prior diffusion models on HumanML3D. Second, it delivered noticeably smoother and more realistic transitions, minimizing sudden acceleration spikes where baseline methods exhibited severe motion glitches. Third, computational efficiency improved substantially, requiring roughly 16% to 47% fewer pose-wise computation steps than existing diffusion-based blending approaches. Finally, the model successfully generalized to cyclic, repetitive movements (such as walking or waving) over extended durations without degrading motion consistency.

These results demonstrate that long, controlled character animations can be produced efficiently in a single processing pass, lowering computational overhead and eliminating manual cleanup. This provides a direct path to reducing animation production costs and accelerating asset generation pipelines in gaming and virtual simulations. The findings also highlight that standard generative evaluation metrics fail to detect physical irregularities like abrupt jerks, underscoring the necessity of motion-derivative metrics for quality control.

Organizations developing interactive avatars or automated animation tools should consider adopting blended encoding architectures to streamline motion generation pipelines without investing in expensive, specialized transition datasets. For production deployment, teams should tune the schedule between absolute and relative encoding steps—allocating roughly 10% of initial steps to absolute positions—to optimally balance text fidelity against transition smoothness.

A recognized limitation of the current architecture is that the initial absolute generation phase models motion segments independently, leaving the system without high-level intent planning across far-apart actions. While confidence in the benchmarked smoothness and accuracy metrics is high, further validation is required before extending the architecture to multi-modal control signals (such as combining text with scene geometry or audio) or applying it to real-time physical robotics.

  • Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Provides the foundational Motion Diffusion Model (MDM) architecture that FlowMDM builds upon and extends for long-sequence composition.
  • Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). Introduces relative and flexible positional conditioning strategies in temporal diffusion models for long sequence generation.
  • Paper: Flow Matching for Generative Modeling, Yaron Lipman et al. (2023). Establishes the core mathematical formulations of flow matching that underpin continuous flow-based generative modeling.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Presents the fundamental denoising diffusion probabilistic framework utilized by diffusion-based motion generation models.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Formulates implicit diffusion sampling techniques that enable efficient deterministic generation in continuous denoising pipelines.
Cover for Seamless Human Motion Composition with Blended Positional Encodings

Abstract

Conditional human motion generation is an important topic with many applications in virtual reality, gaming, and robotics. While prior works have focused on generating motion guided by text, music, or scenes, these typically result in isolated motions confined to short durations. Instead, we address the generation of long, continuous sequences guided by a series of varying textual descriptions. In this context, we introduce FlowMDM, the first diffusion-based model that generates seamless Human Motion Compositions (HMC) without any postprocessing or redundant denoising steps. For this, we introduce the Blended Positional Encodings, a technique that leverages both absolute and relative positional encodings in the denoising chain. More specifically, global motion coherence is recovered at the absolute stage, whereas smooth and realistic transitions are built at the relative stage. As a result, we achieve state-of-the-art results in terms of accuracy, realism, and smoothness on the Babel and HumanML3D datasets. FlowMDM excels when trained with only a single description per motion sequence thanks to its Pose-Centric Cross-ATtention, which makes it robust against varying text descriptions at inference time. Finally, to address the limitations of existing HMC metrics, we propose two new metrics: the Peak Jerk and the Area Under the Jerk, to detect abrupt transitions.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methodology
  • 3.1. Bidirectional diffusion
  • 3.2. Blended positional encodings
  • 3.3. Pose-centric cross-attention
  • 4. Experiments
  • 4.1. Experimental setup
  • 4.2. Quantitative analysis
  • 4.3. Qualitative results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Diffusion-Based Generative Human Motion Composition via FlowMDM

    model/method

    Generative Human Motion Composition (HMC) aims to synthesize a continuous 3D motion sequence of NN frames driven by non-overlapping temporal intervals [0,τ1),[τ1,τ2),…,[τj,N)[0, \tau_1), [\tau_1, \tau_2), \dots, [\tau_j, N), where 0<τ1<⋯<τj<N0 < \tau_1 < \dots < \tau_j < N. Each interval represents a motion subsequence Si={xτi,…,xτi+1−1}S_i = \{x_{\tau_i}, \dots, x_{\tau_{i+1}-1}\} governed by a distinct textual condition cic_i, with a maximum subsequence length of LL.

    FlowMDM formulates HMC as a single, simultaneous denoising process using a bidirectional (encoder-only) Transformer diffusion model parameterized by x0x_0 prediction with mean squared error (L2) reconstruction loss. Unlike autoregressive approaches that accumulate errors over long sequences or diffusion infilling techniques that repeatedly blend overlapping temporal windows, FlowMDM generates the entire multi-action composition simultaneously in a single denoising trajectory without requiring postprocessing, linear interpolations, or predetermined transition durations.

  2. Knowl 2 — Blended Positional Encodings (BPE)

    model/method

    Blended Positional Encodings (BPE) is a positional encoding mechanism for diffusion Transformers that combines Absolute Positional Encodings (APE) and Relative Positional Encodings (RPE) across the diffusion sampling trajectory to enable long-sequence extrapolation while preserving action-level semantics.

    In human motion diffusion, early denoising steps prioritize global low-frequency structural dependencies, whereas later steps resolve local high-frequency details. BPE leverages this property:

    1. Early Denoising Phase: The network uses sinusoidal APE, restricting attention within each subsequence [τi,τi+1)[\tau_i, \tau_{i+1}). This provides temporal anchoring (distance to the start and end of an action), allowing the model to distinguish complex intra-action compositions (e.g., discerning 'walk' from 'sit down' in a compound prompt).
    2. Late Denoising Phase: The network switches to Rotary Position Embeddings (RoPE), spanning all frames up to a local attention horizon H<L<NH < L < N. Given query projection matrix WqW_q, key projection matrix WkW_k, and poses xm,xnx_m, x_n at frame indices mm and nn, RoPE applies dd-dimensional rotation matrices RmdR_m^d and RndR_n^d:

    qmTkn=(RmdWqxm)T(RndWkxn)=xmTWqRn−mdWkxnq_m^T k_n = (R_m^d W_q x_m)^T (R_n^d W_k x_n) = x_m^T W_q R_{n-m}^d W_k x_n

    Because Rn−mdR_{n-m}^d depends exclusively on the relative distance n−mn-m, absolute temporal boundaries are eliminated, making high-frequency denoising translation-invariant and allowing smooth, continuous transitions to emerge naturally between subsequences.

    During training, APE and RPE are alternated randomly with equal probability (p=0.5p = 0.5). During inference, a binary step scheduler executes the first ∼\sim10% to 12.5% of denoising steps (e.g., 60 to 125 steps out of 1000) using APE before switching to RPE for the remaining steps.

  3. Knowl 3 — Pose-Centric Cross-Attention (PCCAT)

    model/method

    When a motion diffusion model is trained on single-action sequences, the textual condition embedding remains constant across all frames. At inference time in Human Motion Composition (HMC), consecutive subsequences have differing conditions (cm≠cnc_m \neq c_n). Standard concatenation or vanilla self-attention across frames computes attention scores using entangled pose-condition embeddings Exm,cmE_{x_m, c_m} and Exn,cnE_{x_n, c_n} via qmTkn=(WqExm,cm)T(WkExn,cn)q_m^T k_n = (W_q E_{x_m, c_m})^T (W_k E_{x_n, c_n}), creating an out-of-distribution input during transition boundaries.

    Pose-Centric Cross-Attention (PCCAT) eliminates this training-inference mismatch by injecting the textual condition cmc_m exclusively into the query representation while keeping keys and values strictly dependent on noisy poses xnx_n:

    qmTkn=(WqExm,cm)T(WkExn)=Exm,cmTWqTWkExnq_m^T k_n = (W_q E_{x_m, c_m})^T (W_k E_{x_n}) = E_{x_m, c_m}^T W_q^T W_k E_{x_n}

    Here, WqW_q and WkW_k are linear projection matrices, Exm,cmE_{x_m, c_m} is the combined embedding of the noisy pose and textual condition for frame mm, and ExnE_{x_n} is the embedding of the noisy pose at frame nn. Consequently, the attention output for frame mm is a weighted combination of value projections of neighboring noisy poses, ensuring that frame mm denoises according to its own text prompt while conditioning smoothly on the kinematics of surrounding frames.

  4. Knowl 4 — Peak Jerk and Area Under the Jerk Metrics for Motion Transitions

    definition

    To evaluate kinetic smoothness and detect abrupt velocity changes or discontinuities at transition boundaries between concatenated motion subsequences Si−1S_{i-1} and SiS_i, transition quality is quantified using Peak Jerk (PJ) and Area Under the Jerk (AUJ). Jerk is the first time derivative of acceleration (third time derivative of position).

    For a transition window of length LtrL_{tr} frames centered at transition boundary τi\tau_i across KK human skeleton joints:

    PJ=max⁡1≤i≤K1≤τ≤Ltr∣ji(τ)∣1\text{PJ} = \max_{\substack{1 \le i \le K \\ 1 \le \tau \le L_{tr}}} |j_i(\tau)|_1

    AUJ=∑τ=1Ltrmax⁡1≤i≤K∣ji(τ)−javg∣1\text{AUJ} = \sum_{\tau=1}^{L_{tr}} \max_{1 \le i \le K} |j_i(\tau) - j_{avg}|_1

    where ji(τ)j_i(\tau) is the instantaneous jerk of joint ii at frame τ\tau, and javgj_{avg} is the dataset-wide average maximum joint jerk across ground truth motions.

    Peak Jerk captures extreme instantaneous kinetic fluctuations. Area Under the Jerk measures the aggregate cumulative L1 deviation from the dataset's natural jerk baseline across the full transition duration, penalizing both unnaturally sharp spikes and artificially extended smoothing artifacts.

  5. Knowl 5 — Evaluation of Motion Composition on the Babel Benchmark

    data/table

    Quantitative performance on sequential motion generation evaluated on the Babel dataset across 32 concatenated textual descriptions (32 subsequences, 31 transitions of length Ltr=30L_{tr}=30 frames). Subsequence quality is evaluated using Top-3 R-precision (R-prec ↑\uparrow), Fréchet Inception Distance (FID ↓\downarrow), Diversity (Div →\rightarrow), and Multimodal Distance (MM-Dist ↓\downarrow). Transition quality is measured via FID ↓\downarrow, Diversity →\rightarrow, Peak Jerk (PJ ↓\downarrow), and Area Under the Jerk (AUJ ↓\downarrow). Values report mean ±\pm 95% confidence intervals over 10 evaluation runs.

    Subsequence Transition
    Method R-prec ↑\uparrow FID ↓\downarrow Div →\rightarrow MM-Dist ↓\downarrow FID ↓\downarrow Div →\rightarrow PJ ↓\downarrow AUJ ↓\downarrow
    GT 0.715±\pm0.003 0.00±\pm0.00 8.42±\pm0.15 3.36±\pm0.00 0.00±\pm0.00 6.20±\pm0.06 0.02±\pm0.00 0.00±\pm0.00
    TEACH_B 0.703±\pm0.002 1.71±\pm0.03 8.18±\pm0.14 3.43±\pm0.01 3.01±\pm0.04 6.23±\pm0.05 1.09±\pm0.00 2.35±\pm0.01
    TEACH 0.655±\pm0.002 1.82±\pm0.02 7.96±\pm0.11 3.72±\pm0.01 3.27±\pm0.04 6.14±\pm0.06 0.07±\pm0.00 0.44±\pm0.00
    DoubleTake* 0.596±\pm0.005 3.16±\pm0.06 7.53±\pm0.11 4.17±\pm0.02 3.33±\pm0.06 6.16±\pm0.05 0.28±\pm0.00 1.04±\pm0.01
    DoubleTake 0.668±\pm0.005 1.33±\pm0.04 7.98±\pm0.12 3.67±\pm0.03 3.15±\pm0.05 6.14±\pm0.07 0.17±\pm0.00 0.64±\pm0.01
    MultiDiffusion 0.702±\pm0.005 1.74±\pm0.04 8.37±\pm0.13 3.43±\pm0.02 6.56±\pm0.12 5.72±\pm0.07 0.18±\pm0.00 0.68±\pm0.00
    DiffCollage 0.671±\pm0.003 1.45±\pm0.05 7.93±\pm0.09 3.71±\pm0.01 4.36±\pm0.09 6.09±\pm0.08 0.19±\pm0.00 0.84±\pm0.01
    FlowMDM 0.702±\pm0.004 0.99±\pm0.04 8.36±\pm0.13 3.45±\pm0.02 2.61±\pm0.06 6.47±\pm0.05 0.06±\pm0.00 0.13±\pm0.00

    FlowMDM achieves the best subsequence FID (0.99) and transition FID (2.61) among all methods while delivering the lowest transition smoothness deviation (PJ of 0.06, AUJ of 0.13), demonstrating smoother and more realistic transitions than autoregressive and multi-stage diffusion baselines.

  6. Knowl 6 — Evaluation of Motion Composition on the HumanML3D Benchmark

    data/table

    Quantitative performance on sequential motion generation evaluated on the HumanML3D dataset across 32 concatenated textual descriptions (transition duration Ltr=60L_{tr}=60 frames). Subsequence quality is evaluated using Top-3 R-precision (R-prec ↑\uparrow), FID ↓\downarrow, Diversity (Div →\rightarrow), and Multimodal Distance (MM-Dist ↓\downarrow). Transition quality is evaluated with FID ↓\downarrow, Diversity →\rightarrow, Peak Jerk (PJ ↓\downarrow), and Area Under the Jerk (AUJ ↓\downarrow). Values report mean ±\pm 95% confidence intervals over 10 evaluation runs. Autoregressive methods (TEACH) cannot be trained on HumanML3D due to the absence of consecutive action pair annotations.

    Subsequence Transition
    Method R-prec ↑\uparrow FID ↓\downarrow Div →\rightarrow MM-Dist ↓\downarrow FID ↓\downarrow Div →\rightarrow PJ ↓\downarrow AUJ ↓\downarrow
    GT 0.796±\pm0.004 0.00±\pm0.00 9.34±\pm0.08 2.97±\pm0.01 0.00±\pm0.00 9.54±\pm0.15 0.04±\pm0.00 0.07±\pm0.00
    DoubleTake* 0.643±\pm0.005 0.80±\pm0.02 9.20±\pm0.11 3.92±\pm0.01 1.71±\pm0.05 8.82±\pm0.13 0.52±\pm0.01 2.10±\pm0.03
    DoubleTake 0.628±\pm0.005 1.25±\pm0.04 9.09±\pm0.12 4.01±\pm0.01 4.19±\pm0.09 8.45±\pm0.09 0.48±\pm0.00 1.83±\pm0.02
    MultiDiffusion 0.629±\pm0.002 1.19±\pm0.03 9.38±\pm0.08 4.02±\pm0.01 4.31±\pm0.06 8.37±\pm0.10 0.17±\pm0.00 1.06±\pm0.01
    DiffCollage 0.615±\pm0.005 1.56±\pm0.04 8.79±\pm0.08 4.13±\pm0.02 4.59±\pm0.10 8.22±\pm0.11 0.26±\pm0.00 2.85±\pm0.09
    FlowMDM 0.685±\pm0.004 0.29±\pm0.01 9.58±\pm0.12 3.61±\pm0.01 1.38±\pm0.05 8.79±\pm0.09 0.06±\pm0.00 0.51±\pm0.01

    FlowMDM outperforms all diffusion baselines in text alignment (R-prec of 0.685 vs. ≤0.643\le 0.643, MM-Dist of 3.61 vs. ≥3.92\ge 3.92), visual realism (subsequence FID of 0.29 vs. ≥0.80\ge 0.80, transition FID of 1.38 vs. ≥1.71\ge 1.71), and transition smoothness (PJ of 0.06 vs. ≥0.17\ge 0.17, AUJ of 0.51 vs. ≥1.06\ge 1.06).

  7. Knowl 7 — Ablation of Positional Encodings and Conditioning Mechanisms in HMC

    empirical result

    Ablation experiments on Babel and HumanML3D evaluate the contributions of positional encoding (Absolute [APE], Relative [RPE], and Blended [BPE]) and text conditioning schemes (Pose-Centric Cross-Attention [PCCAT], Self-Attention [SAT], and Cross-Attention [CAT]):

    1. Positional Encoding Effects:

      • Training and sampling solely with APE achieves strong text-matching (R-prec of 0.699 on Babel, 0.689 on HumanML3D) but produces sharp, unnatural transitions with elevated jerk metrics (AUJ of 3.73 on Babel, 3.40 on HumanML3D; PJ of 1.81 and 1.50).
      • Training and sampling solely with RPE achieves smooth transitions (AUJ of 0.20 on Babel, 0.58 on HumanML3D; PJ of 0.03), but struggles with action-level global structures, dropping R-prec to 0.635 on Babel and 0.531 on HumanML3D.
      • BPE training combined with BPE sampling maintains the high semantic accuracy of APE (R-prec 0.702 on Babel, 0.685 on HumanML3D; subsequence FID 0.99 and 0.29) while preserving the low jerk of RPE (AUJ 0.13 on Babel, 0.51 on HumanML3D).
    2. Conditioning Mechanisms:

      • In HumanML3D (where training samples have only single text descriptions), conditioning via SAT and CAT yields degraded transition FID (3.19 and 3.93, respectively) due to out-of-distribution conditioning mixtures at inference. PCCAT resolves this mismatch, reducing transition FID to 1.38.
      • In Babel (where training samples contain multi-action sequences), models are inherently exposed to varying conditions, resulting in competitive transition FID across conditioning types (SAT: 1.91, CAT: 2.57, PCCAT: 2.61).
  8. Knowl 8 — Computational Efficiency and Periodic Motion Extrapolation of FlowMDM

    empirical result

    FlowMDM provides specific computational efficiency gains and extrapolation capabilities for human motion generation:

    1. Denoising Efficiency: Diffusion blending baselines (DiffCollage, MultiDiffusion) denoise overlapping frames repeatedly, while DoubleTake performs an extra unconditional refinement stage on transitions. In contrast, FlowMDM denoises all frames simultaneously without redundant passes, reducing the total number of pose-wise denoising steps by 47.1% relative to DoubleTake, 28.4% relative to DiffCollage, and 16.5% relative to MultiDiffusion.

    2. Periodic Motion Extrapolation: When evaluated on 32 consecutive repetitions of periodic actions (e.g., 'walk forward', 'jumping', 'playing the guitar') to test duration extrapolation beyond training lengths, baseline methods exhibit large jerk spikes and disrupt periodicity at subsequence seams. FlowMDM generates continuous periodic motion without transition jerk spikes, closely matching ground truth motion dynamics.

  9. Knowl 9 — Independent Subsequence Low-Frequency Generation in FlowMDM

    limitation

    In the initial Absolute Positional Encoding (APE) phase of Blended Positional Encodings (BPE), attention is restricted exclusively within each individual subsequence to establish intra-subsequence action semantics. Consequently, the low-frequency global structures across consecutive subsequences are synthesized independently without long-range cross-subsequence intention planning, leaving inter-subsequence coordination to the subsequent high-frequency RPE phase.

Coverage note — No substantial contributed material was omitted. All primary algorithmic mechanisms (FlowMDM, BPE, PCCAT), evaluation metrics (PJ, AUJ), quantitative benchmark data (Tables 1-4), extrapolation findings, efficiency gains, and limitations are represented.

References

  1. 1.Vida Adeli, Mahsa Ehsanpour, Ian Reid, Juan Carlos Niebles, Silvio Savarese, Ehsan Adeli, and Hamid Rezatofighi. Tripod: Human trajectory and pose dynamics forecasting in the wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13390–13400, 2021.
  2. 2.Simon Alexanderson, Rajmund Nagy, Jonas Beskow, and Gustav Eje Henter. Listen, denoise, action! audio-driven motion synthesis with diffusion models. ACM Transactions on Graphics (TOG), 42(4):1–20, 2023.
  3. 3.Sadegh Aliakbarian, Microsoft Fatemeh Saleh ACRV, Stephen Gould ACRV, and Anu Mathieu Salzmann CVLab. Contextually plausible and diverse 3d human motion prediction. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  4. 4.Nikos Athanasiou, Mathis Petrovich, Michael J Black, and G÷l Varol. Teach: Temporal action composition for 3d humans. In 2022 International Conference on 3D Vision (3DV), pages 414–423. IEEE, 2022.
  5. 5.Sivakumar Balasubramanian, Alejandro Melendez-Calderon, and Etienne Burdet. A robust and sensitive metric for quantifying movement smoothness. IEEE transactions on biomedical engineering, 59(8):2126–2136, 2011.
  6. 6.Sivakumar Balasubramanian, Alejandro Melendez-Calderon, Agnes Roby-Brami, and Etienne Burdet. On the analysis of movement smoothness. Journal of neuroengineering and rehabilitation, 12(1):1–11, 2015.
  7. 7.Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel. Multidiffusion: Fusing diffusion paths for controlled image generation. In International Conference on Machine Learning, pages 1737–1752. PMLR, 2023.
  8. 8.German Barquero, Sergio Escalera, and Cristina Palmero. Belfusion: Latent diffusion for behavior-driven human motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2317–2327, 2023.
  9. 9.German Barquero, Johnny Nunez, Sergio Escalera, Zhen Xu, Wei-Wei Tu, Isabelle Guyon, and Cristina Palmero. Didn't see that coming: a survey on non-verbal social human behavior forecasting. In Understanding Social Behavior in Dyadic and Small Group Interactions, pages 139–178. PMLR, 2022.
  10. 10.German Barquero, Johnny Núñez, Zhen Xu, Sergio Escalera, Wei-Wei Tu, Isabelle Guyon, and Cristina Palmero. Comparison of spatio-temporal models for human motion and pose forecasting in face-to-face interaction scenarios. In Understanding Social Behavior in Dyadic and Small Group Interactions, pages 107–138. PMLR, 2022.
  11. 11.Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  12. 12.Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15935–15946, 2022.
  13. 13.Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it smpl: Automatic estimation of 3d human pose and shape from a single image. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, pages 561–578. Springer, 2016.
  14. 14.Paulo Vinicius Koerich Borges, Nicola Conci, and Andrea Cavallaro. Video-based human behavior understanding: A survey. IEEE transactions on circuits and systems for video technology, 23(11):1993–2008, 2013.
  15. 15.Zhe Cao, Hang Gao, Karttikeya Mangalam, Qi-Zhi Cai, Minh Vo, and Jitendra Malik. Long-term human motion prediction with scene context. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, pages 387–404. Springer, 2020.
  16. 16.Angela Castillo, Maria Escobar, Guillaume Jeanneret, Albert Pumarola, Pablo Arbelaez, Ali Thabet, and Artsiom Sanakoyeu. Bodiffusion: Diffusing sparse observations for full-body human motion synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4221–4231, 2023.
  17. 17.Kang Chen, Zhipeng Tan, Jin Lei, Song-Hai Zhang, Yuan-Chen Guo, Weidong Zhang, and Shi-Min Hu. Choreomaster: choreography-oriented music-driven dance synthesis. ACM Transactions on Graphics (TOG), 40(4):1–13, 2021.
  18. 18.Enric Corona, Albert Pumarola, Guillem Alenya, and Francesc Moreno-Noguer. Context-aware human motion prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6992–7001, 2020.
  19. 19.Antonia Creswell, Tom White, Vincent Dumoulin, Kai Arulkumaran, Biswa Sengupta, and Anil A Bharath. Generative adversarial networks: An overview. IEEE signal processing magazine, 35(1):53–65, 2018.
  20. 20.Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  21. 21.Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt. Mofusion: A framework for denoising-diffusion-based motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9760–9770, 2023.
  22. 22.Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In The Twelfth International Conference on Learning Representations, 2023.
  23. 23.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Re. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359, 2022.
  24. 24.David A Engstrom, JA Scott Kelso, and Tom Holroyd. Reaction-anticipation transitions in human perception-action patterns. Human movement science, 15(6):809–832, 1996.
  25. 25.Philipp Gulde and Joachim Hermsdörfer. Smoothness metrics in complex movement tasks. Frontiers in neurology, 9:615, 2018.
  26. 26.Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5152–5161, 2022.
  27. 27.Chuan Guo, Xinxin Zuo, Sen Wang, and Li Cheng. Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts. In European Conference on Computer Vision, pages 580–597. Springer, 2022.
  28. 28.Wen Guo, Xiaoyu Bie, Xavier Alameda-Pineda, and Francesc Moreno-Noguer. Multi-person extreme motion prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13053–13064, 2022.
  29. 29.Félix G Harvey, Mike Yurick, Derek Nowrouzezahrai, and Christopher Pal. Robust motion in-betweening. ACM Transactions on Graphics (TOG), 39(4):60–1, 2020.
  30. 30.Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black. Stochastic scene-aware motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11374–11384, 2021.
  31. 31.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  32. 32.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  33. 33.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.
  34. 34.Neville Hogan and Dagmar Sternad. Sensitivity of smoothness measures to movement duration, amplitude, and arrests. Journal of motor behavior, 41(6):529–534, 2009.
  35. 35.Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign language. Advances in Neural Information Processing Systems, 36, 2024.
  36. 36.Manuel Kaufmann, Emre Aksan, Jie Song, Fabrizio Pece, Remo Ziegler, and Otmar Hilliges. Convolutional autoencoders for human motion infilling. In 2020 International Conference on 3D Vision (3DV), pages 918–927. IEEE, 2020.
  37. 37.Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. The impact of positional encoding on length generalization in transformers. Advances in Neural Information Processing Systems, 36, 2024.
  38. 38.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186, 2019.
  39. 39.Jihoon Kim, Taehyun Byun, Seungyoun Shin, Jungdam Won, and Sungjoon Choi. Conditional motion in-betweening. Pattern Recognition, 132:108894, 2022.
  40. 40.Jihoon Kim, Jiseob Kim, and Sungjoon Choi. Flame: Free-form language-based motion synthesis & editing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 8255–8263, 2023.
  41. 41.Nilesh Kulkarni, Davis Rempe, Kyle Genova, Abhijit Kundu, Justin Johnson, David Fouhey, and Leonidas Guibas. Nifty: Neural object interaction fields for guided human motion synthesis. arXiv preprint arXiv:2307.07511, 2023.
  42. 42.Wilfried Kunde, Katrin Elsner, and Andrea Kiesel. No anticipation–no action: the role of anticipation in action and perception. Cognitive Processing, 8:71–78, 2007.
  43. 43.Caroline Larboulette and Sylvie Gibet. A review of computable expressive descriptors of human motion. In Proceedings of the 2nd International Workshop on Movement and Computing, pages 21–28, 2015.
  44. 44.Taeryung Lee, Gyeongsik Moon, and Kyoung Mu Lee. Multiact: Long-term 3d human motion generation from multiple action labels. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 1231–1239, 2023.
  45. 45.Jiaman Li, Ruben Villegas, Duygu Ceylan, Jimei Yang, Zhengfei Kuang, Hao Li, and Yajie Zhao. Task-generic hierarchical human motion prior using vaes. In 2021 International Conference on 3D Vision (3DV), pages 771–781. IEEE, 2021.
  46. 46.Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. Ai choreographer: Music conditioned 3d dance generation with aist++. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13401–13412, 2021.
  47. 47.Shuai Li, Sisi Zhuang, Wenfeng Song, Xinyu Zhang, Hejia Chen, and Aimin Hao. Sequential texts driven cohesive motions synthesis with natural transitions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9498–9508, 2023.
  48. 48.Weiyu Li, Xuelin Chen, Peizhuo Li, Olga Sorkine-Hornung, and Baoquan Chen. Example-based motion synthesis via generative motion matching. ACM Transactions on Graphics (TOG), 42(4):1–12, 2023.
  49. 49.Yunhao Li, Zhenbo Yu, Yucheng Zhu, Bingbing Ni, Guangtao Zhai, and Wei Shen. Skeleton2humanoid: Animating simulated characters for physically-plausible motion in-betweening. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1493–1502, 2022.
  50. 50.Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x: A large-scale 3d expressive whole-body human motion dataset. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023.
  51. 51.Zhengyi Luo, Jinkun Cao, Kris Kitani, Weipeng Xu, et al. Perpetual humanoid control for real-time simulated avatars. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10895–10904, 2023.
  52. 52.Hengbo Ma, Jiachen Li, Ramtin Hosseini, Masayoshi Tomizuka, and Chiho Choi. Multi-objective diverse human motion prediction with knowledge distillation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  53. 53.Shugao Ma, Tomas Simon, Jason Saragih, Dawei Wang, Yuecheng Li, Fernando De La Torre, and Yaser Sheikh. Pixel codec avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 64–73, 2021.
  54. 54.Antoine Maiorca, Youngwoo Yoon, and Thierry Dutoit. Evaluating the quality of a synthesized motion with the fréchet motion distance. In ACM SIGGRAPH 2022 Posters, pages 1–2, 2022.
  55. 55.Antoine Maiorca, Youngwoo Yoon, and Thierry Dutoit. Validating objective evaluation metric: Is fréchet motion distance able to capture foot skating artifacts? In Proceedings of the 2023 ACM International Conference on Interactive Media Experiences, pages 242–247, 2023.
  56. 56.Wei Mao, Miaomiao Liu, and Mathieu Salzmann. Generating smooth pose sequences for diverse human motion prediction. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  57. 57.Boris N Oreshkin, Antonios Valkanas, Félix G Harvey, Louis-Simon Ménard, Florent Bocquelet, and Mark J Coates. Motion in-betweening via deep delta-interpolator. IEEE Transactions on Visualization and Computer Graphics, 2023.
  58. 58.Cristina Palmero, German Barquero, Julio CS Jacques Junior, Albert Clapés, Johnny Núñez, David Curto, Sorina Smeureanu, Javier Selva, Zejian Zhang, David Saeteros, et al. Chalearn lap challenges on self-reported personality recognition and non-verbal behavior forecasting during social dyadic interactions: Dataset, design, and results. In Understanding Social Behavior in Dyadic and Small Group Interactions, pages 4–52. PMLR, 2022.
  59. 59.Zizheng Pan, Jianfei Cai, and Bohan Zhuang. Fast vision transformers with hilo attention. Advances in Neural Information Processing Systems, 35:14541–14554, 2022.
  60. 60.Sang-Min Park and Young-Gab Kim. A metaverse: Taxonomy, components, applications, and open challenges. IEEE access, 10:4209–4251, 2022.
  61. 61.Mathis Petrovich, Michael J Black, and G÷l Varol. Temos: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision, pages 480–497. Springer, 2022.
  62. 62.Matthias Plappert, Christian Mandery, and Tamim Asfour. The kit motion-language dataset. Big data, 4(4):236–252, 2016.
  63. 63.Abhinanda R Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J Black. Babel: Bodies, action and behavior with english labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 722–731, 2021.
  64. 64.Yijun Qian, Jack Urbanek, Alexander G Hauptmann, and Jungdam Won. Breaking the limits of text-conditioned 3d motion synthesis with elaborative descriptions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2306–2316, 2023.
  65. 65.Jia Qin, Youyi Zheng, and Kun Zhou. Motion in-betweening via two-stage transformers. ACM Transactions on Graphics (TOG), 41(6):1–16, 2022.
  66. 66.Tianxiang Ren, Jubo Yu, Shihui Guo, Ying Ma, Yutao Ouyang, Zijiao Zeng, Yazhan Zhang, and Yipeng Qin. Diverse motion in-betweening with dual posture stitching. arXiv preprint arXiv:2303.14457, 2023.
  67. 67.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  68. 68.Mehdi SM Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. Assessing generative models via precision and recall. Advances in neural information processing systems, 31, 2018.
  69. 69.Tim Salzmann, Marco Pavone, and Markus Ryll. Motron: Multimodal probabilistic human motion forecasting. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  70. 70.Yoni Shafir, Guy Tevet, Roy Kapon, and Amit Haim Bermano. Human motion diffusion as a generative prior. In The Twelfth International Conference on Learning Representations, 2023.
  71. 71.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  72. 72.Paul Starke, Sebastian Starke, Taku Komura, and Frank Steinicke. Motion in-betweening with phase manifolds. Proceedings of the ACM on Computer Graphics and Interactive Techniques, 6(3):1–17, 2023.
  73. 73.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  74. 74.Guofei Sun, Yongkang Wong, Zhiyong Cheng, Mohan S Kankanhalli, Weidong Geng, and Xiangdong Li. Deepdance: music-to-dance motion choreography with adversarial learning. IEEE Transactions on Multimedia, 23:497–509, 2020.
  75. 75.Jiarui Sun and Girish Chowdhary. Towards globally consistent stochastic human motion prediction via motion diffusion. arXiv preprint arXiv:2305.12554, 2023.
  76. 76.Ryo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati, and Nicolai Marquardt. Augmented reality and robotics: A survey and taxonomy for ar-enhanced human-robot interaction and robotic interfaces. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–33, 2022.
  77. 77.Julian Tanke, Linguang Zhang, Amy Zhao, Chengcheng Tang, Yujun Cai, Lezi Wang, Po-Chen Wu, Juergen Gall, and Cem Keskin. Social diffusion: Long-term multiple human motion anticipation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9601–9611, 2023.
  78. 78.Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. In European Conference on Computer Vision, pages 358–374. Springer, 2022.
  79. 79.Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffusion model. In The Eleventh International Conference on Learning Representations, 2022.
  80. 80.Sibo Tian, Minghui Zheng, and Xiao Liang. Transfusion: A practical and effective transformer-based diffusion model for 3d human motion prediction. arXiv preprint arXiv:2307.16106, 2023.
  81. 81.Jonathan Tseng, Rodrigo Castellon, and Karen Liu. Edge: Editable dance generation from music. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 448–458, 2023.
  82. 82.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  83. 83.Jacob Walker, Kenneth Marino, Abhinav Gupta, and Martial Hebert. The pose knows: Video forecasting by generating pose futures. Proceedings of the IEEE international conference on computer vision, 2017.
  84. 84.Jingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan, Dahua Lin, and Bo Dai. Towards diverse and natural scene-aware 3d human motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20460–20469, 2022.
  85. 85.Jiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu, and Xiaolong Wang. Synthesizing long-term 3d human motion and interaction in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9401–9411, 2021.
  86. 86.Jingbo Wang, Sijie Yan, Bo Dai, and Dahua Lin. Scene-aware generative network for human motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12206–12215, 2021.
  87. 87.Zhisheng Xiao, Karsten Kreis, and Arash Vahdat. Tackling the generative learning trilemma with denoising diffusion GANs. In International Conference on Learning Representations (ICLR), 2022.
  88. 88.Jiachen Xu, Min Wang, Jingyu Gong, Wentao Liu, Chen Qian, Yuan Xie, and Lizhuang Ma. Exploring versatile prior for human motion via motion frequency guidance. In 2021 International Conference on 3D Vision (3DV), pages 606–616. IEEE, 2021.
  89. 89.Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. Interdiff: Generating 3d human-object interactions with physics-informed diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14928–14940, 2023.
  90. 90.Sirui Xu, Yu-Xiong Wang, and Liangyan Gui. Stochastic multi-person 3d motion forecasting. In The Eleventh International Conference on Learning Representations, 2022.
  91. 91.Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys, 2022.
  92. 92.Zhao Yang, Bing Su, and Ji-Rong Wen. Synthesizing long-term human motions with diffusion models via coherent sampling. In Proceedings of the 31st ACM International Conference on Multimedia, pages 3954–3964, 2023.
  93. 93.Zijie Ye, Haozhe Wu, Jia Jia, Yaohua Bu, Wei Chen, Fanbo Meng, and Yanfeng Wang. Choreonet: Towards music to dance synthesis with choreographic action unit. In Proceedings of the 28th ACM International Conference on Multimedia, pages 744–752, 2020.
  94. 94.Hongwei Yi, Chun-Hao P Huang, Shashank Tripathi, Lea Hering, Justus Thies, and Michael J Black. Mime: Human-aware 3d scene generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12965–12976, 2023.
  95. 95.Xinyu Yi, Yuxiao Zhou, and Feng Xu. Transpose: Real-time 3d human translation and pose estimation with six inertial sensors. ACM Transactions on Graphics (TOG), 40(4):1–13, 2021.
  96. 96.Ye Yuan and Kris Kitani. Dlow: Diversifying latent flows for diverse human motion prediction. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16, pages 346–364. Springer, 2020.
  97. 97.Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. Physdiff: Physics-guided human motion diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16010–16021, 2023.
  98. 98.Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang, Hongwei Zhao, Hongtao Lu, Xi Shen, and Ying Shan. Generating human motion from textual descriptions with discrete representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14730–14740, 2023.
  99. 99.Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondiffuse: Text-driven human motion generation with diffusion model. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
  100. 100.Qinsheng Zhang, Jiaming Song, Xun Huang, Yongxin Chen, and Ming-Yu Liu. Diffcollage: Parallel generation of large content with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10188–10198, 2023.
  101. 101.Yan Zhang, Michael J Black, and Siyu Tang. Perpetual motion: Generating unbounded human motion. arXiv preprint arXiv:2007.13886, 2020.
  102. 102.Yan Zhang and Siyu Tang. The wanderings of odysseus in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20481–20491, 2022.
  103. 103.Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5745–5753, 2019.
  104. 104.Yi Zhou, Zimo Li, Shuangjiu Xiao, Chong He, Zeng Huang, and Hao Li. Auto-conditioned recurrent networks for extended complex human motion synthesis. In International Conference on Learning Representations, 2018.
  105. 105.Yi Zhou, Jingwan Lu, Connelly Barnes, Jimei Yang, Sitao Xiang, et al. Generative tweening: Long-term inbetweening of 3d human motions. arXiv preprint arXiv:2005.08891, 2020.
  106. 106.Wentao Zhu, Xiaoxuan Ma, Dongwoo Ro, Hai Ci, Jinlu Zhang, Jiaxin Shi, Feng Gao, Qi Tian, and Yizhou Wang. Human motion generation: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  107. 107.Wenlin Zhuang, Congyi Wang, Jinxiang Chai, Yangang Wang, Ming Shao, and Siyu Xia. Music2dance: Dancenet for music-driven dance generation. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 18(2):1–21, 2022.

Citation

MLA
Barquero, G., et al. “Seamless Human Motion Composition with Blended Positional Encodings”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 457–69, https://doi.org/10.1109/CVPR52733.2024.00051.
APA
Barquero, G., Escalera, S., & Palmero, C. (2024). Seamless Human Motion Composition with Blended Positional Encodings. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 457–469. https://doi.org/10.1109/CVPR52733.2024.00051
Chicago
Barquero, G., S. Escalera, and C. Palmero. 2024. “Seamless Human Motion Composition with Blended Positional Encodings”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 457–69. https://doi.org/10.1109/CVPR52733.2024.00051.
Harvard
Barquero, G., Escalera, S. and Palmero, C. (2024) “Seamless Human Motion Composition with Blended Positional Encodings”, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 457–469. Available at: https://doi.org/10.1109/CVPR52733.2024.00051.
Vancouver
1. Barquero G, Escalera S, Palmero C (2024) Seamless Human Motion Composition with Blended Positional Encodings. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 457–469

BibTeX

@inproceedings{Barquero_2024, title={Seamless Human Motion Composition with Blended Positional Encodings}, url={http://dx.doi.org/10.1109/CVPR52733.2024.00051}, DOI={10.1109/cvpr52733.2024.00051}, booktitle={2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Barquero, German and Escalera, Sergio and Palmero, Cristina}, year={2024}, month=June, pages={457–469} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE