Executing your Commands via Motion Diffusion in Latent Space

Xin ChenBiao JiangWen LiuZilong HuangBin FuTao ChenGang Yu

article2023CVPR732 citations

Proposes a motion latent diffusion framework that generates realistic conditional human motion two orders of magnitude faster than raw-space diffusion models by operating directly on compact, learned variational autoencoder representations.

Listen

Generating realistic 3D human motion from textual prompts, action categories, or unconditioned inputs is vital for digital entertainment, virtual reality, and humanoid robotics. However, existing methods encounter significant hurdles: direct cross-modal alignment often reduces movement diversity, while applying diffusion models directly onto raw motion data results in heavy computational costs, slow processing, and vulnerability to capture noise.

The article demonstrates an effective and efficient conditional human motion generation framework called the Motion Latent-based Diffusion model (MLD). Its primary objective is to evaluate whether executing the denoising diffusion process within a low-dimensional, compressed latent space can achieve superior motion fidelity and diversity while drastically cutting computational overhead.

The approach uses a two-stage architecture evaluated across benchmark datasets, including HumanML3D, KIT, HumanAct12, UESTC, and the AMASS collection. First, a transformer-based variational autoencoder equipped with skip connections compresses variable-length raw motions into a compact, information-dense latent representation. Second, a conditional diffusion model learns the denoising process on this latent representation rather than raw joint coordinates, using pre-trained language models or learnable action embeddings as conditioning signals with classifier-free guidance.

The evaluations reveal four primary findings. First, MLD achieves top-tier text-to-motion generation quality, recording the lowest distribution error (FID of 0.473 on HumanML3D and 0.404 on KIT) and superior prompt matching accuracy. Second, it reduces inference latency to approximately 0.22 seconds per sentence, operating two orders of magnitude faster than prior diffusion methods that required 15 to 25 seconds. Third, the custom variational autoencoder dramatically improves motion reconstruction, lowering position error to 14.7 millimeters compared to over 65 millimeters in baseline architectures. Fourth, the approach achieves state-of-the-art accuracy and diversity in action-to-motion and unconditional generation tasks.

These findings indicate that motion synthesis does not require resource-intensive operations on raw joint sequences. By decoupling universal motion compression from conditional diffusion, organizations can train base autoencoders on vast unannotated datasets while training lighter conditional modules separately. This architecture reduces computing hardware costs, shortens deployment cycles, and enables near-real-time animation pipelines on standard consumer hardware.

Teams developing interactive animation or digital avatar systems should adopt latent-space diffusion to reduce server latency and infrastructure costs. Future technical iterations should focus on generating continuous, non-stop motion beyond fixed maximum dataset sequence lengths and expanding the latent diffusion framework to include coordinated hand, facial, and non-human animal movements.

Confidence in these findings is high due to consistent experimental gains across multiple public benchmarks and tasks. However, users should note that motion duration remains bounded by dataset limits and the framework primarily targets full-body skeletal dynamics without fine-grained finger or facial articulation.

  • Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). This paper establishes the foundational Motion Diffusion Model (MDM) operating directly in raw motion space, providing the direct baseline and motivation for MLD's transition to a latent diffusion architecture.
  • Paper: Tutorial on Variational Autoencoders, Carl Doersch (2016). This tutorial explains the core principles and mathematical formulation of Variational Autoencoders (VAEs), which MLD uses to compress noisy motion data into a representative latent space.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This foundational work introduces Denoising Diffusion Probabilistic Models (DDPMs), detailing the fundamental denoising and reverse diffusion processes adapted by MLD.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). This work introduces Denoising Diffusion Implicit Models (DDIMs), establishing fast deterministic sampling mechanisms essential for accelerating diffusion inference in latent spaces.
  • Paper: DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps, Cheng Lu et al. (2022). This paper formulates fast ODE solvers for diffusion sampling, underpinning the rapid inference capabilities achieved in low-dimensional latent diffusion frameworks.
  • Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). This paper provides key conceptual ground for modeling continuous behavioral and temporal sequences holistically using diffusion models.
Cover for Executing your Commands via Motion Diffusion in Latent Space

Abstract

We study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse and have a property of quite different distribution from conditional modalities, such as textual descriptors in natural languages, it is hard to learn a probabilistic mapping from the desired conditional modality to the human motion sequences. Besides, the raw motion data from the motion capture system might be redundant in sequences and contain noises; directly modeling the joint distribution over the raw motion sequences and conditional modalities would need a heavy computational overhead and might result in artifacts introduced by the captured noises. To learn a better representation of the various human motion sequences, we first design a powerful Variational AutoEncoder (VAE) and arrive at a representative and low-dimensional latent code for a human motion sequence. Then, instead of using a diffusion model to establish the connections between the raw motion sequences and the conditional inputs, we perform a diffusion process on the motion latent space. Our proposed Motion Latent-based Diffusion model (MLD) could produce vivid motion sequences conforming to the given conditional inputs and substantially reduce the computational overhead in both the training and inference stages. Extensive experiments on various human motion generation tasks demonstrate that our MLD achieves significant improvements over the state-of-the-art methods among extensive human motion generation tasks, with two orders of magnitude faster than previous diffusion models on raw motion sequences.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Motion Representation in Latents
  • 3.2. Motion Latent Diffusion Model
  • 3.3. Conditional Motion Latent Diffusion Model
  • 4. Experiments
  • 4.1. Datasets and Evaluation Metrics
  • 4.2. Implementation Details
  • 4.3. Comparisons on Text-to-motion
  • 4.4. Comparisons on Action-to-motion
  • Unconditioned Motion Generation on Part of AMASS Dataset
  • 4.5. Comparisons on Unconditional Generation
  • 5. Ablation Studies
  • 6. Discussion
  • 7. Acknowledgements
  • References

Knowls

  1. Knowl 1 — Motion Latent-based Diffusion Framework

    model/method

    The Motion Latent-based Diffusion (MLD) framework synthesizes high-quality, diverse 3D human motion sequences conforming to conditional inputs (such as natural language descriptions or action categories) by applying diffusion modeling in a compact, learned motion latent space rather than directly on raw, redundant motion sequences.

    MLD operates in a two-stage training paradigm:

    1. Motion Representation in Latents (Stage 1): A motion Variational Autoencoder (VAE) V={E,D}\mathcal{V} = \{\mathcal{E}, \mathcal{D}\} is trained to compress a raw motion sequence x1:L={xi}i=1Lx^{1:L} = \{x^i\}_{i=1}^L of length LL into a low-dimensional latent representation z=E(x1:L)∈Rn×dz = \mathcal{E}(x^{1:L}) \in \mathbb{R}^{n \times d}, and reconstruct it as x^1:L=D(z)\hat{x}^{1:L} = \mathcal{D}(z). The motion encoder E\mathcal{E} and decoder D\mathcal{D} are built using transformer backbones augmented with UNet-like long skip connections.

    2. Conditional Latent Diffusion (Stage 2): With the trained motion encoder E\mathcal{E} kept frozen, a transformer-based denoising network ϵθ\epsilon_\theta learns to reverse a forward diffusion Markov process that gradually adds Gaussian noise to the latent vector z0=E(x1:L)z_0 = \mathcal{E}(x^{1:L}). Conditioning information cc is projected via a domain encoder τθ(c)\tau_\theta(c) and provided to the denoiser.

    During inference, pure Gaussian noise zT∼N(0,I)z_T \sim \mathcal{N}(0, I) is sampled in the latent space, iteratively denoised over TT reverse diffusion steps by ϵθ(zt,t,τθ(c))\epsilon_\theta(z_t, t, \tau_\theta(c)) to yield a clean predicted latent z^0\hat{z}_0, and finally decoded into full-frame human motion x^1:L=D(z^0)\hat{x}^{1:L} = \mathcal{D}(\hat{z}_0) in a single forward pass through the motion decoder.

  2. Knowl 2 — Transformer-based Motion Variational Autoencoder with Skip Connections

    model/method

    The motion Variational Autoencoder (VAE) V={E,D}\mathcal{V} = \{\mathcal{E}, \mathcal{D}\} maps arbitrary-length motion sequences into a compact latent distribution and reconstructs them into continuous motion.

    Each motion frame xix^i is represented as a feature vector combining 3D joint positions, rotations, linear velocities, and binary foot contact indicators. The motion encoder E\mathcal{E} and decoder D\mathcal{D} both consist of 9 transformer layers with 4 attention heads and long skip connections:

    • Encoder E\mathcal{E}: Takes learnable distribution tokens alongside the sequence of frame features x1:L={xi}i=1Lx^{1:L} = \{x^i\}_{i=1}^L. The output states corresponding to the distribution tokens are projected to Gaussian distribution parameters μ,σ∈Rn×d\mu, \sigma \in \mathbb{R}^{n \times d} (with feature dimension d=256d = 256). Latent code z∈Rn×dz \in \mathbb{R}^{n \times d} is sampled using the reparameterization trick: z=μ+σ⊙ϵz = \mu + \sigma \odot \epsilon, where ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I).

    • Decoder D\mathcal{D}: Operates via a cross-attention mechanism, taking LL zero motion tokens as queries and the latent code z∈Rn×dz \in \mathbb{R}^{n \times d} as memory (keys and values) to output reconstructed motion x^1:L=D(z)\hat{x}^{1:L} = \mathcal{D}(z).

    The VAE is trained on motion reconstruction using the Mean Squared Error (MSE) loss and the Kullback-Leibler (KL) divergence loss against a standard Gaussian prior N(0,I)\mathcal{N}(0, I):

    LVAE=∥x1:L−x^1:L∥22+βKLDKL(q(z∣x1:L)∥N(0,I))\mathcal{L}_{\text{VAE}} = \|x^{1:L} - \hat{x}^{1:L}\|_2^2 + \beta_{\text{KL}} D_{\text{KL}}(q(z \mid x^{1:L}) \parallel \mathcal{N}(0, I))

  3. Knowl 3 — Motion Latent Diffusion Forward Process and Training Objective

    equation

    Given a motion latent code z0=E(x1:L)∈Rn×dz_0 = \mathcal{E}(x^{1:L}) \in \mathbb{R}^{n \times d} extracted from a motion sequence x1:Lx^{1:L} by a frozen VAE encoder E\mathcal{E}, the forward diffusion process adds Gaussian noise across TT discrete timesteps according to a Markov transition kernel:

    q(zt∣zt−1)=N(zt;αtzt−1,(1−αt)I)q(z_t \mid z_{t-1}) = \mathcal{N}\left(z_t; \sqrt{\alpha_t} z_{t-1}, (1 - \alpha_t) I\right)

    where αt=1−βt∈(0,1)\alpha_t = 1 - \beta_t \in (0, 1), and the variance schedule βt\beta_t scales linearly from 8.5×10−48.5 \times 10^{-4} to 0.0120.012 over T=1000T = 1000 steps during training.

    The conditional denoising transformer ϵθ(zt,t,τθ(c))\epsilon_\theta(z_t, t, \tau_\theta(c)) predicts the noise vector ϵ\epsilon added to latent z0z_0 at step tt given condition embedding τθ(c)\tau_\theta(c). The model is optimized using the simple latent mean squared error objective:

    LMLD:=Eϵ∼N(0,I), t, c[∥ϵ−ϵθ(zt,t,τθ(c))∥22]\mathcal{L}_{\text{MLD}} := \mathbb{E}_{\epsilon \sim \mathcal{N}(0, I), \, t, \, c} \left[ \left\| \epsilon - \epsilon_\theta(z_t, t, \tau_\theta(c)) \right\|_2^2 \right]

    For unconditional motion generation, where c=∅c = \emptyset, the objective reduces to:

    LMLDuncond:=Eϵ∼N(0,I), t[∥ϵ−ϵθ(zt,t)∥22]\mathcal{L}_{\text{MLD}}^{\text{uncond}} := \mathbb{E}_{\epsilon \sim \mathcal{N}(0, I), \, t} \left[ \left\| \epsilon - \epsilon_\theta(z_t, t) \right\|_2^2 \right]

  4. Knowl 4 — Condition Embedding and Classifier-Free Guidance for Motion Latent Diffusion

    model/method

    The conditional motion latent diffusion model incorporates multimodal conditions cc into the latent denoiser ϵθ\epsilon_\theta using dedicated condition encoders τθ(c)\tau_\theta(c) and classifier-free guidance:

    • Condition Encoders:

      1. Text Prompts: A textual description w1:N={wi}i=1Nw^{1:N} = \{w^i\}_{i=1}^N is encoded by a frozen CLIP-ViT-L/14 text encoder τθw(w1:N)∈R1×d\tau_\theta^w(w^{1:N}) \in \mathbb{R}^{1 \times d} with d=256d = 256.
      2. Action Categories: An action class label a∈Aa \in \mathcal{A} is mapped to an embedding vector using a learnable lookup embedding τθa(a)∈R1×d\tau_\theta^a(a) \in \mathbb{R}^{1 \times d}.
    • Condition Injection: The condition embedding τθ(c)∈R1×256\tau_\theta(c) \in \mathbb{R}^{1 \times 256} is concatenated directly with the noisy motion latent zt∈R1×256z_t \in \mathbb{R}^{1 \times 256} along the token dimension as input to the transformer denoiser ϵθ\epsilon_\theta, which outperforms cross-attention injection.

    • Classifier-Free Guidance: During training of ϵθ\epsilon_\theta, the condition embedding is randomly replaced by an empty condition token ∅\emptyset with a 10%10\% dropout probability. During reverse inference, the guided noise prediction ϵ^θs\hat{\epsilon}_\theta^s is computed via linear extrapolation with guidance scale s>1s > 1:

    ϵ^θs(zt,t,c)=s ϵθ(zt,t,c)+(1−s) ϵθ(zt,t,∅)\hat{\epsilon}_\theta^s(z_t, t, c) = s \, \epsilon_\theta(z_t, t, c) + (1 - s) \, \epsilon_\theta(z_t, t, \emptyset)

  5. Knowl 5 — Text-to-Motion Benchmark on HumanML3D and KIT Datasets

    data/table

    The table compares MLD with existing text-conditional human motion synthesis methods on the HumanML3D and KIT datasets. Evaluation metrics (mean and 95% confidence intervals from 20 runs) are computed using a pretrained motion-text feature extractor: R-Precision (Top-1, Top-2, Top-3 retrieval accuracy), Fréchet Inception Distance (FID), Multimodal Distance (MM Dist), Diversity (DIV), and MultiModality within identical text prompts (MM). →\rightarrow indicates that values closer to real motion are better.

    Dataset / Methods R Precision ↑\uparrow FID ↓\downarrow MM Dist ↓\downarrow Diversity →\rightarrow
    Top 1 Top 2 Top 3
    HumanML3D
    Real 0.511±.0030.511^{\pm .003} 0.703±.0030.703^{\pm .003} 0.797±.0020.797^{\pm .002} 0.002±.0000.002^{\pm .000} 2.974±.0082.974^{\pm .008} 9.503±.0659.503^{\pm .065}
    Seq2Seq 0.180±.0020.180^{\pm .002} 0.300±.0020.300^{\pm .002} 0.396±.0020.396^{\pm .002} 11.75±.03511.75^{\pm .035} 5.529±.0075.529^{\pm .007} 6.223±.0616.223^{\pm .061}
    JL2P 0.246±.0010.246^{\pm .001} 0.387±.0020.387^{\pm .002} 0.486±.0020.486^{\pm .002} 11.02±.04611.02^{\pm .046} 5.296±.0085.296^{\pm .008} 7.676±.0587.676^{\pm .058}
    T2G 0.165±.0010.165^{\pm .001} 0.267±.0020.267^{\pm .002} 0.345±.0020.345^{\pm .002} 7.664±.0307.664^{\pm .030} 6.030±.0086.030^{\pm .008} 6.409±.0716.409^{\pm .071}
    Hier 0.301±.0020.301^{\pm .002} 0.425±.0020.425^{\pm .002} 0.552±.0040.552^{\pm .004} 6.532±.0246.532^{\pm .024} 5.012±.0185.012^{\pm .018} 8.332±.0428.332^{\pm .042}
    TEMOS 0.424±.0020.424^{\pm .002} 0.612±.0020.612^{\pm .002} 0.722±.0020.722^{\pm .002} 3.734±.0283.734^{\pm .028} 3.703±.0083.703^{\pm .008} 8.973±.0718.973^{\pm .071}
    T2M 0.457±.0020.457^{\pm .002} 0.639±.0030.639^{\pm .003} 0.740±.0030.740^{\pm .003} 1.067±.0021.067^{\pm .002} 3.340±.0083.340^{\pm .008} 9.188±.0029.188^{\pm .002}
    MotionDiffuse 0.491±.0010.491^{\pm .001} 0.681±.0010.681^{\pm .001} 0.782±.0010.782^{\pm .001} 0.630±.0010.630^{\pm .001} 3.113±.0013.113^{\pm .001} 9.410±.0499.410^{\pm .049}
    MDM 0.320±.0050.320^{\pm .005} 0.498±.0040.498^{\pm .004} 0.611±.0070.611^{\pm .007} 0.544±.0440.544^{\pm .044} 5.566±.0275.566^{\pm .027} 9.559±.0869.559^{\pm .086}
    MLD (Ours) 0.481±.0030.481^{\pm .003} 0.673±.0030.673^{\pm .003} 0.772±.0020.772^{\pm .002} 0.473±.013\mathbf{0.473}^{\pm .013} 3.196±.0103.196^{\pm .010} 9.724±.0829.724^{\pm .082}
    KIT
    Real 0.424±.0050.424^{\pm .005} 0.649±.0060.649^{\pm .006} 0.779±.0060.779^{\pm .006} 0.031±.0040.031^{\pm .004} 2.788±.0122.788^{\pm .012} 11.08±.09711.08^{\pm .097}
    Seq2Seq 0.103±.0030.103^{\pm .003} 0.178±.0050.178^{\pm .005} 0.241±.0060.241^{\pm .006} 24.86±.34824.86^{\pm .348} 7.960±.0317.960^{\pm .031} 6.744±.1066.744^{\pm .106}
    T2G 0.156±.0040.156^{\pm .004} 0.255±.0040.255^{\pm .004} 0.338±.0050.338^{\pm .005} 12.12±.18312.12^{\pm .183} 6.964±.0296.964^{\pm .029} 9.334±.0799.334^{\pm .079}
    JL2P 0.221±.0050.221^{\pm .005} 0.373±.0040.373^{\pm .004} 0.483±.0050.483^{\pm .005} 6.545±.0726.545^{\pm .072} 5.147±.0305.147^{\pm .030} 9.073±.1009.073^{\pm .100}
    Hier 0.255±.0060.255^{\pm .006} 0.432±.0070.432^{\pm .007} 0.531±.0070.531^{\pm .007} 5.203±.1075.203^{\pm .107} 4.986±.0274.986^{\pm .027} 9.563±.0729.563^{\pm .072}
    TEMOS 0.353±.0060.353^{\pm .006} 0.561±.0070.561^{\pm .007} 0.687±.0050.687^{\pm .005} 3.717±.0513.717^{\pm .051} 3.417±.0193.417^{\pm .019} 10.84±.10010.84^{\pm .100}
    T2M 0.370±.0050.370^{\pm .005} 0.569±.0070.569^{\pm .007} 0.693±.0070.693^{\pm .007} 2.770±.1092.770^{\pm .109} 3.401±.0083.401^{\pm .008} 10.91±.11910.91^{\pm .119}
    MotionDiffuse 0.417±.0040.417^{\pm .004} 0.621±.0040.621^{\pm .004} 0.739±.0040.739^{\pm .004} 1.954±.0621.954^{\pm .062} 2.958±.0052.958^{\pm .005} 11.10±.14311.10^{\pm .143}
    MDM 0.164±.0040.164^{\pm .004} 0.291±.0040.291^{\pm .004} 0.396±.0040.396^{\pm .004} 0.497±.0210.497^{\pm .021} 9.191±.0229.191^{\pm .022} 10.85±.10910.85^{\pm .109}
    MLD (Ours) 0.390±.0080.390^{\pm .008} 0.609±.0080.609^{\pm .008} 0.734±.0070.734^{\pm .007} 0.404±.027\mathbf{0.404}^{\pm .027} 3.204±.0273.204^{\pm .027} 10.80±.11710.80^{\pm .117}

    MLD achieves the best FID score on both HumanML3D (0.4730.473) and KIT (0.4040.404), significantly outperforming joint latent models like TEMOS and raw-space diffusion models like MDM, while generating diverse motion sequences.

  6. Knowl 6 — Action-to-Motion Synthesis Benchmark on UESTC and HumanAct12

    data/table

    The table compares MLD with existing action-to-motion generation models on the UESTC and HumanAct12 datasets using SMPL-based motion representations. Metrics reported include FID on the training and test splits, classification Accuracy (ACC) evaluated with a pretrained action recognition model, Diversity (DIV), and MultiModality (MM) representing intra-action diversity.

    Dataset / Methods FIDtrain↓\text{FID}_{\text{train}} \downarrow FIDtest↓\text{FID}_{\text{test}} \downarrow ACC ↑\uparrow DIV →\rightarrow MM →\rightarrow
    UESTC
    Real 2.92±.262.92^{\pm .26} 2.79±.292.79^{\pm .29} 0.988±.0010.988^{\pm .001} 33.34±.32033.34^{\pm .320} 14.16±.0614.16^{\pm .06}
    ACTOR 20.5±2.320.5^{\pm 2.3} 23.43±2.2023.43^{\pm 2.20} 0.911±.0030.911^{\pm .003} 31.96±.3331.96^{\pm .33} 14.52±.0914.52^{\pm .09}
    INR 9.55±.069.55^{\pm .06} 15.00±.0915.00^{\pm .09} 0.941±.0010.941^{\pm .001} 31.59±.1931.59^{\pm .19} 14.68±.0714.68^{\pm .07}
    MDM 9.98±1.339.98^{\pm 1.33} 12.81±1.4612.81^{\pm 1.46} 0.950±.0000.950^{\pm .000} 33.02±.2833.02^{\pm .28} 14.26±.1214.26^{\pm .12}
    MLD (Ours) 12.89±.10912.89^{\pm .109} 15.79±.07915.79^{\pm .079} 0.954±.001\mathbf{0.954}^{\pm .001} 33.52±.1433.52^{\pm .14} 13.57±.0613.57^{\pm .06}
    HumanAct12
    Real 0.020±.0100.020^{\pm .010} - 0.997±.0010.997^{\pm .001} 6.850±.0506.850^{\pm .050} 2.450±.0402.450^{\pm .040}
    ACTOR 0.120±.0000.120^{\pm .000} - 0.955±.0080.955^{\pm .008} 6.840±.0306.840^{\pm .030} 2.530±.0202.530^{\pm .020}
    INR 0.088±.0040.088^{\pm .004} - 0.973±.0010.973^{\pm .001} 6.881±.0486.881^{\pm .048} 2.569±.0402.569^{\pm .040}
    MDM 0.100±.0000.100^{\pm .000} - 0.990±.0000.990^{\pm .000} 6.680±.0506.680^{\pm .050} 2.520±.0102.520^{\pm .010}
    MLD (Ours) 0.077±.004\mathbf{0.077}^{\pm .004} - 0.964±.0020.964^{\pm .002} 6.831±.0506.831^{\pm .050} 2.824±.0382.824^{\pm .038}

    MLD achieves state-of-the-art action recognition accuracy (0.9540.954) and diversity (33.5233.52) on UESTC, and the lowest FIDtrain\text{FID}_{\text{train}} (0.0770.077) with high multimodality on HumanAct12.

  7. Knowl 7 — Architectural Ablation of Motion Variational Autoencoder

    data/table

    The table presents an ablation study evaluating the motion Variational Autoencoder V\mathcal{V} across different latent token lengths nn for z∈Rn×256z \in \mathbb{R}^{n \times 256}, the inclusion of skip connections, and varying transformer layer depths on the motion split of HumanML3D. Reconstruction metrics include Mean Per-Joint Position Error (MPJPE in mm), Procrustes-Aligned MPJPE (PAMPJPE in mm), and Acceleration Error (ACCL). Unconditional generation from the VAE latent space prior is evaluated via FID and Diversity (DIV).

    Configuration Reconstruction Generation
    MPJPE ↓\downarrow PAMPJPE ↓\downarrow ACCL ↓\downarrow FID ↓\downarrow DIV →\rightarrow
    Real - - - 0.0020.002 9.5039.503
    VPoser-t 75.675.6 48.648.6 9.39.3 1.4301.430 8.3368.336
    ACTOR 65.365.3 41.041.0 7.07.0 0.3410.341 9.5699.569
    Ours-1 (z∈R1×256z \in \mathbb{R}^{1 \times 256}) 54.454.4 41.641.6 8.38.3 0.2470.247 9.6309.630
    Ours-2 (z∈R2×256z \in \mathbb{R}^{2 \times 256}) 51.851.8 37.837.8 8.38.3 0.1660.166 9.6269.626
    Ours-5 (z∈R5×256z \in \mathbb{R}^{5 \times 256}) 24.324.3 14.714.7 5.85.8 0.0430.043 9.5939.593
    Ours-7 (z∈R7×256z \in \mathbb{R}^{7 \times 256}, skip, 9 layers) 14.7\mathbf{14.7} 8.9\mathbf{8.9} 5.1\mathbf{5.1} 0.017\mathbf{0.017} 9.5549.554
    Ours-10 (z∈R10×256z \in \mathbb{R}^{10 \times 256}) 17.317.3 11.511.5 5.85.8 0.0250.025 9.5899.589
    Ours-7 (w/ skip) 14.714.7 8.98.9 5.15.1 0.0170.017 9.5549.554
    Ours-7 (w/o skip) 18.518.5 10.410.4 5.65.6 0.0270.027 9.5289.528
    Ours-7 (7 layers) 16.016.0 10.210.2 5.35.3 0.0220.022 9.5939.593
    Ours-7 (9 layers) 14.714.7 8.98.9 5.15.1 0.0170.017 9.5549.554
    Ours-7 (11 layers) 17.217.2 11.211.2 5.45.4 0.0210.021 9.5339.533

    Setting n=7n = 7 yields the best motion reconstruction quality (MPJPE of 14.7 mm14.7\text{ mm} and PAMPJPE of 8.9 mm8.9\text{ mm}). Adding long skip connections noticeably improves reconstruction accuracy (reducing MPJPE from 18.5 mm18.5\text{ mm} to 14.7 mm14.7\text{ mm}) and generation FID.

  8. Knowl 8 — Ablation of Latent Dimensions and Denoiser Components in Text-to-Motion Diffusion

    data/table

    The table presents an ablation study of the Motion Latent Diffusion denoiser ϵθ\epsilon_\theta on HumanML3D text-to-motion synthesis, examining the effect of latent shape z∈Rn×256z \in \mathbb{R}^{n \times 256}, condition injection mode (concatenation vs. cross-attention), skip connections, and transformer depth.

    Model Variant R Precision Top-3 ↑\uparrow FID ↓\downarrow MM Dist ↓\downarrow Diversity →\rightarrow MModality ↑\uparrow
    Real 0.797±.0020.797^{\pm .002} 0.002±.0000.002^{\pm .000} 2.974±.0082.974^{\pm .008} 9.503±.0659.503^{\pm .065} -
    MLD-1 (z∈R1×256z \in \mathbb{R}^{1 \times 256}) 0.772±.002\mathbf{0.772}^{\pm .002} 0.473±.013\mathbf{0.473}^{\pm .013} 3.196±.010\mathbf{3.196}^{\pm .010} 9.724±.0829.724^{\pm .082} 2.413±.0792.413^{\pm .079}
    MLD-2 (z∈R2×256z \in \mathbb{R}^{2 \times 256}) 0.727±.0030.727^{\pm .003} 0.585±.0150.585^{\pm .015} 3.448±.0113.448^{\pm .011} 9.084±.0819.084^{\pm .081} 2.725±.0932.725^{\pm .093}
    MLD-5 (z∈R5×256z \in \mathbb{R}^{5 \times 256}) 0.722±.0030.722^{\pm .003} 1.554±.0191.554^{\pm .019} 3.511±.0083.511^{\pm .008} 8.424±.0818.424^{\pm .081} 2.542±.0802.542^{\pm .080}
    MLD-7 (z∈R7×256z \in \mathbb{R}^{7 \times 256}) 0.731±.0020.731^{\pm .002} 1.011±.0191.011^{\pm .019} 3.415±.0083.415^{\pm .008} 8.736±.0648.736^{\pm .064} 2.463±.0892.463^{\pm .089}
    MLD-10 (z∈R10×256z \in \mathbb{R}^{10 \times 256}) 0.703±.0030.703^{\pm .003} 1.716±.0271.716^{\pm .027} 3.616±.0123.616^{\pm .012} 8.606±.0678.606^{\pm .067} 2.604±.0872.604^{\pm .087}
    MLD-1 (cross-attention) 0.592±.0040.592^{\pm .004} 1.922±.0411.922^{\pm .041} 4.480±.0154.480^{\pm .015} 8.598±.0888.598^{\pm .088} 3.768±.1263.768^{\pm .126}
    MLD-1 (concatenation) 0.772±.0020.772^{\pm .002} 0.473±.0130.473^{\pm .013} 3.196±.0103.196^{\pm .010} 9.724±.0829.724^{\pm .082} 2.413±.0792.413^{\pm .079}
    MLD-1 (w/o skip) 0.749±.0030.749^{\pm .003} 0.784±.0150.784^{\pm .015} 3.363±.0103.363^{\pm .010} 9.568±.0939.568^{\pm .093} 2.597±.0982.597^{\pm .098}
    MLD-1 (w/ skip) 0.772±.0020.772^{\pm .002} 0.473±.0130.473^{\pm .013} 3.196±.0103.196^{\pm .010} 9.724±.0829.724^{\pm .082} 2.413±.0792.413^{\pm .079}
    MLD-1 (5 layers) 0.760±.0020.760^{\pm .002} 0.314±.0100.314^{\pm .010} 3.259±.0093.259^{\pm .009} 9.706±.0729.706^{\pm .072} 2.635±.0852.635^{\pm .085}
    MLD-1 (7 layers) 0.771±.0030.771^{\pm .003} 0.349±.0120.349^{\pm .012} 3.199±.0123.199^{\pm .012} 9.624±.0629.624^{\pm .062} 2.504±.0882.504^{\pm .088}
    MLD-1 (9 layers) 0.772±.0020.772^{\pm .002} 0.473±.0130.473^{\pm .013} 3.196±.0103.196^{\pm .010} 9.724±.0829.724^{\pm .082} 2.413±.0792.413^{\pm .079}
    MLD-1 (11 layers) 0.771±.0030.771^{\pm .003} 0.402±.0110.402^{\pm .011} 3.203±.0133.203^{\pm .013} 9.876±.0889.876^{\pm .088} 2.478±.0762.478^{\pm .076}

    Key observations:

    1. A 1-token latent vector z∈R1×256z \in \mathbb{R}^{1 \times 256} (MLD-1) achieves the strongest text matching (R Precision Top-3 of 0.7720.772) and lowest FID among latent shapes for diffusion.
    2. Conditioning injection via token concatenation substantially outperforms cross-attention (FID 0.4730.473 vs 1.9221.922).
    3. Incorporating skip connections into the denoiser lowers FID from 0.7840.784 to 0.4730.473.
  9. Knowl 9 — Inference Efficiency of Motion Latent Diffusion vs. Raw-Space Diffusion Models

    empirical result

    Operating the diffusion process in the compact motion latent space R1×256\mathbb{R}^{1 \times 256} rather than across full-length raw motion sequences (L×DjointsL \times D_{\text{joints}}) achieves a two orders of magnitude speedup during inference, while concurrently improving generation quality.

    Average Inference Time per Sentence (AITS, measured in seconds on a single NVIDIA Tesla V100 GPU excluding data loading) and corresponding text-to-motion FID on HumanML3D:

    • MDM (Raw-space diffusion): AITS = 24.74 s24.74\text{ s}, FID = 0.5440.544
    • MotionDiffuse (Raw-space diffusion): AITS = 14.74 s14.74\text{ s}, FID = 0.6300.630
    • T2M (Deterministic model): AITS = 0.038 s0.038\text{ s}, FID = 1.0671.067
    • TEMOS (Joint VAE): AITS = 0.017 s0.017\text{ s}, FID = 3.7343.734
    • MLD-1 (Ours): AITS = 0.217 s0.217\text{ s}, FID = 0.4730.473

    Compared to MDM (24.74 s24.74\text{ s}), MLD-1 requires only 0.217 s0.217\text{ s} per sentence (over 114×114\times faster) because the T=50T = 50 denoising steps execute on low-dimensional latents rather than sequence-length joint tensors, followed by a single-pass decoder invocation.

  10. Knowl 10 — Limitations of Motion Latent-based Diffusion

    limitation

    The MLD model is subject to two main limitations:

    1. Bounded Motion Duration: Although MLD can generate motions of variable frame lengths, the maximum generation duration is bounded by the maximum sequence length present in the training datasets. The model cannot synthesize non-stop or infinite continuous human motions with guaranteed long-term temporal consistency.

    2. Articulated Body Scope: The representation is restricted to full-body skeletal articulation and does not model fine-grained motion components such as facial expressions, finger/hand manipulations, or non-human animal kinematics.

Coverage note — Omitted introductory text, broad surveys of related work (image generation, robotics, motion capture hardware), and qualitative rendering samples from figures/user studies.

References

  1. 1.Hyemin Ahn, Timothy Ha, Yunho Choi, Hwiyeon Yoo, and Songhwai Oh. Text2action: Generative adversarial synthesis from language to action. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 5915–5920. IEEE, 2018.
  2. 2.Chaitanya Ahuja and Louis-Philippe Morency. Language2pose: Natural language grounded pose forecasting. In 2019 International Conference on 3D Vision (3DV), pages 719–728. IEEE, 2019.
  3. 3.Martin Arjovsky, Soumith Chintala, and Leon Bottou. Wasserstein generative adversarial networks. In International conference on machine learning, pages 214–223. PMLR, 2017.
  4. 4.Fan Bao, Chongxuan Li, Yue Cao, and Jun Zhu. All are worth words: a vit backbone for score-based diffusion models. arXiv preprint arXiv:2209.12152, 2022.
  5. 5.Uttaran Bhattacharya, Nicholas Rewkowski, Abhishek Banerjee, Pooja Guhan, Aniket Bera, and Dinesh Manocha. Text2gestures: A transformer-based network for generating emotive body gestures for virtual agents. In 2021 IEEE Virtual Reality and 3D User Interfaces (VR), pages 1–10. IEEE, 2021.
  6. 6.Xuan Cao, Zhang Chen, Anpei Chen, Xin Chen, Shiying Li, and Jingyi Yu. Sparse photometric 3d face reconstruction guided by morphable models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4635–4644, 2018.
  7. 7.Pablo Cervantes, Yusuke Sekikawa, Ikuro Sato, and Koichi Shinoda. Implicit neural representations for variable length human motion generation. In European Conference on Computer Vision, pages 356–372. Springer, 2022.
  8. 8.Xin Chen, Anqi Pang, Wei Yang, Yuexin Ma, Lan Xu, and Jingyi Yu. Sportscap: Monocular 3d human motion capture and fine-grained understanding in challenging sports videos. International Journal of Computer Vision, 129(10):2846–2864, 2021.
  9. 9.Xin Chen, Zhuo Su, Lingbo Yang, Pei Cheng, Lan Xu, Bin Fu, and Gang Yu. Learning variational motion prior for video-based motion capture. arXiv preprint arXiv:2210.15134, 2022.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  11. 11.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  12. 12.Yinglin Duan, Tianyang Shi, Zhengxia Zou, Yenan Lin, Zhehui Qian, Bohan Zhang, and Yi Yuan. Single-shot motion completion with transformer. arXiv preprint arXiv:2103.00776, 2021.
  13. 13.Anindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt, and Philipp Slusallek. Synthesis of compositional animations from textual descriptions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1396–1406, 2021.
  14. 14.John C Gower. Generalized procrustes analysis. Psychometrika, 40(1):33–51, 1975.
  15. 15.Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. Improved training of wasserstein gans. Advances in neural information processing systems, 30, 2017.
  16. 16.Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5152–5161, June 2022.
  17. 17.Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5152–5161, June 2022.
  18. 18.Chuan Guo, Xinxin Zuo, Sen Wang, and Li Cheng. Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts. In ECCV, 2022.
  19. 19.Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng. Action2motion: Conditioned generation of 3d human motions. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2021–2029, 2020.
  20. 20.Felix G Harvey, Mike Yurick, Derek Nowrouzezahrai, and Christopher Pal. Robust motion in-betweening. ACM Transactions on Graphics (TOG), 39(4):60–1, 2020.
  21. 21.Yannan He, Anqi Pang, Xin Chen, Han Liang, Minye Wu, Yuexin Ma, and Lan Xu. Challencap: Monocular 3d capture of challenging human performances using multi-modal references. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11400–11411, 2021.
  22. 22.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022.
  23. 23.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  24. 24.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  25. 25.Daniel Holden, Jun Saito, and Taku Komura. A deep learning framework for character motion synthesis and editing. ACM Transactions on Graphics (TOG), 35(4):1–11, 2016.
  26. 26.Yanli Ji, Feixiang Xu, Yang Yang, Fumin Shen, Heng Tao Shen, and Wei-Shi Zheng. A large-scale rgb-d database for arbitrary-view human action recognition. In Proceedings of the 26th ACM international Conference on Multimedia, pages 1510–1518, 2018.
  27. 27.Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics (TOG), 36(4):1–12, 2017.
  28. 28.Jihoon Kim, Jiseob Kim, and Sungjoon Choi. Flame: Freeform language-based motion synthesis & editing. arXiv preprint arXiv:2209.00349, 2022.
  29. 29.Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. Advances in neural information processing systems, 34:21696–21707, 2021.
  30. 30.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  31. 31.Muhammed Kocabas, Nikos Athanasiou, and Michael J. Black. Vibe: Video inference for human body pose and shape estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  32. 32.Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. Dancing to music. Advances in neural information processing systems, 32, 2019.
  33. 33.Buyu Li, Yongchi Zhao, Shi Zhelun, and Lu Sheng. Danceformer: Music conditioned 3d dance generation with parametric motion transformer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 1272–1279, 2022.
  34. 34.Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. Ai choreographer: Music conditioned 3d dance generation with aist++. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13401–13412, 2021.
  35. 35.Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B Hashimoto. Diffusion-lm improves controllable text generation. 2022.
  36. 36.Yuwei Li, Minye Wu, Yuyao Zhang, Lan Xu, and Jingyi Yu. Piano: A parametric hand bone model from magnetic resonance imaging. arXiv preprint arXiv:2106.10893, 2021.
  37. 37.Yuwei Li, Longwen Zhang, Zesong Qiu, Yingwenqi Jiang, Nianyi Li, Yuexin Ma, Yuyao Zhang, Lan Xu, and Jingyi Yu. Nimble: a non-rigid hand model with bones and muscles. ACM Transactions on Graphics (TOG), 41(4):1–16, 2022.
  38. 38.Xiao Lin and Mohamed R Amer. Human motion modeling using dvgans. arXiv preprint arXiv:1804.10652, 2018.
  39. 39.Hung Yu Ling, Fabio Zinno, George Cheng, and Michiel van de Panne. Character controllers using motion vaes. ACM Trans. Graph., 39(4), 2020.
  40. 40.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. Smpl: A skinned multi-person linear model. ACM Trans. Graph., 34(6):248:1–248:16, Oct. 2015.
  41. 41.Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. Amass: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019.
  42. 42.Christian Mandery, Omer Terlemez, Martin Do, Nikolaus Vahrenkamp, and Tamim Asfour. The kit whole-body human motion database. In 2015 International Conference on Advanced Robotics (ICAR), pages 329–336. IEEE, 2015.
  43. 43.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019.
  44. 44.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 10975–10985, June 2019.
  45. 45.Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa. Amp: Adversarial motion priors for stylized physics-based character control. ACM Trans. Graph., 40(4), July 2021.
  46. 46.Mathis Petrovich, Michael J. Black, and Gul Varol. Action-conditioned 3D human motion synthesis with transformer VAE. In International Conference on Computer Vision (ICCV), 2021.
  47. 47.Mathis Petrovich, Michael J. Black, and Gul Varol. TEMOS: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision (ECCV), 2022.
  48. 48.Matthias Plappert, Christian Mandery, and Tamim Asfour. The kit motion-language dataset. Big Data, 4(4):236–252, dec 2016.
  49. 49.Matthias Plappert, Christian Mandery, and Tamim Asfour. Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks. Robotics and Autonomous Systems, 109:13–26, 2018.
  50. 50.Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J. Black. BABEL: Bodies, action and behavior with english labels. In Proceedings IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pages 722–731, June 2021.
  51. 51.Sigal Raab, Inbal Leibovitch, Peizhuo Li, Kfir Aberman, Olga Sorkine-Hornung, and Daniel Cohen-Or. Modi: Unconditional motion synthesis from diverse data. arXiv preprint arXiv:2206.08010, 2022.
  52. 52.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  53. 53.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  54. 54.Kashif Rasul, Calvin Seward, Ingmar Schuster, and Roland Vollgraf. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. In International Conference on Machine Learning, pages 8857–8868. PMLR, 2021.
  55. 55.Davis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang, Srinath Sridhar, and Leonidas J. Guibas. Humor: 3d human motion model for robust pose estimation. In International Conference on Computer Vision (ICCV), 2021.
  56. 56.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and BjArn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  57. 57.Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6), Nov. 2017.
  58. 58.Javier Romero, Dimitrios Tzionas, and Michael J Black. Embodied hands: Modeling and capturing hands and bodies together. arXiv preprint arXiv:2201.02610, 2022.
  59. 59.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  60. 60.Nadine Rueegg, Silvia Zuffi, Konrad Schindler, and Michael J. Black. BARC: Learning to regress 3D dog shape from images by exploiting breed information. In IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pages 3876–3884, June 2022.
  61. 61.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  62. 62.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  63. 63.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  64. 64.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  65. 65.Sebastian Starke, Ian Mason, and Taku Komura. Deepphase: periodic autoencoders for learning motion phase manifolds. ACM Transactions on Graphics (TOG), 41(4):1–13, 2022.
  66. 66.Sebastian Starke, He Zhang, Taku Komura, and Jun Saito. Neural state machine for character-scene interactions. ACM Trans. Graph., 38(6):209–1, 2019.
  67. 67.Omer Terlemez, Stefan Ulbrich, Christian Mandery, Martin Do, Nikolaus Vahrenkamp, and Tamim Asfour. Master motor map (mmm)—framework and toolkit for capturing, representing, and reproducing human motion on humanoid robots. In 2014 IEEE-RAS International Conference on Humanoid Robots, pages 894–901. IEEE, 2014.
  68. 68.Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. arXiv preprint arXiv:2203.08063, 2022.
  69. 69.Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Amit H Bermano, and Daniel Cohen-Or. Human motion diffusion model. arXiv preprint arXiv:2209.14916, 2022.
  70. 70.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  71. 71.Timo von Marcard, Roberto Henschel, Michael Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In European Conference on Computer Vision (ECCV), sep 2018.
  72. 72.Ziniu Wan, Zhengjia Li, Maoqing Tian, Jianbo Liu, Shuai Yi, and Hongsheng Li. Encoder-decoder with multi-level attention for 3d human shape and pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13033–13042, 2021.
  73. 73.Minkai Xu, Lantao Yu, Yang Song, Chence Shi, Stefano Ermon, and Jian Tang. Geodiff: A geometric diffusion model for molecular conformation generation. In International Conference on Learning Representations, 2022.
  74. 74.Sijie Yan, Zhizhong Li, Yuanjun Xiong, Huahan Yan, and Dahua Lin. Convolutional sequence generation for skeleton-based action synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4394–4402, 2019.
  75. 75.Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001, 2022.
  76. 76.Yan Zhang, Michael J Black, and Siyu Tang. Perpetual motion: Generating unbounded human motion. arXiv preprint arXiv:2007.13886, 2020.
  77. 77.Yan Zhang, Michael J Black, and Siyu Tang. We are more than our joints: Predicting how 3d bodies move. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3372–3382, 2021.
  78. 78.Rui Zhao, Hui Su, and Qiang Ji. Bayesian adversarial human motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6225–6234, 2020.
  79. 79.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 5826–5835, October 2021.
  80. 80.Silvia Zuffi, Angjoo Kanazawa, and Michael J. Black. Lions and tigers and bears: Capturing non-rigid, 3D, articulated shape from images. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3955–3963. IEEE Computer Society, 2018.

Citation

MLA
Chen, X., et al. “Executing Your Commands via Motion Diffusion in Latent Space”. arXiv, 2022, http://arxiv.org/abs/2212.04048v3.
APA
Chen, X., Jiang, B., Liu, W., Huang, Z., Fu, B., Chen, T., Yu, J., & Yu, G. (2022). Executing your Commands via Motion Diffusion in Latent Space. arXiv. http://arxiv.org/abs/2212.04048v3
Chicago
Chen, X., B. Jiang, W. Liu, et al. 2022. “Executing Your Commands via Motion Diffusion in Latent Space”. arXiv. http://arxiv.org/abs/2212.04048v3.
Harvard
Chen, X. et al. (2022) “Executing your Commands via Motion Diffusion in Latent Space”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.04048v3.
Vancouver
1. Chen X, Jiang B, Liu W, Huang Z, Fu B, Chen T, Yu J, Yu G (2022) Executing your Commands via Motion Diffusion in Latent Space. arXiv

BibTeX

@article{chen2022executing,
  title = {Executing your Commands via Motion Diffusion in Latent Space},
  author = {Chen, Xin and Jiang, Biao and Liu, Wen and Huang, Zilong and Fu, Bin and Chen, Tao and Yu, Jingyi and Yu, Gang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.04048v3},
  eprint = {2212.04048}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE