Executing your Commands via Motion Diffusion in Latent Space
Xin ChenBiao JiangWen LiuZilong HuangBin FuTao ChenGang Yu
Proposes a motion latent diffusion framework that generates realistic conditional human motion two orders of magnitude faster than raw-space diffusion models by operating directly on compact, learned variational autoencoder representations.
Generating realistic 3D human motion from textual prompts, action categories, or unconditioned inputs is vital for digital entertainment, virtual reality, and humanoid robotics. However, existing methods encounter significant hurdles: direct cross-modal alignment often reduces movement diversity, while applying diffusion models directly onto raw motion data results in heavy computational costs, slow processing, and vulnerability to capture noise.
The article demonstrates an effective and efficient conditional human motion generation framework called the Motion Latent-based Diffusion model (MLD). Its primary objective is to evaluate whether executing the denoising diffusion process within a low-dimensional, compressed latent space can achieve superior motion fidelity and diversity while drastically cutting computational overhead.
The approach uses a two-stage architecture evaluated across benchmark datasets, including HumanML3D, KIT, HumanAct12, UESTC, and the AMASS collection. First, a transformer-based variational autoencoder equipped with skip connections compresses variable-length raw motions into a compact, information-dense latent representation. Second, a conditional diffusion model learns the denoising process on this latent representation rather than raw joint coordinates, using pre-trained language models or learnable action embeddings as conditioning signals with classifier-free guidance.
The evaluations reveal four primary findings. First, MLD achieves top-tier text-to-motion generation quality, recording the lowest distribution error (FID of 0.473 on HumanML3D and 0.404 on KIT) and superior prompt matching accuracy. Second, it reduces inference latency to approximately 0.22 seconds per sentence, operating two orders of magnitude faster than prior diffusion methods that required 15 to 25 seconds. Third, the custom variational autoencoder dramatically improves motion reconstruction, lowering position error to 14.7 millimeters compared to over 65 millimeters in baseline architectures. Fourth, the approach achieves state-of-the-art accuracy and diversity in action-to-motion and unconditional generation tasks.
These findings indicate that motion synthesis does not require resource-intensive operations on raw joint sequences. By decoupling universal motion compression from conditional diffusion, organizations can train base autoencoders on vast unannotated datasets while training lighter conditional modules separately. This architecture reduces computing hardware costs, shortens deployment cycles, and enables near-real-time animation pipelines on standard consumer hardware.
Teams developing interactive animation or digital avatar systems should adopt latent-space diffusion to reduce server latency and infrastructure costs. Future technical iterations should focus on generating continuous, non-stop motion beyond fixed maximum dataset sequence lengths and expanding the latent diffusion framework to include coordinated hand, facial, and non-human animal movements.
Confidence in these findings is high due to consistent experimental gains across multiple public benchmarks and tasks. However, users should note that motion duration remains bounded by dataset limits and the framework primarily targets full-body skeletal dynamics without fine-grained finger or facial articulation.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). This paper establishes the foundational Motion Diffusion Model (MDM) operating directly in raw motion space, providing the direct baseline and motivation for MLD's transition to a latent diffusion architecture.
- Paper: Tutorial on Variational Autoencoders, Carl Doersch (2016). This tutorial explains the core principles and mathematical formulation of Variational Autoencoders (VAEs), which MLD uses to compress noisy motion data into a representative latent space.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This foundational work introduces Denoising Diffusion Probabilistic Models (DDPMs), detailing the fundamental denoising and reverse diffusion processes adapted by MLD.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). This work introduces Denoising Diffusion Implicit Models (DDIMs), establishing fast deterministic sampling mechanisms essential for accelerating diffusion inference in latent spaces.
- Paper: DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps, Cheng Lu et al. (2022). This paper formulates fast ODE solvers for diffusion sampling, underpinning the rapid inference capabilities achieved in low-dimensional latent diffusion frameworks.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). This paper provides key conceptual ground for modeling continuous behavioral and temporal sequences holistically using diffusion models.
- Paper: Seamless Human Motion Composition with Blended Positional Encodings, Germán Barquero et al. (2024). This paper extends conditional motion diffusion frameworks to long-sequence generation and smooth motion transitions between consecutive text prompts without explicit stitching.
- Paper: Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation, Jiaming Song et al. (2023). This work introduces plug-and-play loss guidance for diffusion models, demonstrating its application to 3D human motion synthesis under path-following and obstacle-avoidance constraints.
