CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control

Guy TevetSigal RaabSetareh CohanDaniele RedaZhengyi LuoXue Bin PengAmit Haim BermanoMichiel van de Panne

article2025ICLR127 citations

Develops a closed-loop framework that integrates real-time motion diffusion planning with reinforcement learning control, enabling physically simulated characters to execute complex sequences of diverse tasks directly from natural language prompts.

Listen

In computer animation, virtual environments, and robotics simulation, generating realistic human movements that seamlessly respond to natural language instructions and physical surroundings remains a fundamental challenge. Traditional data-driven generative models can synthesize a wide variety of human motions from simple text prompts, but they frequently generate physical artifacts such as floating, sliding feet, and unnatural object penetration. Conversely, physics-based reinforcement learning methods enforce realistic physical contacts and balance, yet they typically require complex, labor-intensive reward engineering tailored to individual tasks. The article demonstrates a unified framework called CLoSD (Closing the Loop between Simulation and Diffusion) to evaluate whether pairing a real-time motion diffusion planner with a physics-based reinforcement learning controller in a continuous feedback loop can achieve robust, multi-task character control.

The authors developed an integrated system where a lightweight generative model, termed the Diffusion Planner, creates short-horizon motion plans conditioned on text prompts and 3D target coordinates. This planner operates autoregressively and achieves high computational efficiency, producing 40-frame motion plans at approximately 3,500 frames per second—roughly 175 times faster than real time. A universal reinforcement learning controller executes these planned trajectories inside a physics simulator and feeds the resulting physical state back into the planner to maintain real-time responsiveness. To ensure stability during physical contacts, the tracking controller was fine-tuned in a closed loop across multiple tasks simultaneously using standard reinforcement learning objectives without task-specific reward design.

The experimental findings show that the closed-loop architecture outperforms existing state-of-the-art methods across diverse interactive tasks. In multi-task evaluations, the framework achieved a 100% success rate in navigation, 90% in object striking, 88% in sitting down on furniture, and 98% in standing up. In comparison, prior leading multi-task controllers achieved only a 2% success rate in striking and 8% in standing up because they lacked semantic motion understanding. Furthermore, disabling the closed feedback loop caused the sitting success rate to plunge from 88% to 19% and the get-up success rate from 98% to 23%, confirming that continuous environment feedback is essential. On standard text-to-motion benchmarks, the system matched or exceeded the fidelity of competing physics-based controllers while virtually eliminating physical flaws such as foot skating and body penetration.

These results indicate that generative models can serve as general-purpose kinematic planners for physical controllers, eliminating the need to train bespoke policies for separate movement skills. This offers practical benefits for interactive media, gaming, and virtual avatar development by lowering engineering timelines and reducing runtime system complexity. Decision-makers should consider piloting this architecture for interactive character systems requiring flexible, prompt-driven behaviors. However, because the current design relies on designated joint target coordinates rather than direct vision or elevation maps, future work should integrate visual perception and explore adaptive planning horizons to manage more complex, cluttered environments.

arXiv: 2410.03441
Cover for CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control

Abstract

Motion diffusion models and Reinforcement Learning (RL) based control for physics-based simulations have complementary strengths for human motion generation. The former is capable of generating a wide variety of motions, adhering to intuitive control such as text, while the latter offers physically plausible motion and direct interaction with the environment. In this work, we present a method that combines their respective strengths. CLoSD is a text-driven RL physics-based controller, guided by diffusion generation for various tasks. Our key insight is that motion diffusion can serve as an on-the-fly universal planner for a robust RL controller. To this end, CLoSD maintains a closed-loop interaction between two modules -- a Diffusion Planner (DiP), and a tracking controller. DiP is a fast-responding autoregressive diffusion model, controlled by textual prompts and target locations, and the controller is a simple and robust motion imitator that continuously receives motion plans from DiP and provides feedback from the environment. CLoSD is capable of seamlessly performing a sequence of different tasks, including navigation to a goal location, striking an object with a hand or foot as specified in a text prompt, sitting down, and getting up. this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 CLoSD
  • 4.1 Motion diffusion planner
  • 4.2 Universal Tracking Policy
  • 4.3 Closing the loop
  • 4.4 High-level planning using a state machine
  • 5 Experiments
  • 5.1 Implementation details
  • 5.2 Task evaluation
  • 5.3 Text-to-motion evaluation
  • 6 Conclusions and Limitations
  • References

Knowls

  1. Knowl 1 — CLoSD Closed-Loop Control Architecture

    model/method

    CLoSD (Closing the Loop between Simulation and Diffusion) is a multi-task character control system that integrates an autoregressive kinematic Diffusion Planner (DiP) with a physics-based Reinforcement Learning (RL) tracking controller in a closed feedback loop.

    The framework operates iteratively as follows:

    1. Kinematic Planning: DiP generates a future kinematic motion sequence xpredx^{\text{pred}} spanning NgN_g frames conditioned on a user-provided text instruction, a target joint location/heading constraint, and an observed history prefix xprefixx^{\text{prefix}} of NpN_p frames.
    2. Coordinate Transformation: The predicted trajectory xpredx^{\text{pred}} is transformed from relative HumanML3D motion coordinates to world-space global coordinates xref=R2G(xpred)x^{\text{ref}} = R2G(x^{\text{pred}}).
    3. Tracking in Simulation: A physics-based tracking policy πPHC\pi_{\text{PHC}} receives xrefx^{\text{ref}} frame-by-frame, outputting Proportional-Derivative (PD) targets to simulated humanoid actuators to produce physically valid motion and contact dynamics xsimx^{\text{sim}}.
    4. Feedback Loop Closure: The actual realized physics trajectory of the last NpN_p frames xsim[−Np:]x^{\text{sim}}[-N_p:] is mapped back to the relative representation xprefix=G2R(xsim[−Np:])x^{\text{prefix}} = G2R(x^{\text{sim}}[-N_p:]) and fed back into DiP as the next prefix, enabling continuous replanning in response to simulation dynamics and environmental interactions.
  2. Knowl 2 — Diffusion Planner (DiP) Architecture and Sampling

    model/method

    The Diffusion Planner (DiP) is a real-time, autoregressive motion diffusion model designed for kinematic motion planning. DiP employs an 8-layer transformer decoder architecture with a latent dimension of d=512d = 512 and 4 attention heads.

    At each diffusion timestep t∈[0,T]t \in [0, T]:

    • The input tokens consist of the concatenation of the clean prefix motion sequence xprefix∈RNp×Fx^{\text{prefix}} \in \mathbb{R}^{N_p \times F} (fixed length Np=20N_p = 20) and the noisy motion plan xtpred∈RNg×Fx_t^{\text{pred}} \in \mathbb{R}^{N_g \times F} (fixed length Ng=40N_g = 40), summed with positional embeddings.
    • Conditioning tokens are constructed and fed into each transformer decoder layer via cross-attention. The text prompt is encoded using a frozen DistilBERT model and projected to Ctext∈RNtokens×dC_{\text{text}} \in \mathbb{R}^{N_{\text{tokens}} \times d}. The diffusion step tt and adaptive target condition are mapped through separate shallow fully connected networks to Ct∈RdC_t \in \mathbb{R}^d and Ctarget∈RdC_{\text{target}} \in \mathbb{R}^d, forming the condition sequence (Ct+Ctarget,Ctext)∈R(Ntokens+1)×d(C_t + C_{\text{target}}, C_{\text{text}}) \in \mathbb{R}^{(N_{\text{tokens}}+1) \times d}.
    • The transformer outputs predictions for all frames; the prefix outputs are discarded, and the remaining NgN_g frames directly predict the clean motion x^0pred\hat{x}_0^{\text{pred}}.

    Inference uses DDPM iterative denoising over T=10T = 10 diffusion steps. Generating a 40-frame (2-second) reference plan takes 11.4 ms on an NVIDIA RTX 3090 GPU, operating at 3,500 fps (175×175\times faster than real-time).

  3. Knowl 3 — Adaptive Target Conditioning and Geometric Target Loss

    equation

    To dynamically control arbitrary character joints and body orientations across different tasks, DiP employs adaptive target conditioning combined with a geometric target loss.

    The target condition tuple is defined as ({cj}j∈J,{vj}j∈J,cθ,vθ)(\{c_j\}_{j \in J}, \{v_j\}_{j \in J}, c_\theta, v_\theta), where JJ is the set of selectable joints (including end-effectors and the pelvis), cj∈R3c_j \in \mathbb{R}^3 is the desired global 3D position of joint jj at the final predicted frame x^0[Ng]\hat{x}_0[N_g], vj∈{0,1}v_j \in \{0, 1\} is a boolean validity indicator for joint jj, cθ∈Rc_\theta \in \mathbb{R} is the target body heading angle in the horizontal XYXY-plane, and vθ∈{0,1}v_\theta \in \{0, 1\} is the heading validity indicator.

    The geometric target loss Ltarget\mathcal{L}_{\text{target}} enforces target reaching directly on the predicted clean motion x^0\hat{x}_0 at each diffusion step:

    Ltarget=∑j∈Jvj∥R2G(x^0[Ng])j−cj∥22+vθ∥R2G(x^0[Ng])θ⊖cθ∥22\mathcal{L}_{\text{target}} = \sum_{j \in J} v_j \|R2G(\hat{x}_0[N_g])_j - c_j\|_2^2 + v_\theta \|R2G(\hat{x}_0[N_g])_\theta \ominus c_\theta\|_2^2

    where R2GR2G is the differentiable relative-to-global coordinate converter and ⊖\ominus denotes angular difference.

    The total training objective is:

    L=Lsimple+λtargetLtarget\mathcal{L} = \mathcal{L}_{\text{simple}} + \lambda_{\text{target}} \mathcal{L}_{\text{target}}

    where Lsimple=Ex0∼p(x0∣c),t∼[1,T][∥x0−x^0∥22]\mathcal{L}_{\text{simple}} = \mathbb{E}_{x_0 \sim p(x_0|c), t \sim [1, T]} [\|x_0 - \hat{x}_0\|_2^2] is the standard DDPM x0x_0-prediction loss.

  4. Knowl 4 — Closed-Loop Multi-Task Policy Fine-Tuning

    model/method

    The execution module in CLoSD is a single-primitive, joints-only tracking policy πPHC\pi_{\text{PHC}} based on Perpetual Humanoid Control (PHC). At simulation step nn, πPHC\pi_{\text{PHC}} receives the keypoint-only state s[n]=(xref[n]−xsim[n],xref[n])s[n] = (x^{\text{ref}}[n] - x^{\text{sim}}[n], x^{\text{ref}}[n]) and outputs target angles for proportional-derivative (PD) joint controllers.

    To bridge the gap between idealized motion capture data and kinematic diffusion plans involving physical object interactions, πPHC\pi_{\text{PHC}} is fine-tuned in a closed-loop setting with DiP in-the-loop:

    1. DiP weights are frozen while πPHC\pi_{\text{PHC}} is fine-tuned using Proximal Policy Optimization (PPO) in Isaac Gym across 3,072 parallel environments for 4,000 epochs (after 62,000 pretraining epochs on the AMASS dataset).
    2. Across parallel environments, tasks (such as sitting, getting up, goal navigation, and striking) are sampled simultaneously alongside perturbed object locations and orientations.
    3. Training utilizes standard PHC tracking rewards (the negative exponent of tracking error, Adversarial Motion Prior adversarial reward, and energy penalty) and reset conditions without task-specific reward engineering or task-specific controllers.
  5. Knowl 5 — Dual Motion Representations and Inter-Domain Conversion

    definition

    CLoSD maintains two distinct motion representations to suit diffusion planning and physics simulation:

    • Kinematic HumanML3D Representation (xx): Used by the Diffusion Planner. Each frame nn is defined by a feature vector x[n]=(r˙a,r˙x,r˙z,ry,jp,jr,jv,f)∈RFx[n] = (\dot{r}^a, \dot{r}^x, \dot{r}^z, r^y, j^p, j^r, j^v, f) \in \mathbb{R}^F, where r˙a∈R\dot{r}^a \in \mathbb{R} is the root angular velocity around the vertical axis; r˙x,r˙z∈R\dot{r}^x, \dot{r}^z \in \mathbb{R} are horizontal root linear velocities; ry∈Rr^y \in \mathbb{R} is the root height; jp∈R3(J−1)j^p \in \mathbb{R}^{3(J-1)}, jr∈R6(J−1)j^r \in \mathbb{R}^{6(J-1)}, and jv∈R3Jj^v \in \mathbb{R}^{3J} are local joint positions, 6D rotations, and local velocities relative to the root; and f∈R4f \in \mathbb{R}^4 are binary foot contact labels.
    • Physics State Representation (xsimx^{\text{sim}}): Used by the tracking policy in simulation. Each state is defined by xsim[n]=(jgp,jgv)∈R6Jx^{\text{sim}}[n] = (j^{gp}, j^{gv}) \in \mathbb{R}^{6J}, where jgp∈R3Jj^{gp} \in \mathbb{R}^{3J} are global Cartesian joint positions and jgv∈R3Jj^{gv} \in \mathbb{R}^{3J} are global linear velocities in world coordinates.

    Conversions between representations are handled by two mapping operators:

    • R2G(x)R2G(x): Converts relative HumanML3D kinematic motions to global Cartesian positions and velocities via cumulative forward integration.
    • G2R(xsim)G2R(x^{\text{sim}}): Converts global simulated states to relative HumanML3D format using first-order inverse kinematics for joint angles and a height threshold heuristic for foot contact detection.
  6. Knowl 6 — Task Success Rates on Multi-Task Character Control

    data/table

    Performance comparison of CLoSD against specialized single-task baselines (AMP, InterPhys), the multi-task baseline UniHSI, and ablation variants evaluated over 1,000 episodes per task in simulation.

    Method Goal reaching Object striking Sitting Getting up
    AMP (2021) reach 0.88 - - -
    AMP (2021) strike - 1.0 - -
    InterPhys (2023) sit - - 0.76 -
    UniHSI (2024) 0.96 0.02 0.85 0.08
    CLoSD (Ours) 1.0 0.90 0.88 0.98
    shorter-loop (Ng=10N_g = 10) 0.86 0.71 0.61 0.95
    longer-loop (Ng=80N_g = 80) 0.99 0.86 0.56 0.92
    w.o. fine-tuning 1.0 0.81 0.66 0.53
    open-loop 1.0 0.80 0.19 0.23

    CLoSD matches or exceeds dedicated single-task policies and outperforms the multi-task baseline UniHSI, particularly on Object Striking (0.90 vs. 0.02) and Getting Up (0.98 vs. 0.08). The open-loop baseline and non-fine-tuned controller suffer steep performance drops on contact-rich tasks (Sitting and Getting Up) due to accumulated tracking drift and lack of closed-loop replanning.

  7. Knowl 7 — Text-to-Motion Quality and Physics Plausibility on HumanML3D Benchmark

    data/table

    Evaluation of text-driven motion generation on the HumanML3D test benchmark comparing kinematic generative models, the physics-based MoConVQ controller, and CLoSD (with target conditioning disabled).

    R-precision ↑\uparrow FID ↓\downarrow MultiModal Diversity Penetration ↓\downarrow Floating ↓\downarrow Skating ↓\downarrow
    Method Top 1 Top 2 Top 3 Distance ↓\downarrow [mm] [mm] [mm]
    Ground Truth 0.405 0.632 0.746 0.001 2.95 9.51 0.0 22.9 206⋅10−3206 \cdot 10^{-3}
    MDM (2023) 0.406 0.603 0.719 0.423 3.53 9.52 0.147 28.6 330⋅10−3330 \cdot 10^{-3}
    MoConVQ (2024) 0.309 0.504 0.614 3.279 3.97 8.01 0.249 32.0 294⋅10−3294 \cdot 10^{-3}
    CLoSD (Ours) 0.381 0.566 0.689 1.798 3.65 8.12 0.022 20.0 2⋅10−32 \cdot 10^{-3}
    DiP only 0.464 0.668 0.777 0.283 3.15 9.21 0.083 23.6 629⋅10−3629 \cdot 10^{-3}
    shorter-loop 0.303 0.473 0.586 3.481 4.18 7.93 0.018 21.5 0.3⋅10−30.3 \cdot 10^{-3}
    longer-loop 0.369 0.559 0.674 1.671 3.72 8.20 0.015 19.8 1.2⋅10−31.2 \cdot 10^{-3}
    open-loop 0.367 0.557 0.672 1.445 3.71 9.38 0.007 19.9 1⋅10−31 \cdot 10^{-3}
    w.o. fine-tuning 0.367 0.555 0.670 2.154 3.75 9.34 0.016 20.2 1.6⋅10−31.6 \cdot 10^{-3}

    CLoSD achieves superior text-adherence and distribution metrics compared to the physics-based baseline MoConVQ (FID 1.798 vs. 3.279; Top-1 R-precision 0.381 vs. 0.309). While physics simulation slightly reduces kinematic distribution fidelity compared to pure unconstrained diffusion (DiP only), it virtually eliminates physical motion artifacts, reducing foot skating from 629⋅10−3629 \cdot 10^{-3} mm to 2⋅10−32 \cdot 10^{-3} mm and penetration from 0.083 mm to 0.022 mm.

  8. Knowl 8 — DiP Diffusion Steps and Horizon Ablation

    data/table

    Ablation study evaluating the effect of diffusion denoising steps (TT), prefix length (NpN_p), and generation length (NgN_g) on motion quality and runtime speed on an NVIDIA GeForce RTX 3090 GPU.

    Configuration R-precision (top-3) ↑\uparrow FID ↓\downarrow Runtime [msec] ↓\downarrow Speed [fps] ↑\uparrow
    DiP (Default: Np=20,Ng=40,T=10N_p=20, N_g=40, T=10) 0.78 0.28 11.4 3.5⋅1033.5 \cdot 10^3
    5 Diff-steps (T=5T=5) 0.76 0.32 6.1 6.6⋅1036.6 \cdot 10^3
    20 Diff-steps (T=20T=20) 0.80 0.28 23.0 1.7⋅1031.7 \cdot 10^3
    Np=40N_p = 40 0.74 0.70 13.0 3.1⋅1033.1 \cdot 10^3
    Ng=20N_g = 20 0.78 0.26 11.4 1.7⋅1031.7 \cdot 10^3
    Ng=80N_g = 80 0.74 1.18 16.0 5.0⋅1035.0 \cdot 10^3

    Reducing diffusion steps from 50 (standard in MDM) to 10 in DiP maintains high motion quality (FID 0.28 vs 0.28 at 20 steps) while operating at 3,500 fps. Using Ng=40N_g = 40 frames (2 seconds) and Np=20N_p = 20 frames (1 second) provides the optimal trade-off between autoregressive trajectory consistency and computational overhead.

  9. Knowl 9 — Finite State Machine Task Transitions in CLoSD

    model/method

    CLoSD achieves multi-task composition by executing task sequences using a high-level Finite State Machine (FSM) without requiring dedicated transition policies.

    Each task in the sequence is defined by a pair of condition inputs: a text prompt (specifying motion style and action) and target joint constraints. Task completion criteria are evaluated in the physics simulator:

    • Goal Reaching: Completed when the pelvis joint horizontal distance to the target is ≤0.3\le 0.3 m.
    • Object Striking: Completed when the target bag angle relative to the ground plane drops below 75∘75^\circ.
    • Sitting Down: Completed when the pelvis joint reaches the sitting surface of the sofa.
    • Getting Up: Initialized from sitting; completed when the pelvis reaches a standing target placed 0.3 m in front of the sofa at a height of 0.9 m.

    Upon receiving a completion trigger signal, the FSM updates the conditioning prompt and target location in real time. DiP seamlessly generates transition trajectories from the current prefix state xprefixx^{\text{prefix}} without explicit blending.

  10. Knowl 10 — Limitations of CLoSD

    limitation

    The authors identify several limitations in the CLoSD framework:

    1. Lack of Exteroceptive Perception: Both the Diffusion Planner and the RL tracking controller lack visual or terrain height map inputs, limiting navigation and interaction planning to explicit 3D coordinate targets rather than arbitrary obstacle-cluttered 3D environments.
    2. Mid- and Low-Level Planning Scope: The architecture is specialized for mid- and low-level motion planning and execution, relying on simple state machines or external user input for long-horizon task scheduling.
    3. Fixed Replanning Horizon: The feedback loop operates on a fixed temporal horizon (predicting 40-frame / 2-second segments) rather than adapting its replanning frequency dynamically for agile or high-frequency contact adjustments.

Coverage note — None. All contributed algorithms, mathematical formulations, network designs, experimental tables, and analytical ablations from the paper are represented.

References

  1. 1.Mazen Al Borno, Martin De Lasa, and Aaron Hertzmann. Trajectory optimization for full-body movements with complex contacts. IEEE transactions on visualization and computer graphics, 19(8):1405–1414, 2012.
  2. 2.Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gül Varol. TEACH: Temporal Action Compositions for 3D Humans. In International Conference on 3D Vision (3DV), 2022.
  3. 3.Zhi Cen, Huaijin Pi, Sida Peng, Zehong Shen, Minghui Yang, Shuai Zhu, Hujun Bao, and Xiaowei Zhou. Generating human motion in 3d scenes from text descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1855–1866, 2024.
  4. 4.Rui Chen, Mingyi Shi, Shaoli Huang, Ping Tan, Taku Komura, and Xuelin Chen. Taming diffusion probabilistic models for character control. In SIGGRAPH, 2024.
  5. 5.Nuttapong Chentanez, Matthias Müller, Miles Macklin, Viktor Makoviychuk, and Stefan Jeschke. Physics-based motion capture imitation with deep reinforcement learning. Proceedings - MIG 2018: ACM SIGGRAPH Conference on Motion, Interaction, and Games, 2018.
  6. 6.Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. arXiv preprint arXiv:2303.04137, 2023.
  7. 7.Setareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng, and Michiel van de Panne. Flexible motion in-betweening with diffusion models. In ACM SIGGRAPH 2024 Conference Papers, pp. 1–9, 2024.
  8. 8.Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt. Mofusion: A framework for denoising-diffusion-based motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9760–9770, 2023.
  9. 9.Jacob Devlin. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  10. 10.Levi Fussell, Kevin Bergamin, and Daniel Holden. Supertrack: motion tracking for physically simulated characters using supervised learning. ACM Trans. Graph., 40:1–13, 2021. ISSN 0730-0301.
  11. 11.Purvi Goel, Kuan-Chieh Wang, C Karen Liu, and Kayvon Fatahalian. Iterative motion editing with natural language. In ACM SIGGRAPH 2024 Conference Papers, pp. 1–9, 2024.
  12. 12.Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5152–5161, June 2022.
  13. 13.Mohamed Hassan, Yunrong Guo, Tingwu Wang, Michael Black, Sanja Fidler, and Xue Bin Peng. Synthesizing physical character-scene interactions. In ACM SIGGRAPH 2023 Conference Proceedings, SIGGRAPH ’23, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9798400701597. doi: 10.1145/3588432.3591525. URL https://doi.org/10.1145/3588432.3591525.
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Diffusion Probabilistic Models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  15. 15.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey A. Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen video: High definition video generation with diffusion models. ArXiv, abs/2210.02303, 2022.
  16. 16.Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion-based generation, optimization, and planning in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16750–16761, 2023.
  17. 17.Jordan Juravsky, Yunrong Guo, Sanja Fidler, and Xue Bin Peng. Padl: Language-directed physics-based character control. In SIGGRAPH Asia 2022 Conference Papers, SA ’22, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450394703. doi: 10.1145/3550469.3555391. URL https://doi.org/10.1145/3550469.3555391.
  18. 18.Jordan Juravsky, Yunrong Guo, Sanja Fidler, and Xue Bin Peng. Superpadl: Scaling language-directed physics-based control with progressive supervised distillation. In SIGGRAPH 2024 Conference Papers (SIGGRAPH ’24 Conference Papers),, 2024.
  19. 19.Roy Kapon, Guy Tevet, Daniel Cohen-Or, and Amit H Bermano. Mas: Multi-view ancestral sampling for 3d motion generation using 2d diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1965–1974, 2024.
  20. 20.Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, and Siyu Tang. Optimizing diffusion noise can serve as universal motion priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1334–1345, 2024.
  21. 21.Nilesh Kulkarni, Davis Rempe, Kyle Genova, Abhijit Kundu, Justin Johnson, David Fouhey, and Leonidas Guibas. Nifty: Neural object interaction fields for guided human motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 947–957, 2024.
  22. 22.Lei Li and Angela Dai. Genzi: Zero-shot 3d human-scene interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20465–20474, 2024.
  23. 23.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, October 2015.
  24. 24.Zhengyi Luo, Ryo Hachiuma, Ye Yuan, and Kris Kitani. Dynamics-regulated kinematic policy for egocentric pose estimation. In Advances in Neural Information Processing Systems, 2021.
  25. 25.Zhengyi Luo, Shun Iwase, Ye Yuan, and Kris Kitani. Embodied scene-aware human pose estimation. In Advances in Neural Information Processing Systems, 2022.
  26. 26.Zhengyi Luo, Jinkun Cao, Kris Kitani, Weipeng Xu, et al. Perpetual humanoid control for real-time simulated avatars. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10895–10904, 2023.
  27. 27.Naureen Mahmood, N. Ghorbani, N. Troje, Gerard Pons-Moll, and Michael J. Black. Amass: Archive of motion capture as surface shapes. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5441–5450, 2019.
  28. 28.Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv preprint arXiv:2108.10470, 2021.
  29. 29.Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models. arXiv preprint arXiv:2312.06553, 2023.
  30. 30.Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne. Deepmimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Trans. Graph., 37 (4):143:1–143:14, July 2018. ISSN 0730-0301. doi: 10.1145/3197517.3201311. URL http://doi.acm.org/10.1145/3197517.3201311.
  31. 31.Xue Bin Peng, Michael Chang, Grace Zhang, Pieter Abbeel, and Sergey Levine. Mcp: Learning composable hierarchical control with multiplicative compositional policies. Advances in neural information processing systems, 32, 2019.
  32. 32.Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa. Amp: Adversarial motion priors for stylized physics-based character control. ACM Trans. Graph., 40 (4), July 2021. doi: 10.1145/3450626.3459670. URL http://doi.acm.org/10.1145/3450626.3459670.
  33. 33.Xue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine, and Sanja Fidler. Ase: Large-scale reusable adversarial skill embeddings for physically simulated characters. ACM Trans. Graph., 41(4), July 2022.
  34. 34.Mathis Petrovich, Michael J. Black, and Gül Varol. TEMOS: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision (ECCV), 2022.
  35. 35.Sigal Raab, Inbar Gat, Nathan Sala, Guy Tevet, Rotem Shalev-Arkushin, Ohad Fried, Amit H Bermano, and Daniel Cohen-Or. Monkey see, monkey do: Harnessing self-attention in motion diffusion for zero-shot motion transfer. arXiv preprint arXiv:2406.06508, 2024a.
  36. 36.Sigal Raab, Inbal Leibovitch, Guy Tevet, Moab Arar, Amit H Bermano, and Daniel Cohen-Or. Single motion diffusion. In The Twelfth International Conference on Learning Representations (ICLR), 2024b. URL https://openreview.net/pdf?id=DrhZneqz4n.
  37. 37.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  38. 38.Davis Rempe, Zhengyi Luo, Xue Bin Peng, Ye Yuan, Kris Kitani, Karsten Kreis, Sanja Fidler, and Or Litany. Trace and pace: Controllable pedestrian animation via guided trajectory diffusion. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  39. 39.Jiawei Ren, Mingyuan Zhang, Cunjun Yu, Xiao Ma, Liang Pan, and Ziwei Liu. Insactor: Instruction-driven physics-based characters. NeurIPS, 2023.
  40. 40.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  41. 41.V Sanh. Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  42. 42.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017. URL http://dblp.uni-trier.de/db/journals/corr/corr1707.html#SchulmanWDRK17.
  43. 43.Yoni Shafir, Guy Tevet, Roy Kapon, and Amit Haim Bermano. Human motion diffusion as a generative prior. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=dTpbEdN9kr.
  44. 44.Yi Shi, Jingbo Wang, Xuekun Jiang, Bingkun Lin, Bo Dai, and Xue Bin Peng. Interactive character control with auto-regressive motion diffusion models. ACM Trans. Graph., 43, jul 2024.
  45. 45.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265. PMLR, 2015.
  46. 46.Yang Song and Stefano Ermon. Improved techniques for training score-based generative models. Advances in neural information processing systems, 33:12438–12448, 2020.
  47. 47.Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII, pp. 358–374. Springer, 2022.
  48. 48.Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffusion model. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=SJ1kSyO2jwu.
  49. 49.Takara E Truong, Michael Piseno, Zhaoming Xie, and C Karen Liu. Pdp: Physics-based character animation via diffusion policy. arXiv preprint arXiv:2406.00960, 2024.
  50. 50.Jonathan Tseng, Rodrigo Castellon, and Karen Liu. Edge: Editable dance generation from music. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 448–458, 2023.
  51. 51.Tingwu Wang, Yunrong Guo, Maria Shugrina, and Sanja Fidler. Unicon: Universal neural controller for physics-based character motion. arXiv, 2020. ISSN 2331-8422.
  52. 52.Zan Wang, Yixin Chen, Baoxiong Jia, Puhao Li, Jinlu Zhang, Jingze Zhang, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. Move as you say interact as you can: Language-guided human motion generation with scene affordance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 433–444, 2024.
  53. 53.Jungdam Won, Deepak Gopinath, and Jessica Hodgins. A scalable approach to control diverse behaviors for physically simulated characters. ACM Trans. Graph., 39, 2020. ISSN 0730-0301,1557-7368.
  54. 54.Qianyang Wu, Ye Shi, Xiaoshui Huang, Jingyi Yu, Lan Xu, and Jingya Wang. Thor: Text to human-object interaction diffusion via relation intervention. arXiv preprint arXiv:2403.11208, 2024.
  55. 55.Zeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao, Wenwei Zhang, Bo Dai, Dahua Lin, and Jiangmiao Pang. Unified human-scene interaction via prompted chain-of-contacts. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=1vCnDyQkjg.
  56. 56.Zhaoming Xie, Jonathan Tseng, Sebastian Starke, Michiel van de Panne, and C Karen Liu. Hierarchical planning and control for box loco-manipulation. Proceedings of the ACM on Computer Graphics and Interactive Techniques, 6(3):1–18, 2023.
  57. 57.Heyuan Yao, Zhenhua Song, Yuyang Zhou, Tenglong Ao, Baoquan Chen, and Libin Liu. Moconvq: Unified physics-based motion control via scalable discrete representations. ACM Transactions on Graphics (TOG), 43(4):1–21, 2024.
  58. 58.Hongwei Yi, Justus Thies, Michael J Black, Xue Bin Peng, and Davis Rempe. Generating human interaction motions in scenes with text control. arXiv preprint arXiv:2404.10685, 2024.
  59. 59.Ye Yuan and Kris Kitani. Residual force control for agile human behavior imitation and extended motion synthesis. Advances in Neural Information Processing Systems, 33:21763–21774, 2020.
  60. 60.Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. Physdiff: Physics-guided human motion diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
  61. 61.Wentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu, Wayne Wu, and Yizhou Wang. Motionbert: A unified perspective on learning human motion representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15085–15099, 2023.

Citation

MLA
Tevet, G., et al. “CLoSD: Closing the Loop Between Simulation and Diffusion for Multi-task Character Control”. arXiv, 2024, http://arxiv.org/abs/2410.03441v1.
APA
Tevet, G., Raab, S., Cohan, S., Reda, D., Luo, Z., Peng, X. B., Bermano, A. H., & Panne, M. van . de . (2024). CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control. arXiv. http://arxiv.org/abs/2410.03441v1
Chicago
Tevet, G., S. Raab, S. Cohan, et al. 2024. “CLoSD: Closing the Loop Between Simulation and Diffusion for Multi-task Character Control”. arXiv. http://arxiv.org/abs/2410.03441v1.
Harvard
Tevet, G. et al. (2024) “CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.03441v1.
Vancouver
1. Tevet G, Raab S, Cohan S, Reda D, Luo Z, Peng XB, Bermano AH, Panne M van de (2024) CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control. arXiv

BibTeX

@article{tevet2024closd,
  title = {CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control},
  author = {Tevet, Guy and Raab, Sigal and Cohan, Setareh and Reda, Daniele and Luo, Zhengyi and Peng, Xue Bin and Bermano, Amit H. and Panne, Michiel van de},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.03441v1},
  eprint = {2410.03441}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors