CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control
Guy TevetSigal RaabSetareh CohanDaniele RedaZhengyi LuoXue Bin PengAmit Haim BermanoMichiel van de Panne
Develops a closed-loop framework that integrates real-time motion diffusion planning with reinforcement learning control, enabling physically simulated characters to execute complex sequences of diverse tasks directly from natural language prompts.
In computer animation, virtual environments, and robotics simulation, generating realistic human movements that seamlessly respond to natural language instructions and physical surroundings remains a fundamental challenge. Traditional data-driven generative models can synthesize a wide variety of human motions from simple text prompts, but they frequently generate physical artifacts such as floating, sliding feet, and unnatural object penetration. Conversely, physics-based reinforcement learning methods enforce realistic physical contacts and balance, yet they typically require complex, labor-intensive reward engineering tailored to individual tasks. The article demonstrates a unified framework called CLoSD (Closing the Loop between Simulation and Diffusion) to evaluate whether pairing a real-time motion diffusion planner with a physics-based reinforcement learning controller in a continuous feedback loop can achieve robust, multi-task character control.
The authors developed an integrated system where a lightweight generative model, termed the Diffusion Planner, creates short-horizon motion plans conditioned on text prompts and 3D target coordinates. This planner operates autoregressively and achieves high computational efficiency, producing 40-frame motion plans at approximately 3,500 frames per second—roughly 175 times faster than real time. A universal reinforcement learning controller executes these planned trajectories inside a physics simulator and feeds the resulting physical state back into the planner to maintain real-time responsiveness. To ensure stability during physical contacts, the tracking controller was fine-tuned in a closed loop across multiple tasks simultaneously using standard reinforcement learning objectives without task-specific reward design.
The experimental findings show that the closed-loop architecture outperforms existing state-of-the-art methods across diverse interactive tasks. In multi-task evaluations, the framework achieved a 100% success rate in navigation, 90% in object striking, 88% in sitting down on furniture, and 98% in standing up. In comparison, prior leading multi-task controllers achieved only a 2% success rate in striking and 8% in standing up because they lacked semantic motion understanding. Furthermore, disabling the closed feedback loop caused the sitting success rate to plunge from 88% to 19% and the get-up success rate from 98% to 23%, confirming that continuous environment feedback is essential. On standard text-to-motion benchmarks, the system matched or exceeded the fidelity of competing physics-based controllers while virtually eliminating physical flaws such as foot skating and body penetration.
These results indicate that generative models can serve as general-purpose kinematic planners for physical controllers, eliminating the need to train bespoke policies for separate movement skills. This offers practical benefits for interactive media, gaming, and virtual avatar development by lowering engineering timelines and reducing runtime system complexity. Decision-makers should consider piloting this architecture for interactive character systems requiring flexible, prompt-driven behaviors. However, because the current design relies on designated joint target coordinates rather than direct vision or elevation maps, future work should integrate visual perception and explore adaptive planning horizons to manage more complex, cluttered environments.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Provides the foundational diffusion architecture for conditional 3D human motion generation that underpins kinematic motion planning in character animation.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Establishes the paradigm of formulating trajectory planning and decision-making as conditional sampling from a generative diffusion model.
- Paper: Diffusion policy: Visuomotor policy learning via action diffusion, Cheng Chi et al. (2023). Introduces receding-horizon action generation via diffusion processes for closed-loop physical execution and control.
- Paper: Unified Human-Scene Interaction via Prompted Chain-of-Contacts, Zeqi Xiao et al. (2024). Demonstrates pairing high-level semantic action planners with unified low-level reinforcement learning controllers for physics-based character interactions.
- Paper: Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning, Viktor Makoviychuk et al. (2021). Presents the GPU-accelerated physical simulation framework essential for massively parallel reinforcement learning and character control training.
- Paper: GR00T N1: An Open Foundation Model for Generalist Humanoid Robots, NVIDIA et al. (2025). Scales the coupling of high-level reasoning and high-frequency diffusion action generation into a generalist foundation model across physical humanoid embodiments.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). Extends continuous closed-loop inference and dynamic action streaming to reactive physical manipulation of moving objects.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). Provides a comprehensive benchmark to evaluate whether generative predictive models can serve as actionable, physically grounded closed-loop simulators.
