Learning agile and dynamic motor skills for legged robots
Jemin HwangboJoonho LeeAlexey DosovitskiyDario BellicosoVassilios TsounisVladlen KoltunMarco Hutter
Develops a sim-to-real reinforcement learning pipeline that trains neural network policies in simulation and deploys them on physical quadruped robots, achieving fast running speeds, energy-efficient velocity tracking, and autonomous fall recovery.
Legged robots offer significant potential for operating in complex, unstructured environments where wheeled or tracked machines struggle, such as disaster recovery sites, construction zones, and industrial inspections. However, creating control systems capable of agile and dynamic locomotion remains a primary bottleneck. Traditional model-based control frameworks rely on simplified analytical models and laborious manual tuning that often require months of engineering per behavior, restrict the machine to narrow operating ranges, and require substantial onboard computing power. While reinforcement learning (a machine learning approach based on trial and error) offers an automated alternative, training directly on physical hardware risks severe damage and requires weeks or months of operation. Conversely, controllers trained purely in simulation typically fail on physical hardware because of the reality gap—the discrepancy between simulated physics and real-world dynamics, particularly concerning complex actuator behavior and hardware delays.
The article sets out to demonstrate a simulation-to-reality reinforcement learning methodology that trains neural network control policies entirely in simulation and deploys them directly on a physical, full-scale quadruped robot without real-world fine-tuning. The primary objective was to evaluate whether this framework could surpass existing state-of-the-art model-based controllers in command tracking accuracy, energy efficiency, top operational speed, and dynamic fall recovery on ANYmal, a 32-kilogram, medium-dog-sized quadruped.
To achieve this, the authors created a fast hybrid simulation pipeline. Instead of relying on analytical actuator formulas, they trained a compact neural network (an actuator network) directly on less than four minutes of physical hardware data to capture nonlinear motor dynamics, joint compliance, and communication latencies. This learned actuator model was integrated into an in-house rigid-body contact simulator, allowing simulated training to run approximately 1,000 times faster than real time on standard desktop hardware without specialized computing clusters. The researchers applied a reinforcement learning algorithm (Trust Region Policy Optimization) within a curriculum-learning structure, randomizing physical parameters such as robot mass, inertia, and sensor noise to ensure robust real-world transfer.
The findings show substantial improvements across key performance metrics when tested on physical hardware. The learned locomotion policy significantly outperformed the existing state-of-the-art modular controller, reducing base linear velocity tracking error by approximately 95% and turning rate error by about 60% during random command sequences. Furthermore, the robot broke its previous hardware speed record by 25%, achieving a top forward speed of 1.5 meters per second by fully utilizing the torque and velocity limits of the motors. In terms of efficiency, the learned controller reduced average joint torque by 23% to 36% compared to prior controllers by autonomously discovering a more upright nominal leg posture. Finally, the framework yielded a fully autonomous fall-recovery controller that consistently restored the robot to a standing position from arbitrary upside-down and resting configurations in under three seconds on the physical machine, a feat not previously demonstrated on a quadruped of this complexity.
These results demonstrate that data-driven simulation learning can radically accelerate robotic development timelines while delivering superior operational performance. Shifting the heavy computational workload from runtime planning to offline training reduced onboard processing demands to less than 0.25% of a single processor core, eliminating the need for bulky, external computers. Over three months of rigorous hardware testing, the learned control policies proved resilient against mechanical changes and wear, including structural weight alterations of approximately 2 kilograms and threefold increases in drive stiffness. This confirms that robust simulation-to-reality pipelines can mitigate hardware risks and reduce long-term maintenance costs.
Moving forward, development teams should adopt data-driven actuator modeling and simulation-based training pipelines to minimize manual tuning and deploy performant robotic controllers quickly. Because individual neural network policies remain specialized to single task distributions, future research should integrate hierarchical control structures that enable a robot to transition autonomously among diverse behaviors such as navigation, stair climbing, and recovery. Organizations should establish structured data-logging practices for their hardware platforms to easily train initial actuator networks.
While the results demonstrate high confidence and repeatability on the ANYmal platform, several operational boundaries apply. Actuators with coupled dynamics, such as hydraulic systems sharing a common fluid reservoir, may require unified multi-actuator neural network models rather than the independent joint models used here. In addition, designing task-specific reward functions still requires human engineering judgment to prevent unsafe hardware behaviors, such as high-impact collisions during fall recovery. Finally, the demonstrated locomotion policies relied on flat or minimally disturbed terrain; further validation in highly cluttered or three-dimensional outdoor environments is necessary before deploying the system in critical field missions.
- Paper: Domain randomization for transferring deep neural networks from simulation to the real world, Josh Tobin et al. (2017). Introduces domain randomization techniques essential for bridging the sim-to-real gap when transferring simulated neural network policies to physical robots.
- Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). Develops the core Proximal Policy Optimization (PPO) algorithm heavily relied upon for stable continuous control in legged locomotion learning.
- Paper: High-Dimensional Continuous Control Using Generalized Advantage Estimation, John Schulman et al. (2016). Formulates Generalized Advantage Estimation, which provides critical variance-reduction mechanisms used by policy gradient methods for high-dimensional robotic control.
- Paper: Soft Actor-Critic Algorithms and Applications, Tuomas Haarnoja et al. (2018). Demonstrates practical applications of continuous actor-critic reinforcement learning algorithms on physical legged robots and robotic manipulators.
- Paper: Learning dexterous in-hand manipulation, Marcin Andrychowicz et al. (2018). Establishes large-scale simulation training with dynamics randomization as an effective paradigm for direct policy transfer to complex physical hardware.
- Paper: Trust Region Policy Optimization, John Schulman et al. (2015). Presents Trust Region Policy Optimization, establishing foundational principles of constrained policy updates for high-dimensional continuous locomotion tasks.
- Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). Pioneers deep continuous action-space reinforcement learning for motor control tasks in physics simulation environments.
- Paper: Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning, Viktor Makoviychuk et al. (2021). Accelerates sim-to-real reinforcement learning workflows for systems like the ANYmal quadruped through massively parallel, GPU-native physics simulation.
