Learning quadrupedal locomotion over challenging terrain
Joonho LeeJemin HwangboLorenz WellhausenVladlen KoltunMarco Hutter
Demonstrates that a reinforcement learning controller trained purely in simulation using proprioceptive feedback achieves zero-shot transfer to traverse extreme, unmodeled natural terrains such as mud, snow, and dynamic rubble on physical quadruped robots.
Autonomous legged machines have the potential to traverse extreme environments that remain inaccessible to wheeled vehicles, yet conventional controllers struggle with complex, deformable natural terrains such as mud, snow, and dense vegetation. Traditional control approaches rely on hand-crafted rules and external sensors such as cameras or foot contact sensors, which frequently fail when obscured by environmental debris or water. The article demonstrates a robust control system for four-legged robots operating over challenging, unmapped natural terrain. Its primary objective is to evaluate whether a robot relying strictly on internal bodily motion feedback can achieve reliable locomotion across diverse real-world environments without requiring site-specific tuning.
The authors developed a training approach conducted entirely in simulation using a two-stage learning process. In the initial phase, a simulated teacher controller learned to navigate parameterized rough terrains with full access to ground-truth environmental physics and contact forces. The system then distilled this expertise into a student controller that operated strictly on a rolling history of internal sensor readings, such as joint angles, velocities, and body orientation. An adaptive curriculum continuously adjusted the difficulty of simulated terrains to match the learning progress. The resulting controller was deployed directly onto physical quadruped robots across outdoor forests, mountain trails, and indoor obstacle courses without any prior physical calibration or real-world trial-and-error.
The article demonstrates several significant findings. First, the controller achieved zero-shot transfer from rigid simulations to diverse real-world environments, traversing deep mud, snow, running water, and loose rubble without experiencing catastrophic falls. Second, in direct field comparisons against a state-of-the-art baseline controller, the learned controller achieved more than double the forward speed on moss—0.45 meters per second compared to 0.20 meters per second—while reducing the mechanical cost of transport by approximately 25 to 33 percent, demonstrating superior energetic efficiency. Third, when subjected to an unmodeled 10-kilogram payload representing nearly 23 percent of the robot's total weight, the controller successfully navigated obstacles up to 13.4 centimeters high, whereas the baseline controller failed completely. Fourth, the system autonomously developed protective stepping reflexes, clearing obstacles up to 22.5 centimeters and reducing lateral motion deviation under external pushing forces by 35.5 percent compared to short-memory controllers. Finally, the controller operated with a zero-failure rate across four 60-minute competitive missions in a subterranean challenge.
These findings indicate that complex physical interactions can be managed effectively without laboriously modeling every environmental detail or manually scripting control reflexes. This approach significantly reduces the cost, safety risks, and development timelines associated with field robotics in industrial inspection, disaster response, and defense. Stakeholders should consider adopting this simulation-trained learning architecture as the foundational mobility layer for legged robotic fleets. However, decision-makers should account for current operational boundaries: the system is blind to distant hazards, meaning it cannot detect drop-offs such as cliffs, and it is presently limited to a single trotting gait. Future development should focus on integrating complementary vision systems to prevent hazardous falls while expanding the gait repertoire. Confidence in the core locomotion stability is high given extensive testing across two hardware generations, provided external navigation safeguards are maintained around steep precipices.
- Paper: Learning agile and dynamic motor skills for legged robots, Jemin Hwangbo et al. (2019). This foundational ANYmal locomotion paper establishes the hybrid simulation and learned actuator network framework that the source builds upon for robust sim-to-real terrain traversal.
- Paper: Sim-to-Real Transfer of Robotic Control with Dynamics Randomization, Xue Bin Peng et al. (2017). It introduces the core methodology of training recurrent neural networks with broad dynamics randomization to bridge the simulation-to-reality gap without real-world fine-tuning.
- Paper: Domain randomization for transferring deep neural networks from simulation to the real world, Josh Tobin et al. (2017). It provides the conceptual foundation of domain randomization to achieve zero-shot transfer from simulated physics to physical robotic deployment.
- Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). It details the standard on-policy reinforcement learning optimization algorithm utilized to stably train continuous locomotion policies in simulation.
- Paper: High-Dimensional Continuous Control Using Generalized Advantage Estimation, John Schulman et al. (2016). It presents the Generalized Advantage Estimation method used to control bias and variance in policy gradient estimation for high-dimensional robotic continuous control.
- Paper: Trust Region Policy Optimization, John Schulman et al. (2015). It establishes the underlying trust-region optimization principles that guarantee stable, monotonic policy improvements in complex robotic locomotion tasks.
- Paper: Planning and acting in partially observable stochastic domains, Leslie Pack Kaelbling et al. (1998). It formalizes decision-making under partial observability, framing the theoretical need for policies to infer unobserved environmental states from temporal sensory streams.
- Paper: Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning, Viktor Makoviychuk et al. (2021). It scales up the massively parallel GPU-based simulation and training paradigm pioneered in legged locomotion to achieve orders-of-magnitude faster policy convergence on robots like ANYmal.
