Learning quadrupedal locomotion over challenging terrain

Joonho LeeJemin HwangboLorenz WellhausenVladlen KoltunMarco Hutter

article2020Science Robotics1,758 citations

Demonstrates that a reinforcement learning controller trained purely in simulation using proprioceptive feedback achieves zero-shot transfer to traverse extreme, unmodeled natural terrains such as mud, snow, and dynamic rubble on physical quadruped robots.

Listen

Autonomous legged machines have the potential to traverse extreme environments that remain inaccessible to wheeled vehicles, yet conventional controllers struggle with complex, deformable natural terrains such as mud, snow, and dense vegetation. Traditional control approaches rely on hand-crafted rules and external sensors such as cameras or foot contact sensors, which frequently fail when obscured by environmental debris or water. The article demonstrates a robust control system for four-legged robots operating over challenging, unmapped natural terrain. Its primary objective is to evaluate whether a robot relying strictly on internal bodily motion feedback can achieve reliable locomotion across diverse real-world environments without requiring site-specific tuning.

The authors developed a training approach conducted entirely in simulation using a two-stage learning process. In the initial phase, a simulated teacher controller learned to navigate parameterized rough terrains with full access to ground-truth environmental physics and contact forces. The system then distilled this expertise into a student controller that operated strictly on a rolling history of internal sensor readings, such as joint angles, velocities, and body orientation. An adaptive curriculum continuously adjusted the difficulty of simulated terrains to match the learning progress. The resulting controller was deployed directly onto physical quadruped robots across outdoor forests, mountain trails, and indoor obstacle courses without any prior physical calibration or real-world trial-and-error.

The article demonstrates several significant findings. First, the controller achieved zero-shot transfer from rigid simulations to diverse real-world environments, traversing deep mud, snow, running water, and loose rubble without experiencing catastrophic falls. Second, in direct field comparisons against a state-of-the-art baseline controller, the learned controller achieved more than double the forward speed on moss—0.45 meters per second compared to 0.20 meters per second—while reducing the mechanical cost of transport by approximately 25 to 33 percent, demonstrating superior energetic efficiency. Third, when subjected to an unmodeled 10-kilogram payload representing nearly 23 percent of the robot's total weight, the controller successfully navigated obstacles up to 13.4 centimeters high, whereas the baseline controller failed completely. Fourth, the system autonomously developed protective stepping reflexes, clearing obstacles up to 22.5 centimeters and reducing lateral motion deviation under external pushing forces by 35.5 percent compared to short-memory controllers. Finally, the controller operated with a zero-failure rate across four 60-minute competitive missions in a subterranean challenge.

These findings indicate that complex physical interactions can be managed effectively without laboriously modeling every environmental detail or manually scripting control reflexes. This approach significantly reduces the cost, safety risks, and development timelines associated with field robotics in industrial inspection, disaster response, and defense. Stakeholders should consider adopting this simulation-trained learning architecture as the foundational mobility layer for legged robotic fleets. However, decision-makers should account for current operational boundaries: the system is blind to distant hazards, meaning it cannot detect drop-offs such as cliffs, and it is presently limited to a single trotting gait. Future development should focus on integrating complementary vision systems to prevent hazardous falls while expanding the gait repertoire. Confidence in the core locomotion stability is high given extensive testing across two hardware generations, provided external navigation safeguards are maintained around steep precipices.

Cover for Learning quadrupedal locomotion over challenging terrain

Abstract

Some of the most challenging environments on our planet are accessible to quadrupedal animals but remain out of reach for autonomous machines. Legged locomotion can dramatically expand the operational domains of robotics. However, conventional controllers for legged locomotion are based on elaborate state machines that explicitly trigger the execution of motion primitives and reflexes. These designs have escalated in complexity while falling short of the generality and robustness of animal locomotion. Here we present a radically robust controller for legged locomotion in challenging natural environments. We present a novel solution to incorporating proprioceptive feedback in locomotion control and demonstrate remarkable zero-shot generalization from simulation to natural environments. The controller is trained by reinforcement learning in simulation. It is based on a neural network that acts on a stream of proprioceptive signals. The trained controller has taken two generations of quadrupedal ANYmal robots to a variety of natural environments that are beyond the reach of prior published work in legged locomotion. The controller retains its robustness under conditions that have never been encountered during training: deformable terrain such as mud and snow, dynamic footholds such as rubble, and overground impediments such as thick vegetation and gushing water. The presented work opens new frontiers for robotics and indicates that radical robustness in natural environments can be achieved by training in much simpler domains.

Table of Contents

  • 1 Introduction
  • 2 Results
  • 3 Discussion
  • 4 Materials and Methods
  • 4.1 Overview
  • 5 Acknowledgments
  • References

Knowls

  1. Knowl 1 — Two-Stage Privileged Teacher-Student Learning Framework

    model/method

    To train a blind locomotion policy capable of traversing challenging terrain without requiring exteroceptive perception, a two-stage privileged distillation framework is used.

    In the first stage, a privileged teacher policy πteacher(at∣ot,xt)\pi_{\text{teacher}}(a_t \mid o_t, x_t) is trained via model-free reinforcement learning (Trust Region Policy Optimization, TRPO) in simulation. The state consists of observable measurements ot∈R60o_t \in \mathbb{R}^{60} (command vector, base orientation, base linear and angular velocities, joint positions and velocities, leg phase encodings, joint position error and velocity histories, and previous foot position targets) and privileged simulation state xtx_t (elevation scans of 9 points around each foot within a 10 cm10\text{ cm} radius, contact surface normals, foot contact forces and contact states for feet, shanks, and thighs, foot-ground friction coefficients, and unmodeled external disturbance forces). The teacher architecture uses a multi-layer perceptron (MLP) encoder to map xtx_t into a latent embedding lˉt\bar{l}_t, which is concatenated with oto_t and passed through MLP layers to generate action aˉt\bar{a}_t.

    In the second stage, a student policy πstudent(at∣ot,H)\pi_{\text{student}}(a_t \mid o_t, H) is trained using Dataset Aggregation (DAgger). The student policy receives the observable state oto_t and a temporal buffer of NN previous proprioceptive measurements H={ht−1,ht−2,…,ht−N−1}H = \{h_{t-1}, h_{t-2}, \dots, h_{t-N-1}\}, where ht=ot∖{f0,joint history,previous foot targets}h_t = o_t \setminus \{f_0, \text{joint history}, \text{previous foot targets}\}. The student encodes HH via a Temporal Convolutional Network (TCN) to output latent representation ltl_t, which is merged with oto_t to produce action ata_t. The student is trained via supervised learning using the loss function:

    L:=(aˉt(ot,xt)−at(ot,H))2+(lˉt(ot,xt)−lt(H))2\mathcal{L} := \left(\bar{a}_t(o_t, x_t) - a_t(o_t, H)\right)^2 + \left(\bar{l}_t(o_t, x_t) - l_t(H)\right)^2

    where aˉt\bar{a}_t and lˉt\bar{l}_t are target action and latent representations computed by the frozen teacher. This allows the student to implicitly estimate latent environmental properties solely from proprioceptive history.

  2. Knowl 2 — Adaptive Terrain Curriculum via Particle Filtering

    algorithm

    An adaptive terrain curriculum generates simulated terrains at an optimal difficulty level matching the learning progression of the policy. Terrains are parameterized by a vector cT∈Cc_T \in \mathcal{C} governing geometric features (e.g., roughness, frequency, and amplitude for Perlin hills; step width and height for steps and stairs).

    Terrain traversability for a policy π\pi is defined as the empirical probability that the policy maintains a projected forward base speed exceeding 0.2 m/s0.2\text{ m/s} without terminating:

    Tr(cT,π):=Eξ∼π{ν(st,at,st+1∣cT)}∈[0.0,1.0]\text{Tr}(c_T, \pi) := \mathbb{E}_{\xi \sim \pi} \left\{ \nu(s_t, a_t, s_{t+1} \mid c_T) \right\} \in [0.0, 1.0]

    ν(st,at,st+1):={1if vpr(st+1)>0.20if vpr(st+1)≤0.2∨termination\nu(s_t, a_t, s_{t+1}) := \begin{cases} 1 & \text{if } v_{pr}(s_{t+1}) > 0.2 \\ 0 & \text{if } v_{pr}(s_{t+1}) \le 0.2 \lor \text{termination} \end{cases}

    where vprv_{pr} is the base velocity projected onto the commanded direction. Terrain desirability is defined as mid-range traversability:

    Td(cT,π):=Pr⁡(Tr(cT,π)∈[0.5,0.9])\text{Td}(c_T, \pi) := \Pr(\text{Tr}(c_T, \pi) \in [0.5, 0.9])

    A Sequential Importance Resampling (SIR) particle filter tracks the distribution of desirable parameters cTc_T using NparticleN_{\text{particle}} particles per terrain type.

    Input: Parameter space C\mathcal{C}, number of particles NparticleN_{\text{particle}}, trajectory count per particle NtrajN_{\text{traj}}, evaluation interval NevaluateN_{\text{evaluate}}, transition probability ptransitionp_{\text{transition}}, replay probability preplayp_{\text{replay}}
    Output: Trained policy π\pi
    Initialize replay memory M←∅\mathcal{M} \leftarrow \emptyset
    Sample initial particle parameters {cT,0l}l=1Nparticle∼Uniform(C)\{c_{T,0}^l\}_{l=1}^{N_{\text{particle}}} \sim \text{Uniform}(\mathcal{C})
    i←0,j←0i \leftarrow 0, j \leftarrow 0
    repeat
        for k=0k = 0 to Nevaluate−1N_{\text{evaluate}} - 1 do
            for l=1l = 1 to NparticleN_{\text{particle}} do
                for m=1m = 1 to NtrajN_{\text{traj}} do
                    Generate terrain instance from parameter cT,jlc_{T,j}^l
                    Initialize robot state and execute policy πi\pi_i
                    Compute traversability label ν\nu for each state transition
                    Save trajectories and traversability scores
            Update policy πi+1\pi_{i+1} from collected trajectories using TRPO
            i←i+1i \leftarrow i + 1
        for l=1l = 1 to NparticleN_{\text{particle}} do
            Compute measurement probability Pr⁡(yjl∣cT,jl)≈1NtrajNevaluate∑1(Tr(cT,jl,π)∈[0.5,0.9])\Pr(y_j^l \mid c_{T,j}^l) \approx \frac{1}{N_{\text{traj}} N_{\text{evaluate}}} \sum \mathbf{1}(\text{Tr}(c_{T,j}^l, \pi) \in [0.5, 0.9])
        for l=1l = 1 to NparticleN_{\text{particle}} do
            Compute normalized importance weights wjl=Pr⁡(yjl∣cT,jl)∑mPr⁡(yjm∣cT,jm)w_j^l = \frac{\Pr(y_j^l \mid c_{T,j}^l)}{\sum_m \Pr(y_j^m \mid c_{T,j}^m)}
        Resample NparticleN_{\text{particle}} parameter sets according to weights {wjl}\{w_j^l\}
        Append resampled parameters to replay memory M\mathcal{M}
        for l=1l = 1 to NparticleN_{\text{particle}} do
            With probability preplayp_{\text{replay}}, replace cT,jlc_{T,j}^l with a sample drawn from M\mathcal{M}
            With probability ptransitionp_{\text{transition}}, mutate cT,jlc_{T,j}^l via random walk step to adjacent grid neighbor in C\mathcal{C}
        j←j+1j \leftarrow j + 1
    until policy convergence
  3. Knowl 3 — PMTG-Based Quadrupedal Motion Control Architecture

    model/method

    The locomotion control architecture integrates a Policies Modulating Trajectory Generators (PMTG) scheme with kinematic residuals and high-frequency joint tracking.

    The input command specifies horizontal travel direction and base yaw turning:

    c=⟨(IBv^BT)xy,(ω^T)z⟩=⟨⟨cos⁡(ψT),sin⁡(ψT)⟩,(ω^T)z⟩\mathbf{c} = \left\langle \left({}^B_I \hat{\mathbf{v}}^T_B\right)_{xy}, (\hat{\omega}^T)_z \right\rangle = \left\langle \langle \cos(\psi_T), \sin(\psi_T) \rangle, (\hat{\omega}^T)_z \right\rangle

    where ψT\psi_T is the target yaw heading angle in the base frame and (ω^T)z∈{−1,0,1}(\hat{\omega}^T)_z \in \{-1, 0, 1\} is the turning direction.

    Each leg i∈{1,2,3,4}i \in \{1, 2, 3, 4\} maintains a periodic phase variable ϕi∈[0.0,2π)\phi_i \in [0.0, 2\pi) updated at each time step tt:

    ϕi=(ϕi,0+(f0+fi)t)(mod2π)\phi_i = \left(\phi_{i,0} + (f_0 + f_i)t\right) \pmod{2\pi}

    where f0=1.25 Hzf_0 = 1.25\text{ Hz} is the nominal base trot frequency and fif_i is a policy-generated leg frequency offset. Contact phase corresponds to ϕi∈[0.0,π)\phi_i \in [0.0, \pi) and swing phase corresponds to ϕi∈[π,2π)\phi_i \in [\pi, 2\pi).

    A Foot Trajectory Generator (FTG) produces nominal cyclic foot Cartesian trajectories F(ϕi)∈R3F(\phi_i) \in \mathbb{R}^3 defined in individual horizontal frames HiH_i. The horizontal frame HiH_i is anchored below the ii-th hip joint at nominal leg reach, with its zz-axis (HizH_i z) aligned with gravity ege_g and its xx-axis (HixH_i x) set to the projection of the base xx-axis onto the horizontal plane. Decoupling the roll and pitch attitude of the base from HiH_i isolates foot trajectory generation from body tilt disturbances.

    The policy network outputs a 16-dimensional vector consisting of 4 frequency offsets fif_i and 12 foot position residuals Δrfi,T\Delta r_{f_i, T}. The target foot position in horizontal frame coordinates is:

    rfi,T:=F(ϕi)+Δrfi,Tr_{f_i, T} := F(\phi_i) + \Delta r_{f_i, T}

    Foot target positions are converted to the robot base frame and resolved to target joint positions θ∗∈R12\theta^* \in \mathbb{R}^{12} via analytic inverse kinematics (IK). Joint position targets are tracked at 400 Hz400\text{ Hz} by joint PD controllers, where actuator dynamics of Series Elastic Actuators (SEA) are simulated using an empirical neural network model.

  4. Knowl 4 — Mathematical Formulations for Foot Trajectory Generation and RL Reward

    equation

    The Foot Trajectory Generator (FTG) defines nominal Cartesian foot position F(ϕi)F(\phi_i) along the vertical axis of horizontal frame HiH_i using cubic Hermite splines:

    F(ϕi)={(h(−2k3+3k2)−0.5)Hizk∈[0,1](h(2k3−9k2+12k−4)−0.5)Hizk∈[1,2]−0.5HizotherwiseF(\phi_i) = \begin{cases} \left(h(-2k^3 + 3k^2) - 0.5\right) H_i z & k \in [0, 1] \\ \left(h(2k^3 - 9k^2 + 12k - 4) - 0.5\right) H_i z & k \in [1, 2] \\ -0.5 H_i z & \text{otherwise} \end{cases}

    where k=2(ϕi−π)/πk = 2(\phi_i - \pi)/\pi for leg phase ϕi∈[0,2π)\phi_i \in [0, 2\pi) and maximum foot lift height h=0.2 mh = 0.2\text{ m}.

    The composite reward function for privileged teacher training is formulated as:

    R=0.05rlv+0.05rav+0.04rb+0.01rfc+0.02rbc+0.025rs+2×10−5rτR = 0.05 r_{lv} + 0.05 r_{av} + 0.04 r_b + 0.01 r_{fc} + 0.02 r_{bc} + 0.025 r_s + 2 \times 10^{-5} r_\tau

    Individual terms are defined as follows:

    • Linear velocity reward (vpr=(IBv)xy⋅(IBv^T)xyv_{pr} = ({}^B_I v)_{xy} \cdot ({}^B_I \hat{v}^T)_{xy}): rlv:={exp⁡(−2.0(vpr−0.6)2)vpr<0.61.0vpr≥0.60.0for zero commandr_{lv} := \begin{cases} \exp\left(-2.0(v_{pr} - 0.6)^2\right) & v_{pr} < 0.6 \\ 1.0 & v_{pr} \ge 0.6 \\ 0.0 & \text{for zero command} \end{cases}
    • Angular velocity reward (ωpr=(IBω)z⋅(IBω^T)z\omega_{pr} = ({}^B_I \omega)_z \cdot ({}^B_I \hat{\omega}^T)_z): rav:={exp⁡(−1.5(ωpr−0.6)2)ωpr<0.61.0ωpr≥0.6r_{av} := \begin{cases} \exp\left(-1.5(\omega_{pr} - 0.6)^2\right) & \omega_{pr} < 0.6 \\ 1.0 & \omega_{pr} \ge 0.6 \end{cases}
    • Base motion stability reward with vo=∥(IBv)xy−vpr(IBv^T)xy∥v_o = \|({}^B_I v)_{xy} - v_{pr}({}^B_I \hat{v}^T)_{xy}\| (or vo=∥(IBv)xy∥v_o = \|({}^B_I v)_{xy}\| under zero command): rb:=exp⁡(−1.5vo2)+exp⁡(−1.5∥(IBω)xy∥2)r_b := \exp(-1.5 v_o^2) + \exp\left(-1.5 \|({}^B_I \omega)_{xy}\|^2\right)
    • Foot clearance reward (IswingI_{swing} is swing leg set, Hscan,iH_{scan,i} is elevation scan around foot ii): rfc:=∑i∈Iswing1(rf,i>max⁡(Hscan,i))∣Iswing∣∈[0.0,1.0]r_{fc} := \sum_{i \in I_{swing}} \frac{\mathbf{1}(r_{f,i} > \max(H_{scan,i}))}{|I_{swing}|} \in [0.0, 1.0]
    • Body collision penalty (Ic,bodyI_{c,body} is body contact set, Ic,footI_{c,foot} is foot contact set): rbc:=−∣Ic,body∖Ic,foot∣r_{bc} := -|I_{c,body} \setminus I_{c,foot}|
    • Target foot smoothness penalty: rs:=−∥(rf,d)t−2(rf,d)t−1+(rf,d)t−2∥r_s := -\|(r_{f,d})_t - 2(r_{f,d})_{t-1} + (r_{f,d})_{t-2}\|
    • Joint torque penalty: rτ:=−∑i∈joints∣τi∣r_\tau := -\sum_{i \in joints} |\tau_i|
  5. Knowl 5 — Probing Latent Proprioceptive Representations via Auxiliary Decoders

    model/method

    To evaluate whether a purely proprioceptive student policy implicitly models hidden physical properties of the environment, an auxiliary decoder network is trained on the frozen intermediate representations of the trained Temporal Convolutional Network (TCN).

    The decoder maps intermediate embeddings ⟨ot,lt⟩\langle o_t, l_t \rangle to privileged state features xt∈Xx_t \in X, which include foot contact states, local terrain elevation profile, contact surface normal vectors, foot-ground friction coefficients, and unmodeled external disturbance forces.

    Foot contact states are classified using cross-entropy loss. Continuous states are predicted as Gaussian distributions with mean mim_i and standard deviation σi\sigma_i to capture aleatoric uncertainty, optimized via negative Gaussian log-likelihood loss with weight decay:

    Ldec=∑i∈dim(X∖contact states)((mi−migt)22σi2+log⁡(σi))\mathcal{L}_{dec} = \sum_{i \in \text{dim}(X \setminus \text{contact states})} \left( \frac{(m_i - m_i^{gt})^2}{2\sigma_i^2} + \log(\sigma_i) \right)

    where migtm_i^{gt} is the ground-truth privileged value in simulation. The policy parameters remain completely fixed during decoder training.

    Empirical evaluation of the trained decoder reveals that:

    1. On slippery terrain (such as a moistened whiteboard), the decoded friction coefficient drops immediately upon the first foot slip and remains low until 2 s2\text{ s} after returning to high-traction ground.
    2. In the presence of an unmodeled 10 kg10\text{ kg} payload, the decoder accurately reconstructs a downward external force on the torso.
    3. In thick vegetation, the decoder identifies resistive drag forces opposing the direction of travel, enabling the controller to push through impediments.
  6. Knowl 6 — Emergence of Foot-Trapping Reflexes and Temporal Saliency Attribution

    empirical result

    The proprioceptive controller demonstrates emergent foot-trapping reflexes and obstacle-clearing behaviors without explicit reflex triggers, hand-coded contact thresholds, or discrete finite state machines.

    When encountering an unexpected discrete obstacle (e.g., a 16.8 cm16.8\text{ cm} vertical step):

    1. Foot clearance during normal trotting on flat ground is 12.9 cm12.9\text{ cm} for the left fore (LF) leg and 13.6 cm13.6\text{ cm} for the right fore (RF) leg. Upon foot collision during the swing phase, foot clearance adaptively increases up to 22.5 cm22.5\text{ cm} (LF) and 18.5 cm18.5\text{ cm} (RF) in the subsequent swing phase.
    2. Hind leg clearance adaptively increases from flat-ground baselines of 13.5 cm13.5\text{ cm} (LH) and 9.06 cm9.06\text{ cm} (RH) up to 16.6 cm16.6\text{ cm} (LH) and 15.9 cm15.9\text{ cm} (RH) once the front legs step up.
    3. The policy reacts successfully to mid-shin collisions where foot contact sensors would fail to trigger.

    Gradient saliency mapping over the temporal input buffer H∈R60×NH \in \mathbb{R}^{60 \times N} is evaluated using:

    Mi=∑j∈channels∣∂(rf,T)z∂Hi,j∣∈RM_i = \sum_{j \in channels} \left| \frac{\partial (r_{f,T})_z}{\partial H_{i,j}} \right| \in \mathbb{R}

    where (rf,T)z(r_{f,T})_z is the commanded foot height. The saliency profile displays sharp, sustained peaks focused on the joint position and velocity channels of the colliding leg at the exact time instant of initial impact (t≈2.1 st \approx 2.1\text{ s}), demonstrating that the temporal convolution actively attends to historical collision events across several subsequent gait cycles.

  7. Knowl 7 — Locomotion Performance Comparison in Natural Environments

    data/table

    The learned proprioceptive controller was benchmarked against an optimization-based model-based locomotion baseline on an ANYmal quadruped across three challenging natural environments: wet moss, mud, and dense forest vegetation. Performance was evaluated using average forward speed (m/s\text{m/s}) and dimensionless mechanical Cost of Transport (COT), defined as:

    COT=∑i=112[τiθ˙i]+mgv\text{COT} = \frac{\sum_{i=1}^{12} [\tau_i \dot{\theta}_i]^+}{mgv}

    where [τiθ˙i]+[\tau_i \dot{\theta}_i]^+ is the positive mechanical power exerted by joint actuator ii, mm is the robot mass, gg is gravitational acceleration, and vv is base locomotion speed.

    Quantity Controller Terrain
    Moss Mud Vegetation
    Average speed (m/s) Ours 0.452 0.338 0.248
    Baseline 0.199 0.197 –
    Average mechanical COT Ours 0.423 0.692 1.23
    Baseline 0.625 0.931 –

    The baseline controller suffered frequent foot entrapment and catastrophic falls in thick vegetation and mud, preventing speed and COT measurement in vegetation. The learned controller sustained locomotion without human intervention, achieving higher speeds and lower energy expenditure per unit distance across all terrains. The controller was also deployed on two ANYmal-B quadrupeds in the DARPA Subterranean Challenge Urban Circuit across four 60-minute missions (including stair traversal with 18 cm18\text{ cm} rise at a ∼45∘\sim 45^\circ incline), maintaining a zero failure rate throughout the competition.

  8. Knowl 8 — Ablations on Sequence Modeling, Privileged Distillation, and Curriculum

    empirical result

    Controlled ablation experiments validate the necessity of the sequence model, privileged training, and adaptive curriculum components:

    1. Proprioceptive Memory Length (NN in TCN-NN):

      • Receptive field lengths evaluated: N=1N=1 (0.02 s0.02\text{ s}), N=20N=20 (0.4 s0.4\text{ s}), and N=100N=100 (2.0 s2.0\text{ s}).
      • While slope locomotion speed is comparable across memory lengths, step negotiation success drops dramatically with short memory (TCN-1 fails completely on steps ≥18 cm\ge 18\text{ cm}, largely due to hind-leg trapping).
      • Under a continuous lateral disturbance force of 50 N50\text{ N} applied to the base for 5 s5\text{ s}, TCN-100 exhibits a 35.5%35.5\% lower trajectory deviation compared to TCN-1.
      • A Gated Recurrent Unit (GRU) student achieves performance intermediate between TCN-20 and TCN-100 on steps, but incurs 3×3\times higher stochastic gradient descent update time (0.152 s0.152\text{ s} vs 0.0507 s0.0507\text{ s} for TCN-100).
    2. Privileged Distillation:

      • Directly training a proprioceptive TCN-20 policy via TRPO without the two-stage privileged distillation framework fails to learn locomotion or balance, resulting in near-zero reward and premature episode termination throughout training.
      • Omitting the latent embedding loss term (lˉt−lt)2(\bar{l}_t - l_t)^2 from the student objective (naive imitation learning) significantly reduces step traversal success rates.
    3. Adaptive Curriculum:

      • Training a teacher policy with uniform random terrain sampling causes early falls on excessively difficult profiles, resulting in lower training reward plateaus, shorter mean episode lengths, and lower step/slope success rates compared to training with the particle-filter curriculum.
  9. Knowl 9 — Robustness to Unmodeled Payloads and Low-Friction Slippage

    empirical result

    The controller exhibits zero-shot robustness under severe dynamic and physical model mismatches without fine-tuning:

    1. Unmodeled Payload (10 kg10\text{ kg}, representing 22.7%22.7\% of total robot mass):

      • When commanded to climb vertical steps with the 10 kg10\text{ kg} payload attached, the learned controller successfully traverses step heights up to 13.4 cm13.4\text{ cm}. The model-based baseline fails on all step heights under all commanded speeds with the payload.
      • In flat-ground omnidirectional tracking across 8 commanded directions, the learned controller maintains isotropic speed ( ≈0.4 m/s\,\approx 0.4\text{ m/s}) and an average heading error within 10∘10^\circ both with and without the payload. The baseline exhibits an anisotropic velocity profile with lateral heading errors reaching ∼30∘\sim 30^\circ and experiences catastrophic falls at commanded speeds of 0.6 m/s0.6\text{ m/s}.
    2. Low-Friction Ground (Moistened Whiteboard):

      • On a moistened, slippery whiteboard surface, the model-based baseline violently swings its legs upon slip detection and falls. The learned proprioceptive controller adapts its foot placement and velocity dynamically, preventing loss of balance and successfully tracking the commanded heading.
  10. Knowl 10 — Limitations of Blind Single-Gait Quadrupedal Locomotion

    limitation

    The presented locomotion controller possesses two primary operational limitations:

    1. Gait Pattern Restriction: The controller exclusively converges to a diagonal trot gait. Although the physical ANYmal quadruped hardware is capable of diverse gait patterns (such as pacing, bounding, or walking), the learned policy does not autonomously discover alternative gaits without explicit diversity-inducing training objectives.
    2. Lack of Exteroceptive Anticipation: Operating strictly on proprioception prevents the controller from anticipating non-traversable or catastrophic terrain discontinuities in advance (such as cliffs, deep drop-offs, or impassable obstacles). Consequently, the robot adopts a conservative tactile probing gait, feeling out local terrain compliance and elevation changes only upon bodily contact.

Coverage note — None was omitted; all key contributions—including the two-stage privileged learning framework, adaptive curriculum algorithm, control architecture, mathematical formulations, internal representation probing, emergent reflexes, environmental benchmarks, ablation studies, robustness tests, and limitations—have been fully captured.

References

  1. 1.F. Jenelten, J. Hwangbo, F. Tresoldi, C. D. Bellicoso, M. Hutter, Dynamic locomotion on slippery ground, IEEE Robotics and Automation Letters 4170–4176 (2019).
  2. 2.G. Bledt, P. M. Wensing, S. Ingersoll, S. Kim, Contact model fusion for event-based locomotion in unstructured terrains, 2018 IEEE International Conference on Robotics and Automation (ICRA) (IEEE, 2018).
  3. 3.M. Focchi, R. Orsolino, M. Camurri, V. Barasuol, C. Mastalli, D. G. Caldwell, C. Semini. Heuristic planning for rough terrain locomotion in presence of external disturbances and variable perception quality. Advances in Robotics Research: From Lab to Market (Springer, 2020), 165–209.
  4. 4.J. Reher, W. Ma, A. D. Ames, Dynamic walking with compliance on a Cassie bipedal robot, European Control Conference, 2589–2595 (IEEE, 2019).
  5. 5.Y. Gong, R. Hartley, X. Da, A. Hereid, O. Harib, J. Huang, J. W. Grizzle, Feedback control of a Cassie bipedal robot: Walking, standing, and riding a Segway, American Control Conference, 4559–4566 (IEEE, 2019).
  6. 6.J. Hwangbo, C. D. Bellicoso, P. Fankhauser, M. Huttery, Probabilistic foot contact estimation by fusing information from dynamics and differential/forward kinematics, 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 3872–3878 (IEEE, 2016).
  7. 7.M. Camurri, M. Fallon, S. Bazeille, A. Radulescu, V. Barasuol, D. G. Caldwell, C. Semini, Probabilistic contact estimation and impact detection for state estimation of quadruped robots, IEEE Robotics and Automation Letters 1023–1030 (2017).
  8. 8.M. Focchi, V. Barasuol, M. Frigerio, D. G. Caldwell, C. Semini. Slip detection and recovery for quadruped robots. Robotics Research (Springer, 2018), 185–199.
  9. 9.M. Blösch, C. Gehring, P. Fankhauser, M. Hutter, M. A. Hoepflinger, R. Siegwart, State estimation for legged robots on unstable and slippery terrain, 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, 6058–6064 (IEEE, 2013).
  10. 10.C. Gehring, C. D. Bellicoso, S. Coros, M. Bloesch, P. Fankhauser, M. Hutter, R. Siegwart, Dynamic trotting on slopes for quadrupedal robots, 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 5129–5135 (IEEE, 2015).
  11. 11.R. Hartley, J. Mangelson, L. Gan, M. G. Jadidi, J. M. Walls, R. M. Eustice, J. W. Grizzle, Legged robot state-estimation through combined forward kinematic and preintegrated contact factors, 2018 IEEE International Conference on Robotics and Automation (ICRA), 1–8 (IEEE, 2018).
  12. 12.J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, M. Hutter, Learning agile and dynamic motor skills for legged robots, Science Robotics p. eaau5872 (2019).
  13. 13.T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, S. Levine, Learning to walk via deep reinforcement learning, Robotics: Science and Systems (2019).
  14. 14.Z. Xie, P. Clary, J. Dao, P. Morais, J. Hurst, M. van de Panne, Learning locomotion skills for Cassie: Iterative design and sim-to-real, Conference on Robot Learning (2019).
  15. 15.J. Lee, J. Hwangbo, M. Hutter, Robust recovery controller for a quadrupedal robot using deep reinforcement learning, arXiv:1901.07517 (2019).
  16. 16.J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, V. Vanhoucke, Sim-to-real: Learning agile locomotion for quadruped robots, Robotics: Science and Systems (2018).
  17. 17.Y. Yang, K. Caluwaerts, A. Iscen, T. Zhang, J. Tan, V. Sindhwani, Data efficient reinforcement learning for legged robots, Conference on Robot Learning (2019).
  18. 18.S. Ha, P. Xu, Z. Tan, S. Levine, J. Tan, Learning to walk in the real world with minimal human effort, arXiv:2002.08550 (2020).
  19. 19.X. B. Peng, E. Coumans, T. Zhang, T.-W. Lee, J. Tan, S. Levine, Learning agile robotic locomotion skills by imitating animals, arXiv:2004.00784 (2020).
  20. 20.M. Hutter, C. Gehring, D. Jud, A. Lauber, C. D. Bellicoso, V. Tsounis, J. Hwangbo, K. Bodie, P. Fankhauser, M. Bloesch, R. Diethelm, S. Bachmann, A. Melzer, M. A. Höpflinger, ANYmal - a highly mobile and dynamic quadrupedal robot, IEEE/RSJ International Conference on Intelligent Robots and Systems, 38–44 (IEEE, 2016).
  21. 21.X. B. Peng, M. Andrychowicz, W. Zaremba, P. Abbeel, Sim-to-real transfer of robotic control with dynamics randomization, IEEE International Conference on Robotics and Automation (ICRA) (IEEE, 2018).
  22. 22.S. Bai, J. Z. Kolter, V. Koltun, An empirical evaluation of generic convolutional and recurrent networks for sequence modeling, arXiv:1803.01271 (2018).
  23. 23.D. Chen, B. Zhou, V. Koltun, P. Krähenbühl, Learning by cheating, Conference on Robot Learning (2019).
  24. 24.J. C. Brant, K. O. Stanley, Minimal criterion coevolution: a new approach to open-ended search, Genetic and Evolutionary Computation Conference, 67–74 (2017).
  25. 25.R. Wang, J. Lehman, J. Clune, K. O. Stanley, Paired open-ended trailblazer (poet): Endlessly generating increasingly complex and diverse learning environments and their solutions, arXiv:1901.01753 (2019).
  26. 26.C. D. Bellicoso, F. Jenelten, C. Gehring, M. Hutter, Dynamic locomotion through online nonlinear motion optimization for quadrupedal robots, IEEE Robotics and Automation Letters 2261–2268 (2018).
  27. 27.P. Fankhauser, M. Bloesch, C. Gehring, M. Hutter, R. Siegwart. Robot-centric elevation mapping with uncertainty estimates. Mobile Service Robotics (World Scientific, 2014), 433–440.
  28. 28.S. Collins, A. Ruina, R. Tedrake, M. Wisse, Efficient bipedal robots based on passive-dynamic walkers, Science 1082–1085 (2005).
  29. 29.Ghost Robotics, Vision 60: Latest blind-mode stress testing of V60 legged robot, www.youtube.com/watch?v=tQsLauQWp8M (2019).
  30. 30.J. Hwangbo, J. Lee, M. Hutter, Per-contact iteration method for solving contact dynamics, IEEE Robotics and Automation Letters 895–902 (2018).
  31. 31.E. Coumans, others, Bullet physics library, Open source: bulletphysics.org (2013).
  32. 32.R. Smith, others, Open dynamics engine, Open source: ode.org (2005).
  33. 33.R. M. Alexander, Principles of Animal Locomotion (Princeton University Press, 2003).
  34. 34.A. Iscen, K. Caluwaerts, J. Tan, T. Zhang, E. Coumans, V. Sindhwani, V. Vanhoucke, Policies modulating trajectory generators, Conference on Robot Learning, 916–926 (2018).
  35. 35.V. Barasuol, J. Buchli, C. Semini, M. Frigerio, E. R. De Pieri, D. G. Caldwell, A reactive controller framework for quadrupedal locomotion on challenging terrain, 2013 IEEE International Conference on Robotics and Automation, 2554–2561 (IEEE, 2013).
  36. 36.J. Schulman, S. Levine, P. Abbeel, M. Jordan, P. Moritz, Trust region policy optimization, International Conference on Machine Learning, 1889–1897 (2015).
  37. 37.M. Bloesch, M. Hutter, M. A. Hoepflinger, S. Leutenegger, C. Gehring, C. D. Remy, R. Siegwart, State estimation for legged robots-consistent fusion of leg kinematics and imu, Robotics 17–24 (2013).
  38. 38.S. Ross, G. Gordon, D. Bagnell, A reduction of imitation learning and structured prediction to no-regret online learning, International Conference on Artificial Intelligence and Statistics, 627–635 (2011).
  39. 39.C. Florensa, D. Held, X. Geng, P. Abbeel, Automatic goal generation for reinforcement learning agents, International Conference on Machine Learning, 1514–1523 (2018).
  40. 40.J. Lehman, K. O. Stanley, Revising the evolutionary computation abstraction: minimal criteria novelty search, Genetic and Evolutionary Computation Conference, 103–110 (2010).
  41. 41.T. Matiisen, A. Oliver, T. Cohen, J. Schulman, Teacher-student curriculum learning, IEEE transactions on neural networks and learning systems (2019).
  42. 42.W. Yu, G. Turk, C. K. Liu, Learning symmetric and low-energy locomotion, ACM Transactions on Graphics (TOG) p. 144 (2018).
  43. 43.I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, others, Solving rubik’s cube with a robot hand, arXiv:1910.07113 (2019).
  44. 44.R. O. Chavez-Garcia, J. Guzzi, L. M. Gambardella, A. Giusti, Learning ground traversability from simulations, IEEE Robotics and Automation Letters 1695–1702 (2018).
  45. 45.A. Kendall, Y. Gal, What uncertainties do we need in bayesian deep learning for computer vision?, Advances in neural information processing systems, 5574–5584 (2017).
  46. 46.K. Simonyan, A. Vedaldi, A. Zisserman, Deep inside convolutional networks: Visualising image classification models and saliency maps, arXiv:1312.6034 (2013).
  47. 47.G. A. Pratt, M. M. Williamson, Series elastic actuators, IEEE/RSJ International Conference on Intelligent Robots and Systems, 399–406 (1995).
  48. 48.R. M. Smelik, K. J. De Kraker, T. Tutenel, R. Bidarra, S. A. Groenewegen, A survey of procedural methods for terrain modelling, Proceedings of the CASA Workshop on 3D Advanced Media In Gaming And Simulation (3AMIGAS), 25–34 (2009).
  49. 49.A. Lagae, S. Lefebvre, R. Cook, T. DeRose, G. Drettakis, D. S. Ebert, J. P. Lewis, K. Perlin, M. Zwicker, A survey of procedural noise functions, Computer Graphics Forum, 2579–2600 (Wiley Online Library, 2010).
  50. 50.J. Chung, C. Gulcehre, K. Cho, Y. Bengio, Empirical evaluation of gated recurrent neural networks on sequence modeling, arXiv:1412.3555 (2014).
  51. 51.R. J. Williams, J. Peng, An efficient gradient-based algorithm for on-line training of recurrent network trajectories, Neural Computation 490–501 (1990).
  52. 52.D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, International Conference on Learning Representations (2015).

Citation

MLA
Lee, J., et al. “Learning Quadrupedal Locomotion over Challenging Terrain”. Science Robotics, vol. 5, no. 47, 2020, https://doi.org/10.1126/scirobotics.abc5986.
APA
Lee, J., Hwangbo, J., Wellhausen, L., Koltun, V., & Hutter, M. (2020). Learning quadrupedal locomotion over challenging terrain. Science Robotics, 5(47). https://doi.org/10.1126/scirobotics.abc5986
Chicago
Lee, J., J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter. 2020. “Learning Quadrupedal Locomotion over Challenging Terrain”. Science Robotics 5 (47). https://doi.org/10.1126/scirobotics.abc5986.
Harvard
Lee, J. et al. (2020) “Learning quadrupedal locomotion over challenging terrain”, Science Robotics, 5(47). Available at: https://doi.org/10.1126/scirobotics.abc5986.
Vancouver
1. Lee J, Hwangbo J, Wellhausen L, Koltun V, Hutter M (2020) Learning quadrupedal locomotion over challenging terrain. Science Robotics. https://doi.org/10.1126/scirobotics.abc5986

BibTeX

@article{Lee_2020, title={Learning quadrupedal locomotion over challenging terrain}, volume={5}, ISSN={2470-9476}, url={http://dx.doi.org/10.1126/scirobotics.abc5986}, DOI={10.1126/scirobotics.abc5986}, number={47}, journal={Science Robotics}, publisher={American Association for the Advancement of Science (AAAS)}, author={Lee, Joonho and Hwangbo, Jemin and Wellhausen, Lorenz and Koltun, Vladlen and Hutter, Marco}, year={2020}, month=Oct }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF