Learning agile and dynamic motor skills for legged robots

Jemin HwangboJoonho LeeAlexey DosovitskiyDario BellicosoVassilios TsounisVladlen KoltunMarco Hutter

article2019Science Robotics1,877 citations

Develops a sim-to-real reinforcement learning pipeline that trains neural network policies in simulation and deploys them on physical quadruped robots, achieving fast running speeds, energy-efficient velocity tracking, and autonomous fall recovery.

Listen

Legged robots offer significant potential for operating in complex, unstructured environments where wheeled or tracked machines struggle, such as disaster recovery sites, construction zones, and industrial inspections. However, creating control systems capable of agile and dynamic locomotion remains a primary bottleneck. Traditional model-based control frameworks rely on simplified analytical models and laborious manual tuning that often require months of engineering per behavior, restrict the machine to narrow operating ranges, and require substantial onboard computing power. While reinforcement learning (a machine learning approach based on trial and error) offers an automated alternative, training directly on physical hardware risks severe damage and requires weeks or months of operation. Conversely, controllers trained purely in simulation typically fail on physical hardware because of the reality gap—the discrepancy between simulated physics and real-world dynamics, particularly concerning complex actuator behavior and hardware delays.

The article sets out to demonstrate a simulation-to-reality reinforcement learning methodology that trains neural network control policies entirely in simulation and deploys them directly on a physical, full-scale quadruped robot without real-world fine-tuning. The primary objective was to evaluate whether this framework could surpass existing state-of-the-art model-based controllers in command tracking accuracy, energy efficiency, top operational speed, and dynamic fall recovery on ANYmal, a 32-kilogram, medium-dog-sized quadruped.

To achieve this, the authors created a fast hybrid simulation pipeline. Instead of relying on analytical actuator formulas, they trained a compact neural network (an actuator network) directly on less than four minutes of physical hardware data to capture nonlinear motor dynamics, joint compliance, and communication latencies. This learned actuator model was integrated into an in-house rigid-body contact simulator, allowing simulated training to run approximately 1,000 times faster than real time on standard desktop hardware without specialized computing clusters. The researchers applied a reinforcement learning algorithm (Trust Region Policy Optimization) within a curriculum-learning structure, randomizing physical parameters such as robot mass, inertia, and sensor noise to ensure robust real-world transfer.

The findings show substantial improvements across key performance metrics when tested on physical hardware. The learned locomotion policy significantly outperformed the existing state-of-the-art modular controller, reducing base linear velocity tracking error by approximately 95% and turning rate error by about 60% during random command sequences. Furthermore, the robot broke its previous hardware speed record by 25%, achieving a top forward speed of 1.5 meters per second by fully utilizing the torque and velocity limits of the motors. In terms of efficiency, the learned controller reduced average joint torque by 23% to 36% compared to prior controllers by autonomously discovering a more upright nominal leg posture. Finally, the framework yielded a fully autonomous fall-recovery controller that consistently restored the robot to a standing position from arbitrary upside-down and resting configurations in under three seconds on the physical machine, a feat not previously demonstrated on a quadruped of this complexity.

These results demonstrate that data-driven simulation learning can radically accelerate robotic development timelines while delivering superior operational performance. Shifting the heavy computational workload from runtime planning to offline training reduced onboard processing demands to less than 0.25% of a single processor core, eliminating the need for bulky, external computers. Over three months of rigorous hardware testing, the learned control policies proved resilient against mechanical changes and wear, including structural weight alterations of approximately 2 kilograms and threefold increases in drive stiffness. This confirms that robust simulation-to-reality pipelines can mitigate hardware risks and reduce long-term maintenance costs.

Moving forward, development teams should adopt data-driven actuator modeling and simulation-based training pipelines to minimize manual tuning and deploy performant robotic controllers quickly. Because individual neural network policies remain specialized to single task distributions, future research should integrate hierarchical control structures that enable a robot to transition autonomously among diverse behaviors such as navigation, stair climbing, and recovery. Organizations should establish structured data-logging practices for their hardware platforms to easily train initial actuator networks.

While the results demonstrate high confidence and repeatability on the ANYmal platform, several operational boundaries apply. Actuators with coupled dynamics, such as hydraulic systems sharing a common fluid reservoir, may require unified multi-actuator neural network models rather than the independent joint models used here. In addition, designing task-specific reward functions still requires human engineering judgment to prevent unsafe hardware behaviors, such as high-impact collisions during fall recovery. Finally, the demonstrated locomotion policies relied on flat or minimally disturbed terrain; further validation in highly cluttered or three-dimensional outdoor environments is necessary before deploying the system in critical field missions.

Cover for Learning agile and dynamic motor skills for legged robots

Abstract

Legged robots pose one of the greatest challenges in robotics. Dynamic and agile maneuvers of animals cannot be imitated by existing methods that are crafted by humans. A compelling alternative is reinforcement learning, which requires minimal craftsmanship and promotes the natural evolution of a control policy. However, so far, reinforcement learning research for legged robots is mainly limited to simulation, and only few and comparably simple examples have been deployed on real systems. The primary reason is that training with real robots, particularly with dynamically balancing systems, is complicated and expensive. In the present work, we introduce a method for training a neural network policy in simulation and transferring it to a state-of-the-art legged system, thereby leveraging fast, automated, and cost-effective data generation schemes. The approach is applied to the ANYmal robot, a sophisticated medium-dog-sized quadrupedal system. Using policies trained in simulation, the quadrupedal machine achieves locomotion skills that go beyond what had been achieved with prior methods: ANYmal is capable of precisely and energy-efficiently following high-level body velocity commands, running faster than before, and recovering from falling even in complex configurations.

Table of Contents

  • References

Knowls

  1. Knowl 1 — Data-Driven Actuator Modeling with Neural Networks for Series-Elastic Actuators

    model/method

    To accurately model the non-linear dynamics, hardware/communication delays, cascaded internal feedback loops (such as Field-Oriented Control and joint-level PID loops), and mechanical compliance of Series-Elastic Actuators (SEAs) without manual parameter tuning, the actuator dynamics are approximated via an "actuator network" trained with supervised learning.

    The actuator network is a multi-layer perceptron (MLP) with 3 hidden layers of 32 units each and softsign activation functions: softsign(x)=x1+∣x∣\text{softsign}(x) = \frac{x}{1 + |x|}

    The network takes as input the joint state history consisting of joint position errors ϕ∗−ϕ\phi^* - \phi (commanded position minus measured position) and joint velocities ϕ˙\dot{\phi} at the current time tt and two preceding time steps (t−0.01 st - 0.01\,\text{s} and t−0.02 st - 0.02\,\text{s}), and outputs the predicted joint torque τ∈R\tau \in \mathbb{R}.

    Training data is collected directly on the 12 physical actuators of the ANYmal quadruped at 400 Hz400\,\text{Hz} in under 4 minutes (>106>10^6 samples) using parameterized sinusoidal foot trajectories with frequencies between 1–25 Hz1\text{--}25\,\text{Hz} and amplitudes between 5–10 cm5\text{--}10\,\text{cm}, subjected to manual external disturbances. Evaluating the softsign actuator network for all 12 joints takes 12.2 μs12.2\,\mu\text{s} (compared to 31.6 μs31.6\,\mu\text{s} with tanh⁡\tanh). On a validation dataset, the network achieves an average root-mean-square torque prediction error of 0.740 Nm0.740\,\text{Nm} (close to the 0.2 Nm0.2\,\text{Nm} sensor resolution), compared to 3.55 Nm3.55\,\text{Nm} for an ideal actuator model.

  2. Knowl 2 — Low-Impedance Joint Position Setpoints as Policy Action Space

    model/method

    The control policy outputs desired joint position targets ϕ∗∈R12\phi^* \in \mathbb{R}^{12} rather than direct joint torques. These commanded positions are converted to joint torques τ\tau using an impedance control formulation with fixed low feedback gains and zero target velocity: τ=kp(ϕ∗−ϕ)−kdϕ˙\tau = k_p (\phi^* - \phi) - k_d \dot{\phi} where ϕ\phi and ϕ˙\dot{\phi} are the measured joint positions and velocities, kp=50 Nm/radk_p = 50\,\text{Nm/rad}, and kd=0.1 Nm/(rad/s)k_d = 0.1\,\text{Nm/(rad/s)}.

    The position gain kpk_p is chosen to match the nominal torque range of the robot (pm30 Nm\\pm 30\,\text{Nm}) divided by the nominal range of joint motion (pm0.6 rad\\pm 0.6\,\text{rad}). This parameterization ensures that an initial, untrained random policy behaves as a compliant standing controller rather than causing erratic motions that terminate early in falls, improving exploration and training stability.

  3. Knowl 3 — State Observation Formulation with Projected Gravity and Joint History for Contact Estimation

    model/method

    At each discrete time step tkt_k, the observation vector oko_k fed to the policy network is defined as: ok=⟨ϕg,rz,v,ω,ϕ,ϕ˙,Θ,ak−1,C⟩o_k = \langle \phi^g, r_z, v, \omega, \phi, \dot{\phi}, \Theta, a_{k-1}, C \rangle where:

    • ϕg∈S2\phi^g \in S^2 is the unit vector pointing in the direction of the gravity vector expressed in the robot's Inertial Measurement Unit (IMU) coordinate frame, representing the observable orientation degrees of freedom.
    • rz∈Rr_z \in \mathbb{R} is the estimated base height derived from leg kinematics under a flat-ground assumption using a 1D Kalman filter (omitted when training fall recovery).
    • v∈R3v \in \mathbb{R}^3 and ω∈R3\omega \in \mathbb{R}^3 are the linear and angular velocities of the base body.
    • ϕ∈R12\phi \in \mathbb{R}^{12} and ϕ˙∈R12\dot{\phi} \in \mathbb{R}^{12} are the measured joint positions and velocities.
    • Θ=⟨ϕ(tk−0.01 s),ϕ˙(tk−0.01 s),ϕ(tk−0.02 s),ϕ˙(tk−0.02 s)⟩\Theta = \langle \phi(t_k - 0.01\,\text{s}), \dot{\phi}(t_k - 0.01\,\text{s}), \phi(t_k - 0.02\,\text{s}), \dot{\phi}(t_k - 0.02\,\text{s}) \rangle is the joint state history, which enables the policy to infer ground contact states implicitly without requiring foot force sensors.
    • ak−1∈R12a_{k-1} \in \mathbb{R}^{12} is the previous action (joint position target).
    • CC is the commanded high-level velocity vector (forward velocity, lateral velocity, and yaw rate).
  4. Knowl 4 — Curriculum Factor for Balancing Task Objectives and Control Regularization

    model/method

    To prevent the reinforcement learning agent from converging to a trivial local optimum (such as standing completely still to avoid velocity/torque penalties), a multiplicative curriculum factor kc∈(0,1]k_c \in (0, 1] modulates all regularization, smoothness, and constraint cost terms while leaving the primary task tracking cost unscaled.

    The progression of the curriculum factor across training iterations jj is defined by: kc,0=k0k_{c, 0} = k_0 kc,j+1←(kc,j)kdk_{c, j+1} \leftarrow (k_{c, j})^{k_d} where k0=0.3k_0 = 0.3 is the initial curriculum factor and kd=0.997∈(0,1)k_d = 0.997 \in (0, 1) is the advance rate.

    The sequence {kc,j}\{k_{c, j}\} is strictly monotonically increasing and asymptotically converges to 11. Early in training (kc≈0.3k_c \approx 0.3), the policy focuses primarily on achieving the velocity or reorientation objective; as training progresses and kc→1k_c \to 1, joint speed, torque, slip, and smoothness penalties are steadily enforced to refine energy efficiency and stability.

  5. Knowl 5 — Bounded Logistic Error Kernel for Policy Cost Functions

    equation

    Tracking errors in reinforcement learning cost functions are mapped to a bounded reward using a logistic kernel K:R→[−0.25,0)K: \mathbb{R} \to [-0.25, 0): K(x)=−1ex+2+e−xK(x) = -\frac{1}{e^x + 2 + e^{-x}} where x∈Rx \in \mathbb{R} represents a scaled kinematic or task tracking error (e.g., base linear or angular velocity error).

    Unlike an unbounded Euclidean norm (x2x^2), which assigns large negative penalties early in training and makes early episode termination (such as falling) a locally optimal strategy, the logistic kernel is strictly lower-bounded by −0.25-0.25 at zero error and approaches 00 as ∣x∣→∞|x| \to \infty, ensuring that episode continuation remains preferred over termination.

  6. Knowl 6 — Sim-to-Real Domain Randomization and Observation Noise Injection

    experimental setup

    To transfer policies trained in simulation to the ANYmal quadruped (weight ≈32 kg\approx 32\,\text{kg}, 12 degrees of freedom), the simulation environment combines a hard-contact dynamics solver respecting Coulomb friction cones with stochastic dynamics and noise randomization:

    • Inertial randomization: The policy is trained across 30 randomized ANYmal models where link masses are sampled with uniform variations U(−15,15)%U(-15, 15)\%, center of mass positions are perturbed by U(−2,2) cmU(-2, 2)\,\text{cm}, and joint positions are perturbed by U(−2,2) cmU(-2, 2)\,\text{cm}.
    • Observation noise: Uniform noise is injected during simulation into joint velocity measurements (U(−0.5,0.5) rad/sU(-0.5, 0.5)\,\text{rad/s} to simulate numerical differentiation error), base linear velocity (U(−0.08,0.08) m/sU(-0.08, 0.08)\,\text{m/s}), and base angular velocity (U(−0.16,0.16) rad/sU(-0.16, 0.16)\,\text{rad/s}).
    • Policy training setup: Policies are parameterized as MLPs with two hidden layers (256 and 128 units, tanh⁡\tanh activations) and trained using Trust Region Policy Optimization (TRPO) with a discount factor γ=0.9988\gamma = 0.9988 (half-life 5.77 s5.77\,\text{s}) for locomotion and γ=0.993\gamma = 0.993 (half-life 4.93 s4.93\,\text{s}) for recovery. Training processes ∼2.5×108\sim 2.5 \times 10^8 transitions in ∼4–11\sim 4\text{--}11 hours on a single standard desktop computer (1 CPU, 1 GPU) at ≈500,000\approx 500{,}000 simulation steps per second.
  7. Knowl 7 — Command-Conditioned Locomotion Performance vs. State-of-the-Art Model-Based Control

    empirical result

    When evaluated on the physical ANYmal robot under randomized joystick velocity commands (forward velocity in [−1.0,1.0] m/s[-1.0, 1.0]\,\text{m/s}, lateral velocity in [−0.4,0.4] m/s[-0.4, 0.4]\,\text{m/s}, yaw rate in [−1.2,1.2] rad/s[-1.2, 1.2]\,\text{rad/s}), the learned reinforcement learning policy substantially outperformed the best existing model-based controller (Bellicoso et al., 2018):

    • Tracking accuracy: Average linear velocity error was 0.143 m/s0.143\,\text{m/s} for the learned controller vs. 0.231 m/s0.231\,\text{m/s} for the model-based flying trot (a 95%95\% higher error in the baseline). Average yaw rate error was 0.174 rad/s0.174\,\text{rad/s} vs. 0.278 rad/s0.278\,\text{rad/s} (a 60%60\% higher error in the baseline).
    • Energy efficiency: Average joint torque was reduced from 11.7 Nm11.7\,\text{Nm} to 8.23 Nm8.23\,\text{Nm} (23%ext−−36%23\% ext{--}36\% less torque across speed ranges), and average mechanical power decreased from 97.3 W97.3\,\text{W} to 78.1 W78.1\,\text{W}.
    • Gait emergence and posture: The policy walked with 10∘ext−−15∘10^\circ ext{--}15^\circ straighter nominal knee posture (which cannot be hand-tuned into the baseline controller without causing falls) and autonomously manifested a flying trot at high speeds and walking trot at low speeds without pre-specified contact schedules.
  8. Knowl 8 — Sim-to-Real Transfer Failure of Ideal and Analytical Actuator Models

    empirical result

    Locomotion policies trained in simulation under identical RL procedures but using alternative actuator models failed to transfer to the physical ANYmal robot:

    • Ideal actuator model: Assuming infinite bandwidth and zero latency resulted in policies that could not take a single step on the real robot, causing violent shaking and immediate falls.
    • Analytical actuator model: An analytical model incorporating identified dynamics, CAD parameters, joint friction, damping, low-level PID/FOC code, and hand-tuned latency (tuned over more than a week) also failed to take a single step on hardware, causing severe oscillations due to amplitude-dependent response delays inherent in SEAs.

    Only the policy trained with the data-driven neural network actuator model successfully bridged the reality gap and transferred directly to hardware without real-world fine-tuning.

  9. Knowl 9 — Autonomous High-Speed Locomotion at Hardware Limits

    empirical result

    A high-speed locomotion policy trained with command distributions up to 1.6 m/s1.6\,\text{m/s} forward velocity enabled ANYmal to reach a maximum forward speed of 1.5 m/s1.5\,\text{m/s} on the physical robot (1.58 m/s1.58\,\text{m/s} in simulation), exceeding ANYmal's previous speed record of 1.2 m/s1.2\,\text{m/s} by 25%25\%.

    The learned controller operated at the full hardware capacity of the robot, saturating both the maximum joint torque limit (40 Nm40\,\text{Nm}) and the maximum joint velocity limit (12 rad/s12\,\text{rad/s}). The policy converged to an asymmetric flying trot gait characterized by extended flight phases, an optimization mode distinct from standard animal gaits.

  10. Knowl 10 — Dynamic Fall Recovery and Self-Righting Controller

    model/method

    A dynamic fall recovery controller was trained to restore ANYmal to an upright standing posture from arbitrary fallen orientations. The simulation utilized a 41-body rigid collision model representing the chassis, links, and appendages to resolve complex internal and environmental multi-contact dynamics.

    Initialization for training involved dropping the robot from a 1.0 m1.0\,\text{m} height with randomized orientations and joint configurations, simulating for 1.2 s1.2\,\text{s} until settled. Cost terms penalized base roll orientation deviation, joint acceleration, joint velocities exceeding 8 rad/s8\,\text{rad/s}, contact slip velocity, and non-foot body contact impulses.

    When deployed on the physical robot across 9 challenging initial configurations (including completely upside-down states and leg-entangled configurations), the controller achieved a 100%100\% success rate, dynamically swinging limbs and leveraging momentum to roll over and stand up in less than 3 seconds.

Coverage note — All primary contributions—actuator modeling, state/action formulations, curriculum reward shaping, logistic error kernels, domain randomization, locomotion experiments, actuator ablations, high-speed limits, and fall recovery—are covered; standard implementation subroutines like baseline TRPO updates and specific CAD dimensions were omitted as background.

References

  1. 1.M. Raibert, K. Blankespoor, G. Nelson, R. Playter, Bigdog, the roughterrain quadruped robot, Proceedings of the 2008 International Federation of Automatic Control, 10822–10825 (Elsevier, 2008).
  2. 2.G. Nelson, A. Saunders, N. Neville, B. Swilling, J. Bondaryk, D. Billings, C. Lee, R. Playter, M. Raibert, Petman: A humanoid robot for testing chemical protective clothing, Journal of the Robotics Society of Japan 372–377 (2012).
  3. 3.S. Seok, A. Wang, M. Y. Chuah, D. Otten, J. Lang, S. Kim, Design principles for highly efficient quadrupeds and implementation on the mit cheetah robot, IEEE International Conference on Robotics and Automation, 3307–3312 (IEEE, 2013).
  4. 4.Spotmini autonomous navigation, https://youtu.be/Ve9kWX_KXus. Accessed: 2018-08-11.
  5. 5.M. Hutter, C. Gehring, D. Jud, A. Lauber, C. D. Bellicoso, V. Tsounis, J. Hwangbo, K. Bodie, P. Fankhauser, M. Bloesch, R. Diethelm, S. Bachmann, A. Melzer, M. Hoepflinger, Anymal-a highly mobile and dynamic quadrupedal robot, IEEE/RSJ International Conference on Intelligent Robots and Systems, 38–44 (IEEE, 2016).
  6. 6.R. J. Full, D. E. Koditschek, Templates and anchors: neuromechanical hypotheses of legged locomotion on land, Journal of experimental biology 3325–3332 (1999).
  7. 7.M. H. Raibert, J. J. Craig, Hybrid position/force control of manipulators, Journal of Dynamic Systems, Measurement, and Control 126–133 (1981).
  8. 8.J. Pratt, J. Carff, S. Drakunov, A. Goswami, Capture point: A step toward humanoid push recovery, 2006 6th IEEE-RAS International Conference on Humanoid Robots, 200–207 (IEEE, 2006).
  9. 9.A. Goswami, B. Espiau, A. Keramane, Limit cycles in a passive compass gait biped and passivity-mimicking control laws, Autonomous Robots 273–286 (1997).
  10. 10.W. J. Schwind, Spring loaded inverted pendulum running: A plant model., Ph.D. thesis (1998).
  11. 11.M. Kalakrishnan, J. Buchli, P. Pastor, M. Mistry, S. Schaal, Fast, robust quadruped locomotion over challenging terrain, 2010 IEEE International Conference on Robotics and Automation, 2665–2670 (IEEE, 2010).
  12. 12.C. D. Bellicoso, F. Jenelten, C. Gehring, M. Hutter, Dynamic locomotion through online nonlinear motion optimization for quadrupedal robots, IEEE Robotics and Automation Letters 2261–2268 (2018).
  13. 13.M. Neunert, F. Farshidian, A. W. Winkler, J. Buchli, Trajectory optimization through contacts and automatic gait discovery for quadrupeds, IEEE Robotics and Automation Letters 1502–1509 (2017).
  14. 14.I. Mordatch, E. Todorov, Z. Popovic, Discovery of complex behaviors through contact-invariant optimization, ACM Transactions on Graphics (TOG) p. 43 (2012).
  15. 15.F. Farshidian, M. Neunert, A. W. Winkler, G. Rey, J. Buchli, An efficient optimal planning and control framework for quadrupedal locomotion, 2017 IEEE International Conference on Robotics and Automation, 93–100 (IEEE, 2017).
  16. 16.M. Posa, C. Cantu, R. Tedrake, A direct method for trajectory optimization of rigid bodies through contact, The International Journal of Robotics Research 69–81 (2014).
  17. 17.J. Carius, R. Ranftl, V. Koltun, M. Hutter, Trajectory optimization with implicit hard contacts, IEEE Robotics and Automation Letters (2018).
  18. 18.S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, D. Quillen, Learning handeye coordination for robotic grasping with deep learning and large-scale data collection, The International Journal of Robotics Research 421–436 (2018).
  19. 19.R. Tedrake, T. W. Zhang, H. S. Seung, Stochastic policy gradient reinforcement learning on a simple 3d biped, 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2849–2854 (IEEE, 2004).
  20. 20.J. Yosinski, J. Clune, D. Hidalgo, S. Nguyen, J. C. Zagal, H. Lipson, Evolving robot gaits in hardware: the hyperneat generative encoding vs. parameter optimization., 2011 European Conference on Artificial Life, 890–897 (2011).
  21. 21.S. Levine, V. Koltun, Learning complex neural network policies with trajectory optimization, 2014 International Conference on Machine Learning, 829–837 (2014).
  22. 22.J. Schulman, S. Levine, P. Abbeel, M. Jordan, P. Moritz, Trust region policy optimization, International Conference on Machine Learning, 1889–1897 (2015).
  23. 23.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347 (2017).
  24. 24.N. Heess, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, A. Eslami, M. Riedmiller, others, Emergence of locomotion behaviours in rich environments, arXiv preprint arXiv:1707.02286 (2017).
  25. 25.X. B. Peng, G. Berseth, K. Yin, M. Van De Panne, Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning, ACM Transactions on Graphics (TOG) p. 41 (2017).
  26. 26.Z. Xie, G. Berseth, P. Clary, J. Hurst, M. van de Panne, Feedback control for cassie with deep reinforcement learning, arXiv preprint arXiv:1803.05580 (2018).
  27. 27.T. Lee, F. C. Park, A geometric algorithm for robust multibody inertial parameter identification, IEEE Robotics and Automation Letters 2455–2462 (2018).
  28. 28.M. Neunert, T. Boaventura, J. Buchli. Why off-the-shelf physics simulators fail in evaluating feedback controller performance-a case study for quadrupedal robots. Advances in Cooperative Robotics (World Scientific, 2017), 464–472.
  29. 29.J. Bongard, V. Zykov, H. Lipson, Resilient machines through continuous self-modeling, Science 1118–1121 (2006).
  30. 30.D. Nguyen-Tuong, M. Seeger, J. Peters, Model learning with local gaussian process regression, Advanced Robotics 2015–2034 (2009).
  31. 31.D. Nguyen-Tuong, J. Peters, Learning robot dynamics for computed torque control using local gaussian processes regression, 2008 ECSIS Symposium on Learning and Adaptive Behaviors for Robotic Systems, 59–64 (IEEE, 2008).
  32. 32.K. Nikzad, J. Ghaboussi, S. L. Paul, Actuator dynamics and delay compensation using neurocontrollers, Journal of engineering mechanics 966–975 (1996).
  33. 33.R. S. Sutton, A. G. Barto, Reinforcement learning: An introduction, vol. 1 (MIT press Cambridge, 1998).
  34. 34.I. Mordatch, K. Lowrey, E. Todorov, Ensemble-cio: Full-body dynamic motion planning that transfers to physical humanoids, 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems, 5307–5314 (IEEE, 2015).
  35. 35.J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, V. Vanhoucke, Sim-to-real: Learning agile locomotion for quadruped robots, Proceedings of Robotics: Science and Systems (RSS) (2018).
  36. 36.X. B. Peng, M. Andrychowicz, W. Zaremba, P. Abbeel, Sim-toreal transfer of robotic control with dynamics randomization, CoRR abs/1710.06537 (2017).
  37. 37.N. Jakobi, P. Husbands, I. Harvey, Noise and the reality gap: The use of simulation in evolutionary robotics, European Conference on Artificial Life, 704–720 (Springer, 1995).
  38. 38.A. Dosovitskiy, V. Koltun, Learning to act by predicting the future, International Conference on Learning Representations (ICLR) (2017).
  39. 39.C. Gehring, S. Coros, M. Hutter, C. D. Bellicoso, H. Heijnen, R. Diethelm, M. Bloesch, P. Fankhauser, J. Hwangbo, M. Hoepflinger, others, Practice makes perfect: An optimization-based approach to controlling agile motions for a quadruped robot, IEEE Robotics & Automation Magazine 34–43 (2016).
  40. 40.R. Featherstone, Rigid body dynamics algorithms (Springer, 2014).
  41. 41.J. Hwangbo, J. Lee, M. Hutter, Per-contact iteration method for solving contact dynamics, IEEE Robotics and Automation Letters 895–902 (2018).
  42. 42.R. Smith, Open dynamics engine (2005).
  43. 43.D. A. Winter, The biomechanics and motor control of human gait: normal, elderly and pathological. 2, University of Waterloo, Waterloo (1991).
  44. 44.H.-W. Park, S. Park, S. Kim, Variable-speed quadrupedal bounding using impulse planning: Untethered high-speed 3d running of mit cheetah 2, 2015 IEEE international conference on Robotics and automation, 5163–5170 (IEEE, 2015).
  45. 45.Introducing wildcat, https://youtu.be/wE3fmFTtP9g. Accessed: 2018-08-06.
  46. 46.A. W. Winkler, F. Farshidian, D. Pardo, M. Neunert, J. Buchli, Fast trajectory optimization for legged robots using vertex-based zmp constraints, IEEE Robotics and Automation Letters 2201–2208 (2017).
  47. 47.S. Shamsuddin, L. I. Ismail, H. Yussof, N. I. Zahari, S. Bahari, H. Hashim, A. Jaffar, Humanoid robot nao: Review of control and motion exploration, 2011 IEEE International Conference on Control System, Computing and Engineering, 511–516 (IEEE, 2011).
  48. 48.U. Saranli, M. Buehler, D. E. Koditschek, Rhex: A simple and highly mobile hexapod robot, The International Journal of Robotics Research 616–631 (2001).
  49. 49.E. Ackerman, Boston dynamics sand flea robot demonstrates astonishing jumping skills, IEEE Spectrum Robotics Blog 2 (2012).
  50. 50.J. Morimoto, K. Doya, Acquisition of stand-up behavior by a real robot using hierarchical reinforcement learning, Robotics and Autonomous Systems 37–51 (2001).
  51. 51.C. Semini, N. G. Tsagarakis, E. Guglielmino, M. Focchi, F. Cannella, D. G. Caldwell, Design of hyq–a hydraulically and electrically actuated quadruped robot, Proceedings of the Institution of Mechanical Engineers, Part I: Journal of Systems and Control Engineering 831–849 (2011).
  52. 52.N. G. Tsagarakis, S. Morfey, G. M. Cerda, L. Zhibin, D. G. Caldwell, Compliant humanoid coman: Optimal joint stiffness tuning for modal frequency control, 2013 IEEE International Conference on Robotics and Automation, 673–678 (IEEE, 2013).
  53. 53.J. Bergstra, G. Desjardins, P. Lamblin, Y. Bengio, Quadratic polynomials learn better image features, Technical report, 1337 (2009).
  54. 54.J. Schulman, P. Moritz, S. Levine, M. Jordan, P. Abbeel, Highdimensional continuous control using generalized advantage estimation, Proceedings of the International Conference on Learning Representations (2016).
  55. 55.J. Hwangbo, I. Sa, R. Siegwart, M. Hutter, Control of a quadrotor with reinforcement learning, IEEE Robotics and Automation Letters 2096–2103 (2017).
  56. 56.M. Bloesch, M. Hutter, M. A. Hoepflinger, S. Leutenegger, C. Gehring, C. D. Remy, R. Siegwart, State estimation for legged robots-consistent fusion of leg kinematics and imu, Robotics 17–24 (2013).
  57. 57.X. B. Peng, M. van de Panne, Learning locomotion skills using deeprl: does the choice of action space matter?, Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation, p. 12 (ACM, 2017).
  58. 58.W. Yu, G. Turk, C. K. Liu, Learning symmetric and low-energy locomotion, ACM Transactions on Graphics (TOG) p. 144 (2018).
  59. 59.Y. Bengio, J. Louradour, R. Collobert, J. Weston, Curriculum learning, Proceedings of the 26th annual International Conference on Machine Learning, 41–48 (ACM, 2009).
  60. 60.G. A. Pratt, M. M. Williamson, Series elastic actuators, Intelligent Robots and Systems 95.’Human Robot Interaction and Cooperative Robots’, Proceedings. 1995 IEEE/RSJ International Conference on, 399–406 (IEEE, 1995).
  61. 61.M. Hutter, K. Bodie, A. Lauber, J. Hwangbo, EP16181251 - Joint unit, joint system, robot for manipulation and/or transportation, robotic exoskeleton system and method for manipulation and/or transportation (2016).

Citation

MLA
Hwangbo, J., et al. “Learning Agile and Dynamic Motor Skills for Legged Robots”. Science Robotics, vol. 4, no. 26, 2019, https://doi.org/10.1126/scirobotics.aau5872.
APA
Hwangbo, J., Lee, J., Dosovitskiy, A., Bellicoso, D., Tsounis, V., Koltun, V., & Hutter, M. (2019). Learning agile and dynamic motor skills for legged robots. Science Robotics, 4(26). https://doi.org/10.1126/scirobotics.aau5872
Chicago
Hwangbo, J., J. Lee, A. Dosovitskiy, et al. 2019. “Learning Agile and Dynamic Motor Skills for Legged Robots”. Science Robotics 4 (26). https://doi.org/10.1126/scirobotics.aau5872.
Harvard
Hwangbo, J. et al. (2019) “Learning agile and dynamic motor skills for legged robots”, Science Robotics, 4(26). Available at: https://doi.org/10.1126/scirobotics.aau5872.
Vancouver
1. Hwangbo J, Lee J, Dosovitskiy A, Bellicoso D, Tsounis V, Koltun V, Hutter M (2019) Learning agile and dynamic motor skills for legged robots. Science Robotics. https://doi.org/10.1126/scirobotics.aau5872

BibTeX

@article{Hwangbo_2019, title={Learning agile and dynamic motor skills for legged robots}, volume={4}, ISSN={2470-9476}, url={http://dx.doi.org/10.1126/scirobotics.aau5872}, DOI={10.1126/scirobotics.aau5872}, number={26}, journal={Science Robotics}, publisher={American Association for the Advancement of Science (AAAS)}, author={Hwangbo, Jemin and Lee, Joonho and Dosovitskiy, Alexey and Bellicoso, Dario and Tsounis, Vassilios and Koltun, Vladlen and Hutter, Marco}, year={2019}, month=Jan }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF