Sim-to-Real Transfer of Robotic Control with Dynamics Randomization

Xue Bin PengMarcin AndrychowiczWojciech ZarembaPieter Abbeel

article2017IEEE International Conference on Robotics and Automation1,840 citations

Demonstrates that randomizing physical dynamics during simulation training enables robotic control policies to transfer directly to physical hardware without requiring real-world data or exact model calibration.

Listen

Training advanced robotic control systems using reinforcement learning requires massive amounts of trial-and-error data, which presents major safety hazards, hardware wear, and prohibitive time costs when conducted directly on physical machines. While computer simulations offer a fast and risk-free environment for training, policies developed in simulation often fail when deployed in the real world due to unavoidable modeling inaccuracies and calibration errors—a challenge known as the reality gap.

The article demonstrates that training recurrent neural network policies with dynamics randomization allows robotic controllers to transfer directly from low-fidelity simulations to real-world hardware without requiring real-world training data or precise system calibration.

To evaluate this approach, the authors trained a robotic controller exclusively in simulation to push an object to arbitrary target locations using a physical seven-degree-of-freedom robotic arm. The simulated training process applied broad randomization across 95 physical parameters, including object mass, joint friction, table height, control latency, and sensor noise. Rather than using standard memoryless architectures, the policy employed a recurrent neural network—specifically Long Short-Term Memory units—allowing the controller to implicitly infer environmental physical properties over time from its history of actions and sensory states. The model was trained entirely using sparse binary goal rewards combined with hindsight experience replay techniques.

When deployed directly onto the physical robot, the recurrent policy trained with dynamics randomization achieved an 89% success rate, closely matching its 91% success rate in simulation. In contrast, standard feedforward policies trained without randomization failed completely in the real world (0% success), while feedforward networks with access to recent history reached only a 70% success rate. Ablation analyses revealed that randomizing control latency and observation noise was critical, as omitting either dropped real-world success rates to below 30%. Furthermore, the system showed remarkable physical robustness: when friction dynamics were drastically altered by attaching a snack bag to the target puck, the policy maintained a 91% success rate and autonomously adopted complex correction strategies, such as tilting the puck to slide it into position.

These findings indicate that organizations can bypass costly, labor-intensive system calibration and avoid the hazards of real-world exploratory training. By randomizing dynamic simulation parameters, developers can rely on lower-fidelity simulators to produce adaptive controllers that are resilient to latency, sensor inaccuracies, and changing operating environments. This substantially reduces development timelines, equipment wear, and safety risks in robotic deployments.

Organizations developing autonomous manipulation systems should adopt dynamics randomization alongside recurrent memory architectures rather than investing excessive resources into ultra-high-fidelity simulation calibration. Follow-on research and development should prioritize extending this framework to more complex robotic tasks and integrating additional sensory modalities, such as vision.

While the demonstrated performance is high, readers should note that the evaluation was confined to a single non-grasping manipulation task (tabletop pushing) using motion-capture tracking rather than camera vision. Additional validation is required to ensure these results generalize to multi-stage assembly, grasping, and dynamic multi-agent environments.

Cover for Sim-to-Real Transfer of Robotic Control with Dynamics Randomization

Abstract

Simulations are attractive environments for training agents as they provide an abundant source of data and alleviate certain safety concerns during the training process. But the behaviours developed by agents in simulation are often specific to the characteristics of the simulator. Due to modeling error, strategies that are successful in simulation may not transfer to their real world counterparts. In this paper, we demonstrate a simple method to bridge this "reality gap". By randomizing the dynamics of the simulator during training, we are able to develop policies that are capable of adapting to very different dynamics, including ones that differ significantly from the dynamics on which the policies were trained. This adaptivity enables the policies to generalize to the dynamics of the real world without any training on the physical system. Our approach is demonstrated on an object pushing task using a robotic arm. Despite being trained exclusively in simulation, our policies are able to maintain a similar level of performance when deployed on a real robot, reliably moving an object to a desired location from random initial configurations. We explore the impact of various design decisions and show that the resulting policies are robust to significant calibration error.

Table of Contents

  • I INTRODUCTION
  • II RELATED WORK
  • II-A Domain Adaptation
  • II-B Domain Randomization
  • II-C Non-prehensile Manipulation
  • III BACKGROUND
  • III-A Policy Gradient Methods
  • III-B Hindsight Experience Replay
  • IV METHOD
  • IV-A Tasks
  • IV-B State and Action
  • IV-C Dynamics Randomization
  • IV-D Adaptive Policy
  • IV-E Recurrent Deterministic Policy Gradient
  • IV-F Network Architecture
  • V EXPERIMENTS
  • V-A Comparison of Architectures
  • V-B Ablation
  • V-C Robustness
  • VI Conclusions
  • VII Acknowledgement
  • References

Knowls

  1. Knowl 1 — Dynamics Randomization Objective for Policy Generalization

    model/method

    To transfer control policies directly from simulation to the real world without training on the physical robot, training is conducted across a parameterized distribution of simulated dynamics models. Let p∗(st+1∣st,at)p^*(s_{t+1} \mid s_t, a_t) denote the true dynamics of the physical environment, and let p^(st+1∣st,at,μ)\hat{p}(s_{t+1} \mid s_t, a_t, \mu) denote an approximate simulated transition model parameterized by physical dynamics parameters μ\mu (such as mass, damping, friction, latency, and sensor noise).

    Given a distribution over dynamics parameters ρμ\rho_\mu, the reinforcement learning objective is to optimize policy parameters to maximize the expected discounted return over randomized dynamics:

    Eμ∼ρμ[Eτ∼p(τ∣π,μ)[∑t=0T−1r(st,at)]]\mathbb{E}_{\mu \sim \rho_\mu} \left[ \mathbb{E}_{\tau \sim p(\tau \mid \pi, \mu)} \left[ \sum_{t=0}^{T-1} r(s_t, a_t) \right] \right]

    where τ=(s0,a0,s1,…,sT)\tau = (s_0, a_0, s_1, \dots, s_T) is the trajectory sampled under the simulator dynamics p^(⋅∣⋅,⋅,μ)\hat{p}(\cdot \mid \cdot, \cdot, \mu) governed by policy π\pi, r(st,at)r(s_t, a_t) is the reward at timestep tt, and TT is the episode horizon. By exposing the policy to high variability during training, the policy learns control strategies that generalize to unmodeled real-world physical dynamics.

  2. Knowl 2 — Implicit System Identification via Recurrent Policy Architecture

    model/method

    Rather than using an explicit system identification module to predict estimated physical parameters μ^=ϕ(st,ht)\hat{\mu} = \phi(s_t, h_t) from trajectory history ht=[at−1,st−1,at−2,st−2,… ]h_t = [a_{t-1}, s_{t-1}, a_{t-2}, s_{t-2}, \dots], system identification is implicitly learned end-to-end using a recurrent neural network architecture π(at∣st,zt,g)\pi(a_t \mid s_t, z_t, g), where zt=z(ht)z_t = z(h_t) represents an internal memory state summarizing past observations.

    The policy network consists of two branches:

    1. A recurrent branch that receives the current state sts_t and the previous action at−1a_{t-1} to infer dynamics properties online. It comprises a fully-connected embedding layer of 128 units followed by a recurrent layer of 128 LSTM units.
    2. A feedforward branch that receives the current state sts_t directly (providing direct access without filtering through memory) and the goal vector gg. The goal vector gg is excluded from the recurrent branch because it contains task objective information rather than system dynamics information.

    The feature outputs from both branches are concatenated and passed through 2 fully connected layers of 128 units each. Hidden layers use ReLU activations. The output layer uses tanh⁡\tanh activation units scaled to the control bounds of each action parameter.

  3. Knowl 3 — Omniscient Critic for Recurrent Deterministic Policy Gradients

    model/method

    In the actor-critic framework for recurrent continuous control policies, an asymmetric information structure is used between the policy (actor) and the value function (critic).

    The universal action-value function is modeled as Q(st,at,yt,g,μ)Q(s_t, a_t, y_t, g, \mu), where sts_t is the current state, ata_t is the candidate action, yt=y(ht)y_t = y(h_t) is the value function's recurrent memory over trajectory history hth_t, gg is the goal, and μ\mu represents the exact dynamics parameters sampled in simulation for that episode.

    Because the critic is utilized only during offline simulation training, providing the privileged dynamics parameters μ\mu as an explicit input to the critic (an omniscient critic) reduces variance in the temporal-difference error and policy gradients. Conversely, the deployable policy π(st,zt,g)\pi(s_t, z_t, g) does not receive μ\mu, forcing the policy to rely solely on its recurrent internal memory ztz_t to infer the environment dynamics.

  4. Knowl 4 — Dynamics Randomization Training with HER and RDPG

    algorithm

    The control policy is trained in simulation by combining Recurrent Deterministic Policy Gradients (RDPG), Hindsight Experience Replay (HER), dynamics parameter randomization, and an omniscient critic.

    Input: Policy network πθ\pi_\theta, omniscient critic network QϕQ_\phi, goal distribution ρg\rho_g, dynamics parameter distribution ρμ\rho_\mu, replay buffer MM, HER replay probability kk, discount factor γ\gamma, learning rate α\alpha
    Initialize network weights θ\theta and ϕ\phi randomly
    while training not finished do
        Sample goal g∼ρgg \sim \rho_g
        Sample dynamics parameters μ∼ρμ\mu \sim \rho_\mu
        Generate rollout trajectory τ=(s0,a0,s1,a1,…,sT)\tau = (s_0, a_0, s_1, a_1, \dots, s_T) in simulator configured with μ\mu
        for t=0t = 0 to T−1T-1 do
            Compute sparse reward rt←r(st,g)r_t \leftarrow r(s_t, g)
        end for
        Store episode (τ,{rt}t=0T−1,g,μ)(\tau, \{r_t\}_{t=0}^{T-1}, g, \mu) in replay buffer MM
        Sample an episode (τ,{rt}t=0T−1,g,μ)(\tau, \{r_t\}_{t=0}^{T-1}, g, \mu) from MM
        with probability kk do
            g←g \leftarrow goal satisfied in final state sTs_T
            for t=0t = 0 to T−1T-1 do
                rt←r(st,g)r_t \leftarrow r(s_t, g)
            end for
        end with
        for t=0t = 0 to T−1T-1 do
            Update recurrent memory states zt,zt+1z_t, z_{t+1} for policy and yt,yt+1y_t, y_{t+1} for critic
            a^t+1←πθ(st+1,zt+1,g)\hat{a}_{t+1} \leftarrow \pi_\theta(s_{t+1}, z_{t+1}, g)
            a^t←πθ(st,zt,g)\hat{a}_t \leftarrow \pi_\theta(s_t, z_t, g)
            qt←rt+γQϕ(st+1,a^t+1,yt+1,g,μ)q_t \leftarrow r_t + \gamma Q_\phi(s_{t+1}, \hat{a}_{t+1}, y_{t+1}, g, \mu)
            Δqt←qt−Qϕ(st,at,yt,g,μ)\Delta q_t \leftarrow q_t - Q_\phi(s_t, a_t, y_t, g, \mu)
        end for
        Compute critic gradient ∇ϕ=1T∑t=0T−1Δqt∂Qϕ(st,at,yt,g,μ)∂ϕ\nabla_\phi = \frac{1}{T} \sum_{t=0}^{T-1} \Delta q_t \frac{\partial Q_\phi(s_t, a_t, y_t, g, \mu)}{\partial \phi}
        Compute policy gradient ∇θ=1T∑t=0T−1∂Qϕ(st,a,yt,g,μ)∂a∣a=a^t∂πθ(st,zt,g)∂θ\nabla_\theta = \frac{1}{T} \sum_{t=0}^{T-1} \left. \frac{\partial Q_\phi(s_t, a, y_t, g, \mu)}{\partial a} \right|_{a=\hat{a}_t} \frac{\partial \pi_\theta(s_t, z_t, g)}{\partial \theta}
        Update parameters ϕ\phi and θ\theta via Adam optimizer using ∇ϕ\nabla_\phi and ∇θ\nabla_\theta
    end while

    Training uses Adam optimizer with learning rate 5×10−45 \times 10^{-4}, mini-batches of 128 episodes (100 steps per episode), and HER goal replay probability k=0.8k = 0.8. Training completes in approximately 8000 updates (~100 million environment steps).

  5. Knowl 5 — Dynamics Parameter Randomization Ranges and Distributions

    data/table

    In the simulated environment, 95 dynamics parameters are randomized across physical robot links, contact surfaces, actuator latencies, and sensor noise. Mass, damping, friction, and controller gains are sampled logarithmically from uniform intervals in log-space. Table height is sampled uniformly. The action delay Δt\Delta t varies per timestep according to Δt∼Δt0+Exp(λ)\Delta t \sim \Delta t_0 + \text{Exp}(\lambda), where Δt0=0.04 s\Delta t_0 = 0.04\text{ s} is the base control timestep, and the exponential rate parameter λ\lambda is sampled uniformly once per episode. Independent zero-mean Gaussian noise is added to state features at every timestep with standard deviation equal to 5%5\% of each feature's running standard deviation.

    Parameter Range
    Link Mass [0.25,4]×[0.25, 4] \times default mass of each link
    Joint Damping [0.2,20]×[0.2, 20] \times default damping of each joint
    Puck Mass [0.1,0.4] kg[0.1, 0.4]\text{ kg}
    Puck Friction [0.1,5][0.1, 5]
    Puck Damping [0.01,0.2] Ns/m[0.01, 0.2]\text{ Ns/m}
    Table Height [0.73,0.77] m[0.73, 0.77]\text{ m}
    Controller Gains [0.5,2]×[0.5, 2] \times default gains
    Action Timestep λ\lambda [125,1000] s−1[125, 1000]\text{ s}^{-1}
  6. Knowl 6 — Fetch Robotic Puck-Pushing Experimental Setup

    experimental setup

    The experimental environment consists of a 7-DOF Fetch Robotics arm tasked with pushing a cylindrical puck (mass ≈0.2 kg\approx 0.2\text{ kg}, radius 0.065 m0.065\text{ m}) across a table to a randomly designated target position within a 0.3 m×0.3 m0.3\text{ m} \times 0.3\text{ m} region.

    • State space (52D continuous): Joint positions and velocities of the 7 robot arm joints, 3D position of the gripper, and 3D position, orientation quaternion/Euler angles, linear velocity, and angular velocity of the puck. Real-world puck pose is tracked via an optical motion capture system.
    • Action space (7D continuous): Commanded target joint angles for the robot's low-level joint position controller, expressed as relative angle offsets from current joint positions. Control frequency is approximately 25 Hz25\text{ Hz} (nominal Δt0=0.04 s\Delta t_0 = 0.04\text{ s}).
    • Reward function: Sparse binary reward r(s,g)=0r(s, g) = 0 if the Euclidean distance between the center of the puck and the target position gg is less than or equal to 0.07 m0.07\text{ m}, and r(s,g)=−1r(s, g) = -1 otherwise.
  7. Knowl 7 — Sim-to-Real Transfer Performance Across Policy Architectures

    data/table

    Control policies trained exclusively in simulation were evaluated in simulation (100 test trials with randomized dynamics) and on a physical Fetch robot without real-world fine-tuning. Performance is measured by task success rate (proportion of episodes where the puck is within 0.07 m0.07\text{ m} of the goal at episode termination). Four architectures are compared: LSTM policy, memoryless feedforward trained without dynamics randomization (FF no Rand), memoryless feedforward trained with dynamics randomization (FF), and feedforward augmented with an 8-step history of past states and actions (FF + Hist).

    Model Success (Sim) Success (Real) Trials (Real)
    LSTM 0.91±0.030.91 \pm 0.03 0.89±0.060.89 \pm 0.06 28
    FF no Rand 0.51±0.050.51 \pm 0.05 0.00±0.000.00 \pm 0.00 10
    FF 0.83±0.040.83 \pm 0.04 0.67±0.140.67 \pm 0.14 12
    FF + Hist 0.87±0.030.87 \pm 0.03 0.70±0.100.70 \pm 0.10 20

    The LSTM architecture achieves near-identical success rates in the real world (89%89\%) and simulation (91%91\%). Feedforward policies trained without randomization fail completely on the physical system (0%0\%). While feeding explicit observation histories (FF + Hist) improves real-world transfer to 70%70\%, the recurrent LSTM policy adapts significantly better to physical dynamics.

  8. Knowl 8 — Ablation Analysis of Randomized Dynamics Parameters on Real Robot Transfer

    data/table

    To determine the contribution of individual randomized physical parameters to sim-to-real transfer success, LSTM policies were trained in simulation with specific parameter randomizations disabled and subsequently evaluated on the physical Fetch robot.

    Randomization Condition Success Rate (Real) Trials (Real)
    All parameters randomized 0.89±0.060.89 \pm 0.06 28
    Fixed action timestep 0.29±0.110.29 \pm 0.11 17
    No observation noise 0.25±0.120.25 \pm 0.12 12
    Fixed link mass 0.64±0.100.64 \pm 0.10 22
    Fixed puck friction 0.48±0.100.48 \pm 0.10 27

    Fixing the action timestep (disabling latency randomization) and disabling sensor observation noise cause the largest drops in transfer performance (down to 29%29\% and 25%25\% success, respectively), demonstrating that controller latency variations and sensor noise are the most critical factors in bridging the sim-to-real gap for continuous manipulation.

  9. Knowl 9 — Robustness of Recurrent Policies to Real-World Contact Dynamics Perturbations

    empirical result

    The robustness of the simulation-trained LSTM policy to unmodeled real-world dynamics was evaluated on the physical Fetch robot by modifying the contact surface of the puck. A packet of chips was attached to the bottom of the puck, altering both the sliding friction against the table surface and the internal mass/force distribution.

    Under this contact perturbation, the LSTM policy achieved a real-world task success rate of 0.91±0.040.91 \pm 0.04 over 28 trials, which is statistically indistinguishable from the unmodified baseline puck success rate of 0.89±0.060.89 \pm 0.06. The policy adapted online by executing emergent corrective strategies, such as pressing down on one edge to partially tilt the puck before pushing, modulating push trajectories from sides or top depending on alignment, and actively recovering when the puck overshot the target location.

Coverage note — None was omitted; all key contributed methods, architectures, algorithms, experimental parameters, and empirical results have been captured.

References

  1. 1.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, ‐Human-level control through deep reinforcement learning,‐ Nature, vol. 518, no. 7540, pp. 529–533, 02 2015. [Online]. Available: http://dx.doi.org/10.1038/nature14236
  2. 2.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, ‐Continuous control with deep reinforcement learning,‐ CoRR, vol. abs/1509.02971, 2015. [Online]. Available: http://arxiv.org/abs/1509.02971
  3. 3.Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, ‐Benchmarking deep reinforcement learning for continuous control,‐ CoRR, vol. abs/1604.06778, 2016. [Online]. Available: http://arxiv. org/abs/1604.06778
  4. 4.X. B. Peng, G. Berseth, and M. van de Panne, ‐Terrain-adaptive locomotion skills using deep reinforcement learning,‐ ACM Transactions on Graphics (Proc. SIGGRAPH 2016), vol. 35, no. 4, 2016.
  5. 5.X. B. Peng, G. Berseth, K. Yin, and M. van de Panne, ‐Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning,‐ ACM Transactions on Graphics (Proc. SIGGRAPH 2017), vol. 36, no. 4, 2017.
  6. 6.L. Liu and J. Hodgins, ‐Learning to schedule control fragments for physics-based characters using deep q-learning,‐ ACM Trans. Graph., vol. 36, no. 3, pp. 29:1–29:14, Jun. 2017. [Online]. Available: http://doi.acm.org/10.1145/3083723
  7. 7.N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. A. Riedmiller, and D. Silver, ‐Emergence of locomotion behaviours in rich environments,‐ CoRR, vol. abs/1707.02286, 2017. [Online]. Available: http://arxiv.org/abs/1707.02286
  8. 8.S. Levine, N. Wagener, and P. Abbeel, ‐Learning contact-rich manipulation skills with guided policy search,‐ CoRR, vol. abs/1501.05611, 2015. [Online]. Available: http://arxiv.org/abs/1501.05611
  9. 9.S. Levine, C. Finn, T. Darrell, and P. Abbeel, ‐End-to-end training of deep visuomotor policies,‐ CoRR, vol. abs/1504.00702, 2015. [Online]. Available: http://arxiv.org/abs/1504.00702
  10. 10.S. Levine, P. Pastor, A. Krizhevsky, and D. Quillen, ‐Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection,‐ CoRR, vol. abs/1603.02199, 2016. [Online]. Available: http://arxiv.org/abs/1603.02199
  11. 11.E. Tzeng, C. Devin, J. Hoffman, C. Finn, X. Peng, S. Levine, K. Saenko, and T. Darrell, ‐Adapting deep visuomotor representations with weak pairwise constraints,‐ CoRR, vol. abs/1511.07111, 2015. [Online]. Available: http://arxiv.org/abs/1511.07111
  12. 12.Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, ‐Domain-adversarial training of neural networks,‐ J. Mach. Learn. Res., vol. 17, no. 1, pp. 2096–2030, Jan. 2016. [Online]. Available: http://dl.acm.org/citation.cfm?id=2946645.2946704
  13. 13.A. Gupta, C. Devin, Y. Liu, P. Abbeel, and S. Levine, ‐Learning invariant feature spaces to transfer skills with reinforcement learning,‐ CoRR, vol. abs/1703.02949, 2017. [Online]. Available: http://arxiv.org/abs/1703.02949
  14. 14.S. Daftry, J. A. Bagnell, and M. Hebert, ‐Learning transferable policies for monocular reactive MAV control,‐ CoRR, vol. abs/1608.00627, 2016. [Online]. Available: http://arxiv.org/abs/1608.00627
  15. 15.M. Wulfmeier, I. Posner, and P. Abbeel, ‐Mutual alignment transfer learning,‐ CoRR, vol. abs/1707.07907, 2017. [Online]. Available: http://arxiv.org/abs/1707.07907
  16. 16.A. A. Rusu, M. Vecerik, T. Rothorl, N. Heess, R. Pascanu, ¨ and R. Hadsell, ‐Sim-to-real robot learning from pixels with progressive nets,‐ CoRR, vol. abs/1610.04286, 2016. [Online]. Available: http://arxiv.org/abs/1610.04286
  17. 17.P. Christiano, Z. Shah, I. Mordatch, J. Schneider, T. Blackwell, J. Tobin, P. Abbeel, and W. Zaremba, ‐Transfer from simulation to real world through learning deep inverse dynamics model,‐ CoRR, vol. abs/1610.03518, 2016. [Online]. Available: http://arxiv.org/abs/ 1610.03518
  18. 18.F. Sadeghi and S. Levine, ‐Cad2rl: Real single-image flight without a single real image,‐ CoRR, vol. abs/1611.04201, 2016. [Online]. Available: http://arxiv.org/abs/1611.04201
  19. 19.J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, ‐Domain randomization for transferring deep neural networks from simulation to the real world,‐ CoRR, vol. abs/1703.06907, 2017. [Online]. Available: http://arxiv.org/abs/1703.06907
  20. 20.S. James and E. Johns, ‐3d simulation for robot arm control with deep q-learning,‐ CoRR, vol. abs/1609.03759, 2016. [Online]. Available: http://arxiv.org/abs/1609.03759
  21. 21.I. Mordatch, K. Lowrey, and E. Todorov, ‐Ensemble-cio: Full-body dynamic motion planning that transfers to physical humanoids,‐ in 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2015, Hamburg, Germany, September 28 - October 2, 2015, 2015, pp. 5307–5314. [Online]. Available: https://doi.org/10.1109/IROS.2015.7354126
  22. 22.A. Rajeswaran, S. Ghotra, S. Levine, and B. Ravindran, ‐Epopt: Learning robust neural network policies using model ensembles,‐ CoRR, vol. abs/1610.01283, 2016. [Online]. Available: http://arxiv. org/abs/1610.01283
  23. 23.L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, ‐Robust adversarial reinforcement learning,‐ CoRR, vol. abs/1703.02702, 2017. [Online]. Available: http://arxiv.org/abs/1703.02702
  24. 24.W. Yu, C. K. Liu, and G. Turk, ‐Preparing for the unknown: Learning a universal policy with online system identification,‐ CoRR, vol. abs/1702.02453, 2017. [Online]. Available: http://arxiv.org/abs/1702. 02453
  25. 25.R. Antonova, S. Cruciani, C. Smith, and D. Kragic, ‐Reinforcement learning for pivoting task,‐ CoRR, vol. abs/1703.00472, 2017. [Online]. Available: http://arxiv.org/abs/1703.00472
  26. 26.K. Yu, M. Bauza, N. Fazeli, and A. Rodriguez, ‐More than a million ´ ways to be pushed: A high-fidelity experimental data set of planar pushing,‐ CoRR, vol. abs/1604.04038, 2016. [Online]. Available: http://arxiv.org/abs/1604.04038
  27. 27.K. M. Lynch and M. T. Mason, ‐Stable pushing: Mechanics, controllability, and planning,‐ The International Journal of Robotics Research, vol. 15, no. 6, pp. 533–556, 1996.
  28. 28.M. Dogar and S. Srinivasa, ‐A framework for push-grasping in clutter,‐ in Robotics: Science and Systems VII. Pittsburgh, PA: MIT Press, July 2011.
  29. 29.N. Fazeli, R. Kolbert, R. Tedrake, and A. Rodriguez, ‐Parameter and contact force estimation of planar rigid-bodies undergoing frictional contact,‐ The International Journal of Robotics Research, vol. 0, no. 0, p. 0278364917698749, 2016.
  30. 30.S. Akella and M. T. Mason, ‐Posing polygonal objects in the plane by pushing,‐ The International Journal of Robotics Research, vol. 17, no. 1, pp. 70–88, 1998.
  31. 31.C. Finn, I. J. Goodfellow, and S. Levine, ‐Unsupervised learning for physical interaction through video prediction,‐ CoRR, vol. abs/1605.07157, 2016. [Online]. Available: http://arxiv.org/abs/1605. 07157
  32. 32.D. H. Ignasi Clavera and P. Abbeel, ‐Policy transfer via modularity,‐ in IROS. IEEE, 2017.
  33. 33.R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour, ‐Policy gradient methods for reinforcement learning with function approximation,‐ in In Advances in Neural Information Processing Systems 12. MIT Press, 2000, pp. 1057–1063.
  34. 34.T. Schaul, D. Horgan, K. Gregor, and D. Silver, ‐Universal value function approximators,‐ in Proceedings of the 32nd International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, F. Bach and D. Blei, Eds., vol. 37. Lille, France: PMLR, 07–09 Jul 2015, pp. 1312–1320. [Online]. Available: http://proceedings.mlr.press/v37/schaul15.html
  35. 35.M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba, ‐Hindsight experience replay,‐ in Advances in Neural Information Processing Systems, 2017.
  36. 36.N. Heess, J. J. Hunt, T. P. Lillicrap, and D. Silver, ‐Memory-based control with recurrent neural networks,‐ CoRR, vol. abs/1512.04455, 2015. [Online]. Available: http://arxiv.org/abs/1512.04455
  37. 37.J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson, ‐Learning to communicate with deep multi-agent reinforcement learning,‐ CoRR, vol. abs/1605.06676, 2016. [Online]. Available: http://arxiv.org/abs/1605.06676
  38. 38.R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, ‐Multi-agent actor-critic for mixed cooperative-competitive environments,‐ CoRR, vol. abs/1706.02275, 2017. [Online]. Available: http://arxiv.org/abs/1706.02275
  39. 39.E. Todorov, T. Erez, and Y. Tassa, ‐Mujoco: A physics engine for model-based control.‐ in IROS. IEEE, 2012, pp. 5026–5033.
  40. 40.D. P. Kingma and J. Ba, ‐Adam: A method for stochastic optimization,‐ CoRR, vol. abs/1412.6980, 2014. [Online]. Available: http://arxiv.org/abs/1412.6980

Citation

MLA
Peng, X. B., et al. “Sim-to-Real Transfer of Robotic Control with Dynamics Randomization”. 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, pp. 3803–10, https://doi.org/10.1109/ICRA.2018.8460528.
APA
Peng, X. B., Andrychowicz, M., Zaremba, W., & Abbeel, P. (2018). Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. 2018 IEEE International Conference on Robotics and Automation (ICRA), 3803–3810. https://doi.org/10.1109/ICRA.2018.8460528
Chicago
Peng, X. B., M. Andrychowicz, W. Zaremba, and P. Abbeel. 2018. “Sim-to-Real Transfer of Robotic Control with Dynamics Randomization”. 2018 IEEE International Conference on Robotics and Automation (ICRA), 3803–10. https://doi.org/10.1109/ICRA.2018.8460528.
Harvard
Peng, X.B. et al. (2018) “Sim-to-Real Transfer of Robotic Control with Dynamics Randomization”, 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp. 3803–3810. Available at: https://doi.org/10.1109/ICRA.2018.8460528.
Vancouver
1. Peng XB, Andrychowicz M, Zaremba W, Abbeel P (2018) Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. In: 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp 3803–3810

BibTeX

@inproceedings{Peng_2018, title={Sim-to-Real Transfer of Robotic Control with Dynamics Randomization}, url={http://dx.doi.org/10.1109/ICRA.2018.8460528}, DOI={10.1109/icra.2018.8460528}, booktitle={2018 IEEE International Conference on Robotics and Automation (ICRA)}, publisher={IEEE}, author={Peng, Xue Bin and Andrychowicz, Marcin and Zaremba, Wojciech and Abbeel, Pieter}, year={2018}, month=May, pages={3803–3810} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF