Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates

Shixiang GuEthan HollyTimothy LillicrapSergey Levine

article2016IEEE International Conference on Robotics and Automation1,613 citations

Demonstrates that parallelizing asynchronous off-policy deep Q-learning across multiple physical robots dramatically improves sample efficiency, enabling the successful training of complex 3D manipulation tasks directly on real hardware without human demonstrations.

Listen

Deploying autonomous robots to perform complex manipulation tasks typically requires extensive human engineering, such as manual programming, task-specific representations, or pre-recorded demonstrations. While machine learning algorithms can theoretically allow robots to learn directly through trial and error, applying them to physical systems has historically suffered from prohibitive training times and high sample requirements. The article evaluates whether an off-policy continuous learning method called Normalized Advantage Functions can be parallelized across multiple robots to learn complex three-dimensional manipulation skills from scratch without demonstrations or manual representations.

The authors decoupled the learning process into an asynchronous architecture. A central server thread trains a deep neural network model using data stored in a shared experience buffer, while worker threads running on individual robotic arms independently collect physical interaction data and synchronize their control policies at the start of each trial. The method incorporates practical safety boundary constraints, including velocity limits and spherical workspace projections. The researchers evaluated the framework on simulated tasks (random-target reaching, door pushing and pulling, and pick-and-place) and physical 7-degree-of-freedom robotic arms tasked with reaching and door opening.

The findings demonstrate substantial improvements in learning feasibility and efficiency. Deep neural network representations successfully mastered complex multi-stage manipulation tasks, whereas traditional linear representations completely failed on door tasks. Parallelizing experience collection across multiple robots significantly reduced wall-clock training time; on physical hardware, two robots cooperatively learned a complex door-opening skill from scratch to a 100% success rate in approximately 2.5 hours, whereas a single robot required significantly more than 4 hours. Furthermore, data collection speed relative to training speed proved critical: running only one worker thread caused data starvation that degraded both learning speed and final policy quality.

These results show that collective robotic learning removes the need for manual demonstration data and specialized control representations, directly reducing deployment costs, engineering timelines, and hardware wear. However, the evaluation relies on shaped reward functions that provide continuous distance feedback rather than simple binary success indicators, and the physical setup used a fixed door pose. Decision-makers should prioritize multi-robot data collection architectures when deploying autonomous manipulation systems, while allocating future research to sparse-reward exploration and policy distillation across diverse physical environments.

  • Paper: Asynchronous Methods for Deep Reinforcement Learning, Volodymyr Mnih et al. (2016). Introduces the asynchronous parallel training framework for deep reinforcement learning that the source adapts to scale off-policy manipulation updates across multiple physical robots.
  • Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). Establishes deep deterministic policy gradients for continuous control, providing the foundational continuous off-policy actor-critic formulation built upon by the source.
  • Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). Derives the deterministic policy gradient theorem, laying the theoretical groundwork for off-policy continuous action learning utilized in robotic control.
  • Paper: End-to-End Training of Deep Visuomotor Policies, Sergey Levine et al. (2015). Demonstrates guided policy search for end-to-end visuomotor policies on physical robots, establishing the prior manipulation baselines that the source aims to surpass using pure off-policy RL.
  • Paper: Human-level control through deep reinforcement learning, Volodymyr Mnih et al. (2015). Pioneers deep Q-networks with experience replay and target networks, establishing key value-function stabilization mechanisms adapted in continuous off-policy architectures.
  • Paper: Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection, Sergey Levine et al. (2016). Presents large-scale parallel multi-robot data collection for learning grasping policies, introducing the distributed physical hardware paradigm that informs the source's asynchronous setup.
  • Paper: Trust Region Policy Optimization, John Schulman et al. (2015). Introduces Trust Region Policy Optimization for continuous control benchmarks, serving as a primary benchmark and contrasting on-policy paradigm for the source's sample-efficient off-policy method.
Cover for Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates

Abstract

Reinforcement learning holds the promise of enabling autonomous robots to learn large repertoires of behavioral skills with minimal human intervention. However, robotic applications of reinforcement learning often compromise the autonomy of the learning process in favor of achieving training times that are practical for real physical systems. This typically involves introducing hand-engineered policy representations and human-supplied demonstrations. Deep reinforcement learning alleviates this limitation by training general-purpose neural network policies, but applications of direct deep reinforcement learning algorithms have so far been restricted to simulated settings and relatively simple tasks, due to their apparent high sample complexity. In this paper, we demonstrate that a recent deep reinforcement learning algorithm based on off-policy training of deep Q-functions can scale to complex 3D manipulation tasks and can learn deep neural network policies efficiently enough to train on real physical robots. We demonstrate that the training times can be further reduced by parallelizing the algorithm across multiple robots which pool their policy updates asynchronously. Our experimental evaluation shows that our method can learn a variety of 3D manipulation skills in simulation and a complex door opening skill on real robots without any prior demonstrations or manually designed representations.

Table of Contents

  • I. INTRODUCTION
  • II. RELATED WORK
  • III. BACKGROUND
  • IV. ASYNCHRONOUS TRAINING OF NORMALIZED ADVANTAGE FUNCTIONS
  • A. Asynchronous Learning
  • B. Safety Constraints
  • C. Network Architectures
  • V. SIMULATED EXPERIMENTS
  • A. Simulation Tasks
  • B. Neural Network Policy Representations
  • C. Asynchronous Training
  • VI. REAL-WORLD EXPERIMENTS
  • A. Random Target Reaching
  • B. Door Opening
  • VII. DISCUSSION AND FUTURE WORK
  • ACKNOWLEDGEMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Asynchronous Normalized Advantage Functions Algorithm

    algorithm

    The Asynchronous Normalized Advantage Functions (NAF) algorithm parallelizes experience collection across NN worker threads running on individual physical or simulated robots and off-policy Q-function updates on a single centralized trainer thread. The decoupling allows robots to execute continuous control loops in real time without waiting for neural network backpropagation steps.

    Input: Number of collector threads NN, target network update rate τ∈(0,1]\tau \in (0, 1], discount factor γ∈[0,1)\gamma \in [0, 1), minibatch size mm, episode length TT, total iterations II, episodes per worker MM
    Output: Trained continuous policy μ(x∣θμ)\mu(x|\theta^\mu) and value network parameters θQ={θμ,θP,θV}\theta^Q = \{\theta^\mu, \theta^P, \theta^V\}
    // Trainer Thread (Central Server)
    Initialize normalized Q-network Q(x,u∣θQ)=A(x,u∣θA)+V(x∣θV)Q(x, u|\theta^Q) = A(x, u|\theta^A) + V(x|\theta^V) with parameters θQ={θμ,θP,θV}\theta^Q = \{\theta^\mu, \theta^P, \theta^V\}
    Initialize target network Q′Q' with parameters θQ′←θQ\theta^{Q'} \leftarrow \theta^Q
    Initialize shared replay buffer R←∅R \leftarrow \emptyset
    for iteration = 1 to II do
        Sample a random minibatch of mm transitions (xi,ui,ri,xi+1,ti)(x_i, u_i, r_i, x_{i+1}, t_i) from RR
        for each transition ii do
            if ti<Tt_i < T then
                yi←ri+γV′(xi+1∣θQ′)y_i \leftarrow r_i + \gamma V'(x_{i+1}|\theta^{Q'})
            else
                yi←riy_i \leftarrow r_i
            end if
        end for
        Update θQ\theta^Q by minimizing loss L=1m∑i=1m(yi−Q(xi,ui∣θQ))2L = \frac{1}{m} \sum_{i=1}^m (y_i - Q(x_i, u_i|\theta^Q))^2 using a gradient step
        Update target network parameters: θQ′←τθQ+(1−τ)θQ′\theta^{Q'} \leftarrow \tau \theta^Q + (1 - \tau) \theta^{Q'}
    end for
    // Collector Thread nn (n=1,…,Nn = 1, \dots, N on individual robots)
    Initialize policy network parameters θnμ\theta^\mu_n
    for episode = 1 to MM do
        Synchronize local policy parameters: θnμ←θμ\theta^\mu_n \leftarrow \theta^\mu
        Initialize exploration noise process N\mathcal{N}
        Observe initial state x1∼p(x1)x_1 \sim p(x_1)
        for t=1t = 1 to TT do
            Select action ut=μ(xt∣θnμ)+Ntu_t = \mu(x_t|\theta^\mu_n) + \mathcal{N}_t
            Execute action utu_t on the robot and observe reward rtr_t and next state xt+1x_{t+1}
            Append transition tuple (xt,ut,rt,xt+1,t)(x_t, u_t, r_t, x_{t+1}, t) to shared replay buffer RR
        end for
    end for
  2. Knowl 2 — Normalized Advantage Function Parameterization for Continuous Deep Q-Learning

    model/method

    Normalized Advantage Functions (NAF) represent the action-value function Q(x,u∣θQ)Q(x, u|\theta^Q) for continuous state x∈Rdx \in \mathbb{R}^d and continuous action u∈Rpu \in \mathbb{R}^p such that the greedy action arg⁡max⁡uQ(x,u)\arg\max_u Q(x, u) can be evaluated analytically in closed form without requiring an iterative optimization or a separate actor network.

    The action-value function decomposes into a state-value term V(x∣θV)V(x|\theta^V) and an advantage term A(x,u∣θA)A(x, u|\theta^A):

    Q(x,u∣θQ)=A(x,u∣θA)+V(x∣θV)Q(x, u|\theta^Q) = A(x, u|\theta^A) + V(x|\theta^V)

    A(x,u∣θA)=−12(u−μ(x∣θμ))TP(x∣θP)(u−μ(x∣θμ))A(x, u|\theta^A) = -\frac{1}{2} \left(u - \mu(x|\theta^\mu)\right)^T P(x|\theta^P) \left(u - \mu(x|\theta^\mu)\right)

    where θQ={θμ,θP,θV}\theta^Q = \{\theta^\mu, \theta^P, \theta^V\} represents the combined network weights. The matrix P(x∣θP)P(x|\theta^P) is constrained to be state-dependent, symmetric, and positive-definite by parameterizing it via a Cholesky decomposition:

    P(x∣θP)=L(x∣θP)L(x∣θP)TP(x|\theta^P) = L(x|\theta^P) L(x|\theta^P)^T

    where L(x∣θP)L(x|\theta^P) is a lower triangular matrix whose diagonal entries are strictly positive (e.g., via exponentiation of network outputs). The term μ(x∣θμ)\mu(x|\theta^\mu) is the policy network predicting the optimal action, typically capped with a hyperbolic tangent (tanh⁡\tanh) activation to bound action magnitudes. Because P(x∣θP)P(x|\theta^P) is positive-definite, A(x,u∣θA)≤0A(x, u|\theta^A) \le 0 with A(x,μ(x∣θμ)∣θA)=0A(x, \mu(x|\theta^\mu)|\theta^A) = 0, ensuring max⁡uQ(x,u∣θQ)=V(x∣θV)\max_u Q(x, u|\theta^Q) = V(x|\theta^V) and arg⁡max⁡uQ(x,u∣θQ)=μ(x∣θμ)\arg\max_u Q(x, u|\theta^Q) = \mu(x|\theta^\mu).

  3. Knowl 3 — Kinematic Bounding Sphere and Joint Safety Constraints for Real-World Exploration

    model/method

    To prevent physical damage during high-variance exploration in deep Q-learning, action execution on physical robotic manipulators is filtered using three layers of safety constraints:

    1. Joint Velocity Bounds: A hard maximum commanded velocity threshold is enforced for every individual joint.
    2. Joint Position Limits: Strict angular bounds are enforced on each joint angle.
    3. End-Effector Bounding Sphere: A spherical workspace boundary of radius RsafeR_{safe} centered at a designated base point is defined. If a commanded joint velocity vector q˙\dot{q} produces an end-effector velocity e˙=J(q)q˙\dot{e} = J(q)\dot{q} that would move the end-effector outside the sphere (where J(q)J(q) is the manipulator Jacobian at configuration qq), forward kinematics is used to project the velocity vector onto the tangent plane of the sphere's surface, and an inward correction velocity component directed toward the center of the sphere is added before converting back to joint space control.
  4. Knowl 4 — Linear-NAF Quadratic Action-Value Baseline

    model/method

    Linear-NAF is a simplified baseline representing a globally quadratic action-value function with a linear feedback policy bounded by a non-linearity. The action-value function is defined as:

    Q(x,u)=−12(u−μ(x))TP(u−μ(x))+xTBx+xTb+cQ(x, u) = -\frac{1}{2} \left(u - \mu(x)\right)^T P \left(u - \mu(x)\right) + x^T B x + x^T b + c

    μ(x)=tanh⁡(k+Kx)\mu(x) = \tanh(k + Kx)

    where x∈Rdx \in \mathbb{R}^d is the state vector, u∈Rpu \in \mathbb{R}^p is the control action, P∈Rp×pP \in \mathbb{R}^{p \times p} is a positive-definite learnable matrix, B∈Rd×dB \in \mathbb{R}^{d \times d} is a learnable symmetric matrix, K∈Rp×dK \in \mathbb{R}^{p \times d} is a learnable feedback gain matrix, k∈Rpk \in \mathbb{R}^p and b∈Rdb \in \mathbb{R}^d are learnable bias vectors, and c∈Rc \in \mathbb{R} is a learnable scalar. The hyperbolic tangent function tanh⁡\tanh enforces bounded actions. Unlike standard deep NAF where P(x)P(x) and V(x)V(x) are parameterized by deep neural networks, Linear-NAF uses constant or quadratic state-dependent matrices.

  5. Knowl 5 — Comparative Performance of Deep vs. Linear Policy Representations in Robotic Manipulation

    data/table

    The comparison across simulated 3D robotic manipulation tasks demonstrates that expressive multi-layer neural network policies (NAF and DDPG) succeed on multi-stage discontinuous contact tasks where linear representations fail completely, and converge faster even on smooth tasks.

    Max. Success Rate (%) Episodes to 100% Success (1000s)
    Task DDPG Lin-NAF NAF DDPG Lin-NAF NAF
    Reach 100±0100 \pm 0 100±0100 \pm 0 100±0100 \pm 0 3.2±0.73.2 \pm 0.7 8±38 \pm 3 3.6±1.03.6 \pm 1.0
    Door Pull 100±0100 \pm 0 5±65 \pm 6 100±0100 \pm 0 10±810 \pm 8 N/A 6±36 \pm 3
    Door Push 100±0100 \pm 0 40±1040 \pm 10 100±0100 \pm 0 3.1±1.03.1 \pm 1.0 N/A 4.2±1.04.2 \pm 1.0
    Pick Place 100±0100 \pm 0 100±0100 \pm 0 100±0100 \pm 0 4.4±0.64.4 \pm 0.6 12±312 \pm 3 2.9±0.92.9 \pm 0.9

    For single-motion tasks (reaching and pick-and-place), Linear-NAF eventually achieves a 100% success rate but requires 2 to 4 times more training episodes to converge than deep NAF or DDPG. For complex contact-rich tasks requiring composite motions (hooking, turning a latch past 60 degrees, and pulling/pushing), Linear-NAF fails (achieving 5±6%5 \pm 6\% on Door Pull and 40±10%40 \pm 10\% on Door Push), while deep NAF and DDPG achieve 100%100\% success.

  6. Knowl 6 — Autonomous Real-World Multi-Robot Door Opening from Scratch

    empirical result

    Using asynchronous NAF distributed across two physical 7-DoF robotic arms pooling updates into a central replay buffer, the robots successfully learn to pull open a latched door from scratch without human demonstrations, pre-training, or task-specific representations.

    • Learning Time and Convergence: Two physical robots learning in parallel achieved a 100%100\% success rate (evaluated over 20 consecutive trials) in approximately 2.52.5 hours of wall-clock time (under 500,000 parameter updates). In contrast, a single robot required more than 44 hours to reach 100%100\% success.
    • Learning Stages: The empirical learning curve progresses through three distinct qualitative phases:
      1. Free-space exploration: The arms learn to reach toward the vicinity of the door handle.
      2. Handle contact plateau: The end-effector contacts the handle sporadically and accidentally turns it, but lacks consistency (visible as a plateau near zero reward).
      3. Consistent latch unlatching and opening: The policy reliably coordinates hooking the handle, rotating it past the 60∘60^\circ unlatching threshold, and pulling the door outward.
  7. Knowl 7 — Scaling Dynamics and Training-to-Collection Speed Ratio in Asynchronous Off-Policy RL

    empirical result

    In asynchronous NAF with a fixed trainer update rate (e.g., ≈100 Hz\approx 100\text{ Hz}) and fixed worker control rates (e.g., 20 Hz20\text{ Hz}, setting a speed ratio S=1/5S = 1/5 per worker), varying the number of workers NN reveals three properties of parallel off-policy learning:

    1. Gradient Step Acceleration: Increasing the number of collection workers significantly accelerates learning convergence with respect to the total number of gradient update steps.
    2. Critical Collection-to-Training Ratio: If data collection is too slow relative to the trainer thread (e.g., using only 1 worker with an unthrottled trainer), the policy suffers from severe sample staleness and overfitting on small buffers, impairing both learning speed and asymptotic policy success.
    3. Diminishing Returns Saturation: Increasing the number of collector workers beyond a task-dependent saturation threshold yields diminishing speedup if the neural network training thread cannot ingest transitions faster (bounded at one gradient step per time unit).
  8. Knowl 8 — Manipulation Task State Representations and Reward Formulations

    experimental setup

    The manipulation tasks are formulated in continuous action spaces with the following state vectors xx and reward functions r(x,u)r(x, u), where d(⋅,⋅)d(\cdot, \cdot) denotes the Huber loss function, uu is the joint velocity control action, and ci>0c_i > 0 are weighting constants:

    1. Random-Target Reaching (7-DoF Arm):

      • State (2020 dimensions): 7 joint angles, 7 joint velocities, 3D end-effector Cartesian position e(x)e(x), 3D target position yy.
      • Reward: r(x,u)=−c1d(y,e(x))−c2uTur(x, u) = -c_1 d(y, e(x)) - c_2 u^T u
      • Episode: 150 steps (7.57.5 s at 20 Hz20\text{ Hz}). Success criterion: ∥e(x)−y∥≤5 cm\|e(x) - y\| \le 5\text{ cm}.
    2. Door Opening / Pulling / Pushing (7-DoF Arm):

      • State (2525 dimensions in simulation, 2121 dimensions in real robot): 7 joint angles, 7 joint velocities, 3D end-effector position e(x)e(x), resting handle position hh, door frame position, door angle, handle angle (or VectorNav IMU quaternion q(x)q(x)).
      • Reward: r(x,u)=−c1d(h,e(x))+c2(−d(qo,q(x))+di)−c3uTur(x, u) = -c_1 d(h, e(x)) + c_2 (-d(q_o, q(x)) + d_i) - c_3 u^T u, where qoq_o is the handle quaternion when turned and door opened, and di=d(qo,qneutral)d_i = d(q_o, q_{neutral}) is the offset at resting position.
      • Episode: 300 steps (1515 s at 20 Hz20\text{ Hz}). Success criterion: door opened in correct direction by ≥10∘\ge 10^\circ.
    3. Pick & Place (9-DoF Kinova JACO Arm):

      • State (180180 dimensions): Position and rotation matrices of all environment geometries, target position yy, vector from suspended stick to target.
      • Reward: r(x,u)=−c1d(s(x),g(x))−c2∑i=13d(s(x),fi(x))−c3d(y,s(x))−c4uTur(x, u) = -c_1 d(s(x), g(x)) - c_2 \sum_{i=1}^3 d(s(x), f_i(x)) - c_3 d(y, s(x)) - c_4 u^T u, where s(x)s(x) is stick position, g(x)g(x) is the palm grip center, and fi(x)f_i(x) are the 3 fingertip positions.
      • Episode: 300 steps (33 s at 100 Hz100\text{ Hz}). Success criterion: stick brought within 5 cm5\text{ cm} of target yy.
  9. Knowl 9 — Limitations of Asynchronous Deep Q-Learning for Robotics

    limitation

    The asynchronous deep reinforcement learning framework has two primary stated limitations:

    1. Reliance on Shaped Continuous Rewards: Learning complex skills from scratch relies on shaped distance-based reward functions (e.g., distance from gripper to handle, plus door opening angles). If replaced by a sparse binary success reward signal, exploration complexity increases dramatically and the current algorithm fails to find successful trajectories within practical training durations.
    2. Homogeneous Replay Pooling: The method directly aggregates transitions from all worker robots into a single shared replay buffer, assuming identical kinematics, dynamics, and environmental setups. It does not incorporate explicit multi-task or domain-adaptation mechanisms (such as policy distillation) required when robots operate on varying physical hardware or interact with distinct door types and configurations.

Coverage note — None was omitted; all key algorithmic contributions, mathematical models, experimental comparisons (simulated and real-world), and stated limitations are fully covered.

References

  1. 1.N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in International Conference on Robotics and Automation (IROS), 2004.
  2. 2.G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng, “Learning CPG-based biped locomotion with a policy gradient method: Application to a humanoid robot,” International Journal of Robotic Research, vol. 27, no. 2, pp. 213–228, 2008.
  3. 3.J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural Networks, vol. 21, no. 4, pp. 682–697, 2008.
  4. 4.E. Theodorou, J. Buchli, and S. Schaal, “Reinforcement learning of motor skills in high dimensions,” in International Conference on Robotics and Automation (ICRA), 2010.
  5. 5.J. Peters, K. Mulling, and Y. Altun, “Relative entropy policy search,” in AAAI Conference on Artificial Intelligence, 2010.
  6. 6.M. Kalakrishnan, L. Righetti, P. Pastor, and S. Schaal, “Learning force control policies for compliant manipulation,” in International Conference on Intelligent Robots and Systems (IROS), 2011.
  7. 7.P. Abbeel, A. Coates, M. Quigley, and A. Ng, “An application of reinforcement learning to aerobatic helicopter flight,” in Advances in Neural Information Processing Systems (NIPS), 2006.
  8. 8.J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” International Journal of Robotic Research, vol. 32, no. 11, pp. 1238–1274, 2013.
  9. 9.P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal, “Learning and generalization of motor skills by learning from demonstration,” in International Conference on Robotics and Automation (ICRA), 2009.
  10. 10.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” International Conference on Learning Representations (ICLR), 2016.
  11. 11.S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q-learning with model-based acceleration,” in International Conference on Machine Learning (ICML), 2016.
  12. 12.M. Deisenroth, G. Neumann, and J. Peters, “A survey on policy search for robotics,” Foundations and Trends in Robotics, vol. 2, no. 1-2, pp. 1–142, 2013.
  13. 13.K. J. Hunt, D. Sbarbaro, R. Zbikowski, and P. J. Gawthrop, “Neural networks for control systems: A survey,” Automatica, vol. 28, no. 6, pp. 1083–1112, Nov. 1992.
  14. 14.M. Riedmiller, “Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method,” in European Conference on Machine Learning. Springer, 2005, pp. 317–328.
  15. 15.R. Hafner and M. Riedmiller, “Neural reinforcement learning controllers for a real robot application,” in International Conference on Robotics and Automation (ICRA), 2007.
  16. 16.M. Riedmiller, S. Lange, and A. Voigtlaender, “Autonomous reinforcement learning on raw visual input data in a real world application,” in International Joint Conference on Neural Networks, 2012.
  17. 17.J. Koutn<acute></acute>ık, G. Cuccu, J. Schmidhuber, and F. Gomez, “Evolving large-scale neural networks for vision-based reinforcement learning,” in Conference on Genetic and Evolutionary Computation, ser. GECCO ’13, 2013.
  18. 18.J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel, “Trust region policy optimization,” in International Conference on Machine Learning (ICML), 2015.
  19. 19.S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research (JMLR), vol. 17, 2016.
  20. 20.M. Deisenroth and C. Rasmussen, “PILCO: a model-based and data-efficient approach to policy search,” in International Conference on Machine Learning (ICML), 2011.
  21. 21.T. Moldovan, S. Levine, M. Jordan, and S. Abbeel, “Optimism-driven exploration for nonlinear systems,” in International Conference on Robotics and Automation (ICRA), 2015.
  22. 22.R. Lioutikov, A. Paraschos, G. Neumann, and J. Peters, “Sample-based information-theoretic stochastic optimal control,” in International Conference on Robotics and Automation, 2014.
  23. 23.M. P. Deisenroth, G. Neumann, J. Peters et al., “A survey on policy search for robotics.” Foundations and Trends in Robotics, vol. 2, no. 1-2, pp. 1–142, 2013.
  24. 24.Y. Chebotar, M. Kalakrishnan, A. Yahya, A. Li, S. Schaal, and S. Levine, “Path integral guided policy search,” arXiv preprint arXiv:1610.00529, 2016.
  25. 25.R. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, vol. 8, no. 3-4, pp. 229–256, May 1992.
  26. 26.C. J. Watkins and P. Dayan, “Q-learning,” Machine learning, vol. 8, no. 3-4, pp. 279–292, 1992.
  27. 27.R. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems (NIPS), 1999.
  28. 28.J. Koutn<acute></acuteık, G. Cuccu, J. Schmidhuber, and F. Gomez, “Evolving large-scale neural networks for vision-based reinforcement learning,” in Proceedings of the 15th annual conference on Genetic and evolutionary computation. ACM, 2013, pp. 1061–1068.
  29. 29.V. Mnih et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015.
  30. 30.V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning (ICML), 2016, pp. 1928–1937.
  31. 31.R. Hafner and M. Riedmiller, “Reinforcement learning in feedback control,” Machine learning, vol. 84, no. 1-2, pp. 137–169, 2011.
  32. 32.M. Inaba, S. Kagami, F. Kanehiro, and Y. Hoshino, “A platform for robotics research based on the remote-brained robot approach,” International Journal of Robotics Research, vol. 19, no. 10, 2000.
  33. 33.J. Kuffner, “Cloud-enabled humanoid robots,” in IEEE-RAS International Conference on Humanoid Robotics, 2010.
  34. 34.B. Kehoe, A. Matsukawa, S. Candido, J. Kuffner, and K. Goldberg, “Cloud-based robot grasping with the google object recognition engine,” in IEEE International Conference on Robotics and Automation, 2013.
  35. 35.B. Kehoe, S. Patil, P. Abbeel, and K. Goldberg, “A survey of research on cloud robotics and automation,” IEEE Transactions on Automation Science and Engineering, vol. 12, no. 2, April 2015.
  36. 36.A. Yahya, A. Li, M. Kalakrishnan, Y. Chebotar, and S. Levine, “Collective robot reinforcement learning with distributed asynchronous guided policy search,” arXiv preprint arXiv:1610.00673, 2016.
  37. 37.E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2012, pp. 5026–5033.
  38. 38.J. Ba and D. Kingma, “Adam: A method for stochastic optimization,” 2015.
  39. 39.J. Kober and J. Peters, “Learning motor primitives for robotics,” in International Conference on Robotics and Automation (ICRA), 2009.
  40. 40.R. Tedrake, T. W. Zhang, and H. S. Seung, “Learning to walk in 20 minutes.”
  41. 41.S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” arXiv preprint arXiv:1502.03167, 2015.
  42. 42.J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv preprint arXiv:1607.06450, 2016.
  43. 43.A. Rusu, S. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell, “Policy distillation,” in International Conference on Learning Representations (ICLR), 2016.

Citation

MLA
Gu, S., et al. “Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates”. arXiv, 2016, http://arxiv.org/abs/1610.00633v2.
APA
Gu, S., Holly, E., Lillicrap, T., & Levine, S. (2016). Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates. arXiv. http://arxiv.org/abs/1610.00633v2
Chicago
Gu, S., E. Holly, T. Lillicrap, and S. Levine. 2016. “Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates”. arXiv. http://arxiv.org/abs/1610.00633v2.
Harvard
Gu, S. et al. (2016) “Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1610.00633v2.
Vancouver
1. Gu S, Holly E, Lillicrap T, Levine S (2016) Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates. arXiv

BibTeX

@article{gu2016deep,
  title = {Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates},
  author = {Gu, Shixiang and Holly, Ethan and Lillicrap, Timothy and Levine, Sergey},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1610.00633v2},
  eprint = {1610.00633}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF