Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Aravind RajeswaranVikash KumarAbhishek GuptaGiulia VezzaniJohn SchulmanEmanuel TodorovSergey Levine

article2017Robotics: Science and Systems Conference1,427 citationsBest Paper Award

Demonstrates that combining deep reinforcement learning with a small number of human demonstrations drastically cuts sample complexity, enabling a 24-DoF robotic hand to solve complex dexterous manipulation tasks within just a few hours of simulated experience.

Listen

Deploying multi-fingered robotic hands in human environments is essential for complex everyday tasks, yet controlling these high-dimensional systems remains an open challenge. Traditional physics-based models struggle with complex contact dynamics, while model-free deep reinforcement learning (an approach where agents learn optimal actions through trial and error) typically suffers from extreme sample inefficiency and produces unnatural, fragile movements. The article evaluates whether combining deep reinforcement learning with a small number of human demonstrations can scale to complex, high-degree-of-freedom dexterous manipulation and reduce training sample requirements to practical levels.

To evaluate this capability, the authors designed four simulated manipulation tasks—object relocation, in-hand pen repositioning, latch-door opening, and hammering—using a 24-degree-of-freedom anthropomorphic hand simulated in the MuJoCo physics engine. They captured 25 demonstrations per task in virtual reality and evaluated a novel algorithm called Demonstration Augmented Policy Gradient (DAPG). This approach initializes the control policy by mimicking human demonstrations and fine-tunes it using policy gradients combined with an auxiliary loss that decays over time, contrasting against learning from scratch and alternative algorithms.

The findings show that standard deep reinforcement learning trained from scratch fails under simple task-completion rewards and requires laborious reward design, demanding up to 100 simulated robot hours while producing brittle, unnatural motions. In contrast, DAPG solves all four complex tasks using only basic task-completion signals, reducing sample complexity dramatically. Specifically, it enables policy training within 3.3 to 6.1 hours of simulated robot experience—up to a thirtyfold speedup over training from scratch. Furthermore, policies trained with DAPG significantly outperformed alternative methods, maintained high success rates despite environmental variations in object mass and size, and exhibited natural, human-like motion profiles.

These results demonstrate that incorporating human demonstrations effectively overcomes the sample inefficiency and exploration bottlenecks that previously prevented the adoption of deep reinforcement learning on complex robotic manipulators. By lowering required training times to just a few hours and eliminating manual reward engineering, the method substantially reduces computational and operational costs while producing safer, more robust behaviors suitable for unstructured human environments.

Organizations developing dexterous manipulation platforms should adopt hybrid learning approaches that combine demonstration bootstrapping with policy gradient fine-tuning rather than relying purely on trial-and-error learning from scratch. Next development steps should focus on pilot physical hardware deployments, testing policies directly on real robotic hands, and integrating raw visual inputs and tactile feedback. Readers should note that current findings are established exclusively in high-fidelity simulation environments; until hardware validation is conducted, hardware-transfer dynamics and sensing limitations represent the primary remaining uncertainties.

Cover for Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Abstract

Dexterous multi-fingered hands are extremely versatile and provide a generic way to perform a multitude of tasks in human-centric environments. However, effectively controlling them remains challenging due to their high dimensionality and large number of potential contacts. Deep reinforcement learning (DRL) provides a model-agnostic approach to control complex dynamical systems, but has not been shown to scale to high-dimensional dexterous manipulation. Furthermore, deployment of DRL on physical systems remains challenging due to sample inefficiency. Consequently, the success of DRL in robotics has thus far been limited to simpler manipulators and tasks. In this work, we show that model-free DRL can effectively scale up to complex manipulation tasks with a high-dimensional 24-DoF hand, and solve them from scratch in simulated experiments. Furthermore, with the use of a small number of human demonstrations, the sample complexity can be significantly reduced, which enables learning with sample sizes equivalent to a few hours of robot experience. The use of demonstrations result in policies that exhibit very natural movements and, surprisingly, are also substantially more robust.

Table of Contents

  • I Introduction
  • II Related Work
  • III Dexterous Manipulation Tasks
  • III-A Tasks
  • III-A1 Object relocation (Figure )
  • III-A2 In-hand Manipulation – Repositioning a pen (Figure )
  • III-A3 Manipulating Environmental Props (Figure )
  • III-A4 Tool Use – Hammer (Figure )
  • III-B Experimental setup
  • III-B1 ADROIT hand
  • III-B2 Simulator
  • III-C Demonstrations
  • IV Demo Augmented Policy Gradient (DAPG)
  • IV-A Preliminaries
  • IV-B Natural Policy Gradient
  • IV-C Augmenting RL with demonstrations
  • IV-C1 Pretraining with behavior cloning
  • IV-C2 RL fine-tuning with augmented loss
  • V Results and Discussion
  • V-A Reinforcement Learning from Scratch
  • V-B Reinforcement Learning with Demonstrations
  • VI Conclusion
  • References

Knowls

  1. Knowl 1 — Demo Augmented Policy Gradient Algorithm

    algorithm

    The Demo Augmented Policy Gradient (DAPG) algorithm combines imitation learning and on-policy reinforcement learning to solve high-dimensional continuous control tasks. It initializes the policy parameters using Behavior Cloning (BC) on a small set of demonstrations, and subsequently fine-tunes the policy with Natural Policy Gradient (NPG) using an augmented gradient that continuously incorporates the demonstration data with an advantage-scaled, decaying weight.

    Input: Demonstration dataset D={(st(i),at(i),st+1(i),rt(i))}\mathcal{D} = \{(s_t^{(i)}, a_t^{(i)}, s_{t+1}^{(i)}, r_t^{(i)})\}, initial policy parameter θ0\theta_0, step size δ\delta, weighting hyperparameters λ0,λ1\lambda_0, \lambda_1, trajectory horizon TT, batch size NN
    Output: Optimized policy parameter vector θ\theta
    # Stage 1: Pre-training via Behavior Cloning
    θ←arg⁡max⁡θ∑(s,a)∈Dln⁡πθ(a∣s)\theta \leftarrow \arg\max_\theta \sum_{(s, a) \in \mathcal{D}} \ln \pi_\theta(a \mid s)
    # Stage 2: Policy gradient fine-tuning with augmented objective
    for iteration k=0,1,2,…k = 0, 1, 2, \dots do
        Sample NN trajectories ρπ={(st(i),at(i),rt(i))}i=1,t=1N,T\rho_\pi = \{(s_t^{(i)}, a_t^{(i)}, r_t^{(i)})\}_{i=1, t=1}^{N, T} by executing πθ\pi_\theta in the environment
        Estimate advantages A^π(st(i),at(i))\hat{A}^\pi(s_t^{(i)}, a_t^{(i)}) for all transitions in ρπ\rho_\pi
        Compute demonstration weighting scalar:
            wk←λ0λ1kmax⁡(s′,a′)∈ρπA^π(s′,a′)w_k \leftarrow \lambda_0 \lambda_1^k \max_{(s', a') \in \rho_\pi} \hat{A}^\pi(s', a')
        Compute augmented policy gradient:
            gaug←1NT∑(s,a)∈ρπ∇θln⁡πθ(a∣s)A^π(s,a)+1∣D∣∑(s,a)∈D∇θln⁡πθ(a∣s)wkg_{\text{aug}} \leftarrow \frac{1}{NT} \sum_{(s, a) \in \rho_\pi} \nabla_\theta \ln \pi_\theta(a \mid s) \hat{A}^\pi(s, a) + \frac{1}{|\mathcal{D}|} \sum_{(s, a) \in \mathcal{D}} \nabla_\theta \ln \pi_\theta(a \mid s) w_k
        Compute empirical Fisher Information Matrix:
            Fθ←1NT∑i=1N∑t=1T∇θln⁡πθ(at(i)∣st(i))∇θln⁡πθ(at(i)∣st(i))TF_\theta \leftarrow \frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T \nabla_\theta \ln \pi_\theta(a_t^{(i)} \mid s_t^{(i)}) \nabla_\theta \ln \pi_\theta(a_t^{(i)} \mid s_t^{(i)})^T
        Update parameters via normalized natural gradient step:
            θ←θ+δgaugTFθ−1gaugFθ−1gaug\theta \leftarrow \theta + \sqrt{\frac{\delta}{g_{\text{aug}}^T F_\theta^{-1} g_{\text{aug}}}} F_\theta^{-1} g_{\text{aug}}
    end for
    return θ\theta

    In standard implementations, hyperparameters λ0=0.1\lambda_0 = 0.1 and λ1=0.95\lambda_1 = 0.95 are used.

  2. Knowl 2 — Augmented Policy Gradient with Decaying Demonstration Weighting

    equation

    To exploit demonstration data throughout policy gradient fine-tuning rather than solely during initialization, the standard policy gradient is augmented with an auxiliary gradient evaluated over the demonstration transitions:

    gaug=∑(s,a)∈ρπ∇θln⁡πθ(a∣s)Aπ(s,a)+∑(s,a)∈ρD∇θln⁡πθ(a∣s)w(s,a)g_{\text{aug}} = \sum_{(s,a) \in \rho_\pi} \nabla_\theta \ln \pi_\theta(a \mid s) A^\pi(s, a) + \sum_{(s,a) \in \rho_D} \nabla_\theta \ln \pi_\theta(a \mid s) w(s, a)

    where ρπ\rho_\pi is the dataset of state-action pairs collected by executing the current policy πθ\pi_\theta on the Markov Decision Process, ρD\rho_D is the fixed dataset of expert demonstration transitions, Aπ(s,a)=Qπ(s,a)−Vπ(s)A^\pi(s, a) = Q^\pi(s, a) - V^\pi(s) is the advantage function of policy πθ\pi_\theta, and w(s,a)w(s, a) is a state-action weighting function.

    The weighting function is defined using an empirical heuristic:

    w(s,a)=λ0λ1kmax⁡(s′,a′)∈ρπAπ(s′,a′)∀(s,a)∈ρDw(s, a) = \lambda_0 \lambda_1^k \max_{(s', a') \in \rho_\pi} A^\pi(s', a') \quad \forall (s,a) \in \rho_D

    where kk is the policy iteration counter, λ0>0\lambda_0 > 0 sets the baseline scaling factor relative to the maximum advantage found in the current on-policy batch, and λ1∈(0,1)\lambda_1 \in (0, 1) is an exponential decay rate (λ0=0.1,λ1=0.95 \lambda_0 = 0.1, \lambda_1 = 0.95). This decay ensures that demonstration gradients provide directional guidance early in training, while asymptotically decaying to zero as the policy matures and achieves or surpasses expert performance.

  3. Knowl 3 — Sample and Robot-Time Complexity Comparison Across Dexterous Manipulation Tasks

    data/table

    The sample and wall-clock time complexity required to reach a 90% success rate is evaluated across four 24-DoF dexterous manipulation tasks, comparing Demo Augmented Policy Gradient (DAPG) using sparse task completion rewards against Natural Policy Gradient (NPG) trained from scratch with shaped rewards and sparse rewards. Each iteration consists of 200 trajectories of 2 seconds each (400 seconds of simulation experience per iteration).

    Method DAPG (sparse) RL (shaped) RL (sparse)
    Task NN Hours NN Hours NN Hours
    Relocation 52 5.77 880 98 ∞\infty ∞\infty
    Hammer 55 6.10 448 50 ∞\infty ∞\infty
    Door 42 4.67 146 16.2 ∞\infty ∞\infty
    Pen 30 3.33 864 96 2900 322

    Here NN denotes the number of RL iterations needed to reach a 90% success rate, and Hours indicates the total equivalent robot operational time. Standard RL from scratch completely fails to solve tasks with sparse rewards (except on Pen manipulation, which requires 322 hours). With reward shaping, RL achieves 90% success but requires between 16.2 and 98 hours. DAPG achieves 90% success in 3.33 to 6.10 hours using only sparse rewards, achieving up to an approximate 30×30\times reduction in sample complexity.

  4. Knowl 4 — ADROIT 24-DoF Hand Manipulation Benchmark Suite

    experimental setup

    The dexterous manipulation benchmark suite consists of an anthropomorphic ADROIT hand model simulated in MuJoCo, possessing 24 degrees of freedom: the index, middle, and ring fingers each have 4 DoF, the little finger and thumb each have 5 DoF, and the wrist has 2 DoF. Each joint is position-controlled and equipped with a joint angle sensor. Hand-object contacts model planar friction, and fingertip contacts support rolling and torsional friction, alongside joint dry friction.

    The benchmark comprises four representative continuous control tasks:

    1. Object Relocation: A hand picks up a ball and translates it to a target location randomized across the workspace; success is achieved when the ball is within an ϵ\epsilon-ball of the target.
    2. In-Hand Pen Manipulation: The hand repositions and reorients a pen held in the palm to match a randomized 3D target orientation with the hand base fixed.
    3. Door Opening: The hand grasps and unlatches a door latch subjected to dry friction and a closing bias torque, subsequently swinging the door open until it makes contact with a stopper; door location is randomized.
    4. Tool Use (Hammer): The hand picks up a hammer and strikes a nail to drive it into a board against up to 15 N of dry friction; the nail location is randomized, and success requires the entire nail length to enter the board.

    All environments expose observations consisting of joint angles, object pose (position and orientation), and target pose, and accept 24-dimensional target joint angles as control actions.

  5. Knowl 5 — Demonstration Acquisition and Covariate Shift Perturbation Protocol

    model/method

    Demonstrations for high-DoF dexterous manipulation are gathered using an immersive virtual reality setup via the MuJoCo HAPTIX system. A CyberGlove III measures 24-DoF finger joint configurations, an HTC Vive tracker tracks the 6-DoF base pose of the hand, and an HTC Vive head-mounted display provides stereoscopic visualization.

    For each task, 25 successful demonstration trajectories with randomized task instances are recorded directly in simulation. To broaden the state distribution covered by the demonstration dataset and mitigate distribution drift during behavior cloning, uniform random noise in the range [−0.1,0.1][-0.1, 0.1] radians is added to the recorded actuator commands at each time step during data collection.

  6. Knowl 6 — Normalized Natural Policy Gradient Parameter Update

    equation

    In the Natural Policy Gradient (NPG) framework, the policy parameter vector θ\theta is updated along a normalized natural gradient direction defined by the inverse Fisher Information Matrix FθF_\theta:

    θk+1=θk+δgTFθk−1gFθk−1g\theta_{k+1} = \theta_k + \sqrt{\frac{\delta}{g^T F_{\theta_k}^{-1} g}} F_{\theta_k}^{-1} g

    where δ>0\delta > 0 is a step size hyperparameter constraining the update step in probability distribution space, gg is the empirical policy gradient vector:

    g=1NT∑i=1N∑t=1T∇θln⁡πθ(ati∣sti)A^π(sti,ati,t)g = \frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T \nabla_\theta \ln \pi_\theta(a_t^i \mid s_t^i) \hat{A}^\pi(s_t^i, a_t^i, t)

    and FθkF_{\theta_k} is the empirical Fisher Information Matrix estimated from NN trajectories of horizon TT:

    Fθ=1NT∑i=1N∑t=1T∇θln⁡πθ(ati∣sti)∇θln⁡πθ(ati∣sti)TF_\theta = \frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T \nabla_\theta \ln \pi_\theta(a_t^i \mid s_t^i) \nabla_\theta \ln \pi_\theta(a_t^i \mid s_t^i)^T

  7. Knowl 7 — Policy Initialization via Behavior Cloning

    model/method

    To bootstrap policy exploration in reward-sparse or high-dimensional control settings, the parameterized policy πθ\pi_\theta is initialized by maximizing the log-likelihood over the demonstration dataset ρD\rho_D:

    max⁡θ∑(s,a)∈ρDln⁡πθ(a∣s)\max_\theta \sum_{(s,a) \in \rho_D} \ln \pi_\theta(a \mid s)

    While Behavior Cloning (BC) alone provides an informed policy initialization, it suffers from compounding errors due to distributional shift and fails to complete multi-phase tasks (such as reaching, grasping, and hammering) when trained on limited demonstrations. Initializing with BC restricts initial reinforcement learning exploration to task-relevant regions of the state-action space without requiring manual reward shaping.

  8. Knowl 8 — Empirical Superiority of DAPG over Pure RL and Off-Policy Imitation Baselines

    empirical result

    In simulated evaluations across the 24-DoF manipulation suite:

    1. Deep Deterministic Policy Gradient (DDPG) and its demonstration-augmented extension DDPGfD fail to solve the tasks due to extreme sensitivity to hyperparameters and sample instability in contact-rich, high-dimensional spaces.
    2. Pure Natural Policy Gradient (NPG) without demonstrations fails entirely under sparse reward signals on Object Relocation, Door Opening, and Hammer Use.
    3. Pure NPG with manually shaped rewards solves the tasks but requires extensive training time (16 to 98 robot hours) and yields unnatural motion artifacts (e.g., awkward finger placement, unnatural wrist postures, erratic contact forces).
    4. Standalone Behavior Cloning (BC) achieves near 0% task completion rates due to compounding execution errors.
    5. DAPG reliably solves all four tasks with sparse rewards within 3 to 6 robot hours, producing natural, smooth, and human-like grasping and manipulation strategies.
  9. Knowl 9 — Robustness and Generalization Across Physical Property Variations and Model Ensembles

    empirical result

    When evaluated on object relocation across unobserved physical parameters (object mass varied from 0.2 to 1.4, and object size varied from 0.02 to 0.05):

    1. Policies trained with pure RL and shaped rewards from a single environment instance overfit to nominal dynamics and exhibit near-zero success rates under minor variations in mass or size.
    2. Policies trained with DAPG on a single environment instance retain high success rates across wide variations in mass and size, indicating that human demonstrations impart inherently robust manipulation strategies.
    3. When trained on an ensemble of diverse environment variations (varying mass and size distributions simultaneously), pure RL with shaped rewards fails to converge within practical time frames, whereas DAPG successfully converges to policies achieving nearly 100% success across the entire parameter ensemble.
  10. Knowl 10 — Limitations of Demonstration-Augmented Dexterous Manipulation

    limitation

    The methodology exhibits several key limitations:

    1. All evaluations are restricted to rigid-body physics simulation (MuJoCo) and have not been validated on physical hardware.
    2. Demonstrations for in-hand pen manipulation could not be collected from humans in VR due to the absence of haptic/tactile feedback, requiring a computational RL expert trained with dense rewards to provide demonstrations.
    3. The policies operate on low-dimensional state vectors (ground-truth joint positions and object 6D poses) rather than visual observations (e.g., raw pixels) or tactile sensor streams, despite tactile contacts playing a major role in real-world dexterous manipulation.

Coverage note — None was omitted; all contributed algorithms, mathematical formulations, experimental environments, empirical findings, and stated limitations are fully covered.

References

  1. 1.S. Amari. Natural gradient works efficiently in learning. Neural Computation, 10:251–276, 1998.
  2. 2.H. B. Amor, O. Kroemer, U. Hillenbrand, G. Neumann, and J. Peters. Generalization of human grasping for multi-fingered robot hands. In IROS 2012.
  3. 3.M. Andrychowicz, D. Crow, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba. Hindsight experience replay. In NIPS, 2017.
  4. 4.J. A. Bagnell and J. G. Schneider. Covariant policy search. In IJCAI, 2003.
  5. 5.M. Bojarski, D. D. Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, and K. Zieba. End to end learning for self-driving cars. CoRR, abs/1604.07316, 2016.
  6. 6.G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba. OpenAI Gym, 2016.
  7. 7.T. Brys, A. Harutyunyan, H. B. Suay, S. Chernova, M. E. Taylor, and A. Nowe. Reinforcement learning from demonstration through shaping. In IJCAI, 2015.
  8. 8.R. Deimel and O. Brock. A novel type of compliant and underactuated robotic hand for dexterous grasping. I. J. Robotics Res., 35(1-3):161–185, 2016.
  9. 9.Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel. Benchmarking deep reinforcement learning for continuous control. In ICML, 2016.
  10. 10.T. Erez, Y. Tassa, and E. Todorov. Simulation tools for model-based robotics: Comparison of bullet, havok, mujoco, ode and physx. In ICRA, 2015.
  11. 11.Y. Gao, H. Xu, J. Lin, F. Yu, S. Levine, and T. Darrell. Reinforcement learning from imperfect demonstrations. CoRR, abs/1802.05313, 2018.
  12. 12.A. Ghadirzadeh, A. Maki, D. Kragic, and M. Bjorkman. Deep predictive policy training using reinforcement learning. CoRR, abs/1703.00727, 2017.
  13. 13.S. Gu, E. Holly, T. P. Lillicrap, and S. Levine. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In ICRA, 2017.
  14. 14.A. Gupta, C. Eppner, S. Levine, and P. Abbeel. Learning dexterous manipulation for a soft robotic hand from human demonstrations. In IROS, 2016.
  15. 15.N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. A. Riedmiller, and D. Silver. Emergence of locomotion behaviours in rich environments. CoRR, abs/1707.02286, 2017.
  16. 16.P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger. Deep reinforcement learning that matters. CoRR, abs/1709.06560, 2017.
  17. 17.T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, A. Sendonaris, G. Dulac-Arnold, I. Osband, J. Agapiou, J. Z. Leibo, and A. Gruslys. Learning from demonstrations for real world reinforcement learning. CoRR, abs/1704.03732, 2017.
  18. 18.A. J. Ijspeert, J. Nakanishi, and S. Schaal. Movement imitation with nonlinear dynamical systems in humanoid robots. In ICRA, 2002.
  19. 19.S. Kakade. A natural policy gradient. In NIPS, 2001.
  20. 20.S. Kakade and J. Langford. Approximately optimal approximate reinforcement learning. In ICML, 2002.
  21. 21.J. Kober and J. Peters. Learning motor primitives for robotics. In ICRA, 2009.
  22. 22.J. Kober and J. Peters. Policy search for motor primitives in robotics. Machine Learning, 84(1-2):171–203, 2011.
  23. 23.V. Kumar, A. Gupta, E. Todorov, and S. Levine. Learning Dexterous Manipulation Policies from Experience and Imitation. CoRR, abs/1611.05095, 2016.
  24. 24.V. Kumar, Y. Tassa, T. Erez, and E. Todorov. Real-time behaviour synthesis for dynamic hand-manipulation. In ICRA, 2014.
  25. 25.V. Kumar and E. Todorov. Mujoco haptix: A virtual reality system for hand manipulation. In Humanoids, 2015.
  26. 26.V. Kumar, E. Todorov, and S. Levine. Optimal control with learned local models: Application to dexterous manipulation. In ICRA, 2016.
  27. 27.V. Kumar, Z. Xu, and E. Todorov. Fast, strong and compliant pneumatic actuation for dexterous tendon-driven hands. In ICRA, 2013.
  28. 28.S. Levine, C. Finn, T. Darrell, and P. Abbeel. End-to-end training of deep visuomotor policies. JMLR, 17(39):1–40, 2016.
  29. 29.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra. Continuous control with deep reinforcement learning. CoRR, abs/1509.02971, 2015.
  30. 30.I. Mordatch, K. Lowrey, and E.Todorov. Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids. In IROS, 2015.
  31. 31.I. Mordatch, Z. Popovic, and E. Todorov. Contact-invariant optimization for hand manipulation. In ACM SIGGRAPH/Eurographics symposium on computer animation, pages 137–144. Eurographics Association, 2012.
  32. 32.A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel. Overcoming exploration in reinforcement learning with demonstrations. CoRR, abs/1709.10089, 2017.
  33. 33.X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne. Deepmimic: Example-guided deep reinforcement learning of physics-based character skills. In ACM SIGGRAPH, 2018.
  34. 34.J. Peters. Machine learning of motor skills for robotics. PhD Dissertation, University of Southern California, 2007.
  35. 35.J. Peters and S. Schaal. Natural actor-critic. Neurocomputing, 71:1180–1190, 2007.
  36. 36.J. Peters and S. Schaal. Reinforcement learning of motor skills with policy gradients. Neural Networks, 21(4):682–697, 2008.
  37. 37.D. Pomerleau. ALVINN: an autonomous land vehicle in a neural network. In NIPS 1988], pages 305–313, 1988.
  38. 38.M. Posa, C. Cantu, and R. Tedrake. A direct method for trajectory optimization of rigid bodies through contact. I. J. Robotics Res., 33(1):69–81, 2014.
  39. 39.A. Rajeswaran, S. Ghotra, B. Ravindran, and S. Levine. EPOpt: Learning Robust Neural Network Policies Using Model Ensembles. In ICLR, 2017.
  40. 40.A. Rajeswaran, K. Lowrey, E. Todorov, and S. Kakade. Towards Generalization and Simplicity in Continuous Control. In NIPS, 2017.
  41. 41.S. Ross, G. J. Gordon, and D. Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In AISTATS, 2011.
  42. 42.S. Schaal, A. Ijspeert, and A. Billard. Computational approaches to motor learning by imitation. Philosophical Transactions: Biological Sciences, 358(1431):537–547, 2003.
  43. 43.J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel. Trust region policy optimization. In ICML, 2015.
  44. 44.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017.
  45. 45.K. Subramanian, C. L. I. Jr., and A. L. Thomaz. Exploration from demonstration for interactive reinforcement learning. In AAMAS, pages 447–456. ACM, 2016.
  46. 46.W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell. Deeply aggrevated: Differentiable imitation learning for sequential prediction. In ICML, 2017.
  47. 47.R. S. Sutton and A. G. Barto. Reinforcement learning: An introduction, volume 1.
  48. 48.M. E. Taylor, H. B. Suay, and S. Chernova. Integrating reinforcement learning with human demonstrations of varying ability. In AAMAS, 2011.
  49. 49.E. Theodorou, J. Buchli, and S. Schaal. A generalized path integral control approach to reinforcement learning. Journal of Machine Learning Research, 11:3137–3181, 2010.
  50. 50.E. Theodorou, J. Buchli, and S. Schaal. Reinforcement learning of motor skills in high dimensions: A path integral approach. In ICRA, 2010.
  51. 51.E. Todorov, T. Erez, and Y. Tassa. MuJoCo: A physics engine for model-based control. In ICRA, 2012.
  52. 52.H. van Hoof, T. Hermans, G. Neumann, and J. Peters. Learning robot in-hand manipulation with tactile features. In Humanoids, 2015.
  53. 53.M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothorl, T. Lampe, and M. A. Riedmiller. Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards. CoRR, abs/1707.08817, 2017.
  54. 54.R. J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3):229–256, 1992.
  55. 55.Z. Xu, V. Kumar, and E. Todorov. A low-cost and modular, 20-dof anthropomorphic robotic hand: design, actuation and modeling. In Humanoids, 2013.

Citation

MLA
Rajeswaran, A., et al. “Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations”. arXiv, 2017, http://arxiv.org/abs/1709.10087v2.
APA
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., & Levine, S. (2017). Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. arXiv. http://arxiv.org/abs/1709.10087v2
Chicago
Rajeswaran, A., V. Kumar, A. Gupta, et al. 2017. “Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations”. arXiv. http://arxiv.org/abs/1709.10087v2.
Harvard
Rajeswaran, A. et al. (2017) “Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1709.10087v2.
Vancouver
1. Rajeswaran A, Kumar V, Gupta A, Vezzani G, Schulman J, Todorov E, Levine S (2017) Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. arXiv

BibTeX

@article{rajeswaran2017learning,
  title = {Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations},
  author = {Rajeswaran, Aravind and Kumar, Vikash and Gupta, Abhishek and Vezzani, Giulia and Schulman, John and Todorov, Emanuel and Levine, Sergey},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1709.10087v2},
  eprint = {1709.10087}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF