Benchmarking Deep Reinforcement Learning for Continuous Control

Yan DuanXi ChenRein HouthooftJohn SchulmanPieter Abbeel

article2016ICML1,838 citations

Presents a standardized continuous control benchmark suite and systematically evaluates prominent deep reinforcement learning algorithms across diverse tasks, establishing reproducible performance baselines for high-dimensional control problems.

Listen

Recent advances in deep learning have enabled artificial intelligence to excel at complex tasks with discrete choices, such as playing classic video games from visual inputs. However, evaluating progress in physical continuous control problems—such as robotics, locomotion, and industrial manipulation—has remained difficult due to the lack of a comprehensive, standardized testing suite. Existing benchmarks have largely focused on lower-dimensional problems or discrete actions, leaving a significant gap in understanding how modern algorithms handle complex, high-dimensional physical systems.

The article addresses this gap by establishing an open-source suite of 31 continuous control simulation tasks and systematically evaluating leading reinforcement learning algorithms on deep neural network controllers.

To conduct this evaluation, the researchers implemented 31 simulation environments spanning four main categories: basic control problems, high-dimensional locomotion tasks (such as robotic ants and 3D humanoids), partially observable environments with sensory noise or physical delays, and complex hierarchical navigation tasks. They then implemented and rigorously tested a wide spectrum of algorithms across these tasks using standardized performance metrics and controlled hyperparameter tuning evaluated over multiple random trials.

The benchmark revealed several critical findings regarding algorithm performance. First, batch gradient-based methods with policy update constraints—specifically Truncated Natural Policy Gradient and Trust Region Policy Optimization—consistently outperformed other batch approaches across nearly all complex tasks by ensuring stable, step-by-step improvements. Second, the sample-efficient online method, Deep Deterministic Policy Gradient, learned significantly faster on specific high-dimensional tasks but exhibited substantial training instability and high sensitivity to reward scaling. Third, derivative-free evolutionary approaches proved surprisingly capable on simple tasks with thousands of parameters, but failed or exhausted available computing memory on high-dimensional systems like full humanoid locomotion. Finally, while memory-enhanced recurrent policies successfully mitigated sensory noise and missing velocity data in partially observable settings, every tested algorithm completely failed on hierarchical tasks requiring both low-level motor control and long-term navigation goals.

These findings demonstrate that no single algorithm currently offers an off-the-shelf solution for high-level autonomous physical control. In practical terms, while constrained policy gradient methods provide the most reliable, robust performance for robotic locomotion and motor skills, deploying them into real-world systems with compound, multi-tiered objectives carries high risk until algorithmic capabilities improve. The total failure across hierarchical benchmarks underscores that standard exploration techniques are inadequate for solving complex, multi-stage tasks.

For engineering leaders and researchers evaluating continuous control methods, the article suggests adopting Trust Region Policy Optimization or Truncated Natural Policy Gradient as primary baselines for complex motor control problems, while tuning step sizes conservatively. For operational environments with partial observability or sensor delays, teams should pair recurrent network architectures with constrained policy gradient updates. The primary immediate priority for research and development must be the design of algorithms that can automatically discover, structure, and exploit hierarchical skills, which are required for complex, goal-oriented tasks.

Confidence in these comparative findings is high for simulated environments with continuous physics. However, several limitations remain. All evaluations were conducted entirely within physics engines rather than on physical hardware, leaving potential transfer gaps to real-world mechanical systems. In addition, memory and computational constraints limited the scale of certain gradient-free comparisons on the highest-dimensional systems, and hyperparameter searches were conducted on representative subsets rather than exhaustively across all 31 tasks.

arXiv: 1604.06778rllab/rllab
  • Paper: OpenAI Gym, Greg Brockman et al. (2016). Integrates and standardizes the source paper's rllab continuous control benchmark into the ubiquitous OpenAI Gym framework for broad community adoption.
  • Paper: Deep Reinforcement Learning that Matters, Peter Henderson et al. (2018). Directly critiques and expands upon the reproducibility and evaluation methodology established in continuous control benchmark suites like rllab.
  • Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). Diagnoses value overestimation flaws in DDPG and proposes Twin Delayed DDPG (TD3), validating performance improvements across standard continuous control benchmark environments.
  • Paper: Soft Actor-Critic Algorithms and Applications, Tuomas Haarnoja et al. (2018). Presents Soft Actor-Critic (SAC), a state-of-the-art continuous control algorithm that addresses sample complexity on the benchmark tasks established by the source.
  • Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). Introduces Proximal Policy Optimization (PPO) as a more computationally efficient successor to TRPO, evaluating it on standard continuous locomotion benchmarks.
  • Paper: Deep Reinforcement Learning at the Edge of the Statistical Precipice, Rishabh Agarwal et al. (2021). Examines statistical uncertainty and evaluation protocols across reinforcement learning benchmarks, providing modern rigorous evaluation tools.
  • Paper: Stable-Baselines3: Reliable Reinforcement Learning Implementations, A. Raffin et al. (2021). Provides highly optimized and reproducible open-source reference implementations for the continuous control algorithms benchmarked in the literature.
  • Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). Extends continuous control benchmarking into the offline reinforcement learning paradigm using standardized datasets generated from continuous physics tasks.
  • Paper: Constrained Policy Optimization, Joshua Achiam et al. (2017). Builds upon the trust region optimization framework evaluated in the source to introduce safe reinforcement learning constraints on high-dimensional continuous control tasks.
  • Paper: Generative Adversarial Imitation Learning, Jonathan Ho et al. (2016). Applies trust region policy optimization to generative adversarial imitation learning, testing on the continuous control environments analyzed in the benchmark.
Cover for Benchmarking Deep Reinforcement Learning for Continuous Control

Abstract

Recently, researchers have made significant progress combining the advances in deep learning for learning feature representations with reinforcement learning. Some notable examples include training agents to play Atari games based on raw pixel data and to acquire advanced manipulation skills using raw sensory inputs. However, it has been difficult to quantify progress in the domain of continuous control due to the lack of a commonly adopted benchmark. In this work, we present a benchmark suite of continuous control tasks, including classic tasks like cart-pole swing-up, tasks with very high state and action dimensionality such as 3D humanoid locomotion, tasks with partial observations, and tasks with hierarchical structure. We report novel findings based on the systematic evaluation of a range of implemented reinforcement learning algorithms. Both the benchmark and reference implementations are released at this https URL in order to facilitate experimental reproducibility and to encourage adoption by other researchers.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Tasks
  • 3.1 Basic Tasks
  • 3.2 Locomotion Tasks
  • 3.3 Partially Observable Tasks
  • 3.4 Hierarchical Tasks
  • 4 Algorithms
  • 4.1 Batch Algorithms
  • 4.2 Online Algorithms
  • 4.3 Recurrent Variants
  • 5 Experiment Setup
  • 6 Results and Discussion
  • 7 Related Work
  • 8 Conclusion
  • References
  • 1 Task Specifications
  • 1.1 Basic Tasks
  • 1.2 Locomotion Tasks
  • 1.3 Partially Observable Tasks
  • 1.4 Hierarchical Tasks
  • 2 Experiment Parameters

Knowls

  1. Knowl 1 — Continuous Control Benchmark Task Suite Structure

    experimental setup

    The rllab continuous control benchmark consists of 31 continuous action tasks implemented in physics engines (Box2D for 2D dynamics and MuJoCo for 3D dynamics with complex contacts), divided into four distinct categories:

    1. Basic Tasks (5 tasks): Low-dimensional classical control benchmarks including Cart-Pole Balancing, Cart-Pole Swing Up, Mountain Car, Acrobot Swing Up, and Double Inverted Pendulum Balancing.

    2. Locomotion Tasks (6 tasks): High-dimensional, contact-rich 3D robot locomotion tasks requiring coordinated joint actuation to maximize forward velocity: Swimmer (1313-dim state, 22-dim action), Hopper (2020-dim state, 33-dim action), Walker2D (2121-dim state, 66-dim action), Half-Cheetah (2020-dim state, 66-dim action), Ant (125125-dim state, 88-dim action), Simple Humanoid (102102-dim state, 1010-dim action), and Full Humanoid (142142-dim state, 2828-dim action).

    3. Partially Observable Tasks (15 tasks): Constructed by modifying the five basic tasks under three realistic partial observability conditions:

      • Limited Sensors: Velocities are removed from observations, requiring velocity estimation from positional histories.
      • Noisy Observations and Delayed Actions: Zero-mean Gaussian noise (̂\sigma = 0.1) is added to observations and a physical actuation latency of 0.060.06--0.150.15 seconds (33 discretization frames) is introduced.
      • System Identification: Physical parameters (such as link lengths between 50%50\%--150%150\% or valley widths between 75%75\%--125%125\%) vary randomly across episodes.
    4. Hierarchical Tasks (4 tasks): Environments requiring both low-level continuous joint coordination and high-level sparse-reward goal navigation: Locomotion + Food Collection (Swimmer/Ant Gathering with +1+1 per food unit and −1-1 per bomb unit) and Locomotion + Maze (Swimmer/Ant Maze with +1+1 reward only on reaching the goal region).

  2. Knowl 2 — Empirical Benchmark Results Across Deep Reinforcement Learning Algorithms

    data/table

    The following table reports the performance of nine continuous control algorithms evaluated across representative tasks from the basic, locomotion, partially observable (annotated LS for Limited Sensors, NO for Noisy/Delayed, and SI for System Identification), and hierarchical categories. Performance is measured as the average undiscounted return over all iterations across 5 random seeds (reported as mean±standard deviation\text{mean} \pm \text{standard deviation}):

    Task REINFORCE TNPG RWR REPS TRPO CEM CMA-ES DDPG
    Cart-Pole Balancing 4693.7±14.04693.7 \pm 14.0 3986.4±748.93986.4 \pm 748.9 4861.5±12.34861.5 \pm 12.3 565.6±137.6565.6 \pm 137.6 4869.8±37.64869.8 \pm 37.6 4815.4±4.84815.4 \pm 4.8 2440.4±568.32440.4 \pm 568.3 4634.4±87.84634.4 \pm 87.8
    Inverted Pendulum 13.4±18.013.4 \pm 18.0 209.7±55.5209.7 \pm 55.5 84.7±13.884.7 \pm 13.8 −113.3±4.6-113.3 \pm 4.6 247.2±76.1247.2 \pm 76.1 38.2±25.738.2 \pm 25.7 −40.1±5.7-40.1 \pm 5.7 40.0±244.640.0 \pm 244.6
    Mountain Car −67.1±1.0-67.1 \pm 1.0 −66.5±4.5-66.5 \pm 4.5 −79.4±1.1-79.4 \pm 1.1 −275.6±166.3-275.6 \pm 166.3 −61.7±0.9-61.7 \pm 0.9 −66.0±2.4-66.0 \pm 2.4 −85.0±7.7-85.0 \pm 7.7 −288.4±170.3-288.4 \pm 170.3
    Acrobot −508.1±91.0-508.1 \pm 91.0 −395.8±121.2-395.8 \pm 121.2 −352.7±35.9-352.7 \pm 35.9 −1001.5±10.8-1001.5 \pm 10.8 −326.0±24.4-326.0 \pm 24.4 −436.8±14.7-436.8 \pm 14.7 −785.6±13.1-785.6 \pm 13.1 −223.6±5.8-223.6 \pm 5.8
    Double Inverted Pendulum 4116.5±65.24116.5 \pm 65.2 4455.4±37.64455.4 \pm 37.6 3614.8±368.13614.8 \pm 368.1 446.7±114.8446.7 \pm 114.8 4412.4±50.44412.4 \pm 50.4 2566.2±178.92566.2 \pm 178.9 1576.1±51.31576.1 \pm 51.3 2863.4±154.02863.4 \pm 154.0
    Swimmer 92.3±0.192.3 \pm 0.1 96.0±0.296.0 \pm 0.2 60.7±5.560.7 \pm 5.5 3.8±3.33.8 \pm 3.3 96.0±0.296.0 \pm 0.2 68.8±2.468.8 \pm 2.4 64.9±1.464.9 \pm 1.4 85.8±1.885.8 \pm 1.8
    Hopper 714.0±29.3714.0 \pm 29.3 1155.1±57.91155.1 \pm 57.9 553.2±71.0553.2 \pm 71.0 86.7±17.686.7 \pm 17.6 1183.3±150.01183.3 \pm 150.0 63.1±7.863.1 \pm 7.8 20.3±14.320.3 \pm 14.3 267.1±43.5267.1 \pm 43.5
    2D Walker 506.5±78.8506.5 \pm 78.8 1382.6±108.21382.6 \pm 108.2 136.0±15.9136.0 \pm 15.9 −37.0±38.1-37.0 \pm 38.1 1353.8±85.01353.8 \pm 85.0 84.5±19.284.5 \pm 19.2 77.1±24.377.1 \pm 24.3 318.4±181.6318.4 \pm 181.6
    Half-Cheetah 1183.1±69.21183.1 \pm 69.2 1729.5±184.61729.5 \pm 184.6 376.1±28.2376.1 \pm 28.2 34.5±38.034.5 \pm 38.0 1914.0±120.11914.0 \pm 120.1 330.4±274.8330.4 \pm 274.8 441.3±107.6441.3 \pm 107.6 2148.6±702.72148.6 \pm 702.7
    Ant 548.3±55.5548.3 \pm 55.5 706.0±127.7706.0 \pm 127.7 37.6±3.137.6 \pm 3.1 39.0±9.839.0 \pm 9.8 730.2±61.3730.2 \pm 61.3 49.2±5.949.2 \pm 5.9 17.8±15.517.8 \pm 15.5 326.2±20.8326.2 \pm 20.8
    Simple Humanoid 128.1±34.0128.1 \pm 34.0 255.0±24.5255.0 \pm 24.5 93.3±17.493.3 \pm 17.4 28.3±4.728.3 \pm 4.7 269.7±40.3269.7 \pm 40.3 60.6±12.960.6 \pm 12.9 28.7±3.928.7 \pm 3.9 99.4±28.199.4 \pm 28.1
    Full Humanoid 262.2±10.5262.2 \pm 10.5 288.4±25.2288.4 \pm 25.2 46.7±5.646.7 \pm 5.6 41.7±6.141.7 \pm 6.1 287.0±23.4287.0 \pm 23.4 36.9±2.936.9 \pm 2.9 N/A 119.0±31.2119.0 \pm 31.2
    Cart-Pole Balancing (LS) 420.9±265.5420.9 \pm 265.5 945.1±27.8945.1 \pm 27.8 68.9±1.568.9 \pm 1.5 898.1±22.1898.1 \pm 22.1 960.2±46.0960.2 \pm 46.0 227.0±223.0227.0 \pm 223.0 68.0±1.668.0 \pm 1.6 –
    Cart-Pole Balancing (NO) 616.0±210.8616.0 \pm 210.8 916.3±23.0916.3 \pm 23.0 93.8±1.293.8 \pm 1.2 99.6±7.299.6 \pm 7.2 606.2±122.2606.2 \pm 122.2 181.4±32.1181.4 \pm 32.1 104.4±16.0104.4 \pm 16.0 –
    Cart-Pole Balancing (SI) 431.7±274.1431.7 \pm 274.1 980.5±7.3980.5 \pm 7.3 69.0±2.869.0 \pm 2.8 702.4±196.4702.4 \pm 196.4 980.3±5.1980.3 \pm 5.1 746.6±93.2746.6 \pm 93.2 71.6±2.971.6 \pm 2.9 –
    Ant + Gathering −0.1±0.1-0.1 \pm 0.1 −0.4±0.1-0.4 \pm 0.1 −5.5±0.5-5.5 \pm 0.5 −6.7±0.7-6.7 \pm 0.7 −0.4±0.0-0.4 \pm 0.0 −4.7±0.7-4.7 \pm 0.7 N/A −0.3±0.3-0.3 \pm 0.3
    Ant + Maze 0.0±0.00.0 \pm 0.0 0.0±0.00.0 \pm 0.0 0.0±0.00.0 \pm 0.0 0.0±0.00.0 \pm 0.0 0.0±0.00.0 \pm 0.0 0.0±0.00.0 \pm 0.0 N/A 0.0±0.00.0 \pm 0.0

    The benchmark results demonstrate that TRPO and TNPG consistently obtain the highest cumulative return across continuous locomotion tasks. DDPG demonstrates rapid learning on tasks like Half-Cheetah but shows high variance and degradation on others. CMA-ES fails (runs out of memory) on the highest-dimensional task (Full Humanoid). All tested algorithms fail to achieve positive progress on hierarchical navigation and maze tasks.

  3. Knowl 3 — Continuous-Action Policy Search Formulations in rllab

    model/method

    The benchmark provides implementations of several continuous policy search algorithms operating on parameterized policies πθ\pi_\theta:

    1. REINFORCE: Computes an empirical estimate of the gradient of the expected return η(πθ)=Eτ[∑t=0Tγtr(st,at)]\eta(\pi_\theta) = \mathbb{E}_\tau \left[ \sum_{t=0}^T \gamma^t r(s_t, a_t) \right] using the likelihood ratio with a state-dependent baseline b(sti)b(s_t^i): ∇θη(πθ)≈1NT∑i=1N∑t=0T∇θlog⁡πθ(ati∣sti)(Rti−b(sti))\nabla_\theta \eta(\pi_\theta) \approx \frac{1}{NT} \sum_{i=1}^N \sum_{t=0}^T \nabla_\theta \log \pi_\theta(a_t^i \mid s_t^i) \left( R_t^i - b(s_t^i) \right) where Rti=∑t′=tTγt′−tr(st′i,at′i)R_t^i = \sum_{t'=t}^T \gamma^{t'-t} r(s_{t'}^i, a_{t'}^i), NN is the batch size of trajectories, and TT is the horizon length.

    2. Truncated Natural Policy Gradient (TNPG): Computes the ascent direction I(θ)−1∇θη(πθ)I(\theta)^{-1} \nabla_\theta \eta(\pi_\theta), where I(θ)I(\theta) is the Fisher Information Matrix (FIM), using conjugate gradient descent without forming or inverting the full matrix I(θ)I(\theta), with step size: α=δKL∇θη(πθ)TI(θ)−1∇θη(πθ)\alpha = \sqrt{\frac{\delta_{KL}}{\nabla_\theta \eta(\pi_\theta)^T I(\theta)^{-1} \nabla_\theta \eta(\pi_\theta)}}

    3. Trust Region Policy Optimization (TRPO): Optimizes a surrogate advantage objective constrained by the average KL divergence between the old and new policy: max⁡θEs∼ρθk,a∼πθk[πθ(a∣s)πθk(a∣s)Aθk(s,a)]s.t.Es∼ρθk[DKL(πθk(⋅∣s)∥πθ(⋅∣s))]≤δKL\max_\theta \mathbb{E}_{s \sim \rho_{\theta_k}, a \sim \pi_{\theta_k}} \left[ \frac{\pi_\theta(a \mid s)}{\pi_{\theta_k}(a \mid s)} A_{\theta_k}(s, a) \right] \quad \text{s.t.} \quad \mathbb{E}_{s \sim \rho_{\theta_k}} \left[ D_{KL}\left( \pi_{\theta_k}(\cdot \mid s) \parallel \pi_\theta(\cdot \mid s) \right) \right] \le \delta_{KL} where ρθk\rho_{\theta_k} is the discounted state visitation distribution under πθk\pi_{\theta_k}. It uses the same conjugate gradient direction as TNPG, followed by a backtracking line search ensuring improvement in the surrogate objective and satisfaction of the KL constraint.

    4. Reward-Weighted Regression (RWR): Maximizes a lower bound on log-expected return via Expectation-Maximization: θk+1=arg⁡max⁡θ1NT∑i=1N∑t=0Tlog⁡π(ati∣sti;θ)ρ(Rti−b(sti))\theta_{k+1} = \arg\max_\theta \frac{1}{NT} \sum_{i=1}^N \sum_{t=0}^T \log \pi(a_t^i \mid s_t^i; \theta) \rho\left(R_t^i - b(s_t^i)\right) where ρ(R)=R−Rmin⁡\rho(R) = R - R_{\min} maps returns to non-negative values (Rmin⁡R_{\min} is the minimum return in the batch).

    5. Relative Entropy Policy Search (REPS): Limits information loss per step by solving a dual optimization problem for η>0\eta > 0 and Bellman parameter vector ν\nu: min⁡η>0,νηδKL+ηlog⁡(1M∑i=1Mexp⁡(ri+νT(ϕ(si′)−ϕ(si))η))\min_{\eta > 0, \nu} \eta \delta_{KL} + \eta \log \left( \frac{1}{M} \sum_{i=1}^M \exp\left( \frac{r_i + \nu^T(\phi(s_i') - \phi(s_i))}{\eta} \right) \right) and updating policy parameters θ\theta via weighted maximum likelihood using optimal dual parameters [η∗,ν∗][\eta^*, \nu^*].

  4. Knowl 4 — Learning Stability and Local Optima in TRPO and TNPG vs REINFORCE

    empirical result

    Empirical evaluation on continuous control benchmarks indicates that TRPO and TNPG substantially outperform vanilla REINFORCE, RWR, and REPS across continuous locomotion tasks.

    • Step Size and Policy Shifts: Vanilla REINFORCE exhibits severe policy distribution instability; even when using a small learning rate, gradient steps occasionally induce large KL divergences in policy space, resulting in catastrophic performance drops and premature convergence to poor local optima. For instance, on the 2D Walker task, REINFORCE converges to gaits that jump forward and fall over for immediate reward rather than learning stable, long-term walking gaits.

    • Trust Region Constraints: By explicitly constraining the KL divergence between consecutive policy updates (DKL(πθk∥πθ)≤δKLD_{KL}(\pi_{\theta_k} \parallel \pi_\theta) \le \delta_{KL}), TNPG and TRPO prevent large policy collapses. TRPO further improves on TNPG by performing a line search on both the surrogate objective and the KL constraint, enabling it to robustly use larger step sizes (e.g., δKL=0.1\delta_{KL} = 0.1 for TRPO vs δKL=0.05\delta_{KL} = 0.05 for TNPG on the Swimmer task), which yields faster convergence without policy degradation.

  5. Knowl 5 — Sample Efficiency, Policy Instability, and Reward Scaling in DDPG

    empirical result

    Deep Deterministic Policy Gradient (DDPG) optimizes a deterministic policy μθ:S→A\mu_\theta: S \to A using off-policy replay buffer transitions and policy gradient updates: ∇θη(μθ)=1B∑i=1B∇aQϕ(si,a)∣a=μθ(si)∇θμθ(si)\nabla_\theta \eta(\mu_\theta) = \frac{1}{B} \sum_{i=1}^B \nabla_a Q_\phi(s_i, a)\big|_{a=\mu_\theta(s_i)} \nabla_\theta \mu_\theta(s_i) where QϕQ_\phi is a critic trained via mean-squared Bellman error regression with target networks.

    Empirical comparisons between DDPG and batch on-policy policy gradient algorithms reveal key trade-offs:

    1. Sample Efficiency: DDPG exhibits significantly higher sample efficiency and faster convergence speed during early training iterations on tasks such as Half-Cheetah (2148.6±702.72148.6 \pm 702.7).
    2. Training Instability: DDPG suffers from training instability compared to batch methods; policy performance frequently experiences sharp degradation or collapse during training.
    3. Reward Sensitivity: DDPG is highly sensitive to the scale of environment rewards. Multiplying the reward function by a constant factor of 0.10.1 across all tasks was necessary to achieve stable training.
  6. Knowl 6 — Scaling Characteristics and Limits of Derivative-Free Optimization (CEM and CMA-ES)

    empirical result

    Derivative-free evolutionary methods optimize policy parameters by sampling perturbations directly in parameter space θ∼N(μ,Σ)\theta \sim \mathcal{N}(\mu, \Sigma) and updating the distribution using elite candidates:

    1. Dimensionality vs. Complexity: The Cross-Entropy Method (CEM) successfully optimizes deep neural network policies containing thousands of parameters on low-dimensional basic control tasks (e.g., Cart-Pole Balancing and Mountain Car), demonstrating that parameter count alone does not prohibit derivative-free optimization.
    2. Performance Degradation on Complex Dynamics: CEM performance degrades dramatically on complex locomotion tasks with high degrees of freedom, underperforming policy gradient methods by an order of magnitude.
    3. CMA-ES Memory and Computation Bottleneck: CEM consistently matches or outperforms CMA-ES. CMA-ES, which incrementally updates a full covariance matrix along evolution paths, incurs prohibitive quadratic memory and computational costs as the observation and parameter dimensions grow. On the Full Humanoid task (142142-dimensional observations), CMA-ES runs out of memory (OOM).
  7. Knowl 7 — Recurrent Policies for Continuous Partially Observable Control

    empirical result

    To evaluate reinforcement learning under partial observability, batch algorithms were extended to recurrent policies π(at∣o1:t,a1:t−1)\pi(a_t \mid o_{1:t}, a_{1:t-1}) parameterized as a single-layer LSTM with 32 hidden units.

    Key findings include:

    1. State Recovery: Recurrent policies achieve higher returns than feedforward policies across partially observable variants (Limited Sensors, Noisy Observations + Delayed Actions, and System Identification) by integrating history to infer missing velocities, filter observation noise, and estimate varying physical system parameters.
    2. Increased Optimization Difficulty: Recurrent neural network policies are significantly harder to optimize than feedforward networks. Derivative-free algorithms (CEM and CMA-ES) degrade substantially when applied to recurrent architectures.
    3. Widened Policy Gradient Gap: The performance gap between first-order REINFORCE and natural gradient methods (TNPG and TRPO) is substantially larger in recurrent policies than in feedforward policies. This occurs because small perturbations in recurrent weight matrices cause compounded, long-horizon shifts in the trajectory policy distribution.
  8. Knowl 8 — Failure of Flat Deep Reinforcement Learning on Hierarchical Continuous Tasks

    limitation

    All standard batch and online reinforcement learning algorithms tested in the benchmark suite (REINFORCE, TNPG, TRPO, RWR, REPS, CEM, CMA-ES, and DDPG) fail completely on the hierarchical continuous control tasks (Swimmer Gathering, Ant Gathering, Swimmer Maze, and Ant Maze), yielding zero or near-zero cumulative returns even after extensive hyperparameter grid search and 500 iterations (25 million simulation steps).

    The failure stems from the multi-timescale nature of the tasks: solving them requires concurrently learning low-level motor actuation (coordinating multiple joints to achieve stable locomotion) and high-level exploration/navigation under sparse rewards (locating food/targets across a 2D environment). Standard flat deep RL exploration strategies cannot discover long-horizon sparse navigation rewards while simultaneously searching for low-level continuous locomotion gaits.

  9. Knowl 9 — Experimental Protocol, Baseline Architecture, and Hyperparameter Selection Criterion

    experimental setup

    The rllab benchmarking experiments follow a standardized protocol across tasks and algorithms:

    • Sampling and Horizon Budget: Each training iteration collects 50,00050{,}000 simulation steps. Total iterations are set to 500500 for Basic, Locomotion, and Hierarchical tasks (horizon T=500T = 500), and 300300 for Partially Observable tasks (horizon T=100T = 100). The discount factor is γ=0.99\gamma = 0.99.

    • Time-Varying Feature Linear Baseline: For variance reduction in batch gradient-based algorithms (except REPS), a linear baseline with state- and time-dependent features is subtracted from the empirical return: ϕ(s,t)=[sT, (s⊙s)T, 0.01t, (0.01t)2, (0.01t)3, 1]T\phi(s, t) = \left[ s^T, \,(s \odot s)^T, \, 0.01t, \, (0.01t)^2, \, (0.01t)^3, \, 1 \right]^T where ss is the state vector and ⊙\odot denotes the element-wise product.

    • Network Architectures:

      • Batch Feedforward Policies: 3-hidden-layer MLP (100,50,25100, 50, 25 hidden units with tanh⁡\tanh activations) mapping states to the Gaussian mean, with a state-independent global trainable vector for log-standard deviations.
      • DDPG Policy and Critic: 2-hidden-layer MLP (400,300400, 300 units with ReLU activations).
      • Recurrent Policies: Single-layer LSTM with 32 hidden units.
    • Hyperparameter Selection Metric: Grid search over step sizes and learning rates evaluates performance under 5 random seeds using the robust selection metric: Score=mean(returns)−std(returns)\text{Score} = \text{mean}(\text{returns}) - \text{std}(\text{returns}) which explicitly penalizes algorithms and step sizes that suffer from large variance or catastrophic policy collapse.

  10. Knowl 10 — Continuous Locomotion Task Reward and Termination Formulations

    definition

    The MuJoCo-based locomotion tasks in rllab are defined with specific reward functions r(s,a)r(s, a) and early termination conditions:

    • Swimmer: r(s,a)=vx−0.005∥a∥22r(s, a) = v_x - 0.005 \|a\|_2^2, where vxv_x is forward velocity. No termination condition.
    • Hopper: r(s,a)=vx−0.005∥a∥22+1r(s, a) = v_x - 0.005 \|a\|_2^2 + 1 (alive bonus). Terminates if body height zbody<0.7z_{\text{body}} < 0.7 or forward pitch ∣θy∣<0.2|\theta_y| < 0.2.
    • Walker2D: r(s,a)=vx−0.005∥a∥22r(s, a) = v_x - 0.005 \|a\|_2^2. Terminates if zbody<0.8z_{\text{body}} < 0.8, zbody>2.0z_{\text{body}} > 2.0, or ∣θy∣>1.0|\theta_y| > 1.0.
    • Half-Cheetah: r(s,a)=vx−0.05∥a∥22r(s, a) = v_x - 0.05 \|a\|_2^2. No termination condition.
    • Ant: r(s,a)=vx−0.005∥a∥22−Ccontact+0.05r(s, a) = v_x - 0.005 \|a\|_2^2 - C_{\text{contact}} + 0.05, where Ccontact=5×10−4∥Fcontact∥22C_{\text{contact}} = 5 \times 10^{-4} \|F_{\text{contact}}\|_2^2 (FcontactF_{\text{contact}} is contact force clipped to [−1,1][-1, 1]). Terminates if zbody<0.2z_{\text{body}} < 0.2 or zbody>1.0z_{\text{body}} > 1.0.
    • Simple Humanoid: r(s,a)=vx−5×10−4∥a∥22−Ccontact−Cdeviation+0.2r(s, a) = v_x - 5 \times 10^{-4} \|a\|_2^2 - C_{\text{contact}} - C_{\text{deviation}} + 0.2, where Ccontact=5×10−6∥Fcontact∥2C_{\text{contact}} = 5 \times 10^{-6} \|F_{\text{contact}}\|_2 and Cdeviation=5×10−3(vy2+vz2)C_{\text{deviation}} = 5 \times 10^{-3} (v_y^2 + v_z^2). Terminates if zbody<0.8z_{\text{body}} < 0.8 or zbody>2.0z_{\text{body}} > 2.0.
    • Full Humanoid: Identical reward and termination criteria to Simple Humanoid, expanded to 19 rigid links and 28 actuated joints.

Coverage note — No substantial contributed material was omitted; all 31 benchmark tasks, algorithmic formulations, hyperparameter selection criteria, and experimental findings are covered across the knowls.

References

  1. 1.Abeyruwan, S. RLLib: Lightweight standard and on/off policy reinforcement learning library (C++). http://web.cs.miami.edu/home/saminda/rllib.html, 2013.
  2. 2.Bagnell, J. A. and Schneider, J. Covariant policy search. pp. 1019–1024. IJCAI, 2003.
  3. 3.Bakker, B. Reinforcement learning with long short-term memory. In NIPS, pp. 1475–1482, 2001.
  4. 4.Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. The Arcade Learning Environment: An evaluation platform for general agents. J. Artif. Intell. Res., 47:253–279, 2013.
  5. 5.Bellman, R. Dynamic Programming. Princeton University Press, 1957.
  6. 6.Bertsekas, Dimitri P and Tsitsiklis, John N. Neuro-dynamic programming: an overview. In CDC, pp. 560–564, 1995.
  7. 7.Busoniu, L. ApproxRL: A Matlab toolbox for approximate RL and DP. http://busoniu.net/files/repository/readme-approxrl.html, 2010.
  8. 8.Catto, E. Box2D: A 2D physics engine for games, 2011.
  9. 9.Coulom, Remi. Reinforcement learning using neural networks, with applications to motor control. PhD thesis, Institut National Polytechnique de Grenoble-INPG, 2002.
  10. 10.Dann, C., Neumann, G., and Peters, J. Policy evaluation with temporal differences: A survey and comparison. J. Mach. Learn. Res., 15(1):809–883, 2014.
  11. 11.Degris, T., Bechu, J., White, A., Modayil, J., Pilarski, P. M., and Denk, C. RLPark. http://rlpark.github.io, 2013.
  12. 12.Deisenroth, M. P., Neumann, G., and Peters, J. A survey on policy search for robotics, foundations and trends in robotics. Found. Trends Robotics, 2(1-2):1–142, 2013.
  13. 13.DeJong, G. and Spong, M. W. Swinging up the Acrobot: An example of intelligent control. In ACC, pp. 2158–2162, 1994.
  14. 14.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In CVPR, pp. 248–255, 2009.
  15. 15.Dietterich, T. G. Hierarchical reinforcement learning with the MAXQ value function decomposition. J. Artif. Intell. Res, 13: 227–303, 2000.
  16. 16.Dimitrakakis, C., Tziortziotis, N., and Tossou, A. Beliefbox: A framework for statistical methods in sequential decision making. http://code.google.com/p/beliefbox/, 2007.
  17. 17.Dimitrakakis, Christos, Li, Guangliang, and Tziortziotis, Nikoalos. The reinforcement learning competition 2014. AI Magazine, 35(3):61–65, 2014.
  18. 18.Donaldson, P. E. K. Error decorrelation: a technique for matching a class of functions. In Proc. 3th Intl. Conf. Medical Electronics, pp. 173–178, 1960.
  19. 19.Doya, K. Reinforcement learning in continuous time and space. Neural Comput., 12(1):219–245, 2000.
  20. 20.Dutech, Alain, Edmunds, Timothy, Kok, Jelle, Lagoudakis, Michail, Littman, Michael, Riedmiller, Martin, Russell, Bryan, Scherrer, Bruno, Sutton, Richard, Timmer, Stephan, et al. Reinforcement learning benchmarks and bake-offs ii. Advances in Neural Information Processing Systems (NIPS), 17, 2005.
  21. 21.Erez, Tom, Tassa, Yuval, and Todorov, Emanuel. Infinite horizon model predictive control for nonlinear periodic tasks. Manuscript under review, 4, 2011.
  22. 22.Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A. The pascal visual object classes (VOC) challenge. Int. J. Comput. Vision, 88(2):303–338, 2010.
  23. 23.Fei-Fei, L., Fergus, R., and Perona, P. One-shot learning of object categories. IEEE Trans. Pattern Anal. Mach. Intell., 28(4):594–611, 2006.
  24. 24.Furuta, K., Okutani, T., and Sone, H. Computer control of a double inverted pendulum. Comput. Electr. Eng., 5(1):67–84, 1978.
  25. 25.Garofolo, J. S., Lamel, L. F., Fisher, W. M., Fiscus, J. G., and Pallett, D. S. DARPA TIMIT acoustic-phonetic continuous speech corpus CD-ROM. NIST speech disc 1-1.1. NASA STI/Recon Technical Report N, 93, 1993.
  26. 26.Godfrey, J. J., Holliman, E. C., and McDaniel, J. SWITCHBOARD: Telephone speech corpus for research and development. In ICASSP, pp. 517–520, 1992.
  27. 27.Gomez, F. and Miikkulainen, R. 2-d pole balancing with recurrent evolutionary networks. In ICANN, pp. 425–430. 1998.
  28. 28.Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X. Deep learning for real-time Atari game play using offline montecarlo tree search planning. In NIPS, pp. 3338–3346. 2014.
  29. 29.Hansen, N. and Ostermeier, A. Completely derandomized selfadaptation in evolution strategies. Evol. Comput., 9(2):159–195, 2001.
  30. 30.Heess, N., Hunt, J., Lillicrap, T., and Silver, D. Memory-based control with recurrent neural networks. arXiv:1512.04455, 2015a.
  31. 31.Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, T. Learning continuous control policies by stochastic value gradients. In NIPS, pp. 2926–2934. 2015b.
  32. 32.Hester, T. and Stone, P. The open-source TEXPLORE code release for reinforcement learning on robots. In RoboCup 2013: Robot World Cup XVII, pp. 536–543. 2013.
  33. 33.Hinton, G., Deng, L., Yu, D., Mohamed, A.-R., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Dahl, T. S. G., and Kingsbury, B. Deep neural networks for acoustic modeling in speech recognition. IEEE Signal Process. Mag, 29(6):82–97, 2012.
  34. 34.Hirsch, H.-G. and Pearce, D. The Aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions. In ASR2000-Automatic Speech Recognition: Challenges for the new Millenium ISCA Tutorial and Research Workshop (ITRW), 2000.
  35. 35.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Comput., 9(8):1735–1780, 1997.
  36. 36.Kakade, S. M. A natural policy gradient. In NIPS, pp. 1531–1538. 2002.
  37. 37.Kimura, H. and Kobayashi, S. Stochastic real-valued reinforcement learning to solve a nonlinear control problem. In IEEE SMC, pp. 510–515, 1999.
  38. 38.Kober, J. and Peters, J. Policy search for motor primitives in robotics. In NIPS, pp. 849–856, 2009.
  39. 39.Kochenderfer, M. JRLF: Java reinforcement learning framework. http://mykel.kochenderfer.com/jrlf, 2006.
  40. 40.Krizhevsky, A. and Hinton, G. Learning multiple layers of features from tiny images. Technical report, 2009.
  41. 41.Krizhevsky, A., Sutskever, I., and Hinton, G. ImageNet classification with deep convolutional neural networks. In NIPS, pp. 1097–1105. 2012.
  42. 42.LeCun, Y., Cortes, C., and Burges, C. The MNIST database of handwritten digits, 1998.
  43. 43.Levine, S. and Koltun, V. Guided policy search. In ICML, pp. 1–9, 2013.
  44. 44.Levine, S., Finn, C., Darrell, T., and Abbeel, P. End-to-end training of deep visuomotor policies. arXiv:1504.00702, 2015.
  45. 45.Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. Continuous control with deep reinforcement learning. arXiv:1509.02971, 2015.
  46. 46.Martin, D., C. Fowlkes, D. Tal, and Malik, J. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In ICCV, pp. 416–423, 2001.
  47. 47.Metzen, J. M. and Edgington, M. Maja machine learning framework. http://mloss.org/software/view/220/, 2011.
  48. 48.Michie, D. and Chambers, R. A. BOXES: An experiment in adaptive control. Machine Intelligence, 2:137–152, 1968.
  49. 49.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
  50. 50.Moore, A. Efficient memory-based learning for robot control. Technical report, University of Cambridge, Computer Laboratory, 1990.
  51. 51.Murray, R. M. and Hauser, J. A case study in approximate linearization: The Acrobot example. Technical report, UC Berkeley, EECS Department, 1991.
  52. 52.Murthy, S. S. and Raibert, M. H. 3D balance in legged locomotion: modeling and simulation for the one-legged case. ACM SIGGRAPH Computer Graphics, 18(1):27–27, 1984.
  53. 53.Neumann, G. A reinforcement learning toolbox and RL benchmarks for the control of dynamical systems. Dynamical principles for neuroscience and intelligent biomimetic devices, pp. 113, 2006.
  54. 54.Papis, B. and Wawrzyński, P. dotrl: A platform for rapid reinforcement learning methods development and validation. In FedCSIS, pp. pages 129–136., 2013.
  55. 55.Parr, Ronald and Russell, Stuart. Reinforcement learning with hierarchies of machines. Advances in neural information processing systems, pp. 1043–1049, 1998.
  56. 56.Peters, J. Policy Gradient Toolbox. http://www.ausy.tu-darmstadt.de/Research/PolicyGradientToolbox, 2002.
  57. 57.Peters, J. and Schaal, S. Reinforcement learning by rewardweighted regression for operational space control. In ICML, pp. 745–750, 2007.
  58. 58.Peters, J. and Schaal, S. Reinforcement learning of motor skills with policy gradients. Neural networks, 21(4):682–697, 2008.
  59. 59.Peters, J., Vijaykumar, S., and Schaal, S. Policy gradient methods for robot control. Technical report, 2003.
  60. 60.Peters, J., Mulling, K., and Altun, Y. Relative entropy policy search. In AAAI, pp. 1607–1612, 2010.
  61. 61.Purcell, E. M. Life at low Reynolds number. Am. J. Phys, 45(1): 3–11, 1977.
  62. 62.Raibert, M. H. and Hodgins, J. K. Animation of dynamic legged locomotion. In ACM SIGGRAPH Computer Graphics, volume 25, pp. 349–358, 1991.
  63. 63.Riedmiller, M., Blum, M., and Lampe, T. CLS2: Closed loop simulation system. http://ml.informatik.uni-freiburg.de/research/clsquare, 2012.
  64. 64.Rubinstein, R. The cross-entropy method for combinatorial and continuous optimization. Methodol. Comput. Appl. Probab., 1 (2):127–190, 1999.
  65. 65.Schafer, A. M. and Udluft, S. Solving partially observable reinforcement learning problems with recurrent neural networks. In ECML Workshops, pp. 71–81, 2005.
  66. 66.Schaul, T., Bayer, J., Wierstra, D., Sun, Y., Felder, M., Sehnke, F., Ruckstiess, T., and Schmidhuber, J. PyBrain. J. Mach. Learn. Res., 11:743–746, 2010.
  67. 67.Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P. Trust region policy optimization. In ICML, pp. 1889–1897, 2015a.
  68. 68.Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P. High-dimensional continuous control using generalized advantage estimation. arXiv:1506.02438, 2015b.
  69. 69.Stephenson, A. On induced stability. Philos. Mag., 15(86):233–236, 1908.
  70. 70.Stone, Peter, Kuhlmann, Gregory, Taylor, Matthew E, and Liu, Yaxin. Keepaway soccer: From machine learning testbed to benchmark. In RoboCup 2005: Robot Soccer World Cup IX, pp. 93–105. Springer, 2005.
  71. 71.Sutton, Richard S, Precup, Doina, and Singh, Satinder. Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning. Artificial intelligence, 112(1):181–211, 1999.
  72. 72.Szita, I. and Lorincz, A. Learning Tetris using the noisy crossentropy method. Neural Comput., 18(12):2936–2941, 2006.
  73. 73.Szita, I., Takacs, B., and Lorincz, A. ̄̀́-MDPs: Learning in varying environments. J. Mach. Learn. Res., 3:145–174, 2003.
  74. 74.Tassa, Yuval, Erez, Tom, and Todorov, Emanuel. Synthesis and stabilization of complex behaviors through online trajectory optimization. In Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ International Conference on, pp. 4906–4913. IEEE, 2012.
  75. 75.Tesauro, G. Temporal difference learning and TD-Gammon. Commun. ACM, 38(3):58–68, 1995.
  76. 76.Todorov, E., Erez, T., and Tassa, Y. MuJoCo: A physics engine for model-based control. In IROS, pp. 5026–5033, 2012.
  77. 77.van Hoof, H., Peters, J., and Neumann, G. Learning of nonparametric control policies with high-dimensional state features. In AISTATS, pp. 995–1003, 2015.
  78. 78.Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M. Embed to control: A locally linear latent dynamics model for control from raw images. In NIPS, pp. 2728–2736, 2015.
  79. 79.Wawrzyński, P. Learning to control a 6-degree-of-freedom walking robot. In IEEE EUROCON, pp. 698–705, 2007.
  80. 80.Widrow, B. Pattern recognition and adaptive control. IEEE Trans. Ind. Appl., 83(74):269–277, 1964.
  81. 81.Wierstra, D., Foerster, A., Peters, J., and Schmidhuber, J. Solving deep memory POMDPs with recurrent policy gradients. In ICANN, pp. 697–706. 2007.
  82. 82.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn., 8: 229–256, 1992.
  83. 83.Yamaguchi, A. and Ogasawara, T. SkyAI: Highly modularized reinforcement learning library. In IEEE-RAS Humanoids, pp. 118–123, 2010.
  84. 84.Yu, D., Ju, Y.-C., Wang, Y.-Y., Zweig, G., and Acero, A. Automated directory assistance system - from theory to practice. In Interspeech, pp. 2709–2712, 2007.

Citation

MLA
Duan, Y., et al. “Benchmarking Deep Reinforcement Learning for Continuous Control”. arXiv, 2016, http://arxiv.org/abs/1604.06778v3.
APA
Duan, Y., Chen, X., Houthooft, R., Schulman, J., & Abbeel, P. (2016). Benchmarking Deep Reinforcement Learning for Continuous Control. arXiv. http://arxiv.org/abs/1604.06778v3
Chicago
Duan, Y., X. Chen, R. Houthooft, J. Schulman, and P. Abbeel. 2016. “Benchmarking Deep Reinforcement Learning for Continuous Control”. arXiv. http://arxiv.org/abs/1604.06778v3.
Harvard
Duan, Y. et al. (2016) “Benchmarking Deep Reinforcement Learning for Continuous Control”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1604.06778v3.
Vancouver
1. Duan Y, Chen X, Houthooft R, Schulman J, Abbeel P (2016) Benchmarking Deep Reinforcement Learning for Continuous Control. arXiv

BibTeX

@article{duan2016benchmarking,
  title = {Benchmarking Deep Reinforcement Learning for Continuous Control},
  author = {Duan, Yan and Chen, Xi and Houthooft, Rein and Schulman, John and Abbeel, Pieter},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1604.06778v3},
  eprint = {1604.06778}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission