Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning

Denis YaratsRob FergusAlessandro LazaricLerrel Pinto

article2022ICLR505 citations

Introduces DrQ-v2, a fast and lightweight model-free reinforcement learning algorithm that achieves state-of-the-art visual continuous control on the DeepMind Control Suite, solving complex humanoid locomotion directly from pixels in only eight hours on a single GPU.

Listen

Training autonomous systems to perform continuous control directly from visual camera inputs has long posed a significant technical hurdle in reinforcement learning. Prior approaches have required prohibitive amounts of interaction data and immense computational resources involving multi-GPU clusters, while still failing to master highly complex physical behaviors such as humanoid locomotion.

The article introduces and evaluates DrQ-v2, an improved model-free reinforcement learning algorithm designed to achieve state-of-the-art sample and computational efficiency in vision-based continuous control. The primary objective is to demonstrate that a conceptually simple, data-augmented model-free framework can successfully solve highly intricate visual control benchmarks on modest single-GPU hardware.

To establish its findings, the authors conducted extensive empirical evaluations across 24 continuous control tasks from the DeepMind Control Suite, categorizing them into easy, medium, and hard tiers. The experiments assessed both sample efficiency (the number of environment interactions required) and computational throughput (wall-clock training time) on a single NVIDIA V100 graphics processing unit. The approach replaces earlier base algorithms with deep deterministic policy gradient, incorporates multi-step returns, refines image augmentation, applies an exploration noise decay schedule, and optimizes memory and processing pipelines.

The evaluation yielded several key findings. First, DrQ-v2 achieved state-of-the-art sample efficiency among model-free methods and became the first model-free method to successfully solve complex humanoid locomotion tasks directly from pixels. Second, it delivered a significant computational speedup, achieving 96 frames per second—a 3.4-fold improvement over its predecessor DrQ and a 6-fold increase over contrastive methods. Third, DrQ-v2 allowed most benchmark tasks to train in roughly 8.6 hours on a single processor, with easy tasks finishing in under 3 hours and the hardest humanoid tasks completing in about 86 hours. Fourth, when compared against DreamerV2, a leading model-based alternative, DrQ-v2 matched sample efficiency on many tasks while training approximately four times faster in wall-clock time due to lower computational overhead.

These results demonstrate that complex visual control does not necessarily require computationally heavy world models or massive distributed compute clusters. By drastically lowering wall-clock training times and hardware demands, the algorithm substantially reduces the infrastructure cost and experimentation risk associated with visual control research, making high-performance development accessible on standard hardware.

Based on these findings, the authors recommend that practitioners adopt DrQ-v2 as a strong, computationally efficient baseline for vision-based continuous control. They also advise the research community to retire saturated, easy benchmarks and focus evaluation efforts on medium and hard environments, such as quadruped and humanoid control. Further investigation is recommended to understand why model-based approaches still outperform model-free methods on specific sparse-reward or intricate tasks.

Confidence in these findings is supported by consistent results across 10 random seeds and 24 varied benchmark environments. However, readers should note that the evaluation is confined to simulated physics domains, and real-world robotic deployments may introduce additional physical uncertainties, visual variations, and boundary conditions not captured in the simulation suite.

  • Paper: DeepMind Control Suite, Yuval Tassa et al. (2018). It introduces the standardized DeepMind Control Suite continuous control benchmark tasks and pixel-based baselines directly evaluated and built upon by DrQ-v2.
  • Paper: CURL: Contrastive Unsupervised Representations for Reinforcement Learning, Aravind Srinivas et al. (2020). It establishes foundational data-efficient visual reinforcement learning on the DeepMind Control Suite, providing the key contrastive model-free baseline and image-augmentation context improved upon by DrQ-v2.
  • Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). It introduces the Deep Deterministic Policy Gradient (DDPG) framework that DrQ-v2 adopts as its foundational model-free continuous control base algorithm.
  • Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). It addresses function approximation error and overestimation in continuous actor-critic methods, introducing multi-critic techniques and target noise mechanisms that inform DrQ-v2's policy optimization.
  • Paper: Mastering Atari with Discrete World Models, Danijar Hafner et al. (2021). It presents DreamerV2, the state-of-the-art model-based visual RL framework against which DrQ-v2 directly benchmarks its sample efficiency and computational speedup.
  • Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). It introduces latent imagination for visual continuous control in the DeepMind Control Suite, establishing the model-based paradigm compared against DrQ-v2's model-free approach.
  • Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). It provides the underlying theoretical foundation of deterministic policy gradient methods used in continuous-action visual reinforcement learning.
Cover for Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning

Abstract

We present DrQ-v2, a model-free reinforcement learning (RL) algorithm for visual continuous control. DrQ-v2 builds on DrQ, an off-policy actor-critic approach that uses data augmentation to learn directly from pixels. We introduce several improvements that yield state-of-the-art results on the DeepMind Control Suite. Notably, DrQ-v2 is able to solve complex humanoid locomotion tasks directly from pixel observations, previously unattained by model-free RL. DrQ-v2 is conceptually simple, easy to implement, and provides significantly better computational footprint compared to prior work, with the majority of tasks taking just 8 hours to train on a single GPU. Finally, we publicly release DrQ-v2's implementation to provide RL practitioners with a strong and computationally efficient baseline.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Reinforcement Learning from Images
  • 2.2 Deep Deterministic Policy Gradient
  • 2.3 Data Augmentation in Reinforcement Learning
  • 3 DrQ-v2: Improved Data-Augmented Reinforcement Learning
  • 3.1 Algorithmic Details
  • 3.2 Implementation Details
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Comparison to Model-Free Methods
  • 4.3 Comparison to Model-Based Methods
  • 4.4 Ablation Study
  • 5 Related Work
  • 6 Conclusion
  • 7 Acknowledgements
  • References
  • A Benchmarks
  • B Hyper-parameters

Knowls

  1. Knowl 1 — DrQ-v2 Algorithm for Visual Continuous Control

    algorithm

    DrQ-v2 is an off-policy model-free actor-critic algorithm for image-based continuous control. It integrates deterministic actor updates (DDPG), clipped double Q-learning with nn-step returns, random shift data augmentation with bilinear interpolation, and a scheduled exploration noise mechanism.

    The policy πϕ\pi_\phi and Q-functions Qθ1,Qθ2Q_{\theta_1}, Q_{\theta_2} operate on low-dimensional latent embeddings produced by a convolutional encoder fξf_\xi. The encoder parameters ξ\xi are updated solely via the critic loss, using the most recent encoder weights for both current and future states without an encoder target network.

    Input: Parametric convolutional encoder fξf_\xi, deterministic actor policy πϕ\pi_\phi, critic functions Qθ1,Qθ2Q_{\theta_1}, Q_{\theta_2}, target critic networks Qθˉ1,Qθˉ2Q_{\bar{\theta}_1}, Q_{\bar{\theta}_2} with target weights θˉ1←θ1,θˉ2←θ2\bar{\theta}_1 \leftarrow \theta_1, \bar{\theta}_2 \leftarrow \theta_2, replay buffer D\mathcal{D}, random shift augmentation function aug\text{aug}, exploration noise standard deviation schedule σ(t)\sigma(t), exploration noise clip value cc, discount factor γ\gamma, nn-step horizon nn, learning rate α\alpha, target update soft rate τ\tau, batch size BB, total timesteps TT
    for t=1t = 1 to TT do
        Compute current exploration standard deviation σt←σ(t)\sigma_t \leftarrow \sigma(t)
        Sample noise ϵ∼N(0,σt2)\epsilon \sim \mathcal{N}(0, \sigma_t^2)
        Select action at←πϕ(fξ(xt))+ϵa_t \leftarrow \pi_\phi(f_\xi(x_t)) + \epsilon
        Execute action ata_t, observe reward rtr_t and next stacked frame observation xt+1x_{t+1}
        Store transition (xt,at,rt,xt+1)(x_t, a_t, r_t, x_{t+1}) in replay buffer D\mathcal{D}
        
        // Critic Update
        Sample mini-batch of BB transitions (xt,at,rt:t+n−1,xt+n)(x_t, a_t, r_{t:t+n-1}, x_{t+n}) from D\mathcal{D}
        ht←fξ(aug(xt))h_t \leftarrow f_\xi(\text{aug}(x_t))
        ht+n←fξ(aug(xt+n))h_{t+n} \leftarrow f_\xi(\text{aug}(x_{t+n}))
        Sample target noise ϵt+n∼clip(N(0,σt2),−c,c)\epsilon_{t+n} \sim \text{clip}(\mathcal{N}(0, \sigma_t^2), -c, c)
        at+n←πϕ(ht+n)+ϵt+na_{t+n} \leftarrow \pi_\phi(h_{t+n}) + \epsilon_{t+n}
        Compute nn-step target y←∑i=0n−1γirt+i+γnmin⁡k=1,2Qθˉk(ht+n,at+n)y \leftarrow \sum_{i=0}^{n-1} \gamma^i r_{t+i} + \gamma^n \min_{k=1,2} Q_{\bar{\theta}_k}(h_{t+n}, a_{t+n})
        Compute critic losses Lθk,ξ←1B∑(Qθk(ht,at)−y)2\mathcal{L}_{\theta_k, \xi} \leftarrow \frac{1}{B} \sum (Q_{\theta_k}(h_t, a_t) - y)^2 for k∈{1,2}k \in \{1, 2\}
        Update encoder weights: ξ←ξ−α∇ξ(Lθ1,ξ+Lθ2,ξ)\xi \leftarrow \xi - \alpha \nabla_\xi (\mathcal{L}_{\theta_1, \xi} + \mathcal{L}_{\theta_2, \xi})
        Update critic weights: θk←θk−α∇θkLθk,ξ\theta_k \leftarrow \theta_k - \alpha \nabla_{\theta_k} \mathcal{L}_{\theta_k, \xi} for k∈{1,2}k \in \{1, 2\}
        Update target critic weights: θˉk←(1−τ)θˉk+τθk\bar{\theta}_k \leftarrow (1 - \tau)\bar{\theta}_k + \tau \theta_k for k∈{1,2}k \in \{1, 2\}
        
        // Actor Update
        Sample mini-batch of BB observations (xt)(x_t) from D\mathcal{D}
        ht←fξ(aug(xt))h_t \leftarrow f_\xi(\text{aug}(x_t))
        Sample noise ϵt∼clip(N(0,σt2),−c,c)\epsilon_t \sim \text{clip}(\mathcal{N}(0, \sigma_t^2), -c, c)
        at←πϕ(ht)+ϵta_t \leftarrow \pi_\phi(h_t) + \epsilon_t
        Compute actor loss Lϕ←−1B∑min⁡k=1,2Qθk(ht,at)\mathcal{L}_\phi \leftarrow -\frac{1}{B} \sum \min_{k=1,2} Q_{\theta_k}(h_t, a_t)
        Update actor weights (without propagating gradients into encoder): ϕ←ϕ−α∇ϕLϕ\phi \leftarrow \phi - \alpha \nabla_\phi \mathcal{L}_\phi
    end for
  2. Knowl 2 — Clipped Double Q-Learning with n-Step Returns and DPG Objectives in DrQ-v2

    model/method

    DrQ-v2 utilizes Deep Deterministic Policy Gradient (DDPG) with nn-step temporal difference (TD) learning and clipped double Q-learning to minimize overestimation bias without requiring per-step policy entropy estimation.

    For a transition sequence τ=(xt,at,rt:t+n−1,xt+n)\tau = (x_t, a_t, r_{t:t+n-1}, x_{t+n}) sampled from the replay buffer D\mathcal{D}, the nn-step TD target yy is computed as:

    y=∑i=0n−1γirt+i+γnmin⁡k=1,2Qθˉk(ht+n,at+n)y = \sum_{i=0}^{n-1} \gamma^i r_{t+i} + \gamma^n \min_{k=1,2} Q_{\bar{\theta}_k}(h_{t+n}, a_{t+n})

    where ht=fξ(aug(xt))h_t = f_\xi(\text{aug}(x_t)) and ht+n=fξ(aug(xt+n))h_{t+n} = f_\xi(\text{aug}(x_{t+n})) are latent representations generated by the convolutional encoder fξf_\xi with the current parameter set ξ\xi, θˉ1,θˉ2\bar{\theta}_1, \bar{\theta}_2 are the exponential moving average weights of the target critics, and at+n=πϕ(ht+n)+ϵa_{t+n} = \pi_\phi(h_{t+n}) + \epsilon with regularized noise ϵ∼clip(N(0,σ2),−c,c)\epsilon \sim \text{clip}(\mathcal{N}(0, \sigma^2), -c, c).

    The two critic networks Qθ1Q_{\theta_1} and Qθ2Q_{\theta_2}, along with the encoder fξf_\xi, are trained by minimizing the mean squared Bellman error:

    Lθk,ξ(D)=Eτ∼D[(Qθk(ht,at)−y)2],k∈{1,2}\mathcal{L}_{\theta_k, \xi}(\mathcal{D}) = \mathbb{E}_{\tau \sim \mathcal{D}}\left[ \left( Q_{\theta_k}(h_t, a_t) - y \right)^2 \right], \quad k \in \{1, 2\}

    The encoder parameters ξ\xi are updated via gradient descent on Lθ1,ξ+Lθ2,ξ\mathcal{L}_{\theta_1, \xi} + \mathcal{L}_{\theta_2, \xi}. Target networks are not maintained for the encoder.

    The deterministic policy πϕ\pi_\phi is trained using the Deterministic Policy Gradient (DPG) objective:

    Lϕ(D)=−Ext∼D[min⁡k=1,2Qθk(ht,at)]\mathcal{L}_\phi(\mathcal{D}) = -\mathbb{E}_{x_t \sim \mathcal{D}}\left[ \min_{k=1,2} Q_{\theta_k}(h_t, a_t) \right]

    where ht=fξ(aug(xt))h_t = f_\xi(\text{aug}(x_t)), at=πϕ(ht)+ϵa_t = \pi_\phi(h_t) + \epsilon, and ϵ∼clip(N(0,σ2),−c,c)\epsilon \sim \text{clip}(\mathcal{N}(0, \sigma^2), -c, c). Gradients from the actor loss Lϕ\mathcal{L}_\phi do not propagate back into the encoder parameters ξ\xi.

  3. Knowl 3 — Random Shift Image Augmentation with Bilinear Interpolation

    model/method

    DrQ-v2 regularizes pixel-based policy and critic representations by applying random shifts followed by bilinear interpolation to stacked observation frames.

    Input states consist of 3 consecutive RGB images of resolution 84×8484 \times 84, stacked along the channel dimension (9×84×849 \times 84 \times 84). The augmentation operates as follows:

    1. Each side of the 84×8484 \times 84 frame is padded by 4 boundary-replicated pixels, expanding the spatial dimensions to 92×9292 \times 92.
    2. A random 84×8484 \times 84 crop is extracted from the padded image, simulating a spatial shift of up to ±4\pm 4 pixels along the horizontal and vertical axes.
    3. Bilinear interpolation is applied over the cropped image, replacing each pixel value with the distance-weighted average of its four nearest neighboring pixels.

    To remove computational bottlenecks and optimize GPU pipeline throughput, the augmentation is implemented using native PyTorch flow-field grid sampling (torch.nn.functional.grid_sample), avoiding CPU-to-GPU data transfers and accelerating processing speed by a factor of 2 compared to standard library implementations.

  4. Knowl 4 — Linearly Decaying Exploration Noise Schedule

    equation

    DrQ-v2 modulates action exploration by linearly decaying the standard deviation σ(t)\sigma(t) of Gaussian perturbation noise added to deterministic actions over a fixed training horizon TT:

    σ(t)=σinit+(1−min⁡(tT,1))(σfinal−σinit)\sigma(t) = \sigma_{\text{init}} + \left(1 - \min\left(\frac{t}{T}, 1\right)\right)(\sigma_{\text{final}} - \sigma_{\text{init}})

    where:

    • tt is the current environment step index.
    • TT is the decay horizon parameter (e.g., 100,000100{,}000 for easy tasks, 500,000500{,}000 for medium tasks, and 2,000,0002{,}000{,}000 for hard tasks).
    • σinit\sigma_{\text{init}} is the initial exploration standard deviation, set to 1.01.0.
    • σfinal\sigma_{\text{final}} is the asymptotic exploration standard deviation, set to 0.10.1.

    The exploration noise is added to the deterministic actor output as a=πϕ(h)+ϵa = \pi_\phi(h) + \epsilon, with ϵ∼clip(N(0,σ(t)2),−c,c)\epsilon \sim \text{clip}(\mathcal{N}(0, \sigma(t)^2), -c, c) where c=0.3c = 0.3. This encourages broad exploration during early phases and fine-grained deterministic exploitation during later stages.

  5. Knowl 5 — Visual Continuous Control Benchmark Performance and Humanoid Locomotion

    empirical result

    DrQ-v2 was evaluated across 24 continuous control tasks from the DeepMind Control Suite (DMC) using 3 stacked 84×8484 \times 84 pixel observations and an action repeat factor of 2:

    • Hard locomotion benchmark (Humanoid Stand, Humanoid Walk, Humanoid Run): DrQ-v2 is the first model-free RL algorithm to solve visual humanoid control (21 action degrees of freedom, 54 state dimensions). On Humanoid Stand and Humanoid Walk, DrQ-v2 reaches near-optimal performance (returns approaching 800) within 30 million environment frames, where previous model-free baselines (SAC, CURL, DrQ) fail completely (returns ≈0\approx 0).
    • Medium benchmark (12 complex tasks): DrQ-v2 consistently outperforms prior model-free baselines (CURL, DrQ, SAC-AE) in both sample efficiency and asymptotic return on tasks involving sparse rewards, underactuation, and contact dynamics (e.g., Acrobot Swingup, Quadruped Walk/Run, Finger Turn Hard).
    • Comparison to model-based Dreamer-v2: DrQ-v2 achieves sample efficiency comparable to the model-based method Dreamer-v2 across the majority of DMC benchmarks while maintaining a substantially lower computational footprint.
  6. Knowl 6 — Wall-Clock Training Efficiency and GPU Throughput Comparison

    empirical result

    Benchmarked on a single NVIDIA V100 GPU with a consistent mini-batch size of 256:

    • DrQ-v2 throughput: Achieves an execution speed of 96 frames per second (FPS).
    • Comparison to model-free baselines: DrQ-v2 is 3.4×3.4\times faster than DrQ (28 FPS) and 6.0×6.0\times faster than CURL (16 FPS).
    • Comparison to model-based methods: DrQ-v2 is 4.0×4.0\times faster than Dreamer-v2 (24 FPS) in wall-clock time due to avoiding expensive latent world-model optimization and imagination rollouts.

    Practical wall-clock training times required by DrQ-v2 on a single V100 GPU:

    • Easy tasks (1×1061\times 10^6 frames): ≈2.9\approx 2.9 hours.
    • Medium tasks (3×1063\times 10^6 frames): ≈8.6\approx 8.6 hours.
    • Hard humanoid tasks (30×10630\times 10^6 frames): ≈86\approx 86 hours (compared to 340 hours for Dreamer-v2).
  7. Knowl 7 — Ablation Analysis of Core Algorithmic Components in DrQ-v2

    empirical result

    Ablation experiments conducted on Acrobot Swingup, Quadruped Walk, and Reacher Hard isolate the performance impact of each component in DrQ-v2:

    1. Base Algorithm (DDPG vs. SAC): Replacing SAC with DDPG provides a substantial improvement on challenging exploration tasks. Automatic entropy temperature adjustment in SAC causes premature entropy collapse when learning from pixels, halting exploration, whereas deterministic policy gradients with Gaussian noise maintain steady exploration.
    2. Multi-step returns: Incorporating nn-step returns (n=3n=3 or n=5n=5) accelerates TD target propagation and yields clear sample efficiency gains over standard 1-step returns (n=1n=1). n=3n=3 is chosen as the default balance between variance and reward propagation speed.
    3. Replay buffer capacity: Increasing replay buffer capacity from 10510^5 to 10610^6 transitions mitigates catastrophic forgetting, producing significant score gains on tasks with diverse initial state distributions such as Reacher Hard.
    4. Exploration schedule: Linearly annealing exploration noise standard deviation from σ=1.0\sigma=1.0 down to 0.10.1 outperforms fixed noise (σ=0.2\sigma=0.2), enabling initial wide-area exploration followed by fine-grained policy refinement.
  8. Knowl 8 — DeepMind Control Suite Benchmark Categorization

    data/table

    The 24 visual continuous control tasks from the DeepMind Control Suite are categorized into three difficulty benchmarks based on the environment interaction budget required to achieve near-optimal returns:

    Task Traits Difficulty Allowed Steps dim⁡(S)\dim(\mathcal{S}) / dim⁡(A)\dim(\mathcal{A})
    Cartpole Balance balance, dense easy 1×1061 \times 10^6 4 / 1
    Cartpole Balance Sparse balance, sparse easy 1×1061 \times 10^6 4 / 1
    Cartpole Swingup swing, dense easy 1×1061 \times 10^6 4 / 1
    Cup Catch swing, catch, sparse easy 1×1061 \times 10^6 8 / 2
    Finger Spin rotate, dense easy 1×1061 \times 10^6 6 / 2
    Hopper Stand stand, dense easy 1×1061 \times 10^6 14 / 4
    Pendulum Swingup swing, sparse easy 1×1061 \times 10^6 2 / 1
    Walker Stand stand, dense easy 1×1061 \times 10^6 18 / 6
    Walker Walk walk, dense easy 1×1061 \times 10^6 18 / 6
    Acrobot Swingup diff. balance, dense medium 3×1063 \times 10^6 4 / 1
    Cartpole Swingup Sparse swing, sparse medium 3×1063 \times 10^6 4 / 1
    Cheetah Run run, dense medium 3×1063 \times 10^6 18 / 6
    Finger Turn Easy turn, sparse medium 3×1063 \times 10^6 6 / 2
    Finger Turn Hard turn, sparse medium 3×1063 \times 10^6 6 / 2
    Hopper Hop move, dense medium 3×1063 \times 10^6 14 / 4
    Quadruped Run run, dense medium 3×1063 \times 10^6 56 / 12
    Quadruped Walk walk, dense medium 3×1063 \times 10^6 56 / 12
    Reach Duplo manipulation, sparse medium 3×1063 \times 10^6 55 / 9
    Reacher Easy reach, dense medium 3×1063 \times 10^6 4 / 2
    Reacher Hard reach, dense medium 3×1063 \times 10^6 4 / 2
    Walker Run run, dense medium 3×1063 \times 10^6 18 / 6
    Humanoid Stand stand, dense hard 30×10630 \times 10^6 54 / 21
    Humanoid Walk walk, dense hard 30×10630 \times 10^6 54 / 21
    Humanoid Run run, dense hard 30×10630 \times 10^6 54 / 21

    Observations consist of 3 stacked 84×8484 \times 84 RGB images. Each episode lasts 1000 environment steps, with per-step rewards in [0,1][0, 1], giving a maximum possible episode return of 1000.

  9. Knowl 9 — Default Hyperparameters of DrQ-v2

    data/table

    The standard hyperparameter configuration used across continuous control environments for DrQ-v2 is summarized below:

    Parameter Setting
    Replay buffer capacity 10610^6
    Action repeat 2
    Seed frames 4000
    Exploration steps 2000
    nn-step returns 3
    Mini-batch size 256
    Discount factor γ\gamma 0.99
    Optimizer Adam
    Learning rate 10−410^{-4}
    Agent update frequency 2
    Critic Q-function soft-update rate τ\tau 0.01
    Features dimension 50
    Hidden dimension 1024
    Exploration stddev clip cc 0.3
    Exploration stddev schedule easy: linear(1.0, 0.1, 100000)
    medium: linear(1.0, 0.1, 500000)
    hard: linear(1.0, 0.1, 2000000)

    Specific task deviations:

    • Walker Stand / Walk / Run: mini-batch size =512= 512, nn-step returns =1= 1.
    • Quadruped Run: replay buffer capacity =105= 10^5.
    • Humanoid Stand / Walk: learning rate =8×10−5= 8 \times 10^{-5}, features dimension =100= 100.

Coverage note — None omitted; all core algorithmic mechanisms, mathematical formulations, experimental benchmarks, empirical results, ablations, and hyperparameter configurations are fully covered.

References

  1. 1.Pulkit Agrawal, Ashvin Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine. Learning to poke by poking: experiential learning of intuitive physics. In Proceedings of the 30th International Conference on Neural Information Processing Systems, pages 5092–5100, 2016.
  2. 2.Brandon Amos, Samuel Stanton, Denis Yarats, and Andrew Gordon Wilson. On the model-based stochastic value gradient for continuous reinforcement learning. CoRR, 2020.
  3. 3.Gabriel Barth-Maron, Matthew W. Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva TB, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap. Distributional policy gradients. In International Conference on Learning Representations, 2018.
  4. 4.Richard Bellman. A markovian decision process. Indiana Univ. Math. J., 1957.
  5. 5.Bryan Chen, Alexander Sax, Gene Lewis, Iro Armeni, Silvio Savarese, Amir Roshan Zamir, Jitendra Malik, and Lerrel Pinto. Robust policies via mid-level visual representations: An experimental study in manipulation and navigation. CoRR, 2020.
  6. 6.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In Proceedings of the IEEE International Conference on Computer Vision, pages 1422–1430, 2015.
  7. 7.Jing eng and Ronald J. Williams. Incremental multi-step q-learning. Machine Learning, 1996.
  8. 8.William Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio, Hugo Larochelle, Mark Rowland, and Will Dabney. Revisiting fundamentals of experience replay. CoRR, 2020.
  9. 9.Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel. Learning visual feature spaces for robotic manipulation with deep spatial autoencoders. CoRR, 2015.
  10. 10.Scott Fujimoto, Herke van Hoof, and David Meger. Addressing function approximation error in actor-critic methods. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmassan, Stockholm, Sweden, July 10-15, 2018, 2018.
  11. 11.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations, 2018.
  12. 12.Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta a nd Pieter Abbeel, and Sergey Levine. Soft actor-critic algorithms and applications. CoRR, 2018a.
  13. 13.Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905, 2018b.
  14. 14.Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. Learning latent dynamics for planning from pixels. arXiv preprint arXiv:1811.04551, 2018.
  15. 15.Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603, 2019.
  16. 16.Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world models. CoRR, 2020.
  17. 17.Nicklas Hansen and Xiaolong Wang. Generalization in reinforcement learning by soft data augmentation. In International Conference on Robotics and Automation, 2021.
  18. 18.Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Daniel Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver. Rainbow: Combining improvements in deep reinforcement learning. arXiv preprint arXiv:1710.02298, 2017.
  19. 19.Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas. Acme: A research framework for distributed reinforcement learning, 2020.
  20. 20.Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu. Reinforcement learning with unsupervised auxiliary tasks, 2016.
  21. 21.Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas. Reinforcement learning with augmented data, 2020.
  22. 22.A. X. Lee, A. Nagabandi, P. Abbeel, and S. Levine. Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model. arXiv e-prints, 2019.
  23. 23.Alex X Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine. Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model. arXiv preprint arXiv:1907.00953, 2019.
  24. 24.Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. CoRR, 2015a.
  25. 25.Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. CoRR, abs/1509.02971, 2015b.
  26. 26.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. Playing atari with deep reinforcement learning. arXiv e-prints, 2013.
  27. 27.Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. CoRR, 2016a.
  28. 28.Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning, 2016b.
  29. 29.Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare. Safe and efficient off-policy reinforcement learning. CoRR, 2016.
  30. 30.Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In European Conference on Computer Vision, pages 69–84. Springer, 2016.
  31. 31.Lerrel Pinto, Dhiraj Gandhi, Yuanfeng Han, Yong-Lae Park, and Abhinav Gupta. The curious robot: Learning visual representations via physical interactions. CoRR, 2016.
  32. 32.Roberta Raileanu, Max Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus. Automatic data augmentation for generalization in deep reinforcement learning. CoRR, 2020.
  33. 33.Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman. Data-efficient reinforcement learning with momentum predictive representations. arXiv preprint arXiv:2007.05929, 2020a.
  34. 34.Max Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm, Aaron C. Courville, and Philip Bachman. Data-efficient reinforcement learning with momentum predictive representations. CoRR, 2020b.
  35. 35.David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. Deterministic policy gradient algorithms. In Proceedings of the 31st International Conference on Machine Learning, 2014.
  36. 36.Aravind Srinivas, Michael Laskin, and Pieter Abbeel. Curl: Contrastive unsupervised representations for reinforcement learning. arXiv preprint arXiv:2004.04136, 2020.
  37. 37.Adam Stooke, Kimin Lee, Pieter Abbeel, and Michael Laskin. Decoupling representation learning from reinforcement learning. arXiv preprint arXiv, 2020.
  38. 38.Yuval Tassa, Tom Erez, and Emanuel Todorov. Synthesis and stabilization of complex behaviors through online trajectory optimization. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2012.
  39. 39.Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al. Deepmind control suite. arXiv preprint arXiv:1801.00690, 2018.
  40. 40.Emanuel Todorov, Tom Erez, and Yuval Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2012.
  41. 41.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, pages 1096–1103. ACM, 2008.
  42. 42.Xiaolong Wang and Abhinav Gupta. Unsupervised learning of visual representations using videos. In ICCV, 2015.
  43. 43.Christopher J. C. H. Watkins and Peter Dayan. Q-learning. Machine Learning, 1992.
  44. 44.Christopher John Cornish Hellaby Watkins. Learning from Delayed Rewards. PhD thesis, King’s College, 1989.
  45. 45.Wilson Yan, Ashwin Vangipuram, Pieter Abbeel, and Lerrel Pinto. Learning predictive representations for deformable objects using contrastive estimation. arXiv preprint arXiv:2003.05436, 2020.
  46. 46.Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus. Improving sample efficiency in model-free reinforcement learning from images. arXiv preprint arXiv:1910.01741, 2019.
  47. 47.Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto. Reinforcement learning with prototypical representations. CoRR, 2021a.
  48. 48.Denis Yarats, Ilya Kostrikov, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. In 9th International Conference on Learning Representations, ICLR 2021, 2021b.
  49. 49.Sarah Young, Dhiraj Gandhi, Shubham Tulsiani, Abhinav Gupta, Pieter Abbeel, and Lerrel Pinto. Visual imitation made easy. CoRR, 2020.
  50. 50.Albert Zhan, Philip Zhao, Lerrel Pinto, Pieter Abbeel, and Michael Laskin. A framework for efficient robotic manipulation. CoRR, 2020.
  51. 51.Richard Zhang, Phillip Isola, and Alexei A Efros. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1058–1067, 2017.

Citation

MLA
Yarats, D., et al. “Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning”. arXiv, 2021, http://arxiv.org/abs/2107.09645v1.
APA
Yarats, D., Fergus, R., Lazaric, A., & Pinto, L. (2021). Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning. arXiv. http://arxiv.org/abs/2107.09645v1
Chicago
Yarats, D., R. Fergus, A. Lazaric, and L. Pinto. 2021. “Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning”. arXiv. http://arxiv.org/abs/2107.09645v1.
Harvard
Yarats, D. et al. (2021) “Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2107.09645v1.
Vancouver
1. Yarats D, Fergus R, Lazaric A, Pinto L (2021) Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning. arXiv

BibTeX

@article{yarats2021mastering,
  title = {Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning},
  author = {Yarats, Denis and Fergus, Rob and Lazaric, Alessandro and Pinto, Lerrel},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2107.09645v1},
  eprint = {2107.09645}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors