Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
Denis YaratsRob FergusAlessandro LazaricLerrel Pinto
Introduces DrQ-v2, a fast and lightweight model-free reinforcement learning algorithm that achieves state-of-the-art visual continuous control on the DeepMind Control Suite, solving complex humanoid locomotion directly from pixels in only eight hours on a single GPU.
Training autonomous systems to perform continuous control directly from visual camera inputs has long posed a significant technical hurdle in reinforcement learning. Prior approaches have required prohibitive amounts of interaction data and immense computational resources involving multi-GPU clusters, while still failing to master highly complex physical behaviors such as humanoid locomotion.
The article introduces and evaluates DrQ-v2, an improved model-free reinforcement learning algorithm designed to achieve state-of-the-art sample and computational efficiency in vision-based continuous control. The primary objective is to demonstrate that a conceptually simple, data-augmented model-free framework can successfully solve highly intricate visual control benchmarks on modest single-GPU hardware.
To establish its findings, the authors conducted extensive empirical evaluations across 24 continuous control tasks from the DeepMind Control Suite, categorizing them into easy, medium, and hard tiers. The experiments assessed both sample efficiency (the number of environment interactions required) and computational throughput (wall-clock training time) on a single NVIDIA V100 graphics processing unit. The approach replaces earlier base algorithms with deep deterministic policy gradient, incorporates multi-step returns, refines image augmentation, applies an exploration noise decay schedule, and optimizes memory and processing pipelines.
The evaluation yielded several key findings. First, DrQ-v2 achieved state-of-the-art sample efficiency among model-free methods and became the first model-free method to successfully solve complex humanoid locomotion tasks directly from pixels. Second, it delivered a significant computational speedup, achieving 96 frames per second—a 3.4-fold improvement over its predecessor DrQ and a 6-fold increase over contrastive methods. Third, DrQ-v2 allowed most benchmark tasks to train in roughly 8.6 hours on a single processor, with easy tasks finishing in under 3 hours and the hardest humanoid tasks completing in about 86 hours. Fourth, when compared against DreamerV2, a leading model-based alternative, DrQ-v2 matched sample efficiency on many tasks while training approximately four times faster in wall-clock time due to lower computational overhead.
These results demonstrate that complex visual control does not necessarily require computationally heavy world models or massive distributed compute clusters. By drastically lowering wall-clock training times and hardware demands, the algorithm substantially reduces the infrastructure cost and experimentation risk associated with visual control research, making high-performance development accessible on standard hardware.
Based on these findings, the authors recommend that practitioners adopt DrQ-v2 as a strong, computationally efficient baseline for vision-based continuous control. They also advise the research community to retire saturated, easy benchmarks and focus evaluation efforts on medium and hard environments, such as quadruped and humanoid control. Further investigation is recommended to understand why model-based approaches still outperform model-free methods on specific sparse-reward or intricate tasks.
Confidence in these findings is supported by consistent results across 10 random seeds and 24 varied benchmark environments. However, readers should note that the evaluation is confined to simulated physics domains, and real-world robotic deployments may introduce additional physical uncertainties, visual variations, and boundary conditions not captured in the simulation suite.
- Paper: DeepMind Control Suite, Yuval Tassa et al. (2018). It introduces the standardized DeepMind Control Suite continuous control benchmark tasks and pixel-based baselines directly evaluated and built upon by DrQ-v2.
- Paper: CURL: Contrastive Unsupervised Representations for Reinforcement Learning, Aravind Srinivas et al. (2020). It establishes foundational data-efficient visual reinforcement learning on the DeepMind Control Suite, providing the key contrastive model-free baseline and image-augmentation context improved upon by DrQ-v2.
- Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). It introduces the Deep Deterministic Policy Gradient (DDPG) framework that DrQ-v2 adopts as its foundational model-free continuous control base algorithm.
- Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). It addresses function approximation error and overestimation in continuous actor-critic methods, introducing multi-critic techniques and target noise mechanisms that inform DrQ-v2's policy optimization.
- Paper: Mastering Atari with Discrete World Models, Danijar Hafner et al. (2021). It presents DreamerV2, the state-of-the-art model-based visual RL framework against which DrQ-v2 directly benchmarks its sample efficiency and computational speedup.
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). It introduces latent imagination for visual continuous control in the DeepMind Control Suite, establishing the model-based paradigm compared against DrQ-v2's model-free approach.
- Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). It provides the underlying theoretical foundation of deterministic policy gradient methods used in continuous-action visual reinforcement learning.
- Paper: VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning, Che Wang et al. (2022). It advances visual continuous control beyond pure online model-free RL by integrating visual pretraining and offline demonstrations on the DeepMind Control Suite.
- Paper: Mastering Diverse Domains through World Models, Danijar Hafner et al. (2023). It advances general visual control through a robust world-model framework (DreamerV3) across diverse domains, providing a complementary perspective to model-free visual continuous control.
- Paper: VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training, Yecheng Jason Ma et al. (2023). It extends visual reinforcement learning into self-supervised representation and universal reward learning from passive video data for downstream continuous robotic control.
