VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning

Che WangXufang LuoKeith W. RossDongsheng Li

article2022NeurIPS71 citations

Presents a three-stage visual deep reinforcement learning framework that systematically integrates ImageNet pretraining, offline expert demonstrations, and online fine-tuning to dramatically increase sample efficiency and reduce compute on complex robotic manipulation tasks.

Listen

Deploying visual reinforcement learning in realistic domains like robotic manipulation has long been hindered by poor sample efficiency, sparse reward signals, and high computational costs. While standard benchmarks rely heavily on direct physical state inputs or simplistic visuals, real-world tasks require processing complex camera streams. Existing solutions typically tackle visual pretraining, offline demonstrations, or online exploration in isolation, missing the opportunity to leverage diverse data sources within a single, unified pipeline.

The article introduces and evaluates VRL3, a streamlined, three-stage data-driven framework designed to solve complex visual control tasks efficiently. The core objective is to demonstrate that systematically integrating non-reinforcement learning datasets, small sets of offline demonstrations, and targeted online reinforcement learning dramatically improves learning efficiency, model size, and task success rates compared to existing state-of-the-art approaches.

The framework proceeds across three structured phases: First, it pretrains a lightweight convolutional visual encoder on standard image classification data (ImageNet) to capture general visual features, adapting it to multi-frame inputs via a convolutional channel expansion technique. Second, it initializes an actor-critic agent using offline demonstrations and applies conservative offline reinforcement learning updates to build task-specific representations across the entire network. Third, it fine-tunes the policy online using off-policy updates, stabilized by a Safe Q-target mechanism that caps value estimates to prevent divergence during the offline-to-online transition. The authors evaluate this approach primarily on four simulated dexterous hand manipulation tasks (the Adroit suite) and 24 DeepMind Control Suite tasks, comparing results against established baselines across multiple random seeds.

The evaluation yields several critical findings. On the visual Adroit benchmark, the framework improves average sample efficiency by 780% over the previous state-of-the-art method. On the most difficult hand manipulation task (Relocate), it achieves a 1,220% improvement in sample efficiency—rising to 2,440% when using a wider encoder—while solving the task using only 10% of the total computation. Additionally, the architecture is roughly 50 times smaller in its visual encoder and 3 times smaller across the entire agent compared to prior competitive models. Ablations show that conservative offline reinforcement learning in the second stage significantly outperforms behavioral cloning or contrastive learning by priming both the policy and value networks for subsequent online fine-tuning.

These findings indicate that complex visual control policies can be trained with substantially lower data-collection overhead, shorter training timelines, and minimal compute expenses. By demonstrating that compact encoders initialized on standard computer vision data can outperform much larger architectures, the article establishes that structured multi-stage data integration provides a practical pathway toward affordable real-world robotic learning.

For engineering and research teams deploying visual control, the article recommends adopting multi-stage training pipelines that combine general image pretraining with conservative offline updates before online exploration. Practitioners should prioritize tuning the encoder learning rate scale and utilizing data augmentation, while maintaining stable value targets during online transitions. Future work should focus on testing this pipeline on physical robotic hardware to confirm performance outside simulated environments and investigating performance bounds when task visual styles diverge sharply from standard pretraining datasets.

arXiv: 2202.10324
Cover for VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning

Abstract

We propose VRL3, a powerful data-driven framework with a simple design for solving challenging visual deep reinforcement learning (DRL) tasks. We analyze a number of major obstacles in taking a data-driven approach, and present a suite of design principles, novel findings, and critical insights about data-driven visual DRL. Our framework has three stages: in stage 1, we leverage non-RL datasets (e.g. ImageNet) to learn task-agnostic visual representations; in stage 2, we use offline RL data (e.g. a limited number of expert demonstrations) to convert the task-agnostic representations into more powerful task-specific representations; in stage 3, we fine-tune the agent with online RL. On a set of challenging hand manipulation tasks with sparse reward and realistic visual inputs, compared to the previous SOTA, VRL3 achieves an average of 780% better sample efficiency. And on the hardest task, VRL3 is 1220% more sample efficient (2440% when using a wider encoder) and solves the task with only 10% of the computation. These significant results clearly demonstrate the great potential of data-driven deep reinforcement learning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Learning visual tasks with online RL
  • 2.2 RL with ImageNet pretraining
  • 2.3 Offline RL and data-driven RL
  • 3 VRL3: Visual DRL in 3 stages
  • 4 Results
  • 5 Challenges, Design Decisions, and Insights
  • 5.1 Stage 1: Pretraining with non-RL data
  • 5.2 Stage 2 and Stage 3
  • 5.3 Contribution of each stage
  • 5.4 Additional Studies and Analysis
  • 6 Discussion, Limitations and Future Work
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Three-Stage Architecture of the VRL3 Framework

    model/method

    Visual Deep Reinforcement Learning in 3 Stages (VRL3) is a framework designed to solve complex visual continuous control tasks with sparse rewards by sequentially utilizing non-reinforcement learning (non-RL) data, offline RL data, and online RL interactions.

    The framework consists of three sequential training stages:

    1. Stage 1 (Task-Agnostic Visual Pretraining): A convolutional encoder fξf_\xi is pretrained via supervised learning on a general-domain image classification dataset (such as 1000-class ImageNet), resizing training images to 84×8484 \times 84 pixels to match the downstream reinforcement learning observation resolution.
    2. Stage 2 (Offline RL Task-Specific Adaptation): The actor-critic agent is initialized with the pretrained encoder fξf_\xi, two critic Q-networks Qθ1,Qθ2Q_{\theta_1}, Q_{\theta_2}, target critics Qθˉ1,Qθˉ2Q_{\bar{\theta}_1}, Q_{\bar{\theta}_2}, and a policy network πϕ\pi_\phi. Offline task data (e.g., expert demonstration trajectories) is loaded into a replay buffer D\mathcal{D}. The agent optimizes the actor, critic, and fine-tunes the encoder using an off-policy actor-critic objective augmented with a conservative Q-learning penalty and a Safe Q target technique.
    3. Stage 3 (Online Off-Policy Fine-Tuning): The replay buffer D\mathcal{D} retaining the offline transitions is augmented with newly collected online rollouts. The agent fine-tunes all network parameters online using standard off-policy RL updates with image augmentation and Safe Q target regulation, omitting the conservative Q penalty.

    For an observation comprising visual frames xx and optional proprioceptive sensor vectors zz, the visual input is augmented with random shift augmentations aug(x)\mathrm{aug}(x), processed into visual features h=fξ(aug(x))h = f_\xi(\mathrm{aug}(x)), concatenated into state representation s=cat(h,z)s = \mathrm{cat}(h, z), and supplied to the policy and Q-value networks.

  2. Knowl 2 — Convolutional Channel Expansion (CCE)

    model/method

    Convolutional Channel Expansion (CCE) is an architectural adaptation method that converts a 2D convolutional encoder pretrained on single 3-channel RGB images into an encoder capable of processing mm consecutive stacked video frames or multi-camera inputs (3m3m total channels) without disrupting learned spatial representations.

    Let the first convolutional layer of the pretrained encoder possess a weight tensor W∈RCout×3×Kh×KwW \in \mathbb{R}^{C_{out} \times 3 \times K_h \times K_w}, where CoutC_{out} is the number of output channels and Kh,KwK_h, K_w are the spatial kernel dimensions. CCE replicates the weight tensor mm times along the input channel dimension and scales the resulting values by 1/m1/m:

    WCCE=1m[W  ∥  W  ∥  …  ∥  W]⏟m copies∈RCout×3m×Kh×KwW_{\mathrm{CCE}} = \frac{1}{m} \underbrace{\Big[ W \;\Vert\; W \;\Vert\; \dots \;\Vert\; W \Big]}_{m \text{ copies}} \in \mathbb{R}^{C_{out} \times 3m \times K_h \times K_w}

    where ∥\Vert denotes concatenation along the input channel axis. If the input consists of identical repeated frames, this transformation guarantees that the layer output is mathematically identical to that of the single-image pretrained layer, enabling seamless initialization for temporal observation processing with zero added computational overhead.

  3. Knowl 3 — Safe Q Target Technique for Offline-to-Online Reinforcement Learning

    model/method

    The Safe Q target technique stabilizes critic updates and prevents value overestimation caused by distribution shift during offline RL training and offline-to-online transitions.

    Given discount factor γ∈(0,1)\gamma \in (0, 1) and maximum per-step environmental reward normalized such that rmax⁡=1r_{\max} = 1, the maximum theoretical return achievable by an optimal policy is defined by the infinite geometric series sum:

    Qmax⁡=∑t=0∞γtrmax⁡=rmax⁡1−γQ_{\max} = \sum_{t=0}^\infty \gamma^t r_{\max} = \frac{r_{\max}}{1 - \gamma}

    For γ=0.99\gamma = 0.99, Qmax⁡=100Q_{\max} = 100.

    During temporal-difference updates with nn-step returns, target values are computed as:

    y=∑i=0n−1γirt+i+γnmin⁡k=1,2Qθˉk(st+n,a~t+n)y = \sum_{i=0}^{n-1} \gamma^i r_{t+i} + \gamma^n \min_{k=1, 2} Q_{\bar{\theta}_k}(s_{t+n}, \tilde{a}_{t+n})

    where a~t+n∼πϕ(st+n)\tilde{a}_{t+n} \sim \pi_\phi(s_{t+n}) and QθˉkQ_{\bar{\theta}_k} are target critic networks. Whenever a target value yy exceeds the threshold Qmax⁡+1Q_{\max} + 1, it is compressed via a soft-saturation update:

    y←Qmax⁡+(y−Qmax⁡)ηy \leftarrow Q_{\max} + (y - Q_{\max})^\eta

    where η∈[0,1]\eta \in [0, 1] is the Safe Q factor that attenuates anomalous value spikes.

  4. Knowl 4 — Optimization Objectives and Parameter Updates in Stage 2 Offline RL

    model/method

    In Stage 2 of VRL3, the agent samples mini-batches of transitions (xt,zt,at,rt:t+n−1,xt+n,zt+n)(x_t, z_t, a_t, r_{t:t+n-1}, x_{t+n}, z_{t+n}) from replay buffer D\mathcal{D}. Let st=cat(fξ(aug(xt)),zt)s_t = \mathrm{cat}(f_\xi(\mathrm{aug}(x_t)), z_t) and a~t∼πϕ(st)\tilde{a}_t \sim \pi_\phi(s_t).

    The Q-networks Qθ1,Qθ2Q_{\theta_1}, Q_{\theta_2} and encoder fξf_\xi are updated via gradient descent on the joint loss LQ(D)+LQC(D)\mathcal{L}_Q(\mathcal{D}) + \mathcal{L}_{QC}(\mathcal{D}), where the standard Q-loss for k∈{1,2}k \in \{1, 2\} is:

    LQ(D)=Eτ∼D[(Qθk(st,at)−y)2]\mathcal{L}_Q(\mathcal{D}) = \mathbb{E}_{\tau \sim \mathcal{D}}\left[ \left(Q_{\theta_k}(s_t, a_t) - y\right)^2 \right]

    and the conservative Q-loss penalty is:

    LQC(D)=Eτ∼D[log⁡∑a~texp⁡(Qθk(st,a~t))−Qθk(st,at)]\mathcal{L}_{QC}(\mathcal{D}) = \mathbb{E}_{\tau \sim \mathcal{D}}\left[ \log \sum_{\tilde{a}_t} \exp\left(Q_{\theta_k}(s_t, \tilde{a}_t)\right) - Q_{\theta_k}(s_t, a_t) \right]

    The policy network πϕ\pi_\phi is updated by gradient descent on:

    Lπ(D)=−Eτ∼D[min⁡k=1,2Qθk(st,a~t)]\mathcal{L}_\pi(\mathcal{D}) = -\mathbb{E}_{\tau \sim \mathcal{D}}\left[ \min_{k=1, 2} Q_{\theta_k}(s_t, \tilde{a}_t) \right]

    Policy gradients do not propagate into encoder fξf_\xi. The encoder learning rate is scaled by βenc∈(0,1]\beta_{\mathrm{enc}} \in (0, 1] relative to the base actor-critic learning rate α\alpha:

    αenc=βencα\alpha_{\mathrm{enc}} = \beta_{\mathrm{enc}} \alpha

    Target networks QθˉkQ_{\bar{\theta}_k} are updated via Polyak averaging with parameter τ\tau.

  5. Knowl 5 — Reduced Learning Rate for Visual Encoder Fine-Tuning in Reinforcement Learning

    model/method

    When transferring a visual encoder pretrained on non-RL datasets (e.g., ImageNet) to reinforcement learning tasks, noisy policy and value gradients can destroy visual representations in early training stages.

    VRL3 decouples the optimization speed by setting the encoder learning rate αenc\alpha_{\mathrm{enc}} to a fraction of the learning rate α\alpha used for policy and Q-value MLPs:

    αenc=βencα\alpha_{\mathrm{enc}} = \beta_{\mathrm{enc}} \alpha

    where βenc∈[0.001,0.1]\beta_{\mathrm{enc}} \in [0.001, 0.1]. Setting βenc=0\beta_{\mathrm{enc}} = 0 (completely freezing the encoder) retains useful visual features despite domain gaps but limits performance. Setting βenc=1.0\beta_{\mathrm{enc}} = 1.0 (equal learning rate) leads to catastrophic forgetting of visual features due to gradient instability. An intermediate scale (e.g., βenc≈0.01\beta_{\mathrm{enc}} \approx 0.01) protects pretrained spatial representations while allowing domain adaptation.

  6. Knowl 6 — Empirical Performance on the Visual Adroit Dexterous Manipulation Benchmark

    empirical result

    VRL3 was evaluated across four sparse-reward visual dexterous manipulation tasks in the Adroit suite (Door, Hammer, Pen, Relocate) initialized with 25 human expert demonstration trajectories per task.

    Key empirical findings include:

    • Sample Efficiency: Across all 4 Adroit tasks, VRL3 achieves an average of 780% higher sample efficiency compared to the prior state-of-the-art on visual Adroit, RRL (ResNet as Representation for Reinforcement Learning).
    • Relocate Task Performance: On Relocate (the most challenging task requiring grasping and placing a ball to a target position), VRL3 requires 1220% less online environment data than RRL to reach a 90% success rate, which increases to 2440% higher sample efficiency when utilizing a wider convolutional encoder.
    • Computational and Parameter Efficiency: VRL3 utilizes a lightweight 5-layer convolutional encoder that has 50 times fewer parameters than the ResNet-34 backbone used in RRL and achieves a 3-fold parameter reduction across the full agent. Furthermore, VRL3 solves Relocate with 10% of the total computation required by RRL.
    • Baseline Comparisons: Standard online RL without offline data (DrQv2) and demonstration-buffered variants without Stage 2 offline updates (DrQv2fD) fail completely on the Relocate task.
  7. Knowl 7 — Suboptimality of Behavioral Cloning and Contrastive Representation Learning in Offline Pre-training

    empirical result

    Ablation experiments comparing Stage 2 offline training strategies demonstrate that conservative RL updates outperform Behavioral Cloning (BC) and contrastive unsupervised representation learning (FERM) when leveraging limited expert demonstrations for visual RL:

    1. Behavioral Cloning Failure: Replacing conservative RL in Stage 2 with BC on expert demonstrations yields performance equivalent to completely omitting Stage 2. BC optimizes only policy parameters πϕ\pi_\phi without training critic networks QθQ_\theta, resulting in an actor-critic distribution mismatch that destabilizes early Stage 3 online training.
    2. Contrastive Learning Limitations: Utilizing contrastive representation learning in Stage 2 (as in FERM) updates only visual representations fξf_\xi and leaves the actor and critic uninitialized. In complex sparse-reward environments (Adroit), contrastive pretraining fails to solve the exploration bottleneck and performs substantially worse than conservative RL pre-training.
    3. Joint Agent Training: Training the encoder, critic networks, and actor simultaneously via conservative RL in Stage 2 provides both task-specific representations and calibrated value/policy initializations for Stage 3.
  8. Knowl 8 — Ablation Analysis of VRL3 Training Stages and Encoder Learning

    empirical result

    Empirical ablations isolating individual training stages and encoder update schedules on the Adroit manipulation suite reveal the following behaviors:

    • Stage Inclusion Ablations: Complete execution (S123S123) outperforms removing Stage 1 pretraining (S23S23). Removing Stage 2 offline RL training (S13S13) causes severe performance degradation, confirming that offline demonstration exploitation cannot be compensated for by online RL (S3S3) alone in sparse-reward visual domains.
    • Encoder Update Ablations: Restricting encoder updates exclusively to Stage 3 while freezing it in Stage 2 (S3S3-encoder) achieves substantially higher performance than freezing it across all RL stages (S1S1-encoder or S2S2-encoder), demonstrating that task-specific visual fine-tuning via RL gradients is critical for optimal performance.
    • Non-RL vs. RL Data Regimes: When downstream RL transition data is scarce, Stage 1 ImageNet pretraining provides a superior visual representation despite the large cross-domain gap (S1>S2S1 > S2 encoder performance). When downstream online RL data is abundant, task-specific RL representations outperform frozen out-of-domain representations (S3>S1S3 > S1).
  9. Knowl 9 — Hyperparameter Robustness and Sensitivity Landscape of VRL3

    empirical result

    An evaluation of 16 hyperparameters and architectural components across the four Adroit tasks identified 9 parameters critical to performance, categorized into robust and sensitive groups:

    • Critical and Robust Components (6 parameters):
      1. Stage 1 Pretraining: Essential for difficult manipulation tasks (Pen, Hammer, Relocate); negligible impact on the simple Door task.
      2. Data Augmentation: Random shift augmentation during Stages 2 and 3 is universally critical across all tasks.
      3. Safe Q Target Application: Prevents value blowup without requiring task-specific threshold tuning.
      4. Stage 2 Offline Update Steps: Even small budgets of offline updates (e.g., 5,000–30,000 updates) provide the majority of performance gains.
      5. Batch Size
      6. Target Network Polyak Factor (τ\tau)
    • Critical and Sensitive Hyperparameters (3 parameters):
      1. Encoder Learning Rate Scale (βenc\beta_{\mathrm{enc}}): The only sensitive hyperparameter unique to VRL3, requiring tuning between 0.0010.001 and 0.10.1.
      2. Base Learning Rate (α\alpha): Inherited from the backbone RL optimizer.
      3. Exploration Action Noise Standard Deviation: Inherited from the backbone RL optimizer.
  10. Knowl 10 — Limitations of the VRL3 Framework

    limitation

    The VRL3 framework has several documented limitations:

    1. Evaluation Restricted to Simulation: Experimental validation is conducted entirely within simulated environments (Adroit dexterous hand suite and DeepMind Control Suite). Real-world deployment on physical robotic platforms with physical sensors, latencies, and real camera feeds remains unverified.
    2. Sensitivity to Extreme Domain Gaps: The efficacy of Stage 1 ImageNet pretraining when task visuals drastically diverge from natural image statistics (e.g., extreme non-photorealistic or non-natural sensor modalities) has not been established.
    3. Model-Free Scope: The framework is restricted to model-free off-policy reinforcement learning and does not integrate model-based dynamics pretraining or world-model trajectory rollouts.

Coverage note — No substantial contributed material was omitted. DeepMind Control Suite (DMC) benchmark details from Section 4 are covered comparatively, and all core algorithms, ablations, formulas, and empirical findings are fully captured.

References

  1. 1.Arthur Argenson and Gabriel Dulac-Arnold. Model-based offline planning. arXiv preprint arXiv:2008.05556, 2020.
  2. 2.Chenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhihong Deng, Animesh Garg, Peng Liu, and Zhaoran Wang. Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning. arXiv preprint arXiv:2202.11566, 2022.
  3. 3.Johan Bjorck, Carla P Gomes, and Kilian Q Weinberger. Towards deeper deep reinforcement learning. arXiv preprint arXiv:2106.01151, 2021.
  4. 4.Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym. arXiv preprint arXiv:1606.01540, 2016.
  5. 5.Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee. Sample-efficient reinforcement learning with stochastic ensemble value expansion. In Advances in Neural Information Processing Systems, pages 8224–8234, 2018.
  6. 6.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34, 2021.
  7. 7.Xinyue Chen, Che Wang, Zijian Zhou, and Keith Ross. Randomized ensembled double q-learning: Learning fast without a model. International Conference on Learning Representations, 2021.
  8. 8.Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, and Keith Ross. Bail: Best-action imitation learning for batch deep reinforcement learning. In Advances in Neural Information Processing Systems, volume 33, pages 18353–18363, 2020.
  9. 9.Hyesong Choi, Hunsang Lee, Wonil Song, Sangryul Jeon, Kwanghoon Sohn, and Dongbo Min. Self-supervised structured representations for deep reinforcement learning. 2021.
  10. 10.Ignasi Clavera, Violet Fu, and Pieter Abbeel. Model-augmented actor-critic: Backpropagating through paths. arXiv preprint arXiv:2005.08068, 2020.
  11. 11.Karl W Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman. Phasic policy gradient. In International Conference on Machine Learning, pages 2020–2027. PMLR, 2021.
  12. 12.Yuchen Cui, Scott Niekum, Abhinav Gupta, Vikash Kumar, and Aravind Rajeswaran. Can foundation models perform zero-shot task specification for robot manipulation? In Learning for Dynamics and Control Conference, pages 893–905. PMLR, 2022.
  13. 13.Kefan Dong, Yuping Luo, Tianhe Yu, Chelsea Finn, and Tengyu Ma. On the expressivity of neural networks for deep reinforcement learning. International Conference on Machine Learning (ICML), 2020.
  14. 14.Danny Driess, Ingmar Schubert, Pete Florence, Yunzhu Li, and Marc Toussaint. Reinforcement learning with neural radiance fields. arXiv preprint arXiv:2206.01634, 2022.
  15. 15.Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel. Benchmarking deep reinforcement learning for continuous control. In International Conference on Machine Learning, pages 1329–1338, 2016.
  16. 16.Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel. Deep spatial autoencoders for visuomotor learning. In 2016 IEEE International Conference on Robotics and Automation (ICRA), pages 512–519. IEEE, 2016.
  17. 17.Scott Fujimoto and Shixiang Shane Gu. A minimalist approach to offline reinforcement learning. arXiv preprint arXiv:2106.06860, 2021.
  18. 18.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pages 2052–2062. PMLR, 2019.
  19. 19.Scott Fujimoto, Herke van Hoof, and Dave Meger. Addressing function approximation error in actor-critic methods. arXiv preprint arXiv:1802.09477, 2018.
  20. 20.Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu. Emaq: Expected-max q-learning operator for simple yet effective offline and online rl. In International Conference on Machine Learning, pages 3682–3691. PMLR, 2021.
  21. 21.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33:21271–21284, 2020.
  22. 22.Abhinav Gupta, Adithyavairavan Murali, Dhiraj Prakashchand Gandhi, and Lerrel Pinto. Robot learning in homes: Improving generalization and reducing dataset bias. Advances in neural information processing systems, 31, 2018.
  23. 23.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv preprint arXiv:1801.01290, 2018.
  24. 24.Siddhant Haldar, Vaibhav Mathur, Denis Yarats, and Lerrel Pinto. Watch and match: Supercharging imitation with regularized optimal transport. arXiv preprint arXiv:2206.15469, 2022.
  25. 25.Dongqi Han, Kenji Doya, and Jun Tani. Variational recurrent models for solving partially observable control tasks. arXiv preprint arXiv:1912.10703, 2019.
  26. 26.Matthew Hausknecht and Peter Stone. Deep reinforcement learning in parameterized action space. arXiv preprint arXiv:1511.04143, 2015.
  27. 27.Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  28. 28.Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup. Reproducibility of benchmarked deep reinforcement learning tasks for continuous control. arXiv preprint arXiv:1708.04133, 2017.
  29. 29.Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to trust your model: Model-based policy optimization. In Advances in Neural Information Processing Systems, pages 12519–12530, 2019.
  30. 30.Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. Advances in neural information processing systems, 34, 2021.
  31. 31.Ryan Julian, Benjamin Swanson, Gaurav S Sukhatme, Sergey Levine, Chelsea Finn, and Karol Hausman. Never stop learning: The effectiveness of fine-tuning in robotic reinforcement learning. arXiv preprint arXiv:2004.10190, 2020.
  32. 32.Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims. Morel: Model-based offline reinforcement learning. arXiv preprint arXiv:2005.05951, 2020.
  33. 33.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  34. 34.Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. arXiv preprint arXiv:2110.06169, 2021.
  35. 35.Ilya Kostrikov, Denis Yarats, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. arXiv preprint arXiv:2004.13649, 2020.
  36. 36.Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine. Stabilizing off-policy q-learning via bootstrapping error reduction. arXiv preprint arXiv:1906.00949, 2019.
  37. 37.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. arXiv preprint arXiv:2006.04779, 2020.
  38. 38.Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel. Model-ensemble trust-region policy optimization. arXiv preprint arXiv:1802.10592, 2018.
  39. 39.Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov. Controlling overestimation bias with truncated mixture of continuous distributional quantile critics. arXiv preprint arXiv:2005.04269, 2020.
  40. 40.Hang Lai, Jian Shen, Weinan Zhang, and Yong Yu. Bidirectional model-based policy optimization. arXiv preprint arXiv:2007.01995, 2020.
  41. 41.Qingfeng Lan, Yangchen Pan, Alona Fyshe, and Martha White. Maxmin q-learning: Controlling the estimation bias of q-learning. arXiv preprint arXiv:2002.06487, 2020.
  42. 42.Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba. Benchmarking model-based reinforcement learning. arXiv preprint arXiv:1907.02057, 2019.
  43. 43.Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas. Reinforcement learning with augmented data. arXiv preprint arXiv:2004.14990, 2020.
  44. 44.Michael Laskin, Aravind Srinivas, and Pieter Abbeel. Curl: Contrastive unsupervised representations for reinforcement learning. In International Conference on Machine Learning, pages 5639–5650. PMLR, 2020.
  45. 45.Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel. Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning. arXiv preprint arXiv:2007.04938, 2020.
  46. 46.Seunghyun Lee, Younggyo Seo, Kimin Lee, Pieter Abbeel, and Jinwoo Shin. Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble. In Conference on Robot Learning, pages 1702–1712. PMLR, 2022.
  47. 47.Sergey Levine. Understanding the world through action. In Conference on Robot Learning, pages 1752–1757. PMLR, 2022.
  48. 48.Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies. The Journal of Machine Learning Research, 17(1):1334–1373, 2016.
  49. 49.Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  50. 50.Xiang Li, Jinghuan Shang, Srijan Das, and Michael S Ryoo. Does self-supervised learning really improve reinforcement learning from pixels? arXiv preprint arXiv:2206.05266, 2022.
  51. 51.Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
  52. 52.Nan Lin, Yuxuan Li, Keke Tang, Yujun Zhu, Xiayu Zhang, Ruolin Wang, Jianmin Ji, Xiaoping Chen, and Xinming Zhang. Manipulation planning from demonstration via goal-conditioned prior action primitive decomposition and alignment. IEEE Robotics and Automation Letters, 7(2):1387–1394, 2022.
  53. 53.Clare Lyle, Mark Rowland, and Will Dabney. Understanding and preventing capacity loss in reinforcement learning. arXiv preprint arXiv:2204.09560, 2022.
  54. 54.Lingheng Meng, Rob Gorbet, and Dana Kulić. The effect of multi-step methods on overestimation in deep reinforcement learning. arXiv preprint arXiv:2006.12692, 2020.
  55. 55.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
  56. 56.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529, 2015.
  57. 57.Yao Mu, Shoufa Chen, Mingyu Ding, Jianyu Chen, Runjian Chen, and Ping Luo. Ctrlformer: Learning transferable state representation for visual control via transformer. arXiv preprint arXiv:2206.08883, 2022.
  58. 58.Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine. Awac: Accelerating online reinforcement learning with offline datasets. 2020.
  59. 59.Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. R3m: A universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601, 2022.
  60. 60.Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville. The primacy bias in deep reinforcement learning. In International Conference on Machine Learning, pages 16828–16847. PMLR, 2022.
  61. 61.Kei Ota, Devesh K Jha, and Asako Kanezaki. Training larger networks for deep reinforcement learning. arXiv preprint arXiv:2102.07920, 2021.
  62. 62.Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, and Abhinav Gupta. The unsurprising effectiveness of pre-trained vision models for control. arXiv preprint arXiv:2203.03580, 2022.
  63. 63.Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019.
  64. 64.Lerrel Pinto and Abhinav Gupta. Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours. In 2016 IEEE international conference on robotics and automation (ICRA), pages 3406–3413. IEEE, 2016.
  65. 65.Senthil Purushwalkam Sh. Visual Representation and Recognition without Human Supervision. PhD thesis, Carnegie Mellon University, 2022.
  66. 66.Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine. Learning complex dexterous manipulation with deep reinforcement learning and demonstrations. arXiv preprint arXiv:1709.10087, 2017.
  67. 67.Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar. A game theoretic framework for model based reinforcement learning. arXiv preprint arXiv:2004.07804, 2020.
  68. 68.Machel Reid, Yutaro Yamada, and Shixiang Shane Gu. Can wikipedia help offline reinforcement learning? arXiv preprint arXiv:2201.12122, 2022.
  69. 69.Max Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand, Laurent Charlin, Devon Hjelm, Philip Bachman, and Aaron Courville. Pretraining representations for data-efficient reinforcement learning. arXiv preprint arXiv:2106.04799, 2021.
  70. 70.Younggyo Seo, Danijar Hafner, Hao Liu, Fangchen Liu, Stephen James, Kimin Lee, and Pieter Abbeel. Masked world models for visual control. arXiv preprint arXiv:2206.14244, 2022.
  71. 71.Younggyo Seo, Kimin Lee, Stephen L James, and Pieter Abbeel. Reinforcement learning with action-free pre-training from videos. In International Conference on Machine Learning, pages 19561–19579. PMLR, 2022.
  72. 72.Pierre Sermanet, Kelvin Xu, and Sergey Levine. Unsupervised perceptual rewards for imitation learning. arXiv preprint arXiv:1612.06699, 2016.
  73. 73.Rutav Shah and Vikash Kumar. Rrl: Resnet as representation for reinforcement learning. In Self-Supervision for Reinforcement Learning Workshop-ICLR 2021, 2021.
  74. 74.Wenling Shang, Xiaofei Wang, Aravind Srinivas, Aravind Rajeswaran, Yang Gao, Pieter Abbeel, and Misha Laskin. Reinforcement learning with latent flow. Advances in Neural Information Processing Systems, 34:22171–22183, 2021.
  75. 75.Aravind Sivakumar, Kenneth Shaw, and Deepak Pathak. Robotic telekinesis: learning a robotic hand imitator by watching humans on youtube. arXiv preprint arXiv:2202.10448, 2022.
  76. 76.Maks Sorokin, Jie Tan, C Karen Liu, and Sehoon Ha. Learning to navigate sidewalks in outdoor environments. IEEE Robotics and Automation Letters, 7(2):3906–3913, 2022.
  77. 77.Adam Stooke, Kimin Lee, Pieter Abbeel, and Michael Laskin. Decoupling representation learning from reinforcement learning. arXiv preprint arXiv:2009.08319, 2020.
  78. 78.Richard S Sutton. The bitter lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html, 2021. Accessed: 2022-01-24.
  79. 79.Allison C Tam, Neil C Rabinowitz, Andrew K Lampinen, Nicholas A Roy, Stephanie CY Chan, DJ Strouse, Jane X Wang, Andrea Banino, and Felix Hill. Semantic exploration from language abstractions and pretrained representations. arXiv preprint arXiv:2204.05080, 2022.
  80. 80.Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al. Deepmind control suite. arXiv preprint arXiv:1801.00690, 2018.
  81. 81.Sebastian Thrun and Anton Schwartz. Issues in using function approximation for reinforcement learning. In Proceedings of the 1993 Connectionist Models Summer School Hillsdale, NJ. Lawrence Erlbaum, 1993.
  82. 82.Emanuel Todorov, Tom Erez, and Yuval Tassa. Mujoco: A physics engine for model-based control. In Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ International Conference on, pages 5026–5033. IEEE, 2012.
  83. 83.Ikechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu, Mengyuan Yan, Joséphine Simon, Matthew Bennice, Chuyuan Fu, Cong Ma, Jiantao Jiao, et al. Jump-start reinforcement learning. arXiv preprint arXiv:2204.02372, 2022.
  84. 84.Hado Van Hasselt, Arthur Guez, and David Silver. Deep reinforcement learning with double q-learning. In AAAI, volume 2, page 5. Phoenix, AZ, 2016.
  85. 85.Homer Walke, Jonathan Yang, Albert Yu, Aviral Kumar, Jedrzej Orbik, Avi Singh, and Sergey Levine. Don’t start from scratch: Leveraging prior data to automate robotic reinforcement learning. arXiv preprint arXiv:2207.04703, 2022.
  86. 86.Che Wang, Yanqiu Wu, Quan Vuong, and Keith Ross. Striving for simplicity and performance in off-policy drl: Output normalization and non-uniform sampling. In International Conference on Machine Learning, pages 10070–10080. PMLR, 2020.
  87. 87.Qing Wang, Jiechao Xiong, Lei Han, Han Liu, Tong Zhang, et al. Exponentially weighted imitation learning for batched historical data. In Advances in Neural Information Processing Systems, pages 6288–6297, 2018.
  88. 88.Yifan Wu, George Tucker, and Ofir Nachum. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  89. 89.Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh. Uncertainty weighted actor-critic for offline reinforcement learning. arXiv preprint arXiv:2105.08140, 2021.
  90. 90.Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik. Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173, 2022.
  91. 91.Mengjiao Yang and Ofir Nachum. Representation matters: Offline pretraining for sequential decision making. arXiv preprint arXiv:2102.05815, 2021.
  92. 92.Denis Yarats, David Brandfonbrener, Hao Liu, Michael Laskin, Pieter Abbeel, Alessandro Lazaric, and Lerrel Pinto. Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning. arXiv preprint arXiv:2201.13425, 2022.
  93. 93.Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto. Mastering visual continuous control: Improved data-augmented reinforcement learning. arXiv preprint arXiv:2107.09645, 2021.
  94. 94.Denis Yarats, Ilya Kostrikov, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. In International Conference on Learning Representations, 2021.
  95. 95.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? arXiv preprint arXiv:1411.1792, 2014.
  96. 96.Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma. Mopo: Model-based offline policy optimization. arXiv preprint arXiv:2005.13239, 2020.
  97. 97.Zhecheng Yuan, Zhengrong Xue, Bo Yuan, Xueqian Wang, Yi Wu, Yang Gao, and Huazhe Xu. Pre-trained image encoder for generalizable visual reinforcement learning. In First Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward at ICML 2022.
  98. 98.Albert Zhan, Philip Zhao, Lerrel Pinto, Pieter Abbeel, and Michael Laskin. A framework for efficient robotic manipulation. arXiv preprint arXiv:2012.07975, 2020.
  99. 99.Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine. Learning invariant representations for reinforcement learning without reconstruction. arXiv preprint arXiv:2006.10742, 2020.
  100. 100.Chi Zhang, Sanmukh Rao Kuppannagari, and Viktor K Prasanna. Maximum entropy model rollouts: Fast model based policy optimization without compounding errors. arXiv preprint arXiv:2006.04802, 2020.
  101. 101.Qihang Zhang, Zhenghao Peng, and Bolei Zhou. Action-conditioned contrastive policy pretraining. arXiv preprint arXiv:2204.02393, 2022.
  102. 102.Han Zheng, Xufang Luo, Xuan Song, Dongsheng Li, Jing Jiang, et al. Adaptive q-learning for interaction-limited reinforcement learning. 2021.

Citation

MLA
Wang, C., et al. “VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 32974–88, https://proceedings.neurips.cc/paper_files/paper/2022/file/d4cc7a2d0d70736e29a3b48c3729bc06-Paper-Conference.pdf.
APA
Wang, C., Luo, X., Ross, K., & Li, D. (2022). VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning. Advances in Neural Information Processing Systems, 35, 32974–32988. https://proceedings.neurips.cc/paper_files/paper/2022/file/d4cc7a2d0d70736e29a3b48c3729bc06-Paper-Conference.pdf
Chicago
Wang, C., X. Luo, K. Ross, and D. Li. 2022. “VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning”. Advances in Neural Information Processing Systems 35: 32974–88. https://proceedings.neurips.cc/paper_files/paper/2022/file/d4cc7a2d0d70736e29a3b48c3729bc06-Paper-Conference.pdf.
Harvard
Wang, C. et al. (2022) “VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 32974–32988. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/d4cc7a2d0d70736e29a3b48c3729bc06-Paper-Conference.pdf.
Vancouver
1. Wang C, Luo X, Ross K, Li D (2022) VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 32974–32988

BibTeX

@inproceedings{wang2022vrl3,
  title = {VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning},
  author = {Wang, Che and Luo, Xufang and Ross, Keith and Li, Dongsheng},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {32974-32988},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/d4cc7a2d0d70736e29a3b48c3729bc06-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission