Offline Meta-Reinforcement Learning with Online Self-Supervision

Vitchyr H. PongAshvin NairLaura SmithCatherine HuangSergey Levine

article2022ICML84 citations

Proposes a semi-supervised offline meta-reinforcement learning method that bridges the distribution shift between static training datasets and deployment exploration by autonomously generating reward labels for unlabeled online interactions.

Listen

Meta-reinforcement learning enables autonomous agents to adapt rapidly to new tasks with minimal data, but the initial meta-training phase typically demands vast amounts of online interaction and costly reward supervision. Relying solely on static, pre-collected datasets through offline meta-reinforcement learning offers a practical alternative to reuse multi-task data labeled only once. However, the article identifies a critical failure mode in offline meta-training: a significant distributional shift in the task context variables. Because the agent's meta-learned exploration policy generates online trajectories that systematically differ from the offline data, the learned adaptation procedure encounters unfamiliar inputs and experiences severe performance drops when adapting to new tasks.

The main objective of the article is to demonstrate that a hybrid approach—combining reward-labeled offline data with reward-free online self-supervision—can bridge this distributional shift and allow agents to successfully generalize to unseen tasks.

To evaluate this concept, the authors developed Semi-Supervised Meta Actor-Critic (SMAC). The method trains an initial policy and a generative reward decoder using limited offline data, and then collects additional online experience without ground-truth reward labels. The reward decoder autonomously generates synthetic labels for these online interactions, allowing the agent to update its value function and policy on realistic exploration trajectories. The authors evaluated the approach across six simulated benchmark domains, including standard locomotion benchmarks (such as Ant Direction, Cheetah Velocity, and Humanoid) and a complex robotic arm manipulation environment involving sparse rewards and diverse object interactions. Testing evaluated policies on held-out tasks using highly suboptimal offline datasets comprising up to three orders of magnitude fewer samples than prior offline benchmarks.

The findings show that standard offline meta-reinforcement learning baselines (such as BOReL and MACAW) struggle when trained on limited, suboptimal datasets, often failing to exceed simple meta-imitation. In contrast, SMAC steadily improved with reward-free online interaction, matching or closely approaching an oracle upper bound that received ground-truth rewards throughout online training. In the complex robotic manipulation domain, SMAC delivered substantial performance gains over competing methods while successfully adapting within just a few interaction episodes. Visual trajectory analyses confirmed that reward-free online fine-tuning successfully eliminated context distributional shifts, aligning post-adaptation behavior across both offline and online data sources.

These results demonstrate that online interaction does not require expensive manual reward engineering to yield high-performing adaptive policies. Organizations can dramatically lower data collection costs, reduce deployment risks, and accelerate development timelines by combining static multi-task datasets with autonomous, unsupervised fine-tuning. Practitioners developing adaptive robotic or autonomous control systems should adopt hybrid training frameworks that leverage unsupervised online rollouts, provided that appropriate physical safeguards and automated reset mechanisms are in place.

A primary operational requirement is the ability of the autonomous system to interact safely in the environment without external oversight during the online collection phase. While the findings provide strong confidence across simulated locomotion and manipulation benchmarks, evaluating the framework on physical hardware under real-world sensing noise remains a critical next step before broad deployment.

Cover for Offline Meta-Reinforcement Learning with Online Self-Supervision

Abstract

Meta-reinforcement learning (RL) methods can meta-train policies that adapt to new tasks with orders of magnitude less data than standard RL, but meta-training itself is costly and time-consuming. If we can meta-train on offline data, then we can reuse the same static dataset, labeled once with rewards for different tasks, to meta-train policies that adapt to a variety of new tasks at meta-test time. Although this capability would make meta-RL a practical tool for real-world use, offline meta-RL presents additional challenges beyond online meta-RL or standard offline RL settings. Meta-RL learns an exploration strategy that collects data for adapting, and also meta-trains a policy that quickly adapts to data from a new task. Since this policy was meta-trained on a fixed, offline dataset, it might behave unpredictably when adapting to data collected by the learned exploration strategy, which differs systematically from the offline data and thus induces distributional shift. We propose a hybrid offline meta-RL algorithm, which uses offline data with rewards to meta-train an adaptive policy, and then collects additional unsupervised online data, without any reward labels to bridge this distributional shift. By not requiring reward labels for online collection, this data can be much cheaper to collect. We compare our method to prior work on offline meta-RL on simulated robot locomotion and manipulation tasks and find that using additional unsupervised online data collection leads to a dramatic improvement in the adaptive capabilities of the meta-trained policies, matching the performance of fully online meta-RL on a range of challenging domains that require generalization to new tasks.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Preliminaries
  • 4. The Problem with Naïve Offline Meta-Reinforcement Learning
  • 5. Semi-Supervised Meta Actor-Critic
  • 5.1. Offline Meta-Training
  • 5.2. Self-Supervised Online Meta-Training
  • 5.3. Algorithm Summary and Details
  • 6. Experiments
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Method Pseudo-code
  • B. Additional Experimental Results and Discussion
  • C. Experimental Details
  • C.1. Data Collection Difference from Prior Work
  • C.2. Environment Details
  • C.3. Hyperparameters

Knowls

  1. Knowl 1 — Context Distribution Shift in Offline Meta-Reinforcement Learning

    definition

    In context-based offline meta-reinforcement learning (meta-RL), an agent learns an adaptation encoder qϕe(z∣h)q_{\phi_e}(z \mid h) and a contextual policy πθ(a∣s,z)\pi_\theta(a \mid s, z), where ss is the state, aa is the action, h={(sk,ak,rk,sk′)}k=1Nench = \{(s_k, a_k, r_k, s'_k)\}_{k=1}^{N_{\text{enc}}} is a transition history batch, and zz is a latent task context vector. During offline meta-training, hh is sampled exclusively from a static dataset hofflineh_{\text{offline}} collected by a behavior policy πβ\pi_\beta.

    Context distribution shift (or distribution shift in zz-space) refers to the discrepancy between the offline context posterior p(z∣hoffline)p(z \mid h_{\text{offline}}) observed during training and the test-time context posterior p(z∣honline)p(z \mid h_{\text{online}}) produced when honlineh_{\text{online}} is gathered by the meta-learned exploration policy πθ\pi_\theta. Because πθ\pi_\theta systematically deviates from πβ\pi_\beta due to optimization artifacts and exploration noise, the encoder produces latent context vectors z∼qϕe(z∣honline)z \sim q_{\phi_e}(z \mid h_{\text{online}}) that lie outside the training support of πθ(a∣s,z)\pi_\theta(a \mid s, z). This mismatch degrades post-adaptation performance on unseen tasks, even if the policy adapts effectively to offline trajectories.

  2. Knowl 2 — Semi-Supervised Meta Actor-Critic

    model/method

    Semi-Supervised Meta Actor-Critic (SMAC) is an offline meta-reinforcement learning method designed to overcome context distribution shift by augmenting offline meta-training with self-supervised online interaction without ground-truth reward supervision.

    SMAC operates in two sequential phases:

    1. Offline Meta-Training: Using a multi-task offline dataset D={Di}i=1Nbuff\mathcal{D} = \{\mathcal{D}_i\}_{i=1}^{N_{\text{buff}}}, SMAC jointly trains an inference encoder qϕe(z∣h)q_{\phi_e}(z \mid h) to map interaction history hh to a Gaussian latent context zz, a generative reward decoder rϕd(s,a,z)r_{\phi_d}(s, a, z) to reconstruct offline rewards from latent context, a context-conditioned critic Qw(s,a,z)Q_w(s, a, z) via temporal-difference learning, and an actor πθ(a∣s,z)\pi_\theta(a \mid s, z) via Advantage-Weighted Actor-Critic (AWAC) updates to prevent out-of-distribution action extrapolation.

    2. Self-Supervised Online Meta-Training: The agent collects unlabeled exploratory trajectories τ=(s1,a1,s2,… )\tau = (s_1, a_1, s_2, \dots) online using the exploration policy conditioned on prior samples z∼pz(z)z \sim p_z(z). For each collected trajectory, a task buffer Di\mathcal{D}_i is selected, an offline history hoffline∼Dih_{\text{offline}} \sim \mathcal{D}_i is passed to the frozen encoder to obtain z∼qϕe(z∣hoffline)z \sim q_{\phi_e}(z \mid h_{\text{offline}}), and rewards for τ\tau are generated synthetically via rgenerated=rϕd(s,a,z)r_{\text{generated}} = r_{\phi_d}(s, a, z). The autonomously relabeled trajectories are added to Di\mathcal{D}_i, and policy and critic updates continue. Freezing ϕe\phi_e and ϕd\phi_d prevents synthetic reward error accumulation while exposing πθ\pi_\theta and QwQ_w to online exploration trajectory distributions.

  3. Knowl 3 — SMAC Meta-Training Algorithm

    algorithm

    The complete execution flow for Semi-Supervised Meta Actor-Critic (SMAC) spans an initial offline optimization phase followed by a self-supervised online fine-tuning phase.

    Input: Multi-task offline datasets D={Di}i=1Nbuff\mathcal{D} = \{\mathcal{D}_i\}_{i=1}^{N_{\text{buff}}}, policy πθ\pi_\theta, Q-function QwQ_w, encoder qϕeq_{\phi_e}, reward decoder rϕdr_{\phi_d}, prior pz(z)p_z(z)
    Input: Number of offline iterations NofflineN_{\text{offline}}, online iterations NonlineN_{\text{online}}
    for iteration n=1,2,…,Nofflinen = 1, 2, \dots, N_{\text{offline}} do
        Sample task buffer Di∼D\mathcal{D}_i \sim \mathcal{D}
        Sample two independent history batches h,h′∼Dih, h' \sim \mathcal{D}_i
        Sample latent context z∼qϕe(z∣h)z \sim q_{\phi_e}(z \mid h)
        Update ϕe\phi_e and ϕd\phi_d by minimizing Lreward\mathcal{L}_{\text{reward}} on hh
        Update ww by minimizing Lcritic\mathcal{L}_{\text{critic}} on h′h' with context zz
        Update θ\theta by minimizing Lactor\mathcal{L}_{\text{actor}} on h′h' with context zz
    end for
    for iteration n=1,2,…,Nonlinen = 1, 2, \dots, N_{\text{online}} do
        Sample exploration context zt∼pz(z)z_t \sim p_z(z)
        Roll out policy πθ(a∣s,zt)\pi_\theta(a \mid s, z_t) to collect unlabeled trajectory τ\tau
        Sample task buffer Di∼D\mathcal{D}_i \sim \mathcal{D} and offline history hoffline∼Dih_{\text{offline}} \sim \mathcal{D}_i
        Sample context z∼qϕe(z∣hoffline)z \sim q_{\phi_e}(z \mid h_{\text{offline}})
        Compute synthetic rewards rt=rϕd(st,at,z)r_t = r_{\phi_d}(s_t, a_t, z) for each transition in τ\tau
        Append labeled trajectory τ\tau to task buffer Di\mathcal{D}_i
        Sample two history batches h,h′∼Dih, h' \sim \mathcal{D}_i
        Sample context z∼qϕe(z∣h)z \sim q_{\phi_e}(z \mid h)
        Update ww by minimizing Lcritic\mathcal{L}_{\text{critic}} on h′h' with context zz
        Update θ\theta by minimizing Lactorself-supervised\mathcal{L}_{\text{actor}}^{\text{self-supervised}} on h′h' with context zz
    end for
  4. Knowl 4 — Offline Objectives for SMAC Critic and Advantage-Weighted Actor

    equation

    During offline meta-training, the context-conditioned action-value function Qw(s,a,z)Q_w(s, a, z) with parameters ww is updated by minimizing the Mean Squared Bellman Error:

    Lcritic(w)=E(s,a,r,s′)∼Di, z∼qϕe(z∣h), a′∼πθ(a′∣s′,z)[(Qw(s,a,z)−(r+γQwˉ(s′,a′,z)))2]\mathcal{L}_{\text{critic}}(w) = \mathbb{E}_{(s, a, r, s') \sim \mathcal{D}_i, \, z \sim q_{\phi_e}(z \mid h), \, a' \sim \pi_\theta(a' \mid s', z)} \left[ \left( Q_w(s, a, z) - \left( r + \gamma Q_{\bar{w}}(s', a', z) \right) \right)^2 \right]

    where wˉ\bar{w} denotes target network parameters updated via Polyak averaging with rate η\eta, γ∈[0,1)\gamma \in [0, 1) is the discount factor, and hh is a history batch sampled from task buffer Di\mathcal{D}_i.

    To prevent bootstrapping error from out-of-distribution actions in the offline dataset, the policy πθ(a∣s,z)\pi_\theta(a \mid s, z) with parameters θ\theta is trained via advantage-weighted regression:

    Lactor(θ)=−E(s,a,s′)∼Di, z∼qϕe(z∣h)[log⁡πθ(a∣s,z)exp⁡(Qw(s,a,z)−V(s′,z)λ)]\mathcal{L}_{\text{actor}}(\theta) = - \mathbb{E}_{(s, a, s') \sim \mathcal{D}_i, \, z \sim q_{\phi_e}(z \mid h)} \left[ \log \pi_\theta(a \mid s, z) \exp \left( \frac{Q_w(s, a, z) - V(s', z)}{\lambda} \right) \right]

    where λ>0\lambda > 0 is a temperature hyperparameter (Lagrange multiplier) and V(s′,z)=Ea′′∼πθ(a′′∣s′,z)[Qw(s′,a′′,z)]V(s', z) = \mathbb{E}_{a'' \sim \pi_\theta(a'' \mid s', z)} [ Q_w(s', a'', z) ] is estimated with a single action sample.

  5. Knowl 5 — Generative Reward Decoder and Encoder Loss Objective

    equation

    To learn a generative model over reward functions of meta-training tasks, SMAC optimizes an encoder qϕe(z∣h)q_{\phi_e}(z \mid h) and a reward decoder rϕd(s,a,z)r_{\phi_d}(s, a, z) on offline batches h={(sk,ak,rk,sk′)}k=1Nench = \{(s_k, a_k, r_k, s'_k)\}_{k=1}^{N_{\text{enc}}} using the reconstruction loss regularized by a Kullback-Leibler (KL) divergence bottleneck against a fixed prior pz(z)=N(0,I)p_z(z) = \mathcal{N}(0, I):

    Lreward(ϕd,ϕe,h,z)=∑(s,a,r)∈h∥r−rϕd(s,a,z)∥22+DKL(qϕe(⋅∣h) ∥ pz(⋅))\mathcal{L}_{\text{reward}}(\phi_d, \phi_e, h, z) = \sum_{(s, a, r) \in h} \| r - r_{\phi_d}(s, a, z) \|_2^2 + D_{\text{KL}}\left( q_{\phi_e}(\cdot \mid h) \,\|\, p_z(\cdot) \right)

    where z∼qϕe(z∣h)z \sim q_{\phi_e}(z \mid h). The posterior qϕe(z∣h)q_{\phi_e}(z \mid h) factorizes as a product of Gaussian factors over transitions:

    qϕe(z∣h)∝∏(s,a,r)∈hN(μϕe(s,a,r), diag(σϕe2(s,a,r)))q_{\phi_e}(z \mid h) \propto \prod_{(s, a, r) \in h} \mathcal{N}\left( \mu_{\phi_e}(s, a, r), \, \text{diag}(\sigma_{\phi_e}^2(s, a, r)) \right)

    where μϕe\mu_{\phi_e} and σϕe\sigma_{\phi_e} are computed by a multi-layer perceptron with softplus output activation for variance.

  6. Knowl 6 — Self-Supervised Labeling and Online Actor Optimization Objective

    equation

    During the self-supervised online phase, unlabeled transitions (s,a)(s, a) collected via exploration are assigned synthetic rewards using the reward decoder conditioned on an offline history:

    rgenerated=rϕd(s,a,z),z∼qϕe(z∣hoffline),hoffline∼Dir_{\text{generated}} = r_{\phi_d}(s, a, z), \quad z \sim q_{\phi_e}(z \mid h_{\text{offline}}), \quad h_{\text{offline}} \sim \mathcal{D}_i

    The actor parameter θ\theta is optimized using a weighted combination of the offline advantage-weighted loss Lactor(θ)\mathcal{L}_{\text{actor}}(\theta) and the maximum-entropy PEARL actor objective LactorPEARL(θ)\mathcal{L}_{\text{actor}}^{\text{PEARL}}(\theta):

    Lactorself-supervised(θ)=Lactor(θ)+λpearlLactorPEARL(θ)\mathcal{L}_{\text{actor}}^{\text{self-supervised}}(\theta) = \mathcal{L}_{\text{actor}}(\theta) + \lambda_{\text{pearl}} \mathcal{L}_{\text{actor}}^{\text{PEARL}}(\theta)

    LactorPEARL(θ)=Es∼Di, z∼qϕe(z∣h)[DKL(πθ(a∣s,z)  ∥  exp⁡(Qw(s,a,z))Z(s))]\mathcal{L}_{\text{actor}}^{\text{PEARL}}(\theta) = \mathbb{E}_{s \sim \mathcal{D}_i, \, z \sim q_{\phi_e}(z \mid h)} \left[ D_{\text{KL}}\left( \pi_\theta(a \mid s, z) \;\Big\|\; \frac{\exp(Q_w(s, a, z))}{Z(s)} \right) \right]

    where λpearl≥0\lambda_{\text{pearl}} \ge 0 balances offline behavioral regularized updates with off-policy maximum-entropy actor-critic updates, and Z(s)Z(s) is an intractable partition function handled via reparameterized sampling.

  7. Knowl 7 — Benchmark Adaptation Performance on Unseen Meta-Test Tasks

    empirical result

    SMAC was evaluated on six meta-RL benchmark domains with limited, suboptimal offline data: Cheetah Velocity, Ant Direction, Humanoid (376-dimensional state space), Walker Param, Hopper Param, and Sawyer Manipulation (robotic arm with sparse rewards and 46% offline success rate). Evaluation was conducted across 4 random seeds on held-out test tasks within T=3T = 3 adaptation episodes.

    Key findings include:

    1. Across all six domains, self-supervised online meta-training steadily increased post-adaptation returns over the offline initialization (step 0), matching the performance of an Online Oracle baseline that received fully labeled ground-truth rewards during online interaction.
    2. In the offline-only setting (step 0), SMAC outperformed prior offline meta-RL algorithms BOReL and MACAW on 4 out of 6 domains.
    3. After online self-supervised fine-tuning, SMAC significantly outperformed BOReL on all 6 domains and MACAW on 5 out of 6 domains.
    4. SMAC greatly exceeded the meta behavior cloning baseline across all tasks, demonstrating substantial policy improvement beyond the suboptimal offline data collection policy.
    5. The SAC-based actor ablation (replacing AWAC with Soft Actor-Critic during offline meta-training) failed to learn an accurate offline value function, achieving poor test return across challenging domains.
  8. Knowl 8 — Empirical Validation of Context Distribution Shift and Mitigation in Latent Space

    empirical result

    On the Ant Direction meta-RL task, evaluating the context encoder qϕe(z∣h)q_{\phi_e}(z \mid h) immediately after offline training revealed a severe distribution shift:

    1. KL Divergence to Prior: The distribution of DKL(qϕe(z∣h)∥pz(z))D_{\text{KL}}(q_{\phi_e}(z \mid h) \parallel p_z(z)) was substantially higher when hh was sampled from online trajectories honlineh_{\text{online}} generated by πθ\pi_\theta compared to histories hofflineh_{\text{offline}} sampled from the offline dataset.
    2. Behavioral Mode Collapse: When the post-adaptation policy was conditioned on z∼qϕe(z∣honline)z \sim q_{\phi_e}(z \mid h_{\text{online}}), the ant robot collapsed to moving only in a single direction (up and to the left), causing a sharp decrease in post-adaptation return. When conditioned on z∼qϕe(z∣hoffline)z \sim q_{\phi_e}(z \mid h_{\text{offline}}), the same policy successfully navigated in all directions.
    3. Post-Finetuning Recovery: Following the self-supervised online phase, exploration trajectories generated by πθ\pi_\theta remained qualitatively similar to those pre-finetuning, but conditioning the post-adaptation policy on z∼qϕe(z∣honline)z \sim q_{\phi_e}(z \mid h_{\text{online}}) enabled navigation across the entire circle of directions, proving that self-supervision directly resolves the encoder-policy latent distribution mismatch.
  9. Knowl 9 — Mitigating State-Space Distribution Shift via Test-Task Self-Supervision

    empirical result

    In the Sawyer Manipulation domain configured with 8 distinct potential manipulation behaviors, self-supervised online meta-training was conducted in two different settings: interacting with the meta-training environments vs. interacting with held-out meta-test environments, both without reward supervision.

    Self-supervision conducted directly on the test environments yielded continuous, significant gains in post-adaptation return on unseen tasks, achieving performance competitive with an online oracle trained with ground-truth test-task rewards. In contrast, self-supervised training restricted to the meta-training task environments failed to improve post-adaptation returns on the test tasks. This demonstrates that unsupervised online exploration in test-domain states mitigates state-space distribution shift in meta-RL without requiring test-time reward annotation.

  10. Knowl 10 — SMAC Hyperparameter Specifications

    data/table

    The standard network architectures and hyperparameters for the self-supervised phase of SMAC are structured across shared and environment-specific settings.

    Hyperparameter Value
    RL batch size 256
    Encoder batch size 64
    Meta batch size 4
    Q-network hidden sizes [300, 300, 300]
    Policy network hidden sizes [300, 300, 300]
    Decoder network hidden sizes [64, 64]
    Encoder network hidden sizes [200, 200, 200]
    zz dimensionality (dzd_z) 5
    Hidden activation ReLU
    Critic / Encoder / Decoder output activation Identity
    Policy output activation tanh⁡\tanh
    Discount factor γ\gamma 0.99
    Target network soft target η\eta 0.005
    Optimizer Adam
    Learning rate (all networks) 3×10−43 \times 10^{-4}
    Gradient steps per environment transition 4
    Offline pretraining gradient steps 50000
    Environment Horizon AWR β\beta Reward Scale # Train Tasks # Test Tasks λpearl\lambda_{\text{pearl}}
    Cheetah Velocity 200 100 5 100 30 1
    Ant Direction 200 100 5 100 20 1
    Sawyer Manipulation 50 0.3 1 50 10 0
    Walker Param 200 100 5 50 5 1
    Hopper Param 200 100 5 50 5 1
    Humanoid 200 100 5 50 5 1

    For each gradient update, the critic and actor observe (RL batch size)×(meta batch size)=1024(\text{RL batch size}) \times (\text{meta batch size}) = 1024 transitions, while the encoder observes (encoder batch size)×(meta batch size)=256(\text{encoder batch size}) \times (\text{meta batch size}) = 256 transitions.

  11. Knowl 11 — Encoder Loss Sensitivity for Dynamics Adaptation

    empirical result

    Ablation experiments comparing encoder training objectives across Walker Param, Hopper Param, and Humanoid domains demonstrated that training the encoder qϕe(z∣h)q_{\phi_e}(z \mid h) exclusively with critic Q-loss (as in standard PEARL) leads to severe performance degradation. This failure is pronounced in Walker Param and Hopper Param, which require inferring varying physical parameters (mass, damping, inertia, friction). The Q-loss alone fails to encode sufficient task information to support accurate synthetic reward decoding. In contrast, training the encoder via reward reconstruction loss Lreward\mathcal{L}_{\text{reward}} enables reliable reward generation and robust adaptation.

  12. Knowl 12 — Requirement of Unlabeled Online Environment Interactions

    limitation

    A fundamental operational assumption of SMAC is that the agent can execute autonomous, unlabeled interactions in the target environment with automatic resets or safeguards. In applications where real-world environment interaction is prohibitively dangerous, damaging to hardware, or strictly disallowed prior to deployment, SMAC cannot execute its self-supervised online phase and is limited to its offline meta-training performance.

Coverage note — None was omitted; all key theoretical formulations, algorithmic mechanics, empirical benchmark results, distribution shift analyses, ablation studies, and architectural configurations are covered.

References

  1. 1.Abbeel, P. and Ng, A. Y. Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning, pp. 1, 2004.
  2. 2.Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., Mcgrew, B., Tobin, J., Abbeel, P., and Zaremba, W. Hindsight Experience Replay. In Advances in Neural Information Processing Systems (NIPS), 2017. URL https://arxiv.org/pdf/1707.01495.pdfhttp://arxiv.org/abs/1707.01495.
  3. 3.Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D. Successor features for transfer in reinforcement learning. In Advances in neural information processing systems, pp. 4055–4065, 2017.
  4. 4.Barreto, A., Borsa, D., Quan, J., Schaul, T., Silver, D., Hessel, M., Mankowitz, D., Zˇ´ıdek, A., and Munos, R. Transfer in deep reinforcement learning using successor features and generalised policy improvement. arXiv preprint arXiv:1901.10964, 2019.
  5. 5.Bloesch, M., Humplik, J., Patraucean, V., Hafner, R., Haarnoja, T., Byravan, A., Siegel, N. Y., Tunyasuvunakool, S., Casarini, F., Batchelor, N., et al. Towards real robot learning in the wild: A case study in bipedal locomotion. In Conference on Robot Learning, pp. 1502–1511. PMLR, 2022.
  6. 6.Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. Openai gym. arXiv preprint arXiv:1606.01540, 2016.
  7. 7.Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. arXiv preprint arXiv:1706.03741, 2017.
  8. 8.Colas, C., Sigaud, O., and Oudeyer, P.-Y. Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms. International Conference on Machine Learning (ICML), 2018.
  9. 9.Coumans, E. and Bai, Y. Pybullet, a python module for physics simulation for games, robotics and machine learning. http://pybullet.org, 2016–2021.
  10. 10.Dorfman, R. and Tamar, A. Offline meta reinforcement learning. arXiv preprint arXiv:2008.02598, 2020.
  11. 11.Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. Rl2 : Fast reinforcement learning via slow reinforcement learning. arXiv preprint arXiv:1611.02779, 2016.
  12. 12.Finn, C., Levine, S., and Abbeel, P. Guided cost learning: Deep inverse optimal control via policy optimization. In International conference on machine learning, pp. 49–58. PMLR, 2016.
  13. 13.Finn, C., Abbeel, P., and Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, pp. 1126–1135. PMLR, 2017.
  14. 14.Fu, J., Luo, K., and Levine, S. Learning robust rewards with adversarial inverse reinforcement learning. arXiv preprint arXiv:1710.11248, 2017.
  15. 15.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
  16. 16.Fujimoto, S., Conti, E., Ghavamzadeh, M., and Pineau, J. Benchmarking batch deep reinforcement learning algorithms. arXiv preprint arXiv:1910.01708, 2019a.
  17. 17.Fujimoto, S., Meger, D., and Precup, D. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pp. 2052–2062. PMLR, 2019b.
  18. 18.Grimm, C., Higgins, I., Barreto, A., Teplyashin, D., Wulfmeier, M., Hertweck, T., Hadsell, R., and Singh, S. Disentangled cumulants help successor representations transfer to new tasks. arXiv preprint arXiv:1911.10866, 2019.
  19. 19.Gupta, A., Eysenbach, B., Finn, C., and Levine, S. Unsupervised meta-learning for reinforcement learning. arXiv preprint arXiv:1806.04640, 2018a.
  20. 20.Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S. Meta-reinforcement learning of structured exploration strategies. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, 2018b.
  21. 21.Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905, 2018.
  22. 22.Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S., and Dragan, A. Inverse reward design. arXiv preprint arXiv:1711.02827, 2017.
  23. 23.Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M. Learning an embedding space for transferable robot skills. In International Conference on Learning Representations, 2018.
  24. 24.Ho, J. and Ermon, S. Generative adversarial imitation learning. arXiv preprint arXiv:1606.03476, 2016.
  25. 25.Humplik, J., Galashov, A., Hasenclever, L., Ortega, P. A., Teh, Y. W., and Heess, N. Meta reinforcement learning as task inference. arXiv preprint arXiv:1905.06424, 2019.
  26. 26.Jabri, A., Hsu, K., Eysenbach, B., Gupta, A., Levine, S., and Finn, C. Unsupervised curricula for visual meta-reinforcement learning. arXiv preprint arXiv:1912.04226, 2019.
  27. 27.Kaelbling, L. P. Learning to achieve goals. In International Joint Conference on Artificial Intelligence (IJCAI), volume vol.2, pp. 1094 – 8, 1993.
  28. 28.Kamienny, P.-A., Pirotta, M., Lazaric, A., Lavril, T., Usunier, N., and Denoyer, L. Learning adaptive exploration strategies in dynamic environments through informed policy regularization. arXiv preprint arXiv:2005.02934, 2020.
  29. 29.Khazatsky, A., Nair, A., Jing, D., and Levine, S. What can i do here? learning new skills by imagining visual affordances. In International Conference on Robotics and Automation. IEEE, 2021.
  30. 30.Kirsch, L., van Steenkiste, S., and Schmidhuber, J. Improving generalization in meta reinforcement learning using learned objectives. arXiv preprint arXiv:1910.04098, 2019.
  31. 31.Kofman, J., Wu, X., Luu, T. J., and Verma, S. Teleoperation of a robot manipulator using a vision-based human-robot interface. IEEE transactions on industrial electronics, 52 (5):1206–1219, 2005.
  32. 32.Konyushkova, K., Zolna, K., Aytar, Y., Novikov, A., Reed, S., Cabi, S., and de Freitas, N. Semi-supervised reward learning for offline reinforcement learning. arXiv preprint arXiv:2012.06899, 2020.
  33. 33.Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J. Deep successor reinforcement learning. arXiv preprint arXiv:1606.02396, 2016.
  34. 34.Kumar, A., Fu, J., Tucker, G., and Levine, S. Stabilizing off-policy q-learning via bootstrapping error reduction. arXiv preprint arXiv:1906.00949, 2019.
  35. 35.Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871, 2018.
  36. 36.Levine, S., Pastor, P., Krizhevsky, A., Ibarz, J., and Quillen, D. Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection. The International Journal of Robotics Research, 37(4-5):421–436, 2018.
  37. 37.Levine, S., Kumar, A., Tucker, G., and Fu, J. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  38. 38.Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2016. ISSN 10769757. doi: 10.1613/jair.301. URL https://arxiv.org/pdf/1509.02971.pdf.
  39. 39.Mitchell, E., Rafailov, R., Peng, X. B., Levine, S., and Finn, C. Offline meta-reinforcement learning with advantage weighting. In International Conference on Machine Learning, pp. 7780–7791. PMLR, 2021.
  40. 40.Nair, A., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S. Visual Reinforcement Learning with Imagined Goals. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  41. 41.Nair, A., Gupta, A., Dalal, M., and Levine, S. Awac: Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020.
  42. 42.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  43. 43.Nogueira, R. and Cho, K. End-to-end goal-driven web navigation. Advances in neural information processing systems, 29:1903–1911, 2016.
  44. 44.Peng, X. B., Coumans, E., Zhang, T., Lee, T.-W., Tan, J., and Levine, S. Learning agile robotic locomotion skills by imitating animals. In Robotics: Science and Systems, 2020.
  45. 45.Per´ e, A., Forestier, S., Sigaud, O., and Oudeyer, P.-Y. Un- ´ supervised Learning of Goal Spaces for Intrinsically Motivated Goal Exploration. In International Conference on Learning Representations (ICLR), 2018. URL https://arxiv.org/pdf/1803.00781.pdf.
  46. 46.Pong, V., Gu, S., Dalal, M., and Levine, S. Temporal Difference Models: Model-Free Deep RL For Model-Based Control. In International Conference on Learning Representations (ICLR), 2018. URL https://arxiv.org/pdf/1802.09081.pdf.
  47. 47.Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. Efficient off-policy meta-reinforcement learning via probabilistic context variables. In International conference on machine learning, pp. 5331–5340. PMLR, 2019.
  48. 48.Reddy, S., Dragan, A. D., and Levine, S. Sqil: Imitation learning via reinforcement learning with sparse rewards. arXiv preprint arXiv:1905.11108, 2019.
  49. 49.Ross, S. and Bagnell, D. Efficient reductions for imitation learning. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp. 661–668. JMLR Workshop and Conference Proceedings, 2010.
  50. 50.Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P. Promp: Proximal meta-policy search. In International Conference on Learning Representations, 2018.
  51. 51.Schaal, S. Is imitation learning the route to humanoid robots? Trends in cognitive sciences, 3(6):233–242, 1999.
  52. 52.Schaul, T., Horgan, D., Gregor, K., and Silver, D. Universal Value Function Approximators. In International Conference on Machine Learning (ICML), 2015.
  53. 53.Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., and Hardt, M. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts. In International Conference on Machine Learning (ICML), 2020. URL https://test-time-training.github.io/.
  54. 54.Todorov, E., Erez, T., and Tassa, Y. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 5026–5033. IEEE, 2012.
  55. 55.Torrado, R. R., Bontrager, P., Togelius, J., Liu, J., and Perez-Liebana, D. Deep reinforcement learning for general video game ai. In 2018 IEEE Conference on Computational Intelligence and Games (CIG), pp. 1–8. IEEE, 2018.
  56. 56.Warde-Farley, D., de Wiele, T. V., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V. Unsupervised control through non-parametric discriminative rewards. CoRR, abs/1811.11359, 2018.
  57. 57.Wu, Y., Tucker, G., and Nachum, O. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  58. 58.Xu, D. and Denil, M. Positive-unlabeled reward learning. arXiv preprint arXiv:1911.00459, 2019.
  59. 59.Xu, Z., van Hasselt, H. P., and Silver, D. Meta-gradient reinforcement learning. Advances in neural information processing systems, 31:2396–2407, 2018.
  60. 60.Xu, Z., van Hasselt, H., Hessel, M., Oh, J., Singh, S., and Silver, D. Meta-gradient reinforcement learning with an objective discovered online. arXiv preprint arXiv:2007.08433, 2020.
  61. 61.Yang, B., Zhang, J., Pong, V., Levine, S., and Jayaraman, D. Replab: A reproducible low-cost arm benchmark platform for robotic learning. arXiv preprint arXiv:1905.07447, 2019.
  62. 62.Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning, pp. 1094–1100. PMLR, 2020.
  63. 63.Zhao, T. Z., Nagabandi, A., Rakelly, K., Finn, C., and Levine, S. Meld: Meta-reinforcement learning from images via latent state models. arXiv preprint arXiv:2010.13957, 2020.
  64. 64.Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S. Varibad: a very good method for bayes-adaptive deep rl via meta-learning. Proceedings of ICLR 2020, 2020.
  65. 65.Zolna, K., Novikov, A., Konyushkova, K., Gulcehre, C., Wang, Z., Aytar, Y., Denil, M., de Freitas, N., and Reed, S. Offline learning from demonstrations and unlabeled experience. arXiv preprint arXiv:2011.13885, 2020.

Citation

MLA
Pong, V. H., et al. “Offline Meta-Reinforcement Learning with Online Self-Supervision”. International Conference on Machine Learning, vol. 162, 2022, pp. 17811–29, https://proceedings.mlr.press/v162/pong22a.html.
APA
Pong, V. H., Nair, A. V., Smith, L. M., Huang, C., & Levine, S. (2022). Offline Meta-Reinforcement Learning with Online Self-Supervision. International Conference on Machine Learning, 162, 17811–17829. https://proceedings.mlr.press/v162/pong22a.html
Chicago
Pong, V. H., A. V. Nair, L. M. Smith, C. Huang, and S. Levine. 2022. “Offline Meta-Reinforcement Learning with Online Self-Supervision”. International Conference on Machine Learning 162: 17811–29. https://proceedings.mlr.press/v162/pong22a.html.
Harvard
Pong, V.H. et al. (2022) “Offline Meta-Reinforcement Learning with Online Self-Supervision”, International Conference on Machine Learning. PMLR, pp. 17811–17829. Available at: https://proceedings.mlr.press/v162/pong22a.html.
Vancouver
1. Pong VH, Nair AV, Smith LM, Huang C, Levine S (2022) Offline Meta-Reinforcement Learning with Online Self-Supervision. In: International Conference on Machine Learning. PMLR, pp 17811–17829

BibTeX

@InProceedings{pmlr-v162-pong22a,
  title = 	 {Offline Meta-Reinforcement Learning with Online Self-Supervision},
  author =       {Pong, Vitchyr H and Nair, Ashvin V and Smith, Laura M and Huang, Catherine and Levine, Sergey},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {17811--17829},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/pong22a/pong22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/pong22a.html},
  abstract = 	 {Meta-reinforcement learning (RL) methods can meta-train policies that adapt to new tasks with orders of magnitude less data than standard RL, but meta-training itself is costly and time-consuming. If we can meta-train on offline data, then we can reuse the same static dataset, labeled once with rewards for different tasks, to meta-train policies that adapt to a variety of new tasks at meta-test time. Although this capability would make meta-RL a practical tool for real-world use, offline meta-RL presents additional challenges beyond online meta-RL or standard offline RL settings. Meta-RL learns an exploration strategy that collects data for adapting, and also meta-trains a policy that quickly adapts to data from a new task. Since this policy was meta-trained on a fixed, offline dataset, it might behave unpredictably when adapting to data collected by the learned exploration strategy, which differs systematically from the offline data and thus induces distributional shift. We propose a hybrid offline meta-RL algorithm, which uses offline data with rewards to meta-train an adaptive policy, and then collects additional unsupervised online data, without any reward labels to bridge this distribution shift. By not requiring reward labels for online collection, this data can be much cheaper to collect. We compare our method to prior work on offline meta-RL on simulated robot locomotion and manipulation tasks and find that using additional unsupervised online data collection leads to a dramatic improvement in the adaptive capabilities of the meta-trained policies, matching the performance of fully online meta-RL on a range of challenging domains that require generalization to new tasks.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/