Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning

Haoqi YuanZongqing Lu

article2022ICML58 citations

Presents a contrastive learning framework with a bi-level encoder that decouples task characteristics from behavior policies in offline meta-reinforcement learning, enabling reliable task adaptation even under severe distribution shifts during test time.

Listen

Real-world applications of deep reinforcement learning—such as robotics, recommendation systems, and autonomous control—often face severe data collection costs and safety risks when exploring environments online. Offline meta-reinforcement learning addresses these hurdles by training agents entirely on pre-collected, multi-task datasets so they can quickly adapt to unseen tasks. However, existing context-based approaches struggle when the behavior policy used to collect offline data differs from the exploratory behavior encountered during testing. Because standard models encode full multi-step trajectories, they often memorize features of the data-collection policy instead of learning the true underlying task, causing severe performance drops during deployment.

The article introduces and evaluates CORRO (Contrastive Robust Task Representation Learning), a framework designed to learn task representations that remain stable and accurate even when testing policies deviate significantly from training data. Its main objective is to demonstrate that isolating single-step transitions and optimizing a mutual information contrastive objective enables robust task identification across both varying reward functions and changing environment physics.

To achieve this, the article develops a bi-level task encoder. Instead of processing entire trajectories, the transition encoder processes individual transition steps (state, action, reward, next state) to extract task-relevant signals while stripping out policy-specific patterns. These latent codes are then pooled using an attention-based aggregator to condition the agent’s policy. The encoder is trained via a contrastive objective based on mutual information maximization, differentiating true transitions from synthetically generated negative pairs via generative modeling or reward randomization. The framework was evaluated across multiple continuous control benchmarks involving 20 training and 20 testing tasks across varying physical dynamics and reward conditions.

The experimental findings highlight substantial improvements over existing methods. First, CORRO achieves superior task adaptation performance under standard testing conditions, outperforming prior offline meta-learning approaches and reaching top returns within 20,000 offline training steps. Second, it demonstrates significant robustness against out-of-distribution exploratory behaviors: while baseline methods degraded catastrophically (for example, dropping to returns of -204.1 and -242.7 in velocity tracking tasks), CORRO maintained a stable return of -89.7. Third, it successfully learned structured latent representations that cleanly separate distinct tasks along physical and reward dimensions without requiring access to true task labels during training. Fourth, ablation testing confirms that encoding single-step transitions rather than trajectories is the primary factor driving resistance to policy mismatch.

These findings indicate that offline meta-reinforcement learning can be deployed safely and reliably even when real-world testing environments exhibit unpredictable exploration patterns. By preventing models from relying on spurious behavioral correlations, practitioners can reduce the risk of unexpected deployment failures in safety-critical robotics and control systems. The results also show that high sample diversity in negative pair generation is essential for effective contrastive learning in offline data regimes.

Organizations developing offline reinforcement learning systems should transition from trajectory-level task encoders to transition-level contrastive architectures. When configuring negative sample generation, practitioners should select generative modeling when state-action coverage overlaps across tasks, and switch to reward randomization when task policies explore disjoint state spaces. Moving forward, additional work is recommended to integrate learned exploration policies and explore minimal-interaction hybrid training when offline data is severely limited.

The reported findings are backed by consistent multi-seed experiments across standard continuous control benchmarks. However, confidence should be tempered by the fact that evaluations were conducted in simulated physical environments with pre-collected datasets rather than live physical hardware. Careful pilot testing remains necessary before applying the framework to highly complex, noisy, or unmodeled real-world environments.

No sufficiently relevant recommendations were found.

Cover for Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning

Abstract

We study offline meta-reinforcement learning, a practical reinforcement learning paradigm that learns from offline data to adapt to new tasks. The distribution of offline data is determined jointly by the behavior policy and the task. Existing offline meta-reinforcement learning algorithms cannot distinguish these factors, making task representations unstable to the change of behavior policies. To address this problem, we propose a contrastive learning framework for task representations that are robust to the distribution mismatch of behavior policies in training and test. We design a bi-level encoder structure, use mutual information maximization to formalize task representation learning, derive a contrastive learning objective, and introduce several approaches to approximate the true distribution of negative pairs. Experiments on a variety of offline meta-reinforcement learning benchmarks demonstrate the advantages of our method over prior methods, especially on the generalization to out-of-distribution behavior policies.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 3.1. Problem Formulation
  • 3.2. Context-Based OMRL and FOCAL
  • 3.3. Task Representation Problems in OMRL
  • 4. Method
  • 4.1. Bi-Level Task Encoder
  • 4.2. Contrastive Task Representation Learning
  • 4.3. Negative Pairs Generation
  • 4.4. Algorithm Summary
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Tasks Adaptation Performance
  • 5.3. Robust Task Inference
  • 5.4. Latent Space Visualization
  • 5.5. Ablation: Negative Pairs Generation
  • 5.6. Adaptation with an Exploration Policy
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Contrastive Learning Objective
  • B. Environment Details
  • C. Experimental Details

Knowls

  1. Knowl 1 — CORRO uses transition-level task representations to reduce behavior-policy confounding

    model/method

    In fully offline meta-reinforcement learning, a dataset’s transition distribution reflects both its task and the behavior policy that collected it. An encoder trained on whole trajectories can therefore identify a task from behavior-policy patterns and fail when the context-collection policy changes. CORRO addresses this by learning task features from individual transition tuples (s,a,r,s′)(s,a,r,s') rather than trajectories, then aggregating those features into a representation used by the policy and critic. The transition-level design is intended to retain information about task rewards and dynamics while reducing reliance on trajectory-level behavior signatures.

  2. Knowl 2 — CORRO’s bi-level encoder aggregates transition embeddings with learned attention

    model/method

    For a context c={(si,ai,ri,si′)}i=1kc=\{(s_i,a_i,r_i,s'_i)\}_{i=1}^k, CORRO first applies a transition encoder Eθ1E_{\theta_1} to each tuple, producing zi=Eθ1(si,ai,ri,si′)z_i=E_{\theta_1}(s_i,a_i,r_i,s'_i). An aggregator Eθ2E_{\theta_2} then forms a task representation as a learned weighted sum:

    z=∑j=1ksoftmax⁡({MLP⁡(zi)}i=1k)jzj.z=\sum_{j=1}^{k}\operatorname{softmax}\bigl(\{\operatorname{MLP}(z_i)\}_{i=1}^{k}\bigr)_j z_j.

    Here kk is the number of context transitions, ziz_i is the embedding of transition ii, and the softmax weight for index jj is the normalized attention score assigned to that embedding. The transition encoder is trained with the contrastive objective; the aggregator is trained with the offline reinforcement-learning actor and critic. The policy πϕ(a∣s,z)\pi_\phi(a\mid s,z) and critic Qψ(s,a,z)Q_\psi(s,a,z) are conditioned on the resulting task representation.

  3. Knowl 3 — Task representation learning is formulated as mutual-information maximization

    model/method

    CORRO models the transition encoder as a probabilistic mapping z∼p(z∣x)z\sim p(z\mid x), where x=(s,a,r,s′)x=(s,a,r,s') is a transition generated under a task and a behavior policy. Its intended objective is to make the representation informative about the task while avoiding task-irrelevant information carried by the data-collection policy:

    max⁡I(Z;M)=Ez,M[log⁡p(M∣z)p(M)].\max I(Z;M)=\mathbb{E}_{z,M}\left[\log\frac{p(M\mid z)}{p(M)}\right].

    Here MM is the task random variable, ZZ is the encoded representation, and I(Z;M)I(Z;M) is their mutual information. Because this objective is intractable to optimize directly, CORRO uses a contrastive lower bound. The method’s goal is robustness to changes in behavior policy; the objective does not assume that behavior policies are known.

  4. Knowl 4 — A finite-task contrastive lower bound connects task information to cross-task transitions

    theoretical result

    Let T\mathcal{T} be a finite set of NN tasks drawn from a task distribution, and let M∈TM\in\mathcal{T} be the task generating x=(s,a,r,s′)x=(s,a,r,s'). The state-action pair (s,a)(s,a) may follow an arbitrary distribution; under task MM, r=RM(s,a)r=R_M(s,a) and s′∼TM(⋅∣s,a)s'\sim T_M(\cdot\mid s,a). Draw z∼p(z∣x)z\sim p(z\mid x) and define h(x,z)=p(z∣x)/p(z)h(x,z)=p(z\mid x)/p(z). For each candidate task M∗∈TM^*\in\mathcal{T}, construct x∗=(s,a,r∗,s′∗)x^*=(s,a,r^*,s'^*) using the same (s,a)(s,a), with reward and next state generated under M∗M^*. The paper establishes:

    I(Z;M)−log⁡N  ≥  EM,x,z[log⁡h(x,z)∑M∗∈Th(x∗,z)].I(Z;M)-\log N\;\geq\;\mathbb{E}_{M,x,z}\left[\log\frac{h(x,z)}{\sum_{M^*\in\mathcal{T}}h(x^*,z)}\right].

    The result provides a lower bound on task-representation mutual information using comparisons between a transition and transitions from candidate tasks at a shared state-action pair.

  5. Knowl 5 — CORRO trains the transition encoder with same-task positives and state-action-matched negatives

    equation

    The tractable contrastive objective approximates the density ratio in the finite-task lower bound with the exponential of a similarity score. For each training task MiM_i with offline dataset XiX_i, sample two transitions x,x′∈Xix,x'\in X_i and encode them as z,z′z,z'. For each candidate task M∗M^*, obtain a representation z∗z^* from a transition generated at the same state-action pair as xx; for M∗=MiM^*=M_i, set z∗=z′z^*=z'. CORRO maximizes:

    max⁡θ1  ∑Mi∈T∑x,x′∈Xilog⁡exp⁡(S(z,z′))∑M∗∈Texp⁡(S(z,z∗)).\max_{\theta_1}\;\sum_{M_i\in\mathcal{T}}\sum_{x,x'\in X_i} \log\frac{\exp(S(z,z'))}{\sum_{M^*\in\mathcal{T}}\exp(S(z,z^*))}.

    Here SS is a similarity score between embeddings, and θ1\theta_1 parameterizes the transition encoder. The same-task pair (x,x′)(x,x') is positive; transitions from other tasks at the matched (s,a)(s,a) provide negatives. The paper uses cosine similarity for SS. Comparing tasks at the same state-action pair is intended to make reward and next-state differences, rather than state-action visitation differences, informative for distinguishing tasks.

  6. Knowl 6 — A pooled conditional VAE generates negative transitions when task data overlap

    model/method

    When state-action pairs occur across multiple task datasets, CORRO fits a conditional variational autoencoder to the union of offline datasets. The CVAE has a latent variable uu with Gaussian prior p(u)p(u), an encoder qω(u∣s,a,r,s′)q_\omega(u\mid s,a,r,s'), and a generator pξ(r,s′∣s,a,u)p_\xi(r,s'\mid s,a,u). It is trained by minimizing the negative evidence lower bound over pooled offline transitions:

    LCVAE=−E(s,a,r,s′)∼Dpool[Eu∼qω(⋅∣s,a,r,s′)log⁡pξ(r,s′∣s,a,u)−KL(qω(u∣s,a,r,s′)∥p(u))].\mathcal{L}_{\mathrm{CVAE}}=-\mathbb{E}_{(s,a,r,s')\sim D_{\mathrm{pool}}}\left[ \mathbb{E}_{u\sim q_\omega(\cdot\mid s,a,r,s')}\log p_\xi(r,s'\mid s,a,u) -\mathrm{KL}\bigl(q_\omega(u\mid s,a,r,s')\Vert p(u)\bigr)\right].

    Here DpoolD_{\mathrm{pool}} is the union of the training datasets, and the KL term limits information retained in the latent variable. At contrastive-training time, the generator supplies negative (r,s′)(r,s') outcomes conditioned on a positive transition’s (s,a)(s,a). This pooled modeling can reduce prediction error relative to fitting separate models when state-action coverage overlaps and is similar across tasks. With little overlap, the CVAE may predict only the outcome of the task in which a state-action pair was observed and fail to provide diverse negatives.

  7. Knowl 7 — Reward randomization supplies diverse negatives when tasks differ in reward

    model/method

    If tasks differ in reward functions and their state-action coverage has little overlap, CORRO can generate a negative transition by perturbing the observed reward while retaining the other transition components: r∗=r+νr^*=r+\nu, where ν∼p(ν)\nu\sim p(\nu). The perturbation produces a large supply of diverse reward values, but it is not an approximation to the true cross-task distribution of negative transitions. In the reported experiments, ν∼N(0,0.5)\nu\sim\mathcal{N}(0,0.5) for Point-Robot and Ant-Dir. The approach is suited to reward-varying tasks; it does not provide alternative next states for tasks whose differences lie in transition dynamics.

  8. Knowl 8 — CORRO training separates negative-pair modeling, representation learning, and offline policy learning

    algorithm

    The training procedure takes per-task offline datasets, the selected negative-generation strategy, a transition encoder, an aggregator, a critic, and a policy. It has three stages: (1) if using generative negatives, train the CVAE on pooled transitions with the conditional VAE loss; (2) repeatedly sample a task and same-task positive transitions, generate negative transitions for candidate tasks using the CVAE or reward perturbation, and update the transition encoder to maximize the contrastive objective; (3) sample an offline task dataset and a context, encode its transitions, aggregate them into zz, and train the aggregator, critic, and policy with an offline RL algorithm. At test time, the context is encoded and aggregated once to condition the policy during the episode.

    For the reported meta-training runs, the task batch size was 16, the contrastive batch size 64, the number of negative pairs 16, the offline RL batch size 256, the training budget 2×1052\times10^5 steps, and the learning rate 3×10−43\times10^{-4}. The task-representation dimension was 5 for Point-Robot, Ant-Dir, and Half-Cheetah-Vel; 32 for Walker-Param; and 40 for Hopper-Param. Contexts used in training contained 200 transition tuples. The paper does not specify a separate numerical stopping budget for CVAE pretraining.

  9. Knowl 9 — CORRO improves OOD context-policy returns over prior context-based offline methods

    empirical result

    The evaluation compares mean test return under in-distribution (IID) contexts and out-of-distribution (OOD) contexts. IID contexts come from the same distribution as training data. OOD contexts are collected using saved behavior-policy checkpoints trained on arbitrary tasks, so their context distribution is unseen during meta-training. Values below are mean return ±\pm standard deviation; each method is evaluated on test tasks, using models from the final training epoch.

    Environment Supervised IID Supervised OOD Offline PEARL IID Offline PEARL OOD FOCAL IID FOCAL OOD CORRO IID CORRO OOD
    Point-Robot -4.89±\pm0.10 -5.84±\pm0.14 -5.4±\pm0.17 -6.74±\pm0.19 -6.06±\pm0.42 -7.34±\pm0.20 -5.19±\pm0.05 -6.39±\pm0.05
    Ant-Dir 136±\pm17.6 131.7±\pm11.4 155.4±\pm24.4 141.5±\pm11.3 109.8±\pm12.8 53.5±\pm16.4 156.8±\pm35.2 154.7±\pm25.8
    Half-Cheetah-Vel -31.6±\pm0.7 -32.1±\pm0.9 -31.2±\pm0.5 -242.7±\pm6.0 -38.0±\pm4.0 -204.1±\pm9.5 -33.7±\pm1.1 -89.7±\pm7.4
    Walker-Param 232.7±\pm29.2 221.2±\pm43.4 259.1±\pm48.2 254.7±\pm35.8 225.4±\pm56.4 193.3±\pm151.5 301.5±\pm37.9 284.0±\pm19.3
    Hopper-Param 269.2±\pm20.3 251.9±\pm28.8 244.0±\pm18.5 236.6±\pm18.5 195.6±\pm62.3 199.7±\pm51.9 267.6±\pm25.6 268.0±\pm13.8

    CORRO exceeds both Offline PEARL and FOCAL in OOD return in all five environments. The contrast is especially large in Half-Cheetah-Vel, where Offline PEARL and FOCAL returns fall to −242.7-242.7 and −204.1-204.1, while CORRO reaches −89.7-89.7. The supervised method, which uses ground-truth task descriptions during training, scores higher than CORRO on OOD Point-Robot and Half-Cheetah-Vel; it is not an ordinary label-free meta-RL baseline.

  10. Knowl 10 — Offline task-adaptation benchmarks use five task families and fixed offline SAC datasets

    experimental setup

    Experiments use 20 training tasks and 20 test tasks per environment, with five random seeds. For each task, SAC is trained independently and its replay buffer supplies the offline dataset; meta-training uses no environment interaction. Training contexts are contiguous segments of 200 transitions from the corresponding task dataset, and baseline policy-learning methods and CORRO use SAC with fixed policy-learning hyperparameters.

    Point-Robot varies a goal gg uniformly over [−1,1]2[-1,1]^2, starts at (0,0)(0,0), and rewards negative Euclidean distance to the goal; its episode limit is 20 steps. Ant-Dir varies a direction θ∼U[0,2π]\theta\sim U[0,2\pi] and rewards horizontal velocity projected onto that direction. Half-Cheetah-Vel varies target velocity vg∼U[0,3]v_g\sim U[0,3] and rewards −∣vt−vg∣−12∥at∥22-|v_t-v_g|-\tfrac12\|a_t\|_2^2. Ant-Dir and Half-Cheetah-Vel have 200-step episode limits. Walker-Param varies 32 body-mass and friction parameters; Hopper-Param varies 41 body-mass, inertia, damping, and friction parameters. In both parameterized environments, each parameter is the default value multiplied by 1.5μ1.5^\mu, with μ∼U[−3,3]\mu\sim U[-3,3]. Their rewards are forward velocity minus an action penalty (Walker-Param also adds 1); episodes have a 200-step limit, with Walker-Param terminating early if the walker’s height is below 0.5.

  11. Knowl 11 — IID adaptation favors CORRO on dynamics-varying tasks and Ant-Dir

    empirical result

    For the IID adaptation evaluation, a context trajectory is sampled from the pretrained replay buffer of the same test task, so the context-collection distribution matches meta-training. CORRO and supervised task learning outperform the other methods on Point-Robot. CORRO outperforms all compared methods, including supervised task learning, on Ant-Dir. In Walker-Param and Hopper-Param, where tasks vary in transition dynamics, CORRO has higher adaptation returns than the baselines; the paper also reports more stable policy improvement than FOCAL, whose performance sometimes declines during offline training. In Half-Cheetah-Vel, the plotted IID returns do not clearly separate methods.

  12. Knowl 12 — Negative-pair ablations show that useful negatives must be both diverse and accurate

    data/table

    This ablation compares negative-pair generation methods by contrastive loss and IID/OOD return in Half-Cheetah-Vel and Point-Robot. Returns are mean ±\pm standard deviation. Relabeling fits separate reward and transition models per dataset; None uses transitions from other tasks without matching their state-action pairs. The proposed generators are more effective than these alternatives in the settings where their assumptions fit: generative modeling performs well in Half-Cheetah-Vel, while reward randomization is strongest in Point-Robot. Relabeling has low contrastive loss but poor task performance, which the authors attribute to inaccurate predictions on unseen data. In Point-Robot, behavior policies lead to little state-action overlap, and the CVAE does not provide sufficiently diverse negatives.

    Environment Method Contrastive loss IID return OOD return
    Half-Cheetah-Vel Generative 0.07 -33.7±\pm1.1 -89.7±\pm7.4
    Half-Cheetah-Vel Randomize 0.83 -34.3±\pm1.5 -84.5±\pm1.3
    Half-Cheetah-Vel Relabeling 0.04 -40.8±\pm1.5 -245.3±\pm12.9
    Half-Cheetah-Vel None 1.20 -34.1±\pm2.4 -97.6±\pm3.1
    Point-Robot Generative 2.83 -9.41±\pm0.42 -9.42±\pm0.42
    Point-Robot Randomize 0.54 -5.19±\pm0.05 -6.39±\pm0.05
    Point-Robot Relabeling 0.04 -9.22±\pm0.24 -9.27±\pm0.22
    Point-Robot None 1.46 -5.24±\pm0.27 -6.52±\pm0.08
  13. Knowl 13 — Random-policy contexts preserve CORRO’s advantage on three benchmarks

    empirical result

    The authors also evaluate adaptation with contexts collected by a uniformly random exploration policy. The reported values are mean test returns:

    Environment Supervised Offline PEARL FOCAL CORRO
    Point-Robot -5.32±\pm0.20 -7.06±\pm0.99 -8.64±\pm0.26 -5.59±\pm0.57
    Ant-Dir 149.8±\pm20.5 148.4±\pm35.3 89.8±\pm8.7 163.0±\pm35.8
    Half-Cheetah-Vel -37.6±\pm0.8 -35.4±\pm1.8 -41.6±\pm3.3 -42.9±\pm0.7
    Walker-Param 221.7±\pm91.1 276.6±\pm37.7 245.6±\pm67.8 300.5±\pm34.2
    Hopper-Param 253.5±\pm21.2 245.9±\pm18.9 203.6±\pm46.6 273.3±\pm3.9

    CORRO has the highest reported return among these methods on Ant-Dir, Walker-Param, and Hopper-Param. The supervised baseline is higher on Point-Robot, and CORRO is lower than the other three methods on Half-Cheetah-Vel in this evaluation.

  14. Knowl 14 — The study does not learn an exploration policy or test interaction-assisted training

    limitation

    CORRO uses an arbitrary policy to collect the adaptation context but does not learn an exploration policy. The study also does not evaluate moderate environment interaction as a supplement when the offline training data are extremely limited. These are identified as unaddressed directions, so the experiments establish robustness of task inference under tested context-policy shifts, not the ability to discover an effective exploration strategy or recover from severely limited offline coverage.

Coverage note — The qualitative t-SNE visualization of Half-Cheetah-Vel embeddings is omitted because it is supplementary evidence for task separation and adds no independent quantitative result; background material and proof derivations are also excluded.

References

  1. 1.Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P. Continuous Adaptation Via Meta-Learning in Nonstationary and Competitive Environments. In International Conference on Learning Representations, 2018.
  2. 2.Anand, A., Racah, E., Ozair, S., Bengio, Y., Cotê, M.-A., and Hjelm, R. D. Unsupervised State Representation Learning in Atari. In Advances in neural information processing systems, 2019.
  3. 3.Anonymous. Model-Based Offline Meta-Reinforcement Learning with Regularization. In Submitted to The Tenth International Conference on Learning Representations, 2022. under review.
  4. 4.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A Simple Framework for Contrastive Learning of Visual Representations. In International Conference on Machine Learning, 2020.
  5. 5.Dorfman, R., Shenfeld, I., and Tamar, A. Offline Meta Learning of Exploration. arXiv preprint arXiv:2008.02598, 2020.
  6. 6.Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. RL2: Fast Reinforcement Learning Via Slow Reinforcement Learning. arXiv preprint arXiv:1611.02779, 2016.
  7. 7.Fakoor, R., Chaudhari, P., Soatto, S., and Smola, A. J. Meta-Q-Learning. In International Conference on Learning Representations, 2020.
  8. 8.Finn, C., Abbeel, P., and Levine, S. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In International Conference on Machine Learning, 2017.
  9. 9.Foerster, J., Farquhar, G., Al-Shedivat, M., Rocktaschel, T., Xing, E., and Whiteson, S. DiCE: The Infinitely Differentiable Monte Carlo Estimator. In International Conference on Machine Learning, 2018.
  10. 10.Fujimoto, S., Conti, E., Ghavamzadeh, M., and Pineau, J. Benchmarking Batch Deep Reinforcement Learning Algorithms. arXiv preprint arXiv:1910.01708, 2019a.
  11. 11.Fujimoto, S., Meger, D., and Precup, D. Off-Policy Deep Reinforcement Learning Without Exploration. In International Conference on Machine Learning, 2019b.
  12. 12.Grill, J.-B., Strub, F., Altche, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Michal, V. Bootstrap Your Own Latent: A New Approach To Self-Supervised Learning. In Advances in neural information processing systems, 2020.
  13. 13.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with A Stochastic Actor. In International Conference on Machine Learning, 2018.
  14. 14.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum Contrast for Unsupervised Visual Representation Learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  15. 15.Houthooft, R., Chen, R. Y., Isola, P., Stadie, B. C., Wolski, F., Ho, J., and Abbeel, P. Evolved Policy Gradients. In Advances in neural information processing systems, 2018.
  16. 16.Kumar, A., Fu, J., Tucker, G., and Levine, S. Stabilizing Off-Policy Q-Learning Via Bootstrapping Error Reduction. In Advances in Neural Information Processing Systems, 2019.
  17. 17.Lample, G. and Chaplot, D. S. Playing FPS Games with Deep Reinforcement Learning. In AAAI Conference on Artificial Intelligence, 2017.
  18. 18.Lange, S., Gabel, T., and Riedmiller, M. Batch Reinforcement Learning. In Reinforcement learning, pp. 45–73. Springer, 2012.
  19. 19.Laskin, M., Srinivas, A., and Abbeel, P. Curl: Contrastive Unsupervised Representations for Reinforcement Learning. In International Conference on Machine Learning, 2020.
  20. 20.Levine, S., Kumar, A., Tucker, G., and Fu, J. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. arXiv preprint arXiv:2005.01643, 2020.
  21. 21.Li, J., Vuong, Q., Liu, S., Liu, M., Ciosek, K., Ross, K., Christensen, H. I., and Su, H. Multi-Task Batch Reinforcement Learning with Metric Learning. In International Conference on Learning Representations, 2020.
  22. 22.Li, L., Huang, Y., Chen, M., Luo, S., Luo, D., and Huang, J. Provably Improved Context-Based Offline Meta-RL with Attention and Contrastive Learning. arXiv preprint arXiv:2102.10774, 2021a.
  23. 23.Li, L., Yang, R., and Luo, D. FOCAL: Efficient Fully-Offline Meta-Reinforcement Learning Via Distance Metric Learning and Behavior Regularization. In International Conference on Learning Representations, 2021b.
  24. 24.Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. Continuous Control with Deep Reinforcement Learning. In International Conference on Learning Representations, 2016.
  25. 25.Liu, E. Z., Raghunathan, A., Liang, P., and Finn, C. Explore Then Execute: Adapting Without Rewards Via Factorized Meta-Reinforcement Learning. arXiv preprint arXiv:2008.02790, 2020a.
  26. 26.Liu, Y., Yi, L., Zhang, S., Fan, Q., Funkhouser, T., and Dong, H. P4Contrast: Contrastive Learning with Pairs of Point-Pixel Pairs for RGB-D Scene Understanding. arXiv preprint arXiv:2012.13089, 2020b.
  27. 27.Maaten, L. v. d. and Hinton, G. Visualizing Data Using T-SNE. Journal of Machine Learning Research, 9(11): 2579–2605, 2008.
  28. 28.Mitchell, E., Rafailov, R., Peng, X. B., Levine, S., and Finn, C. Offline Meta-Reinforcement Learning with Advantage Weighting. In International Conference on Machine Learning, 2021.
  29. 29.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. Playing Atari with Deep Reinforcement Learning. arXiv preprint arXiv:1312.5602, 2013.
  30. 30.Nguyen, H. and La, H. Review of Deep Reinforcement Learning for Robot Manipulation. In IEEE International Conference on Robotic Computing, 2019.
  31. 31.Oord, A. v. d., Li, Y., and Vinyals, O. Representation Learning with Contrastive Predictive Coding. arXiv preprint arXiv:1807.03748, 2018.
  32. 32.OroojlooyJadid, A. and Hajinezhad, D. A Review of Cooperative Multi-Agent Deep Reinforcement Learning. arXiv preprint arXiv:1908.03963, 2019.
  33. 33.Parisotto, E., Ghosh, S., Yalamanchi, S. B., Chinnaobireddy, V., Wu, Y., and Salakhutdinov, R. Concurrent Meta Reinforcement Learning. arXiv preprint arXiv:1903.02710, 2019.
  34. 34.Patrick, M., Asano, Y. M., Kuznetsova, P., Fong, R., Henriques, J. F., Zweig, G., and Vedaldi, A. Multi-Modal Self-Supervision From Generalized Data Transformations. arXiv preprint arXiv:2003.04298, 2020.
  35. 35.Peng, X. B., Kumar, A., Zhang, G., and Levine, S. Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning. arXiv preprint arXiv:1910.00177, 2019.
  36. 36.Pong, V. H., Nair, A., Smith, L., Huang, C., and Levine, S. Offline Meta-Reinforcement Learning with Online Self-Supervision. arXiv preprint arXiv:2107.03974, 2021.
  37. 37.Rakelly, K., Zhou, A., Quillen, D., Finn, C., and Levine, S. Efficient Off-Policy Meta-Reinforcement Learning Via Probabilistic Context Variables. In International Conference on Machine Learning, 2019.
  38. 38.Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P. ProMP: Proximal Meta-Policy Search. In International Conference on Learning Representations, 2019.
  39. 39.Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G. Time-Contrastive Networks: Self-Supervised Learning From Video. In IEEE International Conference on Robotics and Automation, 2018.
  40. 40.Sohn, K., Lee, H., and Yan, X. Learning Structured Output Representation using Deep Conditional Generative Models. In Advances in neural information processing systems, 2015.
  41. 41.Stadie, B. C., Yang, G., Houthooft, R., Chen, X., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I. Some Considerations on Learning To Explore Via Meta-Reinforcement Learning. arXiv preprint arXiv:1803.01118, 2018.
  42. 42.Stooke, A., Lee, K., Abbeel, P., and Laskin, M. Decoupling Representation Learning From Reinforcement Learning. In International Conference on Machine Learning, 2021.
  43. 43.Tian, Y., Krishnan, D., and Isola, P. Contrastive Multiview Coding. In European Conference on Computer Vision, 2020.
  44. 44.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is All You Need. In Advances in neural information processing systems, 2017.
  45. 45.Wang, T. and Isola, P. Understanding Contrastive Representation Learning Through Alignment and Uniformity on The Hypersphere. In International Conference on Machine Learning, 2020.
  46. 46.Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N. Dueling Network Architectures for Deep Reinforcement Learning. In International Conference on Machine Learning, 2016.
  47. 47.Wu, Y., Tucker, G., and Nachum, O. Behavior Regularized Offline Reinforcement Learning. arXiv preprint arXiv:1911.11361, 2019.
  48. 48.Zhang, J., Wang, J., Hu, H., Chen, Y., Fan, C., and Zhang, C. Learn to effectively explore in context-based meta-rl. arXiv preprint arXiv:2006.08170, 2020.
  49. 49.Zheng, G., Zhang, F., Zheng, Z., Xiang, Y., Yuan, N. J., Xie, X., and Li, Z. DRN: A Deep Reinforcement Learning Framework for News Recommendation. In World Wide Web Conference, 2018.
  50. 50.Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S. VariBAD: A Very Good Method for Bayes-Adaptive Deep RL Via Meta-Learning. In International Conference on Learning Representations, 2020.
  51. 51.Łukasz Kaiser, Babaeizadeh, M., Miłos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H. Model-Based Reinforcement Learning for Atari. In International Conference on Learning Representations, 2020.

Citation

MLA
Yuan, H., and Z. Lu. “Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning”. International Conference on Machine Learning, vol. 162, 2022, pp. 25747–59, https://proceedings.mlr.press/v162/yuan22a.html.
APA
Yuan, H., & Lu, Z. (2022). Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning. International Conference on Machine Learning, 162, 25747–25759. https://proceedings.mlr.press/v162/yuan22a.html
Chicago
Yuan, H., and Z. Lu. 2022. “Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning”. International Conference on Machine Learning 162: 25747–59. https://proceedings.mlr.press/v162/yuan22a.html.
Harvard
Yuan, H. and Lu, Z. (2022) “Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning”, International Conference on Machine Learning. PMLR, pp. 25747–25759. Available at: https://proceedings.mlr.press/v162/yuan22a.html.
Vancouver
1. Yuan H, Lu Z (2022) Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning. In: International Conference on Machine Learning. PMLR, pp 25747–25759

BibTeX

@InProceedings{pmlr-v162-yuan22a,
  title = 	 {Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning},
  author =       {Yuan, Haoqi and Lu, Zongqing},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {25747--25759},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/yuan22a/yuan22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/yuan22a.html},
  abstract = 	 {We study offline meta-reinforcement learning, a practical reinforcement learning paradigm that learns from offline data to adapt to new tasks. The distribution of offline data is determined jointly by the behavior policy and the task. Existing offline meta-reinforcement learning algorithms cannot distinguish these factors, making task representations unstable to the change of behavior policies. To address this problem, we propose a contrastive learning framework for task representations that are robust to the distribution mismatch of behavior policies in training and test. We design a bi-level encoder structure, use mutual information maximization to formalize task representation learning, derive a contrastive learning objective, and introduce several approaches to approximate the true distribution of negative pairs. Experiments on a variety of offline meta-reinforcement learning benchmarks demonstrate the advantages of our method over prior methods, especially on the generalization to out-of-distribution behavior policies.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/