Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning
Haoqi YuanZongqing Lu
Presents a contrastive learning framework with a bi-level encoder that decouples task characteristics from behavior policies in offline meta-reinforcement learning, enabling reliable task adaptation even under severe distribution shifts during test time.
Real-world applications of deep reinforcement learning—such as robotics, recommendation systems, and autonomous control—often face severe data collection costs and safety risks when exploring environments online. Offline meta-reinforcement learning addresses these hurdles by training agents entirely on pre-collected, multi-task datasets so they can quickly adapt to unseen tasks. However, existing context-based approaches struggle when the behavior policy used to collect offline data differs from the exploratory behavior encountered during testing. Because standard models encode full multi-step trajectories, they often memorize features of the data-collection policy instead of learning the true underlying task, causing severe performance drops during deployment.
The article introduces and evaluates CORRO (Contrastive Robust Task Representation Learning), a framework designed to learn task representations that remain stable and accurate even when testing policies deviate significantly from training data. Its main objective is to demonstrate that isolating single-step transitions and optimizing a mutual information contrastive objective enables robust task identification across both varying reward functions and changing environment physics.
To achieve this, the article develops a bi-level task encoder. Instead of processing entire trajectories, the transition encoder processes individual transition steps (state, action, reward, next state) to extract task-relevant signals while stripping out policy-specific patterns. These latent codes are then pooled using an attention-based aggregator to condition the agent’s policy. The encoder is trained via a contrastive objective based on mutual information maximization, differentiating true transitions from synthetically generated negative pairs via generative modeling or reward randomization. The framework was evaluated across multiple continuous control benchmarks involving 20 training and 20 testing tasks across varying physical dynamics and reward conditions.
The experimental findings highlight substantial improvements over existing methods. First, CORRO achieves superior task adaptation performance under standard testing conditions, outperforming prior offline meta-learning approaches and reaching top returns within 20,000 offline training steps. Second, it demonstrates significant robustness against out-of-distribution exploratory behaviors: while baseline methods degraded catastrophically (for example, dropping to returns of -204.1 and -242.7 in velocity tracking tasks), CORRO maintained a stable return of -89.7. Third, it successfully learned structured latent representations that cleanly separate distinct tasks along physical and reward dimensions without requiring access to true task labels during training. Fourth, ablation testing confirms that encoding single-step transitions rather than trajectories is the primary factor driving resistance to policy mismatch.
These findings indicate that offline meta-reinforcement learning can be deployed safely and reliably even when real-world testing environments exhibit unpredictable exploration patterns. By preventing models from relying on spurious behavioral correlations, practitioners can reduce the risk of unexpected deployment failures in safety-critical robotics and control systems. The results also show that high sample diversity in negative pair generation is essential for effective contrastive learning in offline data regimes.
Organizations developing offline reinforcement learning systems should transition from trajectory-level task encoders to transition-level contrastive architectures. When configuring negative sample generation, practitioners should select generative modeling when state-action coverage overlaps across tasks, and switch to reward randomization when task policies explore disjoint state spaces. Moving forward, additional work is recommended to integrate learned exploration policies and explore minimal-interaction hybrid training when offline data is severely limited.
The reported findings are backed by consistent multi-seed experiments across standard continuous control benchmarks. However, confidence should be tempered by the fact that evaluations were conducted in simulated physical environments with pre-collected datasets rather than live physical hardware. Careful pilot testing remains necessary before applying the framework to highly complex, noisy, or unmodeled real-world environments.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). Its account of offline RL’s distribution-shift and dataset-coverage challenges supplies the setting CORRO addresses.
- Paper: Offline Meta-Reinforcement Learning with Online Self-Supervision, Vitchyr H. Pong et al. (2022). It establishes the offline meta-RL adaptation and context-distribution-shift problem that CORRO tackles with more robust task representations.
- Paper: CURL: Contrastive Unsupervised Representations for Reinforcement Learning, Aravind Srinivas et al. (2020). Its contrastive representation-learning approach for RL provides useful groundwork for understanding CORRO’s transition-level contrastive objective.
- Paper: Contrastive Learning as Goal-Conditioned Reinforcement Learning, Benjamin Eysenbach et al. (2022). It connects contrastive objectives to RL value learning, clarifying the representation-learning principle CORRO repurposes for task inference.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). Its offline-RL benchmark and dataset framework helps explain the evaluation setting CORRO uses to test robustness across data-collection behaviors.
No sufficiently relevant recommendations were found.
