Built independently by an author, for readers. Read the story and support ChapterPal

keyword

context-based meta-RL

Context-based meta-reinforcement learning is a reinforcement learning paradigm in which an agent rapidly adapts to new tasks by encoding past transition history into a latent task representation. Rather than updating policy parameters through gradient descent during evaluation, this approach utilizes a context encoder to map recent experiences, consisting of states, actions, and rewards, into a compact embedding that captures the underlying task identity. A conditioned policy and value network then take this task embedding alongside the current environment state to choose appropriate actions, enabling efficient, feedforward adaptation across different task dynamics and reward functions in both online and offline settings.

1 item

Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning

Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive Learning

Haoqi Yuan, Zongqing Lu

OrganizationsCarnegie Mellon UniversityPeking University

Why you should read this

Presents a contrastive learning framework with a bi-level encoder that decouples task characteristics from behavior policies in offline meta-reinforcement learning, enabling reliable task adaptation even under severe distribution shifts during test time.

We study offline meta-reinforcement learning, a practical reinforcement learning paradigm that learns from offline data to adapt to new tasks. The distribution of offline data is determined jointly by the behavior policy and the task. Existing offline meta-reinforcement learning algorithms cannot distinguish these factors, making task representations unstable to the change of behavior policies. To address this problem, we propose a contrastive learning framework for task representations that are robust to the distribution mismatch of behavior policies in training and test. We design a bi-level encoder structure, use mutual information maximization to formalize task representation learning, derive a contrastive learning objective, and introduce several approaches to approximate the true distribution of negative pairs. Experiments on a variety of offline meta-reinforcement learning benchmarks demonstrate the advantages of our method over prior methods, especially on the generalization to out-of-distribution behavior policies.

Added

2026-10-03