Contrastive Learning as Goal-Conditioned Reinforcement Learning

Benjamin EysenbachTianjun ZhangSergey LevineRuslan Salakhutdinov

article2022NeurIPS314 citations

Proves that contrastive representation learning applied to action-labeled trajectories mathematically equates to learning a goal-conditioned value function, yielding a simpler and more effective reinforcement learning algorithm that outperforms standard approaches on vision-based and offline tasks without auxiliary losses.

Listen

Training automated systems to achieve specific goals via reinforcement learning often requires effective internal representations of the environment. Historically, learning these representations directly alongside decision-making policies has proven fragile and unstable, prompting engineers to rely on auxiliary perceptual objectives, manual reward shaping, or artificial data augmentations. The article addresses this core inefficiency by demonstrating that contrastive representation learning—a technique that maps similar inputs together while separating dissimilar ones—can serve directly as a goal-conditioned reinforcement learning algorithm without requiring separate perceptual machinery.

To evaluate this approach, the authors developed a mathematical framework proving that contrastive learning over action-labeled trajectories estimates a standard goal-conditioned value function. They tested this method across a suite of robotic manipulation and navigation benchmarks, examining both state-based and vision-based inputs across online, offline, and partially observed settings. The experimental setups compared the contrastive approach against standard actor-critic baselines, behavioral cloning, and model-based techniques, measuring success rates and computational throughput.

Across the evaluated tasks, contrastive reinforcement learning consistently matched or exceeded the performance of existing methods. On vision-based tasks, it substantially outperformed traditional actor-critic methods equipped with autoencoders or image augmentations, succeeding on complex manipulation tasks where baseline methods failed entirely. In offline settings where the agent cannot collect new data, the contrastive method exceeded baseline performance on five out of six benchmark environments, achieving a 7% to 9% absolute improvement over top-tier offline baselines on the most difficult navigation tasks. Furthermore, computational training ran nearly four times faster than leading data-augmented reinforcement learning implementations.

These findings indicate that treating representation learning as the decision-making engine itself eliminates the need for ad-hoc perceptual modules, reducing system complexity and computational overhead. Organizations deploying autonomous agents can streamline development pipelines by removing manually engineered reward functions and vision-specific augmentations while achieving higher task reliability.

For engineering teams developing goal-directed autonomous systems, the article supports adopting contrastive architectures to simplify training infrastructure and improve policy success. However, stakeholders should note that the current theoretical proofs and empirical evaluations focus strictly on goal-reaching tasks rather than arbitrary reward-maximization problems. Additional validation in non-goal settings and real-world physical platforms is recommended before broad operational deployment.

arXiv: 2206.07568
Cover for Contrastive Learning as Goal-Conditioned Reinforcement Learning

Abstract

In reinforcement learning (RL), it is easier to solve a task if given a good representation. While deep RL should automatically acquire such good representations, prior work often finds that learning representations in an end-to-end fashion is unstable and instead equip RL algorithms with additional representation learning parts (e.g., auxiliary losses, data augmentation). How can we design RL algorithms that directly acquire good representations? In this paper, instead of adding representation learning parts to an existing RL algorithm, we show (contrastive) representation learning methods can be cast as RL algorithms in their own right. To do this, we build upon prior work and apply contrastive representation learning to action-labeled trajectories, in such a way that the (inner product of) learned representations exactly corresponds to a goal-conditioned value function. We use this idea to reinterpret a prior RL method as performing contrastive learning, and then use the idea to propose a much simpler method that achieves similar performance. Across a range of goal-conditioned RL tasks, we demonstrate that contrastive RL methods achieve higher success rates than prior non-contrastive methods, including in the offline RL setting. We also show that contrastive RL outperforms prior methods on image-based tasks, without using data augmentation or auxiliary objectives. 1

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 Contrastive Learning as an RL Algorithm
  • 4.1 Relating the Q-function to probabilities
  • 4.2 Contrastive Learning Estimates a Q-Function
  • 4.3 Learning the Goal-Conditioned Policy
  • 4.4 A Complete Goal-Conditioned RL Algorithm
  • 4.5 Convergence Guarantees
  • 4.6 C-learning as Contrastive Learning
  • 5 Experiments
  • 5.1 Comparing to prior goal-conditioned RL methods
  • 5.2 Comparing to prior representation learning methods
  • 5.3 Probing the dimensions of contrastive RL
  • 5.4 Partial Observability and Moving Cameras
  • 5.5 Contrastive RL for Offline RL
  • 6 Conclusion
  • References
  • A Additional Related Work
  • B Proofs
  • B.1 Q-function are equivalent to the discounted state occupancy measure
  • B.2 Contrastive RL is Policy Improvement
  • C Contrastive RL (CPC)
  • D Contrastive RL (NCE + C-learning)
  • E Experimental Details
  • E.1 Environments
  • F Additional Experiments
  • F.1 Linear regression with the learned features
  • F.2 When is contrastive learning better than learning a foreward model?
  • F.3 Goals used in the actor loss
  • F.4 Transferring representations to solve new tasks
  • F.5 Robustness to Environment Perturbations
  • F.6 Additional figures
  • G Failed Experiments

Knowls

  1. Knowl 1 — Equivalence of Goal-Conditioned Q-Functions and Discounted State Occupancy Measures

    theoretical result

    In goal-conditioned reinforcement learning, consider an infinite-horizon Markov Decision Process with state space S\mathcal{S}, action space A\mathcal{A}, initial state distribution p0(s)p_0(s), transition dynamics p(st+1∣st,at)p(s_{t+1} \mid s_t, a_t), discount factor γ∈[0,1)\gamma \in [0, 1), and goal distribution pg(sg)p_g(s_g). When the goal-conditioned reward is defined as the transition density of reaching the goal state at the next step:

    rg(st,at)≜(1−γ)p(st+1=sg∣st,at)r_g(s_t, a_t) \triangleq (1 - \gamma) p(s_{t+1} = s_g \mid s_t, a_t)

    with initial state reward rg(s0,a0)=(1−γ)(p(s1=sg∣s0,a0)+p0(s0=sg))r_g(s_0, a_0) = (1 - \gamma)(p(s_1 = s_g \mid s_0, a_0) + p_0(s_0 = s_g)), the goal-conditioned action-value function under policy π(a∣s,sg)\pi(a \mid s, s_g):

    Qsgπ(s,a)≜Eπ(τ∣sg)[∑t′=t∞γt′−trg(st′,at′) ∣ st=s,at=a]Q_{s_g}^\pi(s, a) \triangleq \mathbb{E}_{\pi(\tau \mid s_g)}\left[ \sum_{t'=t}^\infty \gamma^{t'-t} r_g(s_{t'}, a_{t'}) \,\Bigg|\, s_t = s, a_t = a \right]

    is mathematically equivalent to the probability density of reaching state sgs_g under the policy's discounted state occupancy measure:

    Qsgπ(s,a)=pπ(⋅∣⋅,sg)(st+=sg∣s,a)Q_{s_g}^\pi(s, a) = p^{\pi(\cdot \mid \cdot, s_g)}(s_{t+} = s_g \mid s, a)

    where the discounted state occupancy distribution is defined as:

    pπ(⋅∣⋅,sg)(st+=s)≜(1−γ)∑t=0∞γtptπ(⋅∣⋅,sg)(st=s)p^{\pi(\cdot \mid \cdot, s_g)}(s_{t+} = s) \triangleq (1 - \gamma) \sum_{t=0}^\infty \gamma^t p_t^{\pi(\cdot \mid \cdot, s_g)}(s_t = s)

    and ptπ(⋅∣⋅,sg)(s)p_t^{\pi(\cdot \mid \cdot, s_g)}(s) is the state visitation distribution of policy π\pi at time step tt. This equivalence shows that estimating goal-conditioned value functions is equivalent to future state occupancy density estimation.

  2. Knowl 2 — Contrastive Learning of Goal-Conditioned Value Functions

    theoretical result

    Goal-conditioned value functions can be learned directly via binary noise-contrastive estimation (NCE) without temporal difference backups or explicit reward functions. Let state-action pairs (s,a)(s, a) be sampled from a replay buffer distribution p(s,a)p(s, a), and let positive future states sf+∼pπ(⋅∣⋅)(st+∣s,a)s_f^+ \sim p^{\pi(\cdot \mid \cdot)}(s_{t+} \mid s, a) be sampled from the average discounted state occupancy distribution:

    pπ(⋅∣⋅)(st+=s∣s,a)≜∫pπ(⋅∣⋅,sg)(st+=s∣s,a)pπ(sg∣s,a) dsgp^{\pi(\cdot \mid \cdot)}(s_{t+} = s \mid s, a) \triangleq \int p^{\pi(\cdot \mid \cdot, s_g)}(s_{t+} = s \mid s, a) p^\pi(s_g \mid s, a) \, d s_g

    Negative future states sf−∼p(st+)s_f^- \sim p(s_{t+}) are sampled from the marginal state occupancy distribution p(st+)=∫pπ(⋅∣⋅)(st+∣s,a)p(s,a) ds dap(s_{t+}) = \int p^{\pi(\cdot \mid \cdot)}(s_{t+} \mid s, a) p(s, a) \, ds \, da.

    A critic function f(s,a,sf)f(s, a, s_f) trained using the binary classification objective:

    max⁡fE(s,a)∼p(s,a), sf+∼pπ(⋅∣⋅)(st+∣s,a), sf−∼p(st+)[log⁡σ(f(s,a,sf+))+log⁡(1−σ(f(s,a,sf−)))]\max_f \mathbb{E}_{(s, a) \sim p(s, a), \, s_f^+ \sim p^{\pi(\cdot \mid \cdot)}(s_{t+} \mid s, a), \, s_f^- \sim p(s_{t+})} \left[ \log \sigma(f(s, a, s_f^+)) + \log(1 - \sigma(f(s, a, s_f^-))) \right]

    has the Bayes-optimal solution:

    f∗(s,a,sf)=log⁡pπ(⋅∣⋅)(st+=sf∣s,a)p(sf)f^*(s, a, s_f) = \log \frac{p^{\pi(\cdot \mid \cdot)}(s_{t+} = s_f \mid s, a)}{p(s_f)}

    Exponentiating the optimal critic recovers the goal-conditioned Q-function Qsfπ(⋅∣⋅)(s,a)Q_{s_f}^{\pi(\cdot \mid \cdot)}(s, a) up to a state-dependent partition constant:

    exp⁡(f∗(s,a,sf))=1p(sf)Qsfπ(⋅∣⋅)(s,a)\exp(f^*(s, a, s_f)) = \frac{1}{p(s_f)} Q_{s_f}^{\pi(\cdot \mid \cdot)}(s, a)

    Because 1/p(sf)1 / p(s_f) depends solely on the goal state sfs_f and is invariant to the action aa, maximizing f(s,a,sg)f(s, a, s_g) over actions directly maximizes the expected goal-conditioned return.

  3. Knowl 3 — Contrastive RL (NCE) Algorithm

    algorithm

    Contrastive RL (NCE) optimizes goal-conditioned policies and value functions end-to-end by parametrizing the critic as a bilinear representation between state-action pairs and goal states: f(s,a,sg)=ϕ(s,a)Tψ(sg)f(s, a, s_g) = \phi(s, a)^T \psi(s_g), where ϕ\phi is a state-action encoder and ψ\psi is a goal encoder.

    For each transition in a mini-batch, a positive future state sf+s_f^+ is sampled by drawing a geometric time offset t∼GEOM(1−γ)t \sim \text{GEOM}(1 - \gamma) along the same trajectory, while negative future states are formed by the future states of other batch elements. The policy πθ(a∣s,sg)\pi_\theta(a \mid s, s_g) is updated by ascending the gradient of the critic with respect to policy actions.

    import jax.numpy as jnp
    from optax import sigmoid_binary_cross_entropy
    
    def critic_loss(states, actions, future_states, sa_encoder, g_encoder):
        # states: (batch_size, state_dim)
        # actions: (batch_size, action_dim)
        # future_states: (batch_size, state_dim), sampled via geometric offset
        sa_repr = sa_encoder(states, actions)  # (batch_size, repr_dim)
        g_repr = g_encoder(future_states)      # (batch_size, repr_dim)
        logits = jnp.einsum('ik,jk->ij', sa_repr, g_repr)  # (batch_size, batch_size)
        labels = jnp.eye(states.shape[0])
        return sigmoid_binary_cross_entropy(logits=logits, labels=labels).mean()
    
    def actor_loss(states, goals, policy, sa_encoder, g_encoder):
        # states: (batch_size, state_dim)
        # goals: (batch_size, state_dim), random goals sampled from replay buffer
        actions = policy.sample(states, goal=goals)  # (batch_size, action_dim)
        sa_repr = sa_encoder(states, actions)        # (batch_size, repr_dim)
        g_repr = g_encoder(goals)                    # (batch_size, repr_dim)
        logits = jnp.einsum('ik,ik->i', sa_repr, g_repr)  # (batch_size,)
        return -1.0 * jnp.mean(logits)
    

    The bilinear structure allows goal representations ψ(sg)\psi(s_g) to be computed once per batch and shared across all positive and negative comparisons. The method does not require target networks, replay double Q-learning ensembles, or TD backups. On image-based tasks, an action entropy regularization term is added to the actor loss.

  4. Knowl 4 — Approximate Policy Improvement Bound for Filtered Contrastive RL

    theoretical result

    In tabular state and action spaces with a Bayes-optimal contrastive critic, policy improvement is guaranteed when updating a goal-conditioned policy π(a∣s,sg)\pi(a \mid s, s_g) using trajectories filtered according to the ratio of commanded to achieved goal likelihoods.

    Let trajectory segments τi:j=(si,ai,si+1,ai+1,…,sj,aj)\tau_{i:j} = (s_i, a_i, s_{i+1}, a_{i+1}, \dots, s_j, a_j) sampled from policy π(τ∣sg)\pi(\tau \mid s_g) be filtered out if:

    ∣π(τi:j∣sg)π(τi:j∣sj)−1∣>ϵ\left| \frac{\pi(\tau_{i:j} \mid s_g)}{\pi(\tau_{i:j} \mid s_j)} - 1 \right| > \epsilon

    for a tolerance threshold ϵ>0\epsilon > 0, where sjs_j is the state actually reached at step jj.

    Let π′(a∣s,sg)=arg⁡max⁡af∗(s,a,sg)\pi'(a \mid s, s_g) = \arg\max_a f^*(s, a, s_g) be the updated policy obtained after one iteration of contrastive RL on the filtered dataset. Then for all goals sgs_g with pg(sg)>0p_g(s_g) > 0, the expected discounted reward under the goal-conditioned reward rsgr_{s_g} satisfies:

    Eπ′(τ∣sg)[∑t=0∞γtrsg(st,at)]≥Eπ(τ∣sg)[∑t=0∞γtrsg(st,at)]−2γϵ1−γ\mathbb{E}_{\pi'(\tau \mid s_g)}\left[ \sum_{t=0}^\infty \gamma^t r_{s_g}(s_t, a_t) \right] \ge \mathbb{E}_{\pi(\tau \mid s_g)}\left[ \sum_{t=0}^\infty \gamma^t r_{s_g}(s_t, a_t) \right] - \frac{2\gamma \epsilon}{1 - \gamma}

    Iterating data re-collection and filtered contrastive updates corresponds to approximate policy iteration with bounded error.

  5. Knowl 5 — Contrastive RL for Offline Goal-Conditioned Reinforcement Learning

    model/method

    Contrastive RL (NCE) can be adapted to offline goal-conditioned reinforcement learning, where environment interaction is disallowed, by combining the contrastive value objective with a goal-conditioned behavioral cloning regularizer and an ensemble of critics.

    The offline goal-conditioned policy π(a∣s,sg)\pi(a \mid s, s_g) is updated to maximize:

    max⁡π(a∣s,sg)Eπ(a∣s,sg), p(s,aorig,sg)[(1−λ)min⁡k=1,…,Kfk(s,a,sf=sg)+λlog⁡π(aorig∣s,sg)]\max_{\pi(a \mid s, s_g)} \mathbb{E}_{\pi(a \mid s, s_g), \, p(s, a_{\text{orig}}, s_g)} \left[ (1 - \lambda) \min_{k=1,\dots,K} f_k(s, a, s_f = s_g) + \lambda \log \pi(a_{\text{orig}} \mid s, s_g) \right]

    where:

    • λ∈[0,1]\lambda \in [0, 1] controls the trade-off between the contrastive Q-value and the behavioral cloning regularizer (setting λ=1\lambda = 1 yields standard Goal-Conditioned Behavioral Cloning, GCBC).
    • aoriga_{\text{orig}} is the action recorded in the offline dataset transition.
    • {fk}k=1K\{f_k\}_{k=1}^K is an ensemble of KK independently trained contrastive critics (with K∈{2,5}K \in \{2, 5\}), where taking the minimum provides conservative value estimation for the actor update.

    This method performs goal relabeling purely via contrastive representation learning without temporal difference bootstrapping or out-of-distribution value penalties.

  6. Knowl 6 — Offline Goal-Conditioned Benchmark Evaluation on D4RL AntMaze

    data/table

    Contrastive RL with behavioral cloning regularization (Contrastive RL + BC) was evaluated across six goal-conditioned AntMaze benchmark environments from D4RL. It was compared against methods that omit temporal difference (TD) learning (Behavioral Cloning [BC], Decision Transformer [DT], Goal-Conditioned BC [GCBC / RvS-G]) and TD-based offline RL algorithms (TD3+BC, Implicit Q-Learning [IQL]).

    Task BC DT GCBC Contrastive RL + BC (2 nets) Contrastive RL + BC (5 nets) TD3+BC IQL
    umaze-v2 54.6 65.6 65.4 81.9 (±\pm1.7) 79.8 (±\pm1.4) 78.6 87.5
    umaze-diverse-v2 45.6 51.2 60.9 75.4 (±\pm3.5) 77.6 (±\pm2.8) 71.4 62.2
    medium-play-v2 0.0 1.0 58.1 71.5 (±\pm5.2) 72.6 (±\pm2.9) 10.6 71.2
    medium-diverse-v2 0.0 0.6 67.3 72.5 (±\pm2.8) 71.5 (±\pm1.3) 3.0 70.0
    large-play-v2 0.0 0.0 32.4 41.6 (±\pm6.0) 48.6 (±\pm4.4) 0.2 39.6
    large-diverse-v2 0.0 0.2 36.9 49.3 (±\pm6.3) 54.1 (±\pm5.5) 0.0 47.5

    Values represent the normalized success rate percentage (mean ±\pm standard deviation across random seeds). Contrastive RL + BC achieves the highest performance on 5 out of 6 tasks. On the hardest tasks (large-play-v2 and large-diverse-v2), Contrastive RL achieves a 7.0%7.0\% to 9.0%9.0\% absolute improvement over IQL, and achieves a median improvement of 15.0%15.0\% over GCBC across the suite. Scaling the ensemble from 2 to 5 critics improves performance on the large maze environments.

  7. Knowl 7 — Visual Goal-Conditioned RL Performance Without Auxiliary Representation Objectives

    empirical result

    In image-based goal-conditioned control tasks (Fetch Reach, Fetch Push, Sawyer Push, Sawyer Bin, and Point Spiral 11×1111 \times 11), Contrastive RL (NCE) was compared against TD3 with Hindsight Experience Replay (TD3+HER) augmented with standard representation learning methods:

    • Data augmentation via DrQ (averaging Q-values over 4 image augmentations).
    • Autoencoder (AE) auxiliary reconstruction objectives.
    • Contrastive Unsupervised Representations for RL (CURL), using contrastive loss on augmented images.

    Key empirical findings across 5 random seeds include:

    • Contrastive RL (NCE) achieved higher success rates across all tasks than TD3+HER and its representation-augmented variants (DrQ, AE, CURL).
    • On challenging visual manipulation environments (Sawyer Push and Sawyer Bin), baseline methods (TD3+HER, TD3+HER+DrQ, TD3+HER+AE, TD3+HER+CURL) achieved approximately 0%0\% success rate, whereas Contrastive RL (NCE) learned the task, achieving ∼40%\sim 40\% success on Sawyer Push and 10–20%10\text{--}20\% on Sawyer Bin.
    • Contrastive RL achieved these results without using data augmentations, visual reconstruction losses, or auxiliary pre-training objectives.
  8. Knowl 8 — Comparison Across Contrastive RL Objective Variants

    empirical result

    The family of contrastive RL algorithms was analyzed across state-based and image-based manipulation and navigation tasks by comparing four design variants across 5 random seeds:

    1. Contrastive RL (NCE): Uses binary noise-contrastive estimation with a bilinear critic f(s,a,sg)=ϕ(s,a)Tψ(sg)f(s, a, s_g) = \phi(s, a)^T \psi(s_g) and Monte Carlo geometric future sampling.
    2. Contrastive RL (CPC): Replaces binary NCE with multi-class InfoNCE (Contrastive Predictive Coding).
    3. C-learning: Uses temporal difference (TD) learning to train a classifier distinguishing future goals from random goals.
    4. Contrastive RL (NCE + C-learning): Combines Monte Carlo discounted state occupancy estimation with TD-based classifier updates.

    The comparisons show:

    • Contrastive RL (CPC) outperformed Contrastive RL (NCE) on several benchmarks (such as Ant U-Maze and Fetch Push), indicating that swapping mutual information estimators can improve value estimation accuracy.
    • C-learning outperformed NCE on select tasks but underperformed on others.
    • The hybrid Contrastive RL (NCE + C-learning) consistently attained the highest overall performance across both state-based and image-based environments.
  9. Knowl 9 — Robustness of Contrastive RL to Moving Cameras and Partial Observability

    empirical result

    Contrastive RL (NCE) was evaluated on a partially observed robotic manipulation task by modifying the Sawyer Push environment so that the camera is rigidly mounted to the robot's moving wrist rather than statically placed in the environment.

    Because the camera moves with the arm, the target puck is occluded behind a table barrier at the start of each episode, requiring the agent to move the arm into the workspace before the puck becomes visible. Without using recurrent memory architectures or non-Markovian policies, Contrastive RL (NCE) achieved a task success rate of approximately 35%35\% after 5×1065 \times 10^6 environment steps (compared to 75%75\% when evaluated with a static overhead camera), demonstrating that contrastive value estimation can handle moving viewpoints and partial observability.

  10. Knowl 10 — Restriction of Contrastive RL to Goal-Conditioned Settings

    limitation

    The theoretical equivalence between contrastive representation learning and reinforcement learning established in Contrastive RL relies on defining the reward function as the probability density of transitioning into a specified goal state sgs_g, which matches the discounted state occupancy distribution pπ(st+=sg∣s,a)p^\pi(s_{t+} = s_g \mid s, a). Consequently, the method is inherently designed for goal-conditioned RL and multi-goal reaching tasks. Extending contrastive value learning to general RL problems with arbitrary scalar reward functions (such as dense reward shaping or general cost penalties) remains an open challenge.

Coverage note — Deliberately omitted detailed Appendix-only derivations of the CPC and NCE+C-learning objective formulations, summarizing their definitions and results directly in the comparative empirical knowl.

References

  1. 1.Achiam, J., Edwards, H., Amodei, D., and Abbeel, P. (2018). Variational option discovery algorithms. arXiv preprint arXiv:1807.10299.
  2. 2.Achiam, J., Knight, E., and Abbeel, P. (2019). Towards characterizing divergence in deep Q-learning. arXiv preprint arXiv:1903.08894.
  3. 3.Alain, G. and Bengio, Y. (2016). Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644.
  4. 4.Anand, A., Racah, E., Ozair, S., Bengio, Y., Côté, M.-A., and Hjelm, R. D. (2019). Unsupervised state representation learning in Atari. Advances in Neural Information Processing Systems, 32.
  5. 5.Andreas, J., Klein, D., and Levine, S. (2017). Modular multitask reinforcement learning with policy sketches. In International Conference on Machine Learning, pages 166–175. PMLR.
  6. 6.Andrychowicz, M., Crow, D., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W. (2017). Hindsight experience replay. In NeurIPS.
  7. 7.Annasamy, R. M. and Sycara, K. (2019). Towards better interpretability in deep Q-networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 4561–4569.
  8. 8.Authors, I. (2022). Private Communication.
  9. 9.Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D. (2017). Successor features for transfer in reinforcement learning. Advances in neural information processing systems, 30.
  10. 10.Bertsekas, D. P. and Tsitsiklis, J. N. (1996). Neuro-dynamic programming. Athena Scientific.
  11. 11.Blier, L., Tallec, C., and Ollivier, Y. (2021). Learning successor states and goal-dependent values: A mathematical viewpoint. arXiv preprint arXiv:2101.07123.
  12. 12.Borsa, D., Barreto, A., Quan, J., Mankowitz, D., Munos, R., Van Hasselt, H., Silver, D., and Schaul, T. (2018). Universal successor features approximators. arXiv preprint arXiv:1812.07626.
  13. 13.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. (2018). JAX: composable transformations of Python+NumPy programs.
  14. 14.Brown, D., Goo, W., Nagarajan, P., and Niekum, S. (2019). Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations. In International conference on machine learning, pages 783–792. PMLR.
  15. 15.Chane-Sane, E., Schmid, C., and Laptev, I. (2021). Goal-conditioned reinforcement learning with imagined subgoals. In International Conference on Machine Learning, pages 1430–1440. PMLR.
  16. 16.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021). Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34.
  17. 17.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E. (2020). A simple framework for contrastive learning of visual representations. ArXiv, abs/2002.05709.
  18. 18.Chen, X. and He, K. (2021). Exploring simple siamese representation learning. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15745–15753.
  19. 19.Choi, J., Sharma, A., Lee, H., Levine, S., and Gu, S. S. (2021). Variational empowerment as representation learning for goal-conditioned reinforcement learning. In International Conference on Machine Learning, pages 1953–1963. PMLR.
  20. 20.Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. arXiv preprint arXiv:1706.03741.
  21. 21.Dayan, P. (1993). Improving generalization for temporal difference learning: The successor representation. Neural Computation, 5(4):613–624.
  22. 22.Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M. (2019). Goal-conditioned imitation learning. Advances in Neural Information Processing Systems, 32:15324–15335.
  23. 23.Dosovitskiy, A. and Koltun, V. (2016). Learning to act by predicting the future. arXiv preprint arXiv:1611.01779.
  24. 24.Du, Y., Gan, C., and Isola, P. (2021). Curious representation learning for embodied intelligence. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10408–10417.
  25. 25.Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S. (2021). Rvs: What is essential for offline rl via supervised learning? arXiv preprint arXiv:2112.10751.
  26. 26.Eysenbach, B., Geng, X., Levine, S., and Salakhutdinov, R. (2020). Rewriting history with inverse RL: Hindsight inference for policy improvement. ArXiv, abs/2002.11089.
  27. 27.Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S. (2018). Diversity is all you need: Learning skills without a reward function. In International Conference on Learning Representations.
  28. 28.Eysenbach, B., Levine, S., and Salakhutdinov, R. R. (2021a). Replacing rewards with examples: Example-based policy search via recursive classification. Advances in Neural Information Processing Systems, 34.
  29. 29.Eysenbach, B., Salakhutdinov, R., and Levine, S. (2021b). C-learning: Learning to achieve goals via recursive classification. ArXiv, abs/2011.08909.
  30. 30.Eysenbach, B., Salakhutdinov, R. R., and Levine, S. (2019). Search on the replay buffer: Bridging planning and reinforcement learning. Advances in Neural Information Processing Systems, 32.
  31. 31.Eysenbach, B., Udatha, S., Levine, S., and Salakhutdinov, R. (2022). Imitating past successes can be very suboptimal. arXiv preprint arXiv:2206.03378.
  32. 32.Finn, C., Tan, X. Y., Duan, Y., Darrell, T., Levine, S., and Abbeel, P. (2016). Deep spatial autoencoders for visuomotor learning. In 2016 IEEE International Conference on Robotics and Automation (ICRA), pages 512–519. IEEE.
  33. 33.Fischinger, D., Vincze, M., and Jiang, Y. (2013). Learning grasps for unknown objects in cluttered scenes. In 2013 IEEE international conference on robotics and automation, pages 609–616. IEEE.
  34. 34.Florensa, C., Degrave, J., Heess, N., Springenberg, J. T., and Riedmiller, M. (2019). Self-supervised learning of image embedding for continuous control. arXiv preprint arXiv:1901.00943.
  35. 35.Florensa, C., Held, D., Geng, X., and Abbeel, P. (2018). Automatic goal generation for reinforcement learning agents. In International conference on machine learning, pages 1515–1528. PMLR.
  36. 36.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020). D4RL: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219.
  37. 37.Fu, J., Luo, K., and Levine, S. (2017). Learning robust rewards with adversarial inverse reinforcement learning. arXiv preprint arXiv:1710.11248.
  38. 38.Fu, J., Singh, A., Ghosh, D., Yang, L., and Levine, S. (2018). Variational inverse control with events: A general framework for data-driven reward definition. In NeurIPS.
  39. 39.Fujimoto, S. and Gu, S. S. (2021). A minimalist approach to offline reinforcement learning. Advances in Neural Information Processing Systems, 34.
  40. 40.Fujimoto, S., Hoof, H., and Meger, D. (2018). Addressing function approximation error in actor-critic methods. In International conference on machine learning, pages 1587–1596. PMLR.
  41. 41.Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C. M., Eysenbach, B., and Levine, S. (2020). Learning to reach goals via iterated supervised learning. In International Conference on Learning Representations.
  42. 42.Gregor, K., Rezende, D. J., and Wierstra, D. (2016). Variational intrinsic control. arXiv preprint arXiv:1611.07507.
  43. 43.Grill, J.-B., Strub, F., Altch’e, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. Á., Guo, Z. D., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M. (2020). Bootstrap your own latent: A new approach to self-supervised learning. ArXiv, abs/2006.07733.
  44. 44.Guo, Z. D., Azar, M. G., Piot, B., Pires, B. A., and Munos, R. (2018). Neural predictive belief representations. arXiv preprint arXiv:1811.06407.
  45. 45.Guo, Z. D., Pires, B. A., Piot, B., Grill, J.-B., Altché, F., Munos, R., and Azar, M. G. (2020). Bootstrap latent-predictive representations for multitask reinforcement learning. In International Conference on Machine Learning, pages 3875–3886. PMLR.
  46. 46.Gutmann, M. U. and Hyvärinen, A. (2012). Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics. Journal of machine learning research, 13(2).
  47. 47.Ha, D. and Schmidhuber, J. (2018). World models. arXiv preprint arXiv:1803.10122.
  48. 48.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, pages 1861–1870. PMLR.
  49. 49.Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. (2019a). Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603.
  50. 50.Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. (2019b). Learning latent dynamics for planning from pixels. In International conference on machine learning, pages 2555–2565. PMLR.
  51. 51.Han, T., Xie, W., and Zisserman, A. (2020). Self-supervised co-training for video representation learning. Advances in Neural Information Processing Systems, 33:5679–5690.
  52. 52.Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V. (2019). Fast task inference with variational intrinsic successor features. arXiv preprint arXiv:1906.05030.
  53. 53.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738.
  54. 54.Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. (2018). Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670.
  55. 55.Ho, J. and Ermon, S. (2016). Generative adversarial imitation learning. Advances in neural information processing systems, 29:4565–4573.
  56. 56.Hoffer, E. and Ailon, N. (2015). Deep metric learning using triplet network. In International workshop on similarity-based pattern recognition, pages 84–92. Springer.
  57. 57.Hoffman, M., Shahriari, B., Aslanides, J., Barth-Maron, G., Behbahani, F., Norman, T., Abdolmaleki, A., Cassirer, A., Yang, F., Baumli, K., Henderson, S., Novikov, A., Colmenarejo, S. G., Cabi, S., Gulcehre, C., Paine, T. L., Cowie, A., Wang, Z., Piot, B., and de Freitas, N. (2020). Acme: A research framework for distributed reinforcement learning. arXiv preprint arXiv:2006.00979.
  58. 58.Hong, Z.-W., Yang, G., and Agrawal, P. (2022). Bilinear value networks. arXiv preprint arXiv:2204.13695.
  59. 59.Ichter, B., Sermanet, P., and Lynch, C. (2020). Broadly-exploring, local-policy trees for long-horizon task planning. arXiv preprint arXiv:2010.06491.
  60. 60.Janner, M., Mordatch, I., and Levine, S. (2020). gamma-models: Generative temporal difference learning for infinite-horizon prediction. Advances in Neural Information Processing Systems, 33:1724–1735.
  61. 61.Jozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y. (2016). Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410.
  62. 62.Kaelbling, L. P. (1993). Learning to achieve goals. In IJCAI, pages 1094–1099. Citeseer.
  63. 63.Kalashnikov, D., Varley, J., Chebotar, Y., Swanson, B., Jonschkowski, R., Finn, C., Levine, S., and Hausman, K. (2021). Mt-opt: Continuous multi-task robotic reinforcement learning at scale. ArXiv, abs/2104.08212.
  64. 64.Kish, L. (1965). Survey sampling. John Wiley & Sons.
  65. 65.Klingemann, M. (2016). Raster fairy. https://github.com/bmcfee/RasterFairy.
  66. 66.Konda, V. and Tsitsiklis, J. (1999). Actor-critic algorithms. Advances in neural information processing systems, 12.
  67. 67.Konyushkova, K., Zolna, K., Aytar, Y., Novikov, A., Reed, S., Cabi, S., and de Freitas, N. (2020). Semi-supervised reward learning for offline reinforcement learning. arXiv preprint arXiv:2012.06899.
  68. 68.Kostrikov, I., Nair, A., and Levine, S. (2021). Offline reinforcement learning with implicit Q-learning. arXiv preprint arXiv:2110.06169.
  69. 69.Kostrikov, I., Yarats, D., and Fergus, R. (2020). Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. arXiv preprint arXiv:2004.13649.
  70. 70.Kumar, A., Agarwal, R., Ghosh, D., and Levine, S. (2020). Implicit under-parameterization inhibits data-efficient deep reinforcement learning. arXiv preprint arXiv:2010.14498.
  71. 71.Lange, S. and Riedmiller, M. (2010). Deep auto-encoder neural networks in reinforcement learning. In The 2010 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE.
  72. 72.Langford, J. (2010). Specializations of the master problem.
  73. 73.Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A. (2020). Reinforcement learning with augmented data. Advances in Neural Information Processing Systems, 33:19884–19895.
  74. 74.Laskin, M., Liu, H., Peng, X. B., Yarats, D., Rajeswaran, A., and Abbeel, P. (2021). CIC: Contrastive intrinsic control for unsupervised skill discovery. In Deep RL Workshop NeurIPS 2021.
  75. 75.LeCun, Y. (2016). Predictive learning. https://www.youtube.com/watch?v=Ount2Y4qxQo. Keynote Talk.
  76. 76.Levy, A., Konidaris, G., Platt, R., and Saenko, K. (2017). Learning multi-level hierarchies with hindsight. arXiv preprint arXiv:1712.00948.
  77. 77.Levy, O. and Goldberg, Y. (2014). Neural word embedding as implicit matrix factorization. Advances in neural information processing systems, 27.
  78. 78.Li, A., Pinto, L., and Abbeel, P. (2020). Generalized hindsight for reinforcement learning. Advances in neural information processing systems, 33:7754–7767.
  79. 79.Liang, Y., Machado, M. C., Talvitie, E., and Bowling, M. (2015). State of the art control of Atari games using shallow reinforcement learning. arXiv preprint arXiv:1512.01563.
  80. 80.Lin, X., Baweja, H. S., and Held, D. (2019). Reinforcement learning without ground-truth state. ArXiv, abs/1905.07866.
  81. 81.Liu, H. and Abbeel, P. (2021). Aps: Active pretraining with successor features. In International Conference on Machine Learning, pages 6736–6747. PMLR.
  82. 82.Liu, K., Kurutach, T., Tung, C., Abbeel, P., and Tamar, A. (2020). Hallucinative topological memory for zero-shot visual planning. In International Conference on Machine Learning, pages 6259–6270. PMLR.
  83. 83.Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P. (2020). Learning latent plans from play. In Conference on Robot Learning, pages 1113–1132. PMLR.
  84. 84.Ma, Z. and Collins, M. (2018). Noise contrastive estimation and negative sampling for conditional models: Consistency and statistical efficiency. In EMNLP.
  85. 85.Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D. (2021). Discovering and achieving goals via world models. Advances in Neural Information Processing Systems, 34.
  86. 86.Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, 26.
  87. 87.Mnih, A. and Teh, Y. W. (2012). A fast and simple algorithm for training neural probabilistic language models. In ICML.
  88. 88.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602.
  89. 89.Nachum, O., Gu, S., Lee, H., and Levine, S. (2018a). Near-optimal representation learning for hierarchical reinforcement learning. In International Conference on Learning Representations.
  90. 90.Nachum, O., Gu, S. S., Lee, H., and Levine, S. (2018b). Data-efficient hierarchical reinforcement learning. Advances in neural information processing systems, 31.
  91. 91.Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S. (2018). Visual reinforcement learning with imagined goals. Advances in Neural Information Processing Systems, 31:9191–9200.
  92. 92.Nair, S., Mitchell, E., Chen, K., Savarese, S., Finn, C., et al. (2022). Learning language-conditioned robot behavior from offline data and crowd-sourced annotation. In Conference on Robot Learning, pages 1303–1315. PMLR.
  93. 93.Nasiriany, S., Pong, V. H., Lin, S., and Levine, S. (2019). Planning with goal-conditioned policies. In NeurIPS.
  94. 94.Nowozin, S., Cseke, B., and Tomioka, R. (2016). f-GAN: Training generative neural samplers using variational divergence minimization. Advances in neural information processing systems, 29.
  95. 95.Oord, A. v. d., Li, Y., and Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748.
  96. 96.Paster, K., McIlraith, S. A., and Ba, J. (2020). Planning from pixels using inverse dynamics models. arXiv preprint arXiv:2012.02419.
  97. 97.Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al. (2018). Multi-goal reinforcement learning: Challenging robotics environments and request for research. arXiv preprint arXiv:1802.09464.
  98. 98.Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S. (2019). Skew-fit: State-covering self-supervised reinforcement learning. arXiv preprint arXiv:1903.03698.
  99. 99.Poole, B., Ozair, S., Van Den Oord, A., Alemi, A., and Tucker, G. (2019). On variational bounds of mutual information. In International Conference on Machine Learning, pages 5171–5180. PMLR.
  100. 100.Qiu, S., Wang, L., Bai, C., Yang, Z., and Wang, Z. (2022). Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning. In International Conference on Machine Learning, pages 18168–18210. PMLR.
  101. 101.Rakelly, K., Gupta, A., Florensa, C., and Levine, S. (2021). Which mutual-information representation learning objectives are sufficient for control? ArXiv, abs/2106.07278.
  102. 102.Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T. (2018). Learning by playing solving sparse reward tasks from scratch. In International conference on machine learning, pages 4344–4353. PMLR.
  103. 103.Rudner, T. G., Pong, V., McAllister, R., Gal, Y., and Levine, S. (2021). Outcome-driven reinforcement learning via variational inference. Advances in Neural Information Processing Systems, 34.
  104. 104.Rybkin, O., Zhu, C., Nagabandi, A., Daniilidis, K., Mordatch, I., and Levine, S. (2021). Model-based reinforcement learning via latent-space collocation. In International Conference on Machine Learning, pages 9190–9201. PMLR.
  105. 105.Savinov, N., Dosovitskiy, A., and Koltun, V. (2018). Semi-parametric topological memory for navigation. In International Conference on Learning Representations.
  106. 106.Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015). Universal value function approximators. In International conference on machine learning, pages 1312–1320. PMLR.
  107. 107.Schmeckpeper, K., Xie, A., Rybkin, O., Tian, S., Daniilidis, K., Levine, S., and Finn, C. (2020). Learning predictive models from observation and interaction. In European Conference on Computer Vision, pages 708–725. Springer.
  108. 108.Schroff, F., Kalenichenko, D., and Philbin, J. (2015). Facenet: A unified embedding for face recognition and clustering. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 815–823.
  109. 109.Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G. (2018). Time-contrastive networks: Self-supervised learning from video. In 2018 IEEE international conference on robotics and automation (ICRA), pages 1134–1141. IEEE.
  110. 110.Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K. (2019). Dynamics-aware unsupervised discovery of skills. In International Conference on Learning Representations.
  111. 111.Shu, R., Nguyen, T., Chow, Y., Pham, T., Than, K., Ghavamzadeh, M., Ermon, S., and Bui, H. (2020). Predictive coding for locally-linear control. In International Conference on Machine Learning, pages 8862–8871. PMLR.
  112. 112.Silver, D., Singh, S., Precup, D., and Sutton, R. S. (2021). Reward is enough. Artificial Intelligence, 299:103535.
  113. 113.Sohn, K. (2016). Improved deep metric learning with multi-class n-pair loss objective. In NeurIPS.
  114. 114.Srinivas, A. and Abbeel, P. (2021). Unsupervised learning for reinforcement learning. Tutorial.
  115. 115.Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C. (2018). Universal planning networks. ArXiv, abs/1804.00645.
  116. 116.Srinivas, A., Laskin, M., and Abbeel, P. (2020). Curl: Contrastive unsupervised representations for reinforcement learning. arXiv preprint arXiv:2004.04136.
  117. 117.Srivastava, R. K., Shyam, P., Mutz, F., Jaśkowski, W., and Schmidhuber, J. (2019). Training agents using upside-down reinforcement learning. arXiv preprint arXiv:1912.02877.
  118. 118.Stooke, A., Lee, K., Abbeel, P., and Laskin, M. (2021). Decoupling representation learning from reinforcement learning. In International Conference on Machine Learning, pages 9870–9879. PMLR.
  119. 119.Such, F. P., Madhavan, V., Liu, R., Wang, R., Castro, P. S., Li, Y., Zhi, J., Schubert, L., Bellemare, M. G., Clune, J., et al. (2018). An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents. arXiv preprint arXiv:1812.07069.
  120. 120.Sun, H., Li, Z., Liu, X., Zhou, B., and Lin, D. (2019). Policy continuation with hindsight inverse dynamics. Advances in Neural Information Processing Systems, 32:10265–10275.
  121. 121.Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R. (2017). Distral: Robust multitask reinforcement learning. Advances in neural information processing systems, 30.
  122. 122.Tian, Y., Krishnan, D., and Isola, P. (2020). Contrastive multiview coding. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16, pages 776–794. Springer.
  123. 123.Tsai, Y.-H., Zhao, H., Yamada, M., Morency, L.-P., and Salakhutdinov, R. (2020). Neural methods for point-wise dependency estimation. In Proceedings of the Neural Information Processing Systems Conference (Neurips).
  124. 124.Tschannen, M., Djolonga, J., Rubenstein, P. K., Gelly, S., and Lucic, M. (2019). On mutual information maximization for representation learning. arXiv preprint arXiv:1907.13625.
  125. 125.Venkattaramanujam, S., Crawford, E., Doan, T. V., and Precup, D. (2019). Self-supervised learning of distance functions for goal-conditioned reinforcement learning. ArXiv, abs/1907.02998.
  126. 126.Wang, H., Miahi, E., White, M., Machado, M. C., Abbas, Z., Kumaraswamy, R., Liu, V., and White, A. (2022). Investigating the properties of neural network representations in reinforcement learning. arXiv preprint arXiv:2203.15955.
  127. 127.Warde-Farley, D., Van de Wiele, T., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V. (2018). Unsupervised control through non-parametric discriminative rewards. arXiv preprint arXiv:1811.11359.
  128. 128.Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M. (2015). Embed to control: A locally linear latent dynamics model for control from raw images. Advances in neural information processing systems, 28.
  129. 129.Weinberger, K. Q. and Saul, L. K. (2005). Distance metric learning for large margin nearest neighbor classification. In NIPS.
  130. 130.Wilson, A., Fern, A., Ray, S., and Tadepalli, P. (2007). Multi-task reinforcement learning: a hierarchical bayesian approach. In Proceedings of the 24th international conference on Machine learning, pages 1015–1022.
  131. 131.Wu, Y., Tucker, G., and Nachum, O. (2018a). The Laplacian in RL: Learning representations with efficient approximations. arXiv preprint arXiv:1810.04586.
  132. 132.Wu, Z., Xiong, Y., Yu, S. X., and Lin, D. (2018b). Unsupervised feature learning via non-parametric instance discrimination. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3733–3742.
  133. 133.Xie, A., Singh, A., Levine, S., and Finn, C. (2018). Few-shot goal inference for visuomotor learning and planning. In Conference on Robot Learning, pages 40–52. PMLR.
  134. 134.Xu, D. and Denil, M. (2019). Positive-unlabeled reward learning. arXiv preprint arXiv:1911.00459.
  135. 135.Yang, G., Ajay, A., and Agrawal, P. (2021). Overcoming the spectral bias of neural value approximation. In International Conference on Learning Representations.
  136. 136.Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2021a). Mastering visual continuous control: Improved data-augmented reinforcement learning. arXiv preprint arXiv:2107.09645.
  137. 137.Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R. (2021b). Improving sample efficiency in model-free reinforcement learning from images. In AAAI.
  138. 138.Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C. (2020a). Gradient surgery for multi-task learning. Advances in Neural Information Processing Systems, 33:5824–5836.
  139. 139.Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. (2020b). Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning, pages 1094–1100. PMLR.
  140. 140.Zhang, A., McAllister, R. T., Calandra, R., Gal, Y., and Levine, S. (2020a). Learning invariant representations for reinforcement learning without reconstruction. In International Conference on Learning Representations.
  141. 141.Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M., and Levine, S. (2019). Solar: Deep structured representations for model-based reinforcement learning. In International Conference on Machine Learning, pages 7444–7453. PMLR.
  142. 142.Zhang, S., Liu, B., and Whiteson, S. (2020b). Gradientdice: Rethinking generalized offline estimation of stationary values. In International Conference on Machine Learning, pages 11194–11203. PMLR.
  143. 143.Zhang, T., Ren, T., Yang, M., Gonzalez, J., Schuurmans, D., and Dai, B. (2022). Making linear mdps practical via contrastive representation learning. In International Conference on Machine Learning, pages 26447–26466. PMLR.
  144. 144.Zhao, R., Sun, X., and Tresp, V. (2019). Maximum entropy-regularized multi-goal reinforcement learning. In International Conference on Machine Learning, pages 7553–7562. PMLR.
  145. 145.Ziebart, B. D. (2010). Modeling purposeful adaptive behavior with the principle of maximum causal entropy. Carnegie Mellon University.
  146. 146.Zolna, K., Reed, S., Novikov, A., Colmenarejo, S. G., Budden, D., Cabi, S., Denil, M., de Freitas, N., and Wang, Z. (2019). Task-relevant adversarial imitation learning. arXiv preprint arXiv:1910.01077.

Citation

MLA
Eysenbach, B., et al. “Contrastive Learning as Goal-Conditioned Reinforcement Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 35603–20, https://proceedings.neurips.cc/paper_files/paper/2022/file/e7663e974c4ee7a2b475a4775201ce1f-Paper-Conference.pdf.
APA
Eysenbach, B., Zhang, T., Levine, S., & Salakhutdinov, R. R. (2022). Contrastive Learning as Goal-Conditioned Reinforcement Learning. Advances in Neural Information Processing Systems, 35, 35603–35620. https://proceedings.neurips.cc/paper_files/paper/2022/file/e7663e974c4ee7a2b475a4775201ce1f-Paper-Conference.pdf
Chicago
Eysenbach, B., T. Zhang, S. Levine, and R. R. Salakhutdinov. 2022. “Contrastive Learning as Goal-Conditioned Reinforcement Learning”. Advances in Neural Information Processing Systems 35: 35603–20. https://proceedings.neurips.cc/paper_files/paper/2022/file/e7663e974c4ee7a2b475a4775201ce1f-Paper-Conference.pdf.
Harvard
Eysenbach, B. et al. (2022) “Contrastive Learning as Goal-Conditioned Reinforcement Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 35603–35620. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/e7663e974c4ee7a2b475a4775201ce1f-Paper-Conference.pdf.
Vancouver
1. Eysenbach B, Zhang T, Levine S, Salakhutdinov RR (2022) Contrastive Learning as Goal-Conditioned Reinforcement Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 35603–35620

BibTeX

@inproceedings{eysenbach2022contrastive,
  title = {Contrastive Learning as Goal-Conditioned Reinforcement Learning},
  author = {Eysenbach, Benjamin and Zhang, Tianjun and Levine, Sergey and Salakhutdinov, Russ R.},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {35603-35620},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/e7663e974c4ee7a2b475a4775201ce1f-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission