METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Seohong ParkOleh RybkinSergey Levine

article2024ICLR82 citations

Proposes METRA, an unsupervised reinforcement learning objective that maps high-dimensional state spaces into temporal-distance latent metrics, enabling the autonomous discovery of diverse locomotion behaviors directly from pixel inputs in complex environments like Quadruped and Humanoid.

Listen

Unsupervised pre-training has proven highly transformative in areas such as natural language processing and computer vision by allowing models to learn useful representations without human labeling. Reinforcement learning aims to achieve similar success by training autonomous agents to explore and discover diverse behaviors before facing specific downstream tasks. However, existing unsupervised reinforcement learning techniques struggle to scale to complex, high-dimensional environments. Pure exploration approaches attempt to map every state or transition, which becomes computationally infeasible as environments grow larger. Meanwhile, skill discovery methods based on mutual information frequently fail to explore broadly because they focus only on making behaviors distinguishable rather than maximizing physical coverage, often stalling in static behaviors or failing entirely when operating directly from image pixels.

The article introduces and evaluates Metric-Aware Abstraction, termed METRA, a scalable objective designed to autonomously discover diverse and useful behaviors without human supervision or rewards. The primary objective is to demonstrate that an agent can achieve broad, practical environment coverage by abstracting the state space into a compact latent space governed by temporal distances—defined as the minimum number of environment steps required to transition between states.

To evaluate this framework, the authors conducted extensive simulated experiments across five robotic locomotion and manipulation benchmarks, spanning both state-based inputs (Ant and HalfCheetah) and raw pixel observations (Quadruped, Humanoid, and Kitchen). The study compared METRA against eleven prior algorithms spanning pure exploration, mutual information skill learning, and unsupervised goal-reaching baselines. The methodology evaluated performance based on state space coverage, zero-shot goal-reaching capability, and downstream task adaptation using hierarchical controllers.

The analysis produced several key findings. First, METRA demonstrated superior state coverage across all tested domains, substantially outperforming prior skill discovery and pure exploration techniques. Second, METRA is the first unsupervised reinforcement learning method to successfully discover diverse locomotion behaviors in pixel-based Quadruped and Humanoid environments, where all competing methods failed to explore. Third, METRA achieved the highest performance in zero-shot goal-reaching tasks, surpassing leading goal-conditioned methods like LEXA across all five tested domains. Finally, when evaluating downstream utility, high-level controllers trained on top of METRA's pre-trained behaviors achieved the fastest adaptation and highest overall task returns.

These findings indicate that scaling unsupervised reinforcement learning requires moving away from exhaustive state coverage toward metric-aware behavioral abstractions. In practice, METRA provides a mechanism to pre-train general-purpose agents that can immediately adapt to specific objectives without task-specific engineering. This significantly lowers the computational and operational costs of downstream task learning, reduces training timelines, and enables zero-shot deployment for goal-reaching applications directly from raw sensory data such as camera streams.

Organizations developing complex autonomous systems should consider adopting temporal distance metrics for unsupervised pre-training rather than relying on pure exploration or metric-agnostic mutual information. Looking forward, further development is recommended to combine METRA with advanced model-based frameworks to improve sample efficiency and to incorporate asymmetric quasimetrics for environments with irreversible dynamics. Decision-makers should note that while results are statistically robust across the tested continuous control benchmarks, the evaluation was limited to stationary, fully observable simulations and has not yet been extended to non-Markovian settings, discrete game environments, or physical hardware deployments.

No sufficiently relevant recommendations were found.

Cover for METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Abstract

Unsupervised pre-training strategies have proven to be highly effective in natural language processing and computer vision. Likewise, unsupervised reinforcement learning (RL) holds the promise of discovering a variety of potentially useful behaviors that can accelerate the learning of a wide array of downstream tasks. Previous unsupervised RL approaches have mainly focused on pure exploration and mutual information skill learning. However, despite the previous attempts, making unsupervised RL truly scalable still remains a major open challenge: pure exploration approaches might struggle in complex environments with large state spaces, where covering every possible transition is infeasible, and mutual information skill learning approaches might completely fail to explore the environment due to the lack of incentives. To make unsupervised RL scalable to complex, high-dimensional environments, we propose a novel unsupervised RL objective, which we call Metric-Aware Abstraction (METRA). Our main idea is, instead of directly covering the entire state space, to only cover a compact latent space ZZ that is metrically connected to the state space SS by temporal distances. By learning to move in every direction in the latent space, METRA obtains a tractable set of diverse behaviors that approximately cover the state space, being scalable to high-dimensional environments. Through our experiments in five locomotion and manipulation environments, we demonstrate that METRA can discover a variety of useful behaviors even in complex, pixel-based environments, being the first unsupervised RL method that discovers diverse locomotion behaviors in pixel-based Quadruped and Humanoid. Our code and videos are available at this https URL

Table of Contents

  • 1 Introduction
  • 2 Why Might Previous Unsupervised RL Methods Fail To Scale?
  • 3 Preliminaries and Problem Setting
  • 4 A Scalable Objective for Unsupervised RL
  • 4.1 Tractable Optimization
  • 4.2 Full Objective: Metric-Aware Abstraction (METRA)
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Qualitative Comparison
  • 5.3 Quantitative Comparison
  • 6 Conclusion
  • References
  • A Extended Related Work
  • B Theoretical results
  • B.1 Universality of Inner Product Decomposition
  • B.2 Lipschitz Constraint under the Temporal Distance Metric
  • C A Connection between METRA and PCA
  • D Connections between WDM and DIAYN, DADS, and CIC
  • D.1 DIAYN
  • D.2 DADS
  • D.3 CIC
  • E Additional Results
  • E.1 Full qualitative results
  • E.2 Latent space visualization
  • E.3 Ablation Study of Latent Space Sizes
  • F Experimental Details
  • F.1 Environments
  • F.2 Implementation Details

Knowls

  1. Knowl 1 — Metric-Aware Abstraction (METRA) Objective

    model/method

    Metric-Aware Abstraction (METRA) is an unsupervised reinforcement learning objective designed to learn diverse, exploratory skills in high-dimensional state and observation spaces by maximizing the Wasserstein dependency measure (WDM) between states and skill latent vectors under a temporal distance metric.

    Let M=(S,A,μ,p)\mathcal{M} = (\mathcal{S}, \mathcal{A}, \mu, p) be a controlled Markov process with state space S\mathcal{S}, action space A\mathcal{A}, initial state distribution μ\mu, and dynamics kernel pp. Let Z\mathcal{Z} denote a compact latent skill space with a zero-mean prior p(z)p(z) (such that Ep(z)[z]=0\mathbb{E}_{p(z)}[z] = 0), and let π(a∣s,z)\pi(a|s, z) be a latent-conditioned policy.

    METRA defines the dependency between trajectory termination states and skill latent variables using the 1-Wasserstein distance WW on the metric space (S×Z,d)(\mathcal{S} \times \mathcal{Z}, d):

    IW(ST;Z)=W(p(sT,z),p(sT)p(z))I_W(S_T; Z) = W(p(s_T, z), p(s_T)p(z))

    By leveraging the Kantorovich-Rubinstein duality, setting the dual witness function to an inner product form f(s,z)=ϕ(s)⊤zf(s, z) = \phi(s)^\top z with representation mapping ϕ:S→RD\phi: \mathcal{S} \to \mathbb{R}^D, and choosing the underlying metric dd as the temporal distance dtemp(s1,s2)d_{\text{temp}}(s_1, s_2) (the minimum expected environment steps to transition from s1s_1 to s2s_2), the unconstrained variational objective decomposes via a telescoping sum across trajectory steps t=0,…,T−1t = 0, \dots, T-1 into the constrained optimization problem:

    sup⁡π,ϕEτ∼π(⋅∣z),z∼p(z)[∑t=0T−1(ϕ(st+1)−ϕ(st))⊤z]s.t.∥ϕ(s)−ϕ(s′)∥2≤1,∀(s,s′)∈Sadj\sup_{\pi, \phi} \mathbb{E}_{\tau \sim \pi(\cdot|z), z \sim p(z)} \left[ \sum_{t=0}^{T-1} (\phi(s_{t+1}) - \phi(s_t))^\top z \right] \quad \text{s.t.} \quad \|\phi(s) - \phi(s')\|_2 \le 1, \quad \forall (s, s') \in \mathcal{S}_{\text{adj}}

    where Sadj\mathcal{S}_{\text{adj}} is the set of all adjacent (single-step transition) state pairs in the MDP. The policy π(a∣s,z)\pi(a|s, z) is optimized to maximize the intrinsic step reward r(s,z,s′)=(ϕ(s′)−ϕ(s))⊤zr(s, z, s') = (\phi(s') - \phi(s))^\top z while the representation ϕ(s)\phi(s) is trained to preserve temporal distances.

  2. Knowl 2 — METRA Optimization Algorithm

    algorithm

    METRA jointly optimizes the state representation function ϕ(s)\phi(s), the dual Lagrange multiplier λ≥0\lambda \ge 0, and the latent-conditioned policy π(a∣s,z)\pi(a|s, z) via dual gradient descent and off-policy actor-critic reinforcement learning.

    Input: Replay buffer D\mathcal{D}, initial policy π(a∣s,z)\pi(a|s, z), initial state encoder ϕ(s)\phi(s), initial Lagrange multiplier λ\lambda, relaxation constant ε>0\varepsilon > 0, prior distribution p(z)p(z)
    Output: Trained skill policy π(a∣s,z)\pi(a|s, z), metric abstraction function ϕ(s)\phi(s)
    for each epoch do
        for each episode in epoch do
            Sample latent skill z∼p(z)z \sim p(z)
            Roll out trajectory τ=(s0,a0,s1,…,sT)\tau = (s_0, a_0, s_1, \dots, s_T) using policy π(a∣s,z)\pi(a|s, z)
            Store transitions (st,at,st+1,z)(s_t, a_t, s_{t+1}, z) in replay buffer D\mathcal{D}
        end for
        for each gradient step do
            Sample minibatch of transitions (s,z,s′)(s, z, s') from D\mathcal{D}
            Update ϕ\phi via gradient ascent to maximize E(s,z,s′)[(ϕ(s′)−ϕ(s))⊤z+λ⋅min⁡(ε,1−∥ϕ(s)−ϕ(s′)∥22)]\mathbb{E}_{(s, z, s')}[(\phi(s') - \phi(s))^\top z + \lambda \cdot \min(\varepsilon, 1 - \|\phi(s) - \phi(s')\|_2^2)]
            Update λ\lambda via gradient descent to minimize E(s,z,s′)[λ⋅min⁡(ε,1−∥ϕ(s)−ϕ(s′)∥22)]\mathbb{E}_{(s, z, s')}[\lambda \cdot \min(\varepsilon, 1 - \|\phi(s) - \phi(s')\|_2^2)]
            Update π(a∣s,z)\pi(a|s, z) using Soft Actor-Critic (SAC) with intrinsic reward r(s,z,s′)=(ϕ(s′)−ϕ(s))⊤zr(s, z, s') = (\phi(s') - \phi(s))^\top z
        end for
    end for

    In standard implementations, the relaxation constant is set to ε=10−3\varepsilon = 10^{-3} and the initial Lagrange multiplier is set to λ=30\lambda = 30. Continuous skills are normalized such that z~=z/∥z∥2\tilde{z} = z / \|z\|_2 with z∼N(0,Id)z \sim \mathcal{N}(0, I_d).

  3. Knowl 3 — Equivalence of Global Temporal Metric Lipschitz Constraint to Local Transition Constraint

    theoretical result

    Let M=(S,A,μ,p)\mathcal{M} = (\mathcal{S}, \mathcal{A}, \mu, p) be a controlled Markov process, and let dtemp(u,v)d_{\text{temp}}(u, v) denote the temporal distance from state u∈Su \in \mathcal{S} to state v∈Sv \in \mathcal{S}, defined as the minimum number of environment transitions required to reach vv starting from uu. Let Sadj={(s,s′)∈S×S:p(s′∣s,a)>0 for some a∈A}∪{(s,s):s∈S}\mathcal{S}_{\text{adj}} = \{(s, s') \in \mathcal{S} \times \mathcal{S} : p(s'|s, a) > 0 \text{ for some } a \in \mathcal{A}\} \cup \{(s, s) : s \in \mathcal{S}\} be the set of valid 1-step adjacent state pairs.

    For any state embedding function ϕ:S→RD\phi: \mathcal{S} \to \mathbb{R}^D, the following two statements are mathematically equivalent:

    1. Global 1-Lipschitz condition under temporal distance: ∥ϕ(u)−ϕ(v)∥2≤dtemp(u,v),∀u,v∈S\|\phi(u) - \phi(v)\|_2 \le d_{\text{temp}}(u, v), \quad \forall u, v \in \mathcal{S}

    2. Local 1-step adjacent transition constraint: ∥ϕ(s)−ϕ(s′)∥2≤1,∀(s,s′)∈Sadj\|\phi(s) - \phi(s')\|_2 \le 1, \quad \forall (s, s') \in \mathcal{S}_{\text{adj}}

    This equivalence guarantees that enforcing a unit bound on adjacent observed transition pairs in the replay buffer is necessary and sufficient to guarantee that latent Euclidean distances never exceed the shortest path transition distances across the entire state space.

  4. Knowl 4 — Equivalence of Linear Squared METRA and Principal Component Analysis in Temporal Embedding Space

    theoretical result

    Let M\mathcal{M} be an MDP admitting a temporally consistent embedding ψ:S→Rm\psi: \mathcal{S} \to \mathbb{R}^m such that the temporal distance equals the Euclidean distance between embeddings:

    dtemp(u,v)=∥ψ(u)−ψ(v)∥2,∀u,v∈Sd_{\text{temp}}(u, v) = \|\psi(u) - \psi(v)\|_2, \quad \forall u, v \in \mathcal{S}

    Assume that all states are reachable within TT steps, the prior skill distribution is isotropic Gaussian p(z)=N(0,Id)p(z) = \mathcal{N}(0, I_d), the embedding space Ψ={ψ(s):s∈S}\Psi = \{\psi(s) : s \in \mathcal{S}\} forms an ellipsoid Ψ={x∈Rm:x⊤A−1x≤1}\Psi = \{x \in \mathbb{R}^m : x^\top A^{-1} x \le 1\} for some symmetric positive definite matrix A∈S++mA \in \mathcal{S}_{++}^m, and the abstraction function ϕ:S→Rd\phi: \mathcal{S} \to \mathbb{R}^d is a linear mapping of the temporal embedding: ϕ(s)=W⊤ψ(s)\phi(s) = W^\top \psi(s) with W∈Rm×dW \in \mathbb{R}^{m \times d}.

    Under these conditions, maximizing the squared METRA objective:

    sup⁡π,ϕEτ∼π,z∼p(z)[(ϕ(sT)⊤z)2]s.t.∥ϕ(u)−ϕ(v)∥2≤dtemp(u,v),∀u,v∈S\sup_{\pi, \phi} \mathbb{E}_{\tau \sim \pi, z \sim p(z)} [(\phi(s_T)^\top z)^2] \quad \text{s.t.} \quad \|\phi(u) - \phi(v)\|_2 \le d_{\text{temp}}(u, v), \quad \forall u, v \in \mathcal{S}

    is mathematically equivalent to performing Principal Component Analysis (PCA) in the temporal embedding space. The optimal projection matrix is given by:

    W∗=[a1,a2,…,ad]W^* = [a_1, a_2, \dots, a_d]

    where a1,a2,…,ad∈Rma_1, a_2, \dots, a_d \in \mathbb{R}^m are the top dd orthonormal eigenvectors corresponding to the dd largest eigenvalues of AA.

  5. Knowl 5 — Zero-Shot Goal Reaching via Metric-Preserving State Abstraction

    model/method

    Because the METRA representation function ϕ:S→RD\phi: \mathcal{S} \to \mathbb{R}^D is trained to preserve directional temporal distances between states, the vector difference ϕ(g)−ϕ(s)\phi(g) - \phi(s) between the current state ss and a target goal state gg specifies the precise skill vector zz required to navigate from ss to gg.

    Without training a goal-conditioned policy or modifying the pre-trained skill policy π(a∣s,z)\pi(a|s, z), goal reaching is performed zero-shot by dynamically computing the skill command at runtime:

    • For continuous skill latent spaces Z=RD\mathcal{Z} = \mathbb{R}^D: z=ϕ(g)−ϕ(s)∥ϕ(g)−ϕ(s)∥2z = \frac{\phi(g) - \phi(s)}{\|\phi(g) - \phi(s)\|_2}

    • For discrete skill latent spaces Z={e1,e2,…,eK}\mathcal{Z} = \{e_1, e_2, \dots, e_K\}: z=arg⁡max⁡k(ϕ(g)−ϕ(s))kz = \arg\max_{k} \left( \phi(g) - \phi(s) \right)_k

    In continuous locomotion environments, zz is re-computed at every control step to provide closed-loop goal navigation. In multi-stage manipulation environments, zz is evaluated at the initial state and executed across the trajectory.

  6. Knowl 6 — Unified Wasserstein Framework and Connections to DIAYN, DADS, and CIC

    theoretical result

    Standard mutual information (MI) skill discovery objectives are metric-agnostic approximations of specific parametrizations of the Wasserstein Dependency Measure (WDM).

    Let ϕL\phi_L and ψL\psi_L denote 1-Lipschitz-constrained representation mappings. The general WDM objectives:

    IW(S;Z)≈sup⁡ϕL,ψL∑t=0T−1(Ep(τ,z)[ϕL(st)⊤ψL(z)]−Ep(τ)[ϕL(st)]⊤Ep(z)[ψL(z)])I_W(S; Z) \approx \sup_{\phi_L, \psi_L} \sum_{t=0}^{T-1} \left( \mathbb{E}_{p(\tau, z)} [\phi_L(s_t)^\top \psi_L(z)] - \mathbb{E}_{p(\tau)}[\phi_L(s_t)]^\top \mathbb{E}_{p(z)}[\psi_L(z)] \right)

    IW(ST;Z)≈sup⁡ϕL,ψL∑t=0T−1(Ep(τ,z)[(ϕL(st+1)−ϕL(st))⊤ψL(z)]−Ep(τ)[ϕL(st+1)−ϕL(st)]⊤Ep(z)[ψL(z)])I_W(S_T; Z) \approx \sup_{\phi_L, \psi_L} \sum_{t=0}^{T-1} \left( \mathbb{E}_{p(\tau, z)} [(\phi_L(s_{t+1}) - \phi_L(s_t))^\top \psi_L(z)] - \mathbb{E}_{p(\tau)}[\phi_L(s_{t+1}) - \phi_L(s_t)]^\top \mathbb{E}_{p(z)}[\psi_L(z)] \right)

    yield distinct previous methods under specific structural reductions:

    1. DIAYN: Setting ψL(z)=z\psi_L(z) = z in IW(S;Z)I_W(S; Z) with Gaussian assumption p(z)=N(0,I)p(z) = \mathcal{N}(0, I) produces the intrinsic reward rt=ϕL(st)⊤zr_t = \phi_L(s_t)^\top z, which is the Lipschitz inner-product counterpart of DIAYN's reward rtDIAYN=−∥ϕ(st)−z∥22r_t^{\text{DIAYN}} = -\|\phi(s_t) - z\|_2^2.
    2. DADS: Setting ϕL(s)=s\phi_L(s) = s in IW(ST;Z)I_W(S_T; Z) yields rt≈(st+1−st)⊤ψL(z)−1L∑i=1L(st+1−st)⊤ψL(zi)r_t \approx (s_{t+1} - s_t)^\top \psi_L(z) - \frac{1}{L} \sum_{i=1}^L (s_{t+1} - s_t)^\top \psi_L(z_i), matching DADS's transition prediction reward rtDADS=−∥(st+1−st)−ψ(st,z)∥22+1L∑i=1L∥(st+1−st)−ψ(st,zi)∥22r_t^{\text{DADS}} = -\|(s_{t+1} - s_t) - \psi(s_t, z)\|_2^2 + \frac{1}{L} \sum_{i=1}^L \|(s_{t+1} - s_t) - \psi(s_t, z_i)\|_2^2.
    3. CIC: Applying Jensen's inequality to unconstrained IW(S;Z)I_W(S; Z) yields rt=ϕL(st)⊤ψL(z)−log⁡1L∑i=1Lexp⁡(ϕL(st)⊤ψL(zi))r_t = \phi_L(s_t)^\top \psi_L(z) - \log \frac{1}{L} \sum_{i=1}^L \exp(\phi_L(s_t)^\top \psi_L(z_i)), which forms the metric-constrained version of CIC's noise-contrastive estimation reward.
  7. Knowl 7 — Universality of Inner Product Decomposition for Joint Bivariate Functions

    theoretical result

    Let X\mathcal{X} and Y\mathcal{Y} be compact Hausdorff spaces (e.g., compact subsets in RN\mathbb{R}^N), and let C(X×Y)C(\mathcal{X} \times \mathcal{Y}) be the Banach space of real-valued continuous functions on X×Y\mathcal{X} \times \mathcal{Y} endowed with the supremum norm.

    For every target function f(x,y)∈C(X×Y)f(x, y) \in C(\mathcal{X} \times \mathcal{Y}) and any approximation error ε>0\varepsilon > 0, there exists a dimension D∈ND \in \mathbb{N} and continuous vector functions ϕ:X→RD\phi: \mathcal{X} \to \mathbb{R}^D and ψ:Y→RD\psi: \mathcal{Y} \to \mathbb{R}^D such that:

    sup⁡x∈X,y∈Y∣f(x,y)−ϕ(x)⊤ψ(y)∣<ε\sup_{x \in \mathcal{X}, y \in \mathcal{Y}} |f(x, y) - \phi(x)^\top \psi(y)| < \varepsilon

    Furthermore, if Φ⊂C(X)\Phi \subset C(\mathcal{X}) and Ψ⊂C(Y)\Psi \subset C(\mathcal{Y}) are dense function families (such as feedforward neural network parameterizations), the function family T={∑i=1Dϕi(x)ψi(y):D∈N,ϕi∈Φ,ψi∈Ψ}\mathcal{T} = \left\{ \sum_{i=1}^D \phi_i(x)\psi_i(y) : D \in \mathbb{N}, \phi_i \in \Phi, \psi_i \in \Psi \right\} is dense in C(X×Y)C(\mathcal{X} \times \mathcal{Y}). Therefore, the bilinear decomposition f(s,z)=ϕ(s)⊤ψ(z)f(s, z) = \phi(s)^\top \psi(z) retains universal expressiveness as the latent dimension D→∞D \to \infty.

  8. Knowl 8 — Unsupervised Exploration and Skill Discovery in Pixel-Based Locomotion

    empirical result

    METRA was evaluated on state-based locomotion (Gym HalfCheetah, 18-D state; Gym Ant, 29-D state) and pixel-based continuous control (DeepMind Control Suite Quadruped and Humanoid, 64×64×364 \times 64 \times 3 RGB inputs; Franka Kitchen, 64×64×364 \times 64 \times 3 RGB inputs) against 11 baseline methods spanning unsupervised skill discovery (DIAYN, DADS, LSD, CIC), pure exploration (ICM, RND, Plan2Explore/Disagreement, APT, LBS, APS), and unsupervised goal-reaching (LEXA).

    In pixel-based Quadruped and Humanoid, all prior skill discovery methods (DIAYN, DADS, LSD, CIC) completely fail to learn locomotion, remaining near the starting origin. Pure exploration methods (ICM, RND, APT, APS, Plan2Explore, LBS) struggle to cover the state space in complex high-dimensional morphologies (such as Ant and Humanoid) because they exhaust capacity on local joint configurations.

    METRA is the only unsupervised RL method that successfully acquires diverse directional locomotion behaviors purely from raw pixel observations without proprioceptive access, reward signals, or privileged knowledge.

  9. Knowl 9 — Downstream Hierarchical Task Adaptation and Zero-Shot Goal Reaching

    empirical result

    The downstream utility of skills discovered by METRA was evaluated across two paradigms:

    1. Hierarchical Downstream Reinforcement Learning: A high-level policy πh(z∣s,stask)\pi^h(z|s, s_{\text{task}}) was trained on top of frozen skill policies to select a skill zz every KK steps (K=25K=25 for Ant/HalfCheetah, K=50K=50 for Quadruped/Humanoid). Across five downstream benchmarks (AntMultiGoals, HalfCheetahGoal, HalfCheetahHurdle, QuadrupedGoal, HumanoidGoal), METRA achieved superior or near-best cumulative returns compared to skills learned by LSD, CIC, DIAYN, and DADS.

    2. Zero-Shot Goal Navigation: Pre-trained agents were commanded to reach arbitrary goal states by selecting skills via z=(ϕ(g)−ϕ(s))/∥ϕ(g)−ϕ(s)∥2z = (\phi(g) - \phi(s)) / \|\phi(g) - \phi(s)\|_2 without any downstream policy learning. METRA outperformed LEXA, DIAYN, and LSD across all five goal-conditioned environments (HalfCheetah, Ant, Quadruped, Humanoid, and Kitchen), achieving the highest negative goal distance and highest task success counts.

  10. Knowl 10 — Limitations of Metric-Aware Abstraction (METRA)

    limitation

    METRA possesses several key theoretical and practical limitations:

    1. Symmetric Latent Metric Assumption: Embedding temporal distances into a symmetric Euclidean latent space enforces an effective bound ∥ϕ(s1)−ϕ(s2)∥2≤min⁡(dtemp(s1,s2),dtemp(s2,s1))\|\phi(s_1) - \phi(s_2)\|_2 \le \min(d_{\text{temp}}(s_1, s_2), d_{\text{temp}}(s_2, s_1)), which makes distance abstractions overly conservative in environments with highly asymmetric or irreversible transition dynamics.
    2. Linear Latent Trajectory Restriction: Setting the skill representation to the identity ψ(z)=z\psi(z) = z restricts policy optimization to behaviors that progress linearly in the latent space Z\mathcal{Z}, which may limit the diversity of non-linear state trajectories.
    3. Sample Efficiency and Low Update-to-Data Ratio: In practice, METRA utilizes model-free Soft Actor-Critic (SAC) with low update-to-data (UTD) ratios (1/41/4 for Kitchen and 1/161/16 for Quadruped and Humanoid), which reduces sample efficiency relative to model-based or sample-efficient actor-critic alternatives.
    4. Markovian and Stationary Dynamics Assumption: The formulation assumes a stationary, fully observable Markov decision process and does not account for non-Markovian or non-stationary environments.

Coverage note — All substantial contributions—including the WDM objective, algorithmic formulation, Lipschitz constraints under temporal metrics, connection to PCA, connections to prior MI-based methods (DIAYN, DADS, CIC), universality of inner-product decomposition, zero-shot goal reaching, empirical evaluations, and limitations—have been included as standalone knowls.

References

  1. 1.Joshua Achiam, Harrison Edwards, Dario Amodei, and Pieter Abbeel. Variational option discovery algorithms. ArXiv, abs/1807.10299, 2018.
  2. 2.Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andew Bolt, and Charles Blundell. Never give up: Learning directed exploration strategies. In International Conference on Learning Representations (ICLR), 2020.
  3. 3.Richard F Bass. Real analysis for graduate students. Createspace Ind Pub, 2013.
  4. 4.Kate Baumli, David Warde-Farley, Steven Stenberg Hansen, and Volodymyr Mnih. Relative variational intrinsic control. In AAAI Conference on Artificial Intelligence (AAAI), 2021.
  5. 5.Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos. Unifying count-based exploration and intrinsic motivation. In Neural Information Processing Systems (NeurIPS), 2016.
  6. 6.Glen Berseth, Daniel Geng, Coline Devin, Nicholas Rhinehart, Chelsea Finn, Dinesh Jayaraman, and Sergey Levine. Smirl: Surprise minimizing reinforcement learning in unstable environments. In International Conference on Learning Representations (ICLR), 2021.
  7. 7.Rajendra Bhatia. Matrix analysis. Springer Science & Business Media, 2013.
  8. 8.G. Brockman, Vicki Cheung, Ludwig Pettersson, J. Schneider, John Schulman, Jie Tang, and W. Zaremba. OpenAI Gym. ArXiv, abs/1606.01540, 2016.
  9. 9.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Neural Information Processing Systems (NeurIPS), 2020.
  10. 10.Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. Exploration by random network distillation. In International Conference on Learning Representations (ICLR), 2019.
  11. 11.Víctor Campos Camúñez, Alex Trott, Caiming Xiong, Richard Socher, Xavier Giró Nieto, and Jordi Torres Viñals. Explore, discover and learn: unsupervised discovery of state-covering skills. In International Conference on Machine Learning (ICML), 2020.
  12. 12.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML), 2020.
  13. 13.Xinyue Chen, Che Wang, Zijian Zhou, and Keith W. Ross. Randomized ensembled double q-learning: Learning fast without a model. In International Conference on Learning Representations (ICLR), 2021.
  14. 14.Jongwook Choi, Archit Sharma, Honglak Lee, Sergey Levine, and Shixiang Gu. Variational empowerment as representation learning for goal-conditioned reinforcement learning. In International Conference on Machine Learning (ICML), 2021.
  15. 15.John D. Co-Reyes, Yuxuan Liu, Abhishek Gupta, Benjamin Eysenbach, P. Abbeel, and Sergey Levine. Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings. In International Conference on Machine Learning (ICML), 2018.
  16. 16.Ishan Durugkar, Mauricio Tec, Scott Niekum, and Peter Stone. Adversarial intrinsic motivation for reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2021.
  17. 17.Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune. First return, then explore. Nature, 590:580–586, 2020.
  18. 18.Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. Diversity is all you need: Learning skills without a reward function. In International Conference on Learning Representations (ICLR), 2019a.
  19. 19.Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine. Search on the replay buffer: Bridging planning and reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2019b.
  20. 20.Carlos Florensa, Yan Duan, and P. Abbeel. Stochastic neural networks for hierarchical reinforcement learning. In International Conference on Learning Representations (ICLR), 2017.
  21. 21.Carlos Florensa, Jonas Degrave, Nicolas Manfred Otto Heess, Jost Tobias Springenberg, and Martin A. Riedmiller. Self-supervised learning of image embedding for continuous control. ArXiv, abs/1901.00943, 2019.
  22. 22.Justin Fu, John D. Co-Reyes, and Sergey Levine. Ex2: Exploration with exemplar models for deep reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2017.
  23. 23.Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra. Variational intrinsic control. ArXiv, abs/1611.07507, 2016.
  24. 24.Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman. Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning. In Conference on Robot Learning (CoRL), 2019.
  25. 25.Michael U Gutmann and Aapo Hyvärinen. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2010.
  26. 26.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning (ICML), 2018a.
  27. 27.Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, G. Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. Soft actor-critic algorithms and applications. ArXiv, abs/1812.05905, 2018b.
  28. 28.Danijar Hafner, Timothy P. Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. In International Conference on Learning Representations (ICLR), 2020.
  29. 29.Danijar Hafner, Kuang-Huei Lee, Ian S. Fischer, and P. Abbeel. Deep hierarchical planning from pixels. In Neural Information Processing Systems (NeurIPS), 2022.
  30. 30.Nicklas Hansen, Xiaolong Wang, and Hao Su. Temporal difference learning for model predictive control. In International Conference on Machine Learning (ICML), 2022.
  31. 31.S. Hansen, Will Dabney, André Barreto, T. Wiele, David Warde-Farley, and V. Mnih. Fast task inference with variational intrinsic successor features. In International Conference on Learning Representations (ICLR), 2020.
  32. 32.Kristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, and Sergey Levine. Dynamical distance learning for semi-supervised and unsupervised skill discovery. In International Conference on Learning Representations (ICLR), 2020.
  33. 33.Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest. Provably efficient maximum entropy exploration. In International Conference on Machine Learning (ICML), 2019.
  34. 34.Shuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao, and Xiangyang Ji. Wasserstein unsupervised reinforcement learning. In AAAI Conference on Artificial Intelligence (AAAI), 2022.
  35. 35.Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, and Yoshimasa Tsuruoka. Dropout q-functions for doubly efficient reinforcement learning. In International Conference on Learning Representations (ICLR), 2022.
  36. 36.Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and P. Abbeel. Vime: Variational information maximizing exploration. In Neural Information Processing Systems (NeurIPS), 2016.
  37. 37.Edward S. Hu, Richard Chang, Oleh Rybkin, and Dinesh Jayaraman. Planning goals for exploration. In International Conference on Learning Representations (ICLR), 2023.
  38. 38.Zheyuan Jiang, Jingyue Gao, and Jianyu Chen. Unsupervised skill discovery via recurrent skill training. In Neural Information Processing Systems (NeurIPS), 2022.
  39. 39.Leslie Pack Kaelbling. Learning to achieve goals. In International Joint Conference on Artificial Intelligence (IJCAI), 1993.
  40. 40.Pierre-Alexandre Kamienny, Jean Tarbouriech, Alessandro Lazaric, and Ludovic Denoyer. Direct then diffuse: Incremental unsupervised skill discovery for state covering and goal reaching. In International Conference on Learning Representations (ICLR), 2022.
  41. 41.Jaekyeom Kim, Seohong Park, and Gunhee Kim. Unsupervised skill discovery with bottleneck option learning. In International Conference on Machine Learning (ICML), 2021.
  42. 42.Seongun Kim, Kyowoon Lee, and Jaesik Choi. Variational curriculum reinforcement learning for unsupervised discovery of skills. In International Conference on Machine Learning (ICML), 2023.
  43. 43.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  44. 44.Martin Klissarov and Marlos C. Machado. Deep laplacian-based options for temporally-extended exploration. In International Conference on Machine Learning (ICML), 2023.
  45. 45.Ilya Kostrikov, Denis Yarats, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. In International Conference on Learning Representations (ICLR), 2021.
  46. 46.Michael Laskin, Denis Yarats, Hao Liu, Kimin Lee, Albert Zhan, Kevin Lu, Catherine Cang, Lerrel Pinto, and P. Abbeel. Urlb: Unsupervised reinforcement learning benchmark. In Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2021.
  47. 47.Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and P. Abbeel. Unsupervised reinforcement learning with contrastive intrinsic control. In Neural Information Processing Systems (NeurIPS), 2022.
  48. 48.Yann LeCun, Bernhard E. Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne E. Hubbard, and Lawrence D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1:541–551, 1989.
  49. 49.Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric P. Xing, Sergey Levine, and Ruslan Salakhutdinov. Efficient exploration via state marginal matching. ArXiv, abs/1906.05274, 2019.
  50. 50.Mengdi Li, Xufeng Zhao, Jae Hee Lee, Cornelius Weber, and Stefan Wermter. Internally rewarded reinforcement learning. In International Conference on Machine Learning (ICML), 2023.
  51. 51.Hao Liu and Pieter Abbeel. APS: Active pretraining with successor features. In International Conference on Machine Learning (ICML), 2021a.
  52. 52.Hao Liu and Pieter Abbeel. Behavior from the void: Unsupervised active pre-training. In Neural Information Processing Systems (NeurIPS), 2021b.
  53. 53.Marlos C. Machado, Marc G. Bellemare, and Michael Bowling. A laplacian framework for option discovery in reinforcement learning. In International Conference on Machine Learning (ICML), 2017.
  54. 54.Marlos C. Machado, Clemens Rosenbaum, Xiaoxiao Guo, Miao Liu, Gerald Tesauro, and Murray Campbell. Eigenoption discovery through the deep successor representation. In International Conference on Learning Representations (ICLR), 2018.
  55. 55.Pietro Mazzaglia, Ozan Çatal, Tim Verbelen, and B. Dhoedt. Curiosity-driven exploration via latent bayesian surprise. In AAAI Conference on Artificial Intelligence (AAAI), 2022.
  56. 56.Pietro Mazzaglia, Tim Verbelen, B. Dhoedt, Alexandre Lacoste, and Sai Rajeswar. Choreographer: Learning and adapting skills in imagination. In International Conference on Learning Representations (ICLR), 2023.
  57. 57.Russell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner, and Deepak Pathak. Discovering and achieving goals via world models. In Neural Information Processing Systems (NeurIPS), 2021.
  58. 58.Shakir Mohamed and Danilo J. Rezende. Variational information maximisation for intrinsically motivated reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2015.
  59. 59.Mirco Mutti, Lorenzo Pratissoli, and Marcello Restelli. Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate. In AAAI Conference on Artificial Intelligence (AAAI), 2021.
  60. 60.OpenAI OpenAI, Matthias Plappert, Raul Sampedro, Tao Xu, Ilge Akkaya, Vineet Kosaraju, Peter Welinder, Ruben D’Sa, Arthur Petron, Henrique Ponde de Oliveira Pinto, Alex Paino, Hyeonwoo Noh, Lilian Weng, Qiming Yuan, Casey Chu, and Wojciech Zaremba. Asymmetric self-play for automatic goal discovery in robotic manipulation. ArXiv, abs/2101.04882, 2021.
  61. 61.Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos. Count-based exploration with neural density models. In International Conference on Machine Learning (ICML), 2017.
  62. 62.Sherjil Ozair, Corey Lynch, Yoshua Bengio, Aäron van den Oord, Sergey Levine, and Pierre Sermanet. Wasserstein dependency measure for representation learning. In Neural Information Processing Systems (NeurIPS), 2019.
  63. 63.Seohong Park and Sergey Levine. Predictable mdp abstraction for unsupervised model-based rl. In International Conference on Machine Learning (ICML), 2023.
  64. 64.Seohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee, and Gunhee Kim. Lipschitz-constrained unsupervised skill discovery. In International Conference on Learning Representations (ICLR), 2022.
  65. 65.Seohong Park, Dibya Ghosh, Benjamin Eysenbach, and Sergey Levine. Hiql: Offline goal-conditioned rl with latent states as actions. In Neural Information Processing Systems (NeurIPS), 2023a.
  66. 66.Seohong Park, Kimin Lee, Youngwoon Lee, and P. Abbeel. Controllability-aware unsupervised skill discovery. In International Conference on Machine Learning (ICML), 2023b.
  67. 67.Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction. In International Conference on Machine Learning (ICML), 2017.
  68. 68.Deepak Pathak, Dhiraj Gandhi, and Abhinav Kumar Gupta. Self-supervised exploration via disagreement. In International Conference on Machine Learning (ICML), 2019.
  69. 69.Silviu Pitis, Harris Chan, S. Zhao, Bradly C. Stadie, and Jimmy Ba. Maximum entropy gain exploration for long horizon multi-goal reinforcement learning. In International Conference on Machine Learning (ICML), 2020.
  70. 70.Vitchyr H. Pong, Murtaza Dalal, S. Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine. Skew-Fit: State-covering self-supervised reinforcement learning. In International Conference on Machine Learning (ICML), 2020.
  71. 71.Ben Poole, Sherjil Ozair, Aäron van den Oord, Alexander A. Alemi, and G. Tucker. On variational bounds of mutual information. In International Conference on Machine Learning (ICML), 2019.
  72. 72.A. H. Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson, Byron Boots, and Michael C. Yip. Composing task-agnostic policies with deep reinforcement learning. In International Conference on Learning Representations (ICLR), 2020.
  73. 73.Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Pich’e, B. Dhoedt, Aaron C. Courville, and Alexandre Lacoste. Mastering the unsupervised reinforcement learning benchmark from pixels. In International Conference on Machine Learning (ICML), 2023.
  74. 74.Nick Rhinehart, Jenny Wang, Glen Berseth, John D. Co-Reyes, Danijar Hafner, Chelsea Finn, and Sergey Levine. Information is power: Intrinsic control via information capture. In Neural Information Processing Systems (NeurIPS), 2021.
  75. 75.Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun. Semi-parametric topological memory for navigation. In International Conference on Learning Representations (ICLR), 2018.
  76. 76.Tom Schaul, Dan Horgan, Karol Gregor, and David Silver. Universal value function approximators. In International Conference on Machine Learning (ICML), 2015.
  77. 77.John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and P. Abbeel. High-dimensional continuous control using generalized advantage estimation. In International Conference on Learning Representations (ICLR), 2016.
  78. 78.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. ArXiv, abs/1707.06347, 2017.
  79. 79.Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, P. Abbeel, Danijar Hafner, and Deepak Pathak. Planning to explore via self-supervised world models. In International Conference on Machine Learning (ICML), 2020.
  80. 80.Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, P. Abbeel, and Kimin Lee. State entropy maximization with random encoders for efficient exploration. In International Conference on Machine Learning (ICML), 2021.
  81. 81.Nur Muhammad (Mahi) Shafiullah and Lerrel Pinto. One after another: Learning incremental skills for a changing world. In International Conference on Learning Representations (ICLR), 2022.
  82. 82.Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman. Dynamics-aware unsupervised discovery of skills. In International Conference on Learning Representations (ICLR), 2020.
  83. 83.Pranav Shyam, Wojciech Jaśkowski, and Faustino J. Gomez. Model-based active exploration. In International Conference on Machine Learning (ICML), 2019.
  84. 84.DJ Strouse, Kate Baumli, David Warde-Farley, Vlad Mnih, and Steven Stenberg Hansen. Learning more skills through optimistic exploration. In International Conference on Learning Representations (ICLR), 2022.
  85. 85.Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur D. Szlam, and Rob Fergus. Intrinsic motivation and automatic curricula via asymmetric self-play. In International Conference on Learning Representations (ICLR), 2018.
  86. 86.Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and P. Abbeel. #exploration: A study of count-based exploration for deep reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2017.
  87. 87.Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller. Deepmind control suite. ArXiv, abs/1801.00690, 2018.
  88. 88.Emanuel Todorov, Tom Erez, and Yuval Tassa. Mujoco: A physics engine for model-based control. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2012.
  89. 89.Ahmed Touati and Yann Ollivier. Learning one representation to optimize all rewards. In Neural Information Processing Systems (NeurIPS), 2021.
  90. 90.Ahmed Touati, Jérémy Rapin, and Yann Ollivier. Does zero-shot reinforcement learning exist? In International Conference on Learning Representations (ICLR), 2023.
  91. 91.Aäron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. ArXiv, abs/1807.03748, 2018.
  92. 92.Cédric Villani et al. Optimal transport: old and new. Springer, 2009.
  93. 93.Tongzhou Wang, Antonio Torralba, Phillip Isola, and Amy Zhang. Optimal goal-reaching reinforcement learning via quasimetric learning. In International Conference on Machine Learning (ICML), 2023.
  94. 94.David Warde-Farley, Tom Van de Wiele, Tejas Kulkarni, Catalin Ionescu, Steven Hansen, and Volodymyr Mnih. Unsupervised control through non-parametric discriminative rewards. In International Conference on Learning Representations (ICLR), 2019.
  95. 95.Rushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li, Bin Zhao, Zhen Wang, Peng Liu, and Xuelong Li. Behavior contrastive learning for unsupervised skill discovery. In International Conference on Machine Learning (ICML), 2023.
  96. 96.Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto. Reinforcement learning with prototypical representations. In International Conference on Machine Learning (ICML), 2021.
  97. 97.Jesse Zhang, Haonan Yu, and Wei Xu. Hierarchical reinforcement learning by discovering intrinsic options. In International Conference on Learning Representations (ICLR), 2021.
  98. 98.Andrew Zhao, Matthieu Lin, Yangguang Li, Y. Liu, and Gao Huang. A mixture of surprises for unsupervised reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2022.

Citation

MLA
Park, S., et al. “METRA: Scalable Unsupervised RL with Metric-Aware Abstraction”. arXiv, 2023, http://arxiv.org/abs/2310.08887v2.
APA
Park, S., Rybkin, O., & Levine, S. (2023). METRA: Scalable Unsupervised RL with Metric-Aware Abstraction. arXiv. http://arxiv.org/abs/2310.08887v2
Chicago
Park, S., O. Rybkin, and S. Levine. 2023. “METRA: Scalable Unsupervised RL with Metric-Aware Abstraction”. arXiv. http://arxiv.org/abs/2310.08887v2.
Harvard
Park, S., Rybkin, O. and Levine, S. (2023) “METRA: Scalable Unsupervised RL with Metric-Aware Abstraction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.08887v2.
Vancouver
1. Park S, Rybkin O, Levine S (2023) METRA: Scalable Unsupervised RL with Metric-Aware Abstraction. arXiv

BibTeX

@article{park2023metra,
  title = {METRA: Scalable Unsupervised RL with Metric-Aware Abstraction},
  author = {Park, Seohong and Rybkin, Oleh and Levine, Sergey},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.08887v2},
  eprint = {2310.08887}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors