Offline Reinforcement Learning with Value-based Episodic Memory

Xiaoteng MaYiqin YangHao HuJun YangChongjie ZhangQianchuan ZhaoBin LiangQihan Liu

article2022ICLR51 citations

Develops Value-based Episodic Memory, an offline reinforcement learning framework that combines expectile state-value learning with trajectory-based implicit planning to avoid out-of-distribution extrapolation errors and achieve superior performance on sparse-reward tasks.

Listen

Deploying machine learning to optimize decision-making in safety-critical and high-risk environments often requires learning exclusively from pre-collected, offline data rather than live trial and error. A primary barrier in offline learning is extrapolation error, where algorithms severely overestimate the potential rewards of actions not present in the historical dataset. Most current approaches attempt to counteract this instability using complex constraints, behavioral models, or artificial penalties. The article introduces and evaluates Value-based Episodic Memory, a novel framework designed to learn effective decision policies strictly within the bounds of historical data without requiring auxiliary behavioral or dynamic models.

The approach combines two core concepts: Expectile V-Learning and implicit memory-based planning. Instead of evaluating action-specific values that risk extrapolation errors on unseen actions, the framework learns direct state values. An expectile parameter balances conservative imitation of past behavior against the pursuit of optimal performance. The method then performs recursive planning along historical trajectories to enhance advantage estimations, training the final decision policy through standard regression techniques. The authors tested this method across continuous control benchmarks, including the D4RL suite spanning sparse-reward robotic navigation, complex robotic manipulation, and standard locomotion tasks.

The evaluation produced several key findings regarding algorithm performance and stability. First, the proposed framework achieved superior or competitive results compared to leading baseline algorithms across the majority of benchmark tasks. Second, performance gains were especially pronounced in complex, sparse-reward environments; in challenging maze navigation and robotic manipulation tasks, the method achieved success rates substantially higher than existing methods, many of which failed entirely. Third, the framework avoided the catastrophic value overestimation and training collapse that compromised competing action-value models on narrow human demonstration datasets. Finally, theoretical analysis confirmed that the approach is provably convergent and accelerates value learning without introducing systematic bias.

These findings indicate that decision-making models can be reliably trained on offline operational data with lower computational complexity and greater stability. By eliminating the need for complex behavioral generative models or penalty tuning, the framework reduces implementation risk and engineering overhead for data-driven optimization in physical systems. Organizations evaluating autonomous systems can achieve higher operational performance from imperfect or sparse demonstration logs without risking unstable behavior.

To apply these insights, technical teams should consider adopting state-value expectile architectures when training models from limited operational logs. Practitioners must balance the expectile parameter, setting conservative values for narrow or noisy datasets and higher values when data coverage is extensive. Further work should explore deploying this framework in real-world physical pilots beyond simulated benchmarks and evaluating its performance in highly stochastic environments, as the theoretical guarantees assume deterministic system dynamics.

Cover for Offline Reinforcement Learning with Value-based Episodic Memory

Abstract

Offline reinforcement learning (RL) shows promise of applying RL to real-world problems by effectively utilizing previously collected data. Most existing offline RL algorithms use regularization or constraints to suppress extrapolation error for actions outside the dataset. In this paper, we adopt a different framework, which learns the V-function instead of the Q-function to naturally keep the learning procedure within the support of an offline dataset. To enable effective generalization while maintaining proper conservatism in offline learning, we propose Expectile V-Learning (EVL), which smoothly interpolates between the optimal value learning and behavior cloning. Further, we introduce implicit planning along offline trajectories to enhance learned V-values and accelerate convergence. Together, we present a new offline method called Value-based Episodic Memory (VEM). We provide theoretical analysis for the convergence properties of our proposed VEM method, and empirical results in the D4RL benchmark show that our method achieves superior performance in most tasks, particularly in sparse-reward tasks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Preliminaries
  • 2.2 Value-based Offline Reinforcement Learning Methods
  • 3 Method
  • 3.1 Expectile V-Learning
  • 3.2 Implicit Memory-Based Planning
  • 3.3 Generalized Advantage-Weighted Learning
  • 4 Theoretical Analysis
  • 4.1 Convergence Property of the Expectile V-Learning
  • 4.2 Value-based Episodic Memory
  • 4.3 Toy Example
  • 5 Related Work
  • 6 Experiments
  • 6.1 Evaluation environments
  • 6.2 Performance on D4RL tasks
  • 6.3 Analysis of Value Estimation
  • 6.4 Ablations
  • 7 Conclusion
  • 8 Reproducibility
  • References
  • A Algorithm
  • A.1 Value-based Episodic Memory Control
  • B Theoretical Analysis
  • B.1 Complete derivation.
  • B.2 Proof of Lemma
  • B.3 Proof of Lemma
  • B.4 Proof of Lemma
  • B.5 Proof of Lemma
  • B.6 Proof of Lemma
  • B.7 Proof of Proposition
  • C Detailed Implementation
  • C.1 Generalized Advantage-Weighted Learning
  • C.2 BCQ-EM
  • C.3 Hyper-Parameter and Network Structure
  • D Additional Experiments on D4RL
  • D.1 Ablation Study
  • D.2 Complete training curves and value estimation error

Knowls

  1. Knowl 1 — Expectile V-Learning for Offline Reinforcement Learning

    model/method

    Expectile VV-Learning (EVL) is an offline value estimation method that learns state value functions V(s)V(s) entirely within the state support of an offline dataset D\mathcal{D}, avoiding the out-of-distribution action extrapolation errors common in QQ-learning methods without needing explicit policy constraints or actor regularization.

    Given an offline dataset of transition tuples (st,at,rt,st+1)∼D(s_t, a_t, r_t, s_{t+1}) \sim \mathcal{D}, parameterized value network Vθ(s)V_\theta(s), and target network Vθ′(s)V_{\theta'}(s), EVL updates the value network parameters θ\theta by minimizing the expectile regression loss:

    JV(θ)=E(st,at,st+1)∼D[(V^(st)−Vθ(st))2]J_V(\theta) = \mathbb{E}_{(s_t, a_t, s_{t+1}) \sim \mathcal{D}} \left[ \left( \hat{V}(s_t) - V_\theta(s_t) \right)^2 \right]

    where the target value V^(s)\hat{V}(s) is computed using a one-step gradient expectile update:

    V^(s)=Vθ′(s)+2α(τ[δ(s,a,s′)]++(1−τ)[δ(s,a,s′)]−)\hat{V}(s) = V_{\theta'}(s) + 2\alpha \left( \tau [\delta(s, a, s')]_{+} + (1 - \tau) [\delta(s, a, s')]_{-} \right)

    Here, δ(s,a,s′)=r(s,a)+γVθ′(s′)−Vθ′(s)\delta(s, a, s') = r(s, a) + \gamma V_{\theta'}(s') - V_{\theta'}(s) is the one-step temporal-difference (TD) error, γ∈[0,1)\gamma \in [0, 1) is the discount factor, α>0\alpha > 0 is a step-size parameter, τ∈(0,1)\tau \in (0, 1) is the expectile hyperparameter controlling conservatism versus optimality, and the operators are defined as [x]+=max⁡(x,0)[x]_{+} = \max(x, 0) and [x]−=min⁡(x,0)[x]_{-} = \min(x, 0).

  2. Knowl 2 — Implicit Memory-Based Trajectory Planning

    model/method

    Implicit memory-based trajectory planning is a non-parametric value aggregation mechanism that bootstraps state values along offline trajectories to accelerate value propagation and improve advantage estimates in long-horizon and sparse-reward settings without querying out-of-distribution transitions.

    Along an offline trajectory of length TT with steps t∈{1,…,T}t \in \{1, \dots, T\}, rewards rtr_t, and states sts_t, the enhanced target return R^t\hat{R}_t is computed recursively backwards from the final step TT to the first step:

    R^t={rt+γmax⁡(R^t+1,V^(st+1))if t<T,rtif t=T\hat{R}_t = \begin{cases} r_t + \gamma \max\left( \hat{R}_{t+1}, \hat{V}(s_{t+1}) \right) & \text{if } t < T, \\ r_t & \text{if } t = T \end{cases}

    where V^(st+1)\hat{V}(s_{t+1}) is the expectile-generalized value estimate at state st+1s_{t+1} and γ∈[0,1)\gamma \in [0, 1) is the discount factor. Equivalently, this recurrence unrolls to a multi-step maximum over rollout horizons n∈{1,…,nmax⁡}n \in \{1, \dots, n_{\max}\}:

    R^t=max⁡0<n≤nmax⁡V^t,n,where V^t,n={rt+γV^t+1,n−1if n>0,V^(st)if n=0\hat{R}_t = \max_{0 < n \le n_{\max}} \hat{V}_{t, n}, \quad \text{where } \hat{V}_{t, n} = \begin{cases} r_t + \gamma \hat{V}_{t+1, n-1} & \text{if } n > 0, \\ \hat{V}(s_t) & \text{if } n = 0 \end{cases}

    and V^t,n=0\hat{V}_{t, n} = 0 for n>Tn > T. This recursive maximization connects disjoint trajectories in the offline dataset through the learned value function V^\hat{V}, creating an implicit graph-like planning step confined strictly to observed transitions.

  3. Knowl 3 — Value-based Episodic Memory Control Algorithm

    algorithm

    Value-based Episodic Memory (VEM) integrates Expectile VV-Learning, implicit trajectory-based planning, and advantage-weighted policy regression for offline reinforcement learning. It maintains an ensemble of critic networks to mitigate overestimation bias, periodically updates trajectory returns in an episodic memory buffer, and updates the actor network via advantage regression.

    Input: Offline dataset M containing trajectories of transitions (s_t, a_t, r_t, s_{t+1})
    Input: Expectile parameter tau in (0, 1), step-size alpha, discount factor gamma in [0, 1)
    Input: Target update rate kappa, memory update frequency u, batch size N, total iterations T_iter
    Output: Trained policy network pi_phi
    Initialize critic parameters theta_1, theta_2, actor parameters phi, target parameters theta'_1 <- theta_1, theta'_2 <- theta_2
    Initialize trajectory returns R_t^(1), R_t^(2) in episodic memory buffer M using Equation 6
    for t = 1 to T_iter do
        for i in {1, 2} do
            Sample mini-batch of N transitions (s_t, a_t, r_t, s_{t+1}, R_t^(i)) from M
            Update critic parameters:
                theta_i <- theta_i - lr * grad_theta_i ( (1/N) * sum_batch (R_t^(i) - V_theta_i(s_t))^2 )
            Compute advantage:
                A_hat(s_t, a_t) = min_{j in {1,2}} R_t^(j) - (1/2) * sum_{j=1}^2 V_theta_j(s_t)
            Update actor parameters:
                phi <- phi + lr * grad_phi ( (1/N) * sum_batch ( grad_phi log pi_phi(a_t | s_t) * f(A_hat(s_t, a_t)) ) )
        end for
        if t mod u == 0 then
            for i in {1, 2} do
                theta'_i <- kappa * theta_i + (1 - kappa) * theta'_i
            end for
            for each trajectory in buffer M do
                for each transition (s_t, a_t, r_t, s_{t+1}) in reversed trajectory do
                    for i in {1, 2} do
                        Compute target value V_hat(s_{t+1}) using V_theta'_i and Expectile operator
                        Compute R_t^(i) <- r_t + gamma * max(R_{t+1}^(i), V_hat(s_{t+1})) (or r_t if terminal)
                        Store updated R_t^(i) into M
                    end for
                end for
            end for
        end if
    end for
    return pi_phi
  4. Knowl 4 — Generalized Advantage-Weighted Policy Learning

    model/method

    Policy extraction in Value-based Episodic Memory is conducted via generalized advantage-weighted regression over the offline dataset D\mathcal{D}. Given the enhanced trajectory returns R^t\hat{R}_t obtained through implicit memory-based planning and the baseline value estimates V^(st)\hat{V}(s_t), the policy network πϕ(a∣s)\pi_\phi(a|s) is optimized according to:

    max⁡ϕJπ(ϕ)=E(st,at)∼D[log⁡πϕ(at∣st)⋅f(A^(st,at))]\max_\phi J_\pi(\phi) = \mathbb{E}_{(s_t, a_t) \sim \mathcal{D}} \left[ \log \pi_\phi(a_t \mid s_t) \cdot f\left( \hat{A}(s_t, a_t) \right) \right]

    where A^(st,at)=R^t−V^(st)\hat{A}(s_t, a_t) = \hat{R}_t - \hat{V}(s_t) is the estimated advantage and f(⋅)f(\cdot) is a non-negative, monotonically increasing weighting function.

    Two practical functional forms for f(A^)f(\hat{A}) are utilized:

    1. Leaky-ReLU weighting with slope hyperparameter αlr>0\alpha_{lr} > 0:

    f(A^(s,a))={A^(s,a)if A^(s,a)>0A^(s,a)αlrif A^(s,a)≤0f(\hat{A}(s, a)) = \begin{cases} \hat{A}(s, a) & \text{if } \hat{A}(s, a) > 0 \\ \frac{\hat{A}(s, a)}{\alpha_{lr}} & \text{if } \hat{A}(s, a) \le 0 \end{cases}

    1. Softmax weighting with temperature α>0\alpha > 0 normalized across the mini-batch:

    f(A^(sj,aj))=exp⁡(1αA^(sj,aj))∑(si,ai)∈Batchexp⁡(1αA^(si,ai))f(\hat{A}(s_j, a_j)) = \frac{\exp\left( \frac{1}{\alpha} \hat{A}(s_j, a_j) \right)}{\sum_{(s_i, a_i) \in \text{Batch}} \exp\left( \frac{1}{\alpha} \hat{A}(s_i, a_i) \right)}

  5. Knowl 5 — Bellman Expectile Operator and Gradient Expectile Operator

    equation

    The Bellman expectile operator Tτμ\mathcal{T}^\mu_\tau for behavior policy μ\mu and expectile parameter τ∈(0,1)\tau \in (0, 1) is defined by asymmetric squared loss minimization:

    (TτμV)(s):=arg⁡min⁡vEa∼μ(⋅∣s)[τ[δ]+2+(1−τ)[−δ]+2](\mathcal{T}^\mu_\tau V)(s) := \arg\min_v \mathbb{E}_{a \sim \mu(\cdot \mid s)} \left[ \tau [\delta]_+^2 + (1 - \tau) [-\delta]_+^2 \right]

    where δ=Es′∼P(⋅∣s,a)[r(s,a)+γV(s′)−v]\delta = \mathbb{E}_{s' \sim P(\cdot \mid s, a)} [r(s, a) + \gamma V(s') - v], γ∈[0,1)\gamma \in [0, 1), and [x]+=max⁡(x,0)[x]_+ = \max(x, 0). Setting τ=1/2\tau = 1/2 reduces Tτμ\mathcal{T}^\mu_\tau to the standard Bellman expectation operator, while τ→1\tau \to 1 causes it to approach the Bellman optimality operator.

    Because the implicit equation lacks a closed form, the one-step gradient expectile operator ((Tg)τμV)(s)((\mathcal{T}_g)^\mu_\tau V)(s) with step-size α>0\alpha > 0 is defined as:

    ((Tg)τμV)(s)=V(s)+2αEa∼μ(⋅∣s)[τ[δ(s,a)]++(1−τ)[δ(s,a)]−]((\mathcal{T}_g)^\mu_\tau V)(s) = V(s) + 2\alpha \mathbb{E}_{a \sim \mu(\cdot \mid s)} \left[ \tau [\delta(s, a)]_+ + (1 - \tau) [\delta(s, a)]_- \right]

    where δ(s,a)=Es′∼P(⋅∣s,a)[r(s,a)+γV(s′)−V(s)]\delta(s, a) = \mathbb{E}_{s' \sim P(\cdot \mid s, a)} [r(s, a) + \gamma V(s') - V(s)], [x]+=max⁡(x,0)[x]_+ = \max(x, 0), and [x]−=min⁡(x,0)[x]_- = \min(x, 0).

  6. Knowl 6 — Contraction and Asymptotic Optimality of Expectile V-Learning

    theoretical result

    In a deterministic Markov Decision Process (S,A,P,r,γ)(S, A, P, r, \gamma) with discount factor γ∈[0,1)\gamma \in [0, 1), the one-step gradient expectile operator Tτμ\mathcal{T}^\mu_\tau satisfies the following theoretical properties:

    1. Contraction property: For any τ∈[0,1)\tau \in [0, 1) and step-size α≤12max⁡{τ,1−τ}\alpha \le \frac{1}{2 \max\{\tau, 1 - \tau\}}, Tτμ\mathcal{T}^\mu_\tau is a γτ\gamma_\tau-contraction under the supremum norm ∥⋅∥∞\|\cdot\|_\infty:

    ∥TτμV1−TτμV2∥∞≤γτ∥V1−V2∥∞\|\mathcal{T}^\mu_\tau V_1 - \mathcal{T}^\mu_\tau V_2\|_\infty \le \gamma_\tau \|V_1 - V_2\|_\infty

    with contraction factor:

    γτ=1−2α(1−γ)min⁡{τ,1−τ}\gamma_\tau = 1 - 2\alpha (1 - \gamma) \min\{\tau, 1 - \tau\}

    When α=12max⁡{τ,1−τ}\alpha = \frac{1}{2 \max\{\tau, 1 - \tau\}}, the optimal contraction rate is γτ=1−min⁡{τ,1−τ}max⁡{τ,1−τ}(1−γ)\gamma_\tau = 1 - \frac{\min\{\tau, 1 - \tau\}}{\max\{\tau, 1 - \tau\}} (1 - \gamma).

    1. Asymptotic optimality: Let Vτ∗V^*_\tau be the unique fixed point of Tτμ\mathcal{T}^\mu_\tau, and let V∗V^* be the fixed point of the Bellman optimality operator T∗\mathcal{T}^*. In deterministic environments:

    lim⁡τ→1Vτ∗=V∗\lim_{\tau \to 1} V^*_\tau = V^*

  7. Knowl 7 — Contraction and Fixed Point Invariance of the VEM Planning Operator

    theoretical result

    The Value-based Episodic Memory operator Tvem\mathcal{T}_{vem} with maximum rollout horizon nmax⁡∈N+n_{\max} \in \mathbb{N}^+ is defined on state value functions VV as:

    (TvemV)(s)=max⁡1≤n≤nmax⁡{(Tμ)n−1TτμV(s)}(\mathcal{T}_{vem} V)(s) = \max_{1 \le n \le n_{\max}} \left\{ (\mathcal{T}^\mu)^{n-1} \mathcal{T}^\mu_\tau V(s) \right\}

    where Tμ\mathcal{T}^\mu is the standard Bellman expectation operator and Tτμ\mathcal{T}^\mu_\tau is the gradient expectile operator.

    Theoretical guarantees for Tvem\mathcal{T}_{vem} include:

    1. Contraction: For any τ∈(0,1)\tau \in (0, 1) and nmax⁡∈N+n_{\max} \in \mathbb{N}^+, Tvem\mathcal{T}_{vem} is a γτ\gamma_\tau-contraction under the supremum norm, where γτ=1−2α(1−γ)min⁡{τ,1−τ}\gamma_\tau = 1 - 2\alpha (1 - \gamma) \min\{\tau, 1 - \tau\}.
    2. Fixed point invariance: If τ>1/2\tau > 1/2, Tvem\mathcal{T}_{vem} has the exact same fixed point Vτ∗V^*_\tau as Tτμ\mathcal{T}^\mu_\tau.
    3. Optimistic convergence bound: When current value estimates V(s)V(s) are lower than the behavior policy value, Tvem\mathcal{T}_{vem} provides an accelerated update bounded by:

    ∣TvemV(s)−Vτ∗(s)∣≤γn∗(s)−1γτ∥V−Vn∗,τμ∥∞+∥Vn∗,τμ−Vτ∗∥∞,∀s∈S|\mathcal{T}_{vem} V(s) - V^*_\tau(s)| \le \gamma^{n^*(s)-1} \gamma_\tau \|V - V^\mu_{n^*, \tau}\|_\infty + \|V^\mu_{n^*, \tau} - V^*_\tau\|_\infty, \quad \forall s \in S

    where n∗(s)=arg⁡max⁡1≤n≤nmax⁡{(Tμ)n−1TτμV(s)}n^*(s) = \arg\max_{1 \le n \le n_{\max}} \{(\mathcal{T}^\mu)^{n-1} \mathcal{T}^\mu_\tau V(s)\}, and Vn∗,τμV^\mu_{n^*, \tau} is the fixed point of (Tμ)n∗(s)−1Tτμ(\mathcal{T}^\mu)^{n^*(s)-1} \mathcal{T}^\mu_\tau.

  8. Knowl 8 — Monotonicity of Expectile Value Fixed Points with Respect to Expectile Level

    theoretical result

    Let Tτμ\mathcal{T}^\mu_\tau be the one-step gradient expectile operator with expectile parameter τ∈(0,1)\tau \in (0, 1), and let Vτ∗V^*_\tau denote its unique fixed point satisfying TτμVτ∗=Vτ∗\mathcal{T}^\mu_\tau V^*_\tau = V^*_\tau.

    1. Operator monotonicity: For any τ,τ′∈(0,1)\tau, \tau' \in (0, 1) such that τ′≥τ\tau' \ge \tau and any value function VV:

    Tτ′μV(s)≥TτμV(s),∀s∈S\mathcal{T}^\mu_{\tau'} V(s) \ge \mathcal{T}^\mu_\tau V(s), \quad \forall s \in S

    1. Fixed-point monotonicity: For any τ,τ′∈(0,1)\tau, \tau' \in (0, 1) such that τ′≥τ\tau' \ge \tau:

    Vτ′∗(s)≥Vτ∗(s),∀s∈SV^*_{\tau'}(s) \ge V^*_\tau(s), \quad \forall s \in S

    This confirms that increasing τ\tau monotonically shifts the fixed-point value estimates upward from the behavior policy value VμV^\mu toward the optimal value V∗V^*.

  9. Knowl 9 — D4RL Benchmark Evaluation of Value-based Episodic Memory

    data/table

    Evaluation of Value-based Episodic Memory (VEM) on the D4RL benchmark across AntMaze navigation (sparse reward), Adroit robotic hand manipulation, and MuJoCo locomotion domains. Scores are normalized so that 0 represents random policy performance and 100 represents expert demonstration performance, averaged over 3 random seeds with standard deviation.

    Dataset Type Environments VEM (Ours) BAIL BCQ CQL AWR
    fixed antmaze-umaze 87.5 ±\pm 1.1 62.5 ±\pm 2.3 78.9 74.0 56.0
    play antmaze-medium 78.0 ±\pm 3.1 40.0 ±\pm 15.0 0.0 61.2 0.0
    play antmaze-large 57.0 ±\pm 5.0 23.0 ±\pm 5.0 6.7 11.8 0.0
    diverse antmaze-umaze 78.0 ±\pm 1.1 75.0 ±\pm 1.0 55.0 84.0 70.3
    diverse antmaze-medium 77.0 ±\pm 2.2 50.0 ±\pm 10.0 0.0 53.7 0.0
    diverse antmaze-large 58.0 ±\pm 2.1 30.0 ±\pm 5.0 2.2 14.9 0.0
    human adroit-door 11.2 ±\pm 4.2 0.0 ±\pm 0.1 -0.0 9.1 0.4
    human adroit-hammer 3.6 ±\pm 1.0 0.0 ±\pm 0.1 0.5 2.1 1.2
    human adroit-relocate 1.3 ±\pm 0.2 0.0 ±\pm 0.1 0.5 2.1 -0.0
    human adroit-pen 65.0 ±\pm 2.1 32.5 ±\pm 1.5 68.9 55.8 12.3
    cloned adroit-door 3.6 ±\pm 0.3 0.0 ±\pm 0.1 0.0 3.5 0.0
    cloned adroit-hammer 2.7 ±\pm 1.5 0.1 ±\pm 0.1 0.4 5.7 0.4
    cloned adroit-pen 48.7 ±\pm 3.2 46.5 ±\pm 3.5 44.0 40.3 28.0
    expert adroit-door 105.5 ±\pm 0.2 104.7 ±\pm 0.3 99.0 - 102.9
    expert adroit-hammer 128.3 ±\pm 1.1 123.5 ±\pm 3.1 114.9 - 39.0
    expert adroit-relocate 109.8 ±\pm 0.2 94.4 ±\pm 2.7 41.6 - 91.5
    expert adroit-pen 111.7 ±\pm 2.6 126.7 ±\pm 0.3 114.9 - 111.0
    random mujoco-walker2d 6.2 ±\pm 4.7 3.9 ±\pm 2.5 4.9 7.0 1.5
    random mujoco-hopper 11.1 ±\pm 1.0 9.8 ±\pm 0.1 10.6 10.8 10.2
    random mujoco-halfcheetah 16.4 ±\pm 3.6 0.0 ±\pm 0.1 2.2 35.4 2.5
    medium mujoco-walker2d 74.0 ±\pm 1.2 73.0 ±\pm 1.0 53.1 79.2 17.4
    medium mujoco-hopper 56.6 ±\pm 2.3 58.2 ±\pm 1.0 54.5 58.0 35.9
    medium mujoco-halfcheetah 47.4 ±\pm 0.2 42.6 ±\pm 1.2 40.7 44.4 37.4

    VEM achieves superior performance over competing baselines on almost all sparse-reward tasks (AntMaze and Adroit human/cloned), achieving scores such as 58.0 on antmaze-large-diverse (versus 14.9 for CQL and 2.2 for BCQ) and 77.0 on antmaze-medium-diverse (versus 53.7 for CQL).

  10. Knowl 10 — Value Estimation Stability of EVL versus BCQ Action Sampling

    empirical result

    Replacing Expectile VV-Learning in VEM with Batch Constrained QQ-learning's generative action sampling model (BCQ-EM) leads to severe value explosion and degraded policy performance on high-dimensional robotic control tasks (such as Adroit human demonstrations).

    In Adroit tasks with a 24-DoF action space and narrow demonstration distributions, the generative action model in BCQ fails to completely eliminate out-of-distribution actions. Bootstrapping on these generated actions causes the QQ-value estimation errors to explode up to magnitudes between 101310^{13} and 101410^{14} on door-human, hammer-human, relocate-human, and pen-human. In contrast, VEM learns state value functions V(s)V(s) strictly within the offline dataset's support, maintaining bounded estimation errors and consistent training stability.

Coverage note — None was omitted; all key theoretical assertions, core algorithmic formulations, primary benchmark datasets, and diagnostic empirical ablation results were captured.

References

  1. 1.Arthur Argenson and Gabriel Dulac-Arnold. Model-based offline planning. arXiv preprint arXiv:2008.05556, 2020.
  2. 2.Charles Blundell, Benigno Uria, Alexander Pritzel, Yazhe Li, Avraham Ruderman, Joel Z Leibo, Jack Rae, Daan Wierstra, and Demis Hassabis. Model-free episodic control. arXiv preprint arXiv:1606.04460, 2016.
  3. 3.Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, and Keith Ross. BAIL: Best-action imitation learning for batch deep reinforcement learning. Advances in Neural Information Processing Systems, 33, 2020.
  4. 4.Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos. Distributional reinforcement learning with quantile regression. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  5. 5.Robert Dadashi, Shideh Rezaeifar, Nino Vieillard, Léonard Hussenot, Olivier Pietquin, and Matthieu Geist. Offline reinforcement learning with pseudometric learning. arXiv preprint arXiv:2103.01948, 2021.
  6. 6.Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. D4RL: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
  7. 7.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pp. 2052–2062. PMLR, 2019.
  8. 8.Matthieu Geist, Bruno Scherrer, and Olivier Pietquin. A theory of regularized markov decision processes. In International Conference on Machine Learning, pp. 2160–2169. PMLR, 2019.
  9. 9.Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu. EMaQ: Expected-max Q-learning operator for simple yet effective offline and online RL. In International Conference on Machine Learning, pp. 3682–3691. PMLR, 2021.
  10. 10.Hao Hu, Jianing Ye, Zhizhou Ren, Guangxiang Zhu, and Chongjie Zhang. Generalizable episodic memory for deep reinforcement learning. arXiv preprint arXiv:2103.06469, 2021.
  11. 11.Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to trust your model: Model-based policy optimization. Advances in Neural Information Processing Systems, 32:12519–12530, 2019.
  12. 12.Ying Jin, Zhuoran Yang, and Zhaoran Wang. Is pessimism provably efficient for offline RL? In International Conference on Machine Learning, pp. 5084–5096. PMLR, 2021.
  13. 13.Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims. Morel: Model-based offline reinforcement learning. arXiv preprint arXiv:2005.05951, 2020.
  14. 14.Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum. Offline reinforcement learning with fisher divergence critic regularization. In International Conference on Machine Learning, pp. 5774–5783. PMLR, 2021.
  15. 15.Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. Stabilizing off-policy Q-learning via bootstrapping error reduction. Advances in Neural Information Processing Systems, 32:11784–11794, 2019.
  16. 16.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative Q-learning for offline reinforcement learning. arXiv preprint arXiv:2006.04779, 2020.
  17. 17.Jongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau, and Kee-Eung Kim. OptiDICE: Offline policy optimization via stationary distribution correction estimation. arXiv preprint arXiv:2106.10783, 2021.
  18. 18.Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  19. 19.Zichuan Lin, Tianqi Zhao, Guangwen Yang, and Lintao Zhang. Episodic memory deep Q-networks. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, pp. 2433–2439, 2018.
  20. 20.Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine. Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020.
  21. 21.Whitney K Newey and James L Powell. Asymmetric least squares estimation and testing. Econometrica: Journal of the Econometric Society, pp. 819–847, 1987.
  22. 22.Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019.
  23. 23.Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adria Puigdomenech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell. Neural episodic control. In International Conference on Machine Learning, pp. 2827–2836. PMLR, 2017.
  24. 24.Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell. Bridging offline reinforcement learning and imitation learning: A tale of pessimism. arXiv preprint arXiv:2103.12021, 2021.
  25. 25.Mark Rowland, Robert Dadashi, Saurabh Kumar, Rémi Munos, Marc G Bellemare, and Will Dabney. Statistics and samples in distributional reinforcement learning. In International Conference on Machine Learning, pp. 5528–5536. PMLR, 2019.
  26. 26.Mark Rowland, Will Dabney, and Rémi Munos. Adaptive trade-offs in off-policy learning. In International Conference on Artificial Intelligence and Statistics, pp. 34–44. PMLR, 2020.
  27. 27.Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller. Keep doing what worked: Behavioral modelling priors for offline reinforcement learning. arXiv preprint arXiv:2002.08396, 2020.
  28. 28.Yunhao Tang. Self-imitation learning via generalized lower bound Q-learning. Advances in Neural Information Processing Systems, 33, 2020.
  29. 29.Qing Wang, Jiechao Xiong, Lei Han, Peng Sun, Han Liu, and Tong Zhang. Exponentially weighted imitation learning for batched historical data. Advances in Neural Information Processing Systems, 31:6288, 2018.
  30. 30.Ruosong Wang, Yifan Wu, Ruslan Salakhutdinov, and Sham M Kakade. Instabilities of offline rl with pre-trained neural representation. arXiv preprint arXiv:2103.04947, 2021.
  31. 31.Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh. Uncertainty weighted Actor-Critic for offline reinforcement learning. arXiv preprint arXiv:2105.08140, 2021.
  32. 32.Yiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng, Qiyuan Zhang, Gao Huang, Jun Yang, and Qianchuan Zhao. Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning. arXiv preprint arXiv:2106.03400, 2021.
  33. 33.Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma. MOPO: Model-based offline policy optimization. Advances in Neural Information Processing Systems, 33:14129–14142, 2020.
  34. 34.Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans. GenDICE: Generalized offline estimation of stationary values. arXiv preprint arXiv:2002.09072, 2020a.
  35. 35.Shangtong Zhang, Bo Liu, and Shimon Whiteson. GradientDICE: Rethinking generalized offline estimation of stationary values. In International Conference on Machine Learning, pp. 11194–11203. PMLR, 2020b.

Citation

MLA
Ma, X., et al. “Offline Reinforcement Learning with Value-based Episodic Memory”. arXiv, 2021, http://arxiv.org/abs/2110.09796v1.
APA
Ma, X., Yang, Y., Hu, H., Liu, Q., Yang, J., Zhang, C., Zhao, Q., & Liang, B. (2021). Offline Reinforcement Learning with Value-based Episodic Memory. arXiv. http://arxiv.org/abs/2110.09796v1
Chicago
Ma, X., Y. Yang, H. Hu, et al. 2021. “Offline Reinforcement Learning with Value-based Episodic Memory”. arXiv. http://arxiv.org/abs/2110.09796v1.
Harvard
Ma, X. et al. (2021) “Offline Reinforcement Learning with Value-based Episodic Memory”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.09796v1.
Vancouver
1. Ma X, Yang Y, Hu H, Liu Q, Yang J, Zhang C, Zhao Q, Liang B (2021) Offline Reinforcement Learning with Value-based Episodic Memory. arXiv

BibTeX

@article{ma2021offline,
  title = {Offline Reinforcement Learning with Value-based Episodic Memory},
  author = {Ma, Xiaoteng and Yang, Yiqin and Hu, Hao and Liu, Qihan and Yang, Jun and Zhang, Chongjie and Zhao, Qianchuan and Liang, Bin},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.09796v1},
  eprint = {2110.09796}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors