Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction

Aviral KumarJustin FuGeorge TuckerSergey Levine

article2019NeurIPS1,358 citations

Proposes the BEAR algorithm to stabilize offline reinforcement learning by theoretically characterizing out-of-distribution bootstrapping error and constraining policy optimization to the support of static datasets.

Listen

Real-world deployment of reinforcement learning is frequently limited by the high cost and safety risks of active, real-time data collection in operational environments. While industries such as autonomous driving and robotics possess large repositories of pre-collected, static operational logs, existing off-policy algorithms struggle to learn effectively from these fixed datasets without collecting additional live data. When standard algorithms attempt to evaluate choices absent from the historical records, they introduce severe out-of-distribution estimation errors. These errors compound rapidly across training steps, causing value estimates to explode and resulting in policy failure even when the static datasets are large or generated by experts.

The article analyzes the mechanisms behind this error propagation and develops an algorithm that enables artificial intelligence agents to learn high-performing, reliable policies purely from fixed historical datasets without environmental interaction.

To address this challenge, the authors mathematically formulated the error propagation dynamics and introduced Bootstrapping Error Accumulation Reduction, or BEAR. Unlike previous approaches that strictly force the new policy to mirror the exact probability distribution of the data collection policy, BEAR constrains updates to match only the underlying support set—the set of valid, plausible actions—of the historical behavior. The algorithm implements this principle in a continuous-action actor-critic framework using Maximum Mean Discrepancy, a statistical distance metric estimated directly from data samples. The approach was evaluated across simulated continuous-control environments, including standard robotics benchmarks and the complex Humanoid control task, using fixed datasets of one million transitions gathered under random, mediocre, and expert collection policies.

The experimental findings show that BEAR consistently matches or outperforms existing methods across diverse data conditions. On medium-quality demonstration data—the most representative setting for real-world enterprise applications—BEAR achieved substantially higher performance than standard off-policy baselines and existing constrained methods such as Batch-Constrained Q-learning, which tended to merely copy suboptimal behavior. On random datasets, BEAR successfully extracted high-performing policies that significantly exceeded the average quality of the training data, whereas strict distribution-matching methods failed completely. On expert datasets, BEAR matched optimal performance, demonstrating robust generalization regardless of the initial data quality.

These results demonstrate that offline reinforcement learning can be successfully stabilized without resorting to overly conservative imitation. For enterprise applications, this shifts the paradigm from expensive, trial-and-error live testing to data-driven training on legacy logs, substantially lowering operational risk, financial costs, and deployment timelines. Furthermore, ablation experiments established that constraining the support set via sample-based Maximum Mean Discrepancy provides a more stable optimization target than standard probability density constraints like Kullback-Leibler divergence.

Organizations aiming to deploy reinforcement learning on static data assets should adopt support-constrained learning frameworks rather than pure behavioral cloning or unconstrained off-policy algorithms. Practitioners are advised to implement sample-based support constraints with low sample counts, which provide the best balance between safety and policy improvement. However, decision-makers should note certain limitations: the method can still exhibit performance degradation during prolonged training runs due to the lack of an established early-stopping criterion, and action-support constraints may become conservative when applied to complex mixtures of multiple behavioral policies. Future efforts should prioritize developing validation-based early stopping mechanisms and scaling the framework to large-scale, real-world industrial deployments.

Cover for Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction

Abstract

Off-policy reinforcement learning aims to leverage experience collected from prior policies for sample-efficient learning. However, in practice, commonly used off-policy approximate dynamic programming methods based on Q-learning and actor-critic methods are highly sensitive to the data distribution, and can make only limited progress without collecting additional on-policy data. As a step towards more robust off-policy algorithms, we study the setting where the off-policy experience is fixed and there is no further interaction with the environment. We identify bootstrapping error as a key source of instability in current methods. Bootstrapping error is due to bootstrapping from actions that lie outside of the training data distribution, and it accumulates via the Bellman backup operator. We theoretically analyze bootstrapping error, and demonstrate how carefully constraining action selection in the backup can mitigate it. Based on our analysis, we propose a practical algorithm, bootstrapping error accumulation reduction (BEAR). We demonstrate that BEAR is able to learn robustly from different off-policy distributions, including random and suboptimal demonstrations, on a range of continuous control tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background
  • 4 Out-of-Distribution Actions in Q-Learning
  • 4.1 Distribution-Constrained Backups
  • 5 Bootstrapping Error Accumulation Reduction (BEAR)
  • 6 Experiments
  • 6.1 Performance on Medium-Quality Data
  • 6.2 Performance on Random and Optimal Datasets
  • 6.3 Analysis of BEAR-QL
  • 7 Discussion and Future Work
  • References
  • A Distribution-Constrained Backup Operator
  • B Error Propagation
  • C Additional Details Regarding BEAR-QL
  • C.1 Why can we choose actions from Πϵ{\Pi_{\epsilon}}, the support of the training distribution, and need not restrict action selection to the policy distribution?
  • C.2 Details on connection between BEAR-QL and distribution-constrained backups
  • C.3 How effective is the MMD\operatorname{MMD} constraint in constraining supports of distributions?
  • D Additional Experimental Details
  • E Additional Experimental Results

Knowls

  1. Knowl 1 — Bootstrapping Error Accumulation Reduction Q-Learning (BEAR-QL)

    algorithm

    Bootstrapping Error Accumulation Reduction Q-Learning (BEAR-QL) is an offline actor-critic algorithm designed to prevent bootstrapping error accumulation when training on static datasets without environment interaction.

    BEAR-QL maintains an ensemble of KK Q-functions parameterized by {θi}i=1K\{\theta_i\}_{i=1}^K, an actor policy πϕ(a∣s)\pi_\phi(a|s), target Q-networks {θi′}i=1K\{\theta'_i\}_{i=1}^K, a target actor πϕ′(a∣s)\pi_{\phi'}(a|s), and a dual Lagrange multiplier α\alpha for the support constraint. Target networks are updated via exponential moving average with rate τ\tau.

    Input: Static dataset D\mathcal{D}, target update rate τ\tau, batch size NN, MMD sample size nn, target action sample size pp, ensemble weighting factor λ∈[0,1]\lambda \in [0, 1], MMD threshold ε\varepsilon
    Initialize: Q-ensemble parameters {θi}i=1K\{\theta_i\}_{i=1}^K, actor parameters ϕ\phi, Lagrange multiplier α\alpha, target parameters θi′←θi\theta'_i \leftarrow \theta_i for all i∈{1,…,K}i \in \{1, \dots, K\}, target actor ϕ′←ϕ\phi' \leftarrow \phi
    for t=1,…,Niterationst = 1, \dots, N_{\text{iterations}} do
        Sample mini-batch of transitions (s,a,r,s′)∼D(s, a, r, s') \sim \mathcal{D}
        Sample pp candidate target actions {ai′∼πϕ′(⋅∣s′)}i=1p\{a'_i \sim \pi_{\phi'}(\cdot|s')\}_{i=1}^p
        Compute target value for each candidate: y(s,a)=max⁡ai′[λmin⁡j=1,…,KQθj′(s′,ai′)+(1−λ)max⁡j=1,…,KQθj′(s′,ai′)]y(s, a) = \max_{a'_i} [\lambda \min_{j=1,\dots,K} Q_{\theta'_j}(s', a'_i) + (1 - \lambda) \max_{j=1,\dots,K} Q_{\theta'_j}(s', a'_i)]
        for i=1,…,Ki = 1, \dots, K do
            Update θi←arg⁡min⁡θi(Qθi(s,a)−(r+γy(s,a)))2\theta_i \leftarrow \arg\min_{\theta_i} (Q_{\theta_i}(s, a) - (r + \gamma y(s, a)))^2
        end for
        Sample actions {a^i∼πϕ(⋅∣s)}i=1m\{\hat{a}_i \sim \pi_\phi(\cdot|s)\}_{i=1}^m and {aj∼D(s)}j=1n\{a_j \sim \mathcal{D}(s)\}_{j=1}^n with n∈[1,10]n \in [1, 10]
        Update ϕ,α\phi, \alpha by dual gradient descent on the objective: max⁡ϕEs∼D,a∼πϕ(⋅∣s)[min⁡j=1,…,KQθj(s,a)]s.t.Es∼D[MMD2(D(s),πϕ(⋅∣s))]≤ε\max_\phi \mathbb{E}_{s \sim \mathcal{D}, a \sim \pi_\phi(\cdot|s)} [\min_{j=1,\dots,K} Q_{\theta_j}(s, a)] \quad \text{s.t.} \quad \mathbb{E}_{s \sim \mathcal{D}} [\text{MMD}^2(\mathcal{D}(s), \pi_\phi(\cdot|s))] \le \varepsilon
        Update target networks: θi′←τθi+(1−τ)θi′\theta'_i \leftarrow \tau \theta_i + (1 - \tau) \theta'_i for all ii, ϕ′←τϕ+(1−τ)ϕ′\phi' \leftarrow \tau \phi + (1 - \tau) \phi'
    end for

    At evaluation time, policy action selection can be executed greedily by sampling pp candidate actions from πϕ(⋅∣s)\pi_\phi(\cdot|s) and selecting the action a=arg⁡max⁡aimin⁡j=1,…,KQθj(s,ai)a = \arg\max_{a_i} \min_{j=1,\dots,K} Q_{\theta_j}(s, a_i).

  2. Knowl 2 — Performance Bound for Approximate Distribution-Constrained Q-Iteration

    theoretical result

    Let a Markov decision process (MDP) be defined by (S,A,P,R,ρ0,γ)(S, A, P, R, \rho_0, \gamma), where SS is the state space, AA is the action space, P(s′∣s,a)P(s'|s, a) is the transition dynamics, R(s,a)R(s, a) is the reward function, ρ0(s)\rho_0(s) is the initial state distribution, and γ∈(0,1)\gamma \in (0, 1) is the discount factor. Let μ(s,a)\mu(s, a) denote the static training data distribution with state marginal μ(s)\mu(s).

    Let Π\Pi be a restricted policy set used in the distribution-constrained Bellman backup operator TΠQ(s,a)=R(s,a)+γmax⁡π∈ΠEs′∼P(⋅∣s,a)[max⁡π′∈ΠEa′∼π′(⋅∣s′)[Q(s′,a′)]]\mathcal{T}^\Pi Q(s, a) = R(s, a) + \gamma \max_{\pi \in \Pi} \mathbb{E}_{s' \sim P(\cdot|s, a)} [\max_{\pi' \in \Pi} \mathbb{E}_{a' \sim \pi'(\cdot|s')}[Q(s', a')]].

    Define the suboptimality constant α(Π)\alpha(\Pi) measuring the suboptimality bias induced by restricting policies to Π\Pi: α(Π)=max⁡s,a∣TΠQ∗(s,a)−TQ∗(s,a)∣\alpha(\Pi) = \max_{s, a} |\mathcal{T}^\Pi Q^*(s, a) - \mathcal{T} Q^*(s, a)| where Q∗Q^* is the optimal state-action value function and T\mathcal{T} is the unconstrained Bellman optimality operator.

    Assume there exist coefficients c(k)c(k) such that for any sequence π1,…,πk∈Π\pi_1, \dots, \pi_k \in \Pi and s∈Ss \in S, ρ0Pπ1…Pπk(s)≤c(k)μ(s)\rho_0 P^{\pi_1} \dots P^{\pi_k}(s) \le c(k) \mu(s), and define the concentrability coefficient: C(Π)=(1−γ)2∑k=1∞kγk−1c(k)C(\Pi) = (1 - \gamma)^2 \sum_{k=1}^\infty k \gamma^{k-1} c(k)

    If approximate distribution-constrained value iteration is run such that the Bellman error satisfies δ(s,a)≥max⁡k∣Qk(s,a)−TΠQk−1(s,a)∣\delta(s, a) \ge \max_k |Q_k(s, a) - \mathcal{T}^\Pi Q_{k-1}(s, a)|, then the suboptimality of the resulting policy sequence πk\pi_k is bounded by: lim⁡k→∞Eρ0[∣Vπk(s)−V∗(s)∣]≤γ(1−γ)2[C(Π)Eμ[max⁡π∈ΠEπ[δ(s,a)]]+1−γγα(Π)]\lim_{k\to\infty} \mathbb{E}_{\rho_0}[|V^{\pi_k}(s) - V^*(s)|] \le \frac{\gamma}{(1 - \gamma)^2} \left[ C(\Pi) \mathbb{E}_\mu \left[ \max_{\pi \in \Pi} \mathbb{E}_\pi [\delta(s, a)] \right] + \frac{1 - \gamma}{\gamma} \alpha(\Pi) \right]

    This bound characterizes the fundamental trade-off in offline reinforcement learning: enlarging Π\Pi reduces the suboptimality bias α(Π)\alpha(\Pi) but increases distribution shift and error amplification through the concentrability coefficient C(Π)C(\Pi).

  3. Knowl 3 — Concentrability Coefficient Bound for Support-Constrained Policy Sets

    theoretical result

    Let the training data distribution μ(s,a)\mu(s, a) be generated by a behavior policy β(a∣s)\beta(a|s) with state marginal distribution μ(s)\mu(s) in an MDP with discount factor γ∈(0,1)\gamma \in (0, 1) and initial state distribution ρ0\rho_0.

    Define the ϵ\epsilon-support constrained policy class Πϵ\Pi_\epsilon as the set of policies whose action support is contained within the probable regions of β\beta: Πϵ={π∣π(a∣s)=0 whenever β(a∣s)<ϵ}\Pi_\epsilon = \{ \pi \mid \pi(a|s) = 0 \text{ whenever } \beta(a|s) < \epsilon \}

    Let μΠϵ\mu_{\Pi_\epsilon} be the highest discounted marginal state distribution starting from ρ0\rho_0 and following policies π∈Πϵ\pi \in \Pi_\epsilon at every timestep. Define f(ϵ)=min⁡s∈S,μΠϵ(s)>0[μ(s)]>0f(\epsilon) = \min_{s \in S, \mu_{\Pi_\epsilon}(s) > 0} [\mu(s)] > 0 as the minimum state visitation density under the behavior policy on states reachable under Πϵ\Pi_\epsilon.

    Then the concentrability coefficient C(Πϵ)C(\Pi_\epsilon) is bounded relative to the concentrability coefficient of the behavior policy C(β)C(\beta) by: C(Πϵ)≤C(β)(1+γ(1−γ)f(ϵ)(1−ϵ))C(\Pi_\epsilon) \le C(\beta) \left( 1 + \frac{\gamma}{(1 - \gamma) f(\epsilon)} (1 - \epsilon) \right)

    This confirms that constraining the learned policy to the support of the behavior distribution guarantees bounded concentrability without requiring the learned policy distribution π(a∣s)\pi(a|s) to strictly match the behavior policy density β(a∣s)\beta(a|s).

  4. Knowl 4 — Distribution-Constrained Bellman Backup Operator

    definition

    Given a subset of policy distributions Π\Pi, the distribution-constrained Bellman backup operator TΠ\mathcal{T}^\Pi acting on a state-action value function Q:S×A→RQ: S \times A \to \mathbb{R} is defined as: TΠQ(s,a)=R(s,a)+γmax⁡π∈ΠEs′∼P(⋅∣s,a)[V(s′)]\mathcal{T}^\Pi Q(s, a) = R(s, a) + \gamma \max_{\pi \in \Pi} \mathbb{E}_{s' \sim P(\cdot|s, a)} [V(s')] where V(s)=max⁡π∈ΠEa∼π(⋅∣s)[Q(s,a)]V(s) = \max_{\pi \in \Pi} \mathbb{E}_{a \sim \pi(\cdot|s)} [Q(s, a)] and R(s,a)R(s, a) is the reward function, P(s′∣s,a)P(s'|s, a) is the transition distribution, and γ∈(0,1)\gamma \in (0, 1) is the discount factor.

    This operator is mathematically equivalent to the standard Bellman optimality backup in a modified MDP M′=(S,A′,P′,R,ρ0,γ)M' = (S, \mathcal{A}', P', R, \rho_0, \gamma) where the action space is replaced by the policy class A′=Π\mathcal{A}' = \Pi and transition probabilities are given by P′(s′∣s,π)=Ea∼π(⋅∣s)[P(s′∣s,a)]P'(s'|s, \pi) = \mathbb{E}_{a \sim \pi(\cdot|s)}[P(s'|s, a)]. As a result, TΠ\mathcal{T}^\Pi is a γ\gamma-contraction in the ℓ∞\ell_\infty-norm and possesses a unique fixed point QΠQ^\Pi.

  5. Knowl 5 — Support-Constrained Policy Optimization via Sampled Maximum Mean Discrepancy

    model/method

    In Bootstrapping Error Accumulation Reduction (BEAR), the policy improvement step constrains the actor πϕ\pi_\phi to share the support of the behavior dataset D\mathcal{D} rather than matching its density. The policy is optimized by solving: πϕ:=arg⁡max⁡πEs∼D[Ea∼π(⋅∣s)[min⁡j=1,…,KQ^j(s,a)]]s.t.Es∼D[MMD(D(s),π(⋅∣s))]≤ε\pi_\phi := \arg\max_\pi \mathbb{E}_{s \sim \mathcal{D}} \left[ \mathbb{E}_{a \sim \pi(\cdot|s)} \left[ \min_{j=1,\dots,K} \hat{Q}_j(s, a) \right] \right] \quad \text{s.t.} \quad \mathbb{E}_{s \sim \mathcal{D}} [\text{MMD}(\mathcal{D}(s), \pi(\cdot|s))] \le \varepsilon where ε\varepsilon is a tolerance threshold (chosen as ε=0.05\varepsilon = 0.05), and {Q^j}j=1K\{\hat{Q}_j\}_{j=1}^K is an ensemble of learned Q-functions.

    Given samples {xi}i=1n∼P\{x_i\}_{i=1}^n \sim P and {yj}j=1m∼Q\{y_j\}_{j=1}^m \sim Q, the empirical Maximum Mean Discrepancy (MMD) with universal kernel k(⋅,⋅)k(\cdot, \cdot) is: MMD2({x1,…,xn},{y1,…,ym})=1n2∑i,i′k(xi,xi′)−2nm∑i,jk(xi,yj)+1m2∑j,j′k(yj,yj′)\text{MMD}^2(\{x_1, \dots, x_n\}, \{y_1, \dots, y_m\}) = \frac{1}{n^2} \sum_{i,i'} k(x_i, x_{i'}) - \frac{2}{nm} \sum_{i,j} k(x_i, y_j) + \frac{1}{m^2} \sum_{j,j'} k(y_j, y_{j'}) Common choices for k(x,y)k(x, y) are the Laplacian kernel k(x,y)=exp⁡(−∥x−y∥σ)k(x, y) = \exp\left(-\frac{\|x - y\|}{\sigma}\right) or Gaussian kernel k(x,y)=exp⁡(−∥x−y∥22σ2)k(x, y) = \exp\left(-\frac{\|x - y\|^2}{2\sigma^2}\right).

    When estimated with a small-to-intermediate number of samples (n∈[1,10]n \in [1, 10], typically n=4n=4 or 55), the sampled MMD between PP and QQ behaves similarly to the MMD between a uniform distribution over PP's support and QQ. This penalizes placing probability mass outside the support of PP without forcing QQ to match the specific probability density of PP within that support.

  6. Knowl 6 — Bootstrapping Error Accumulation in Offline Q-Learning

    definition

    In offline value-based reinforcement learning with a static dataset D={(s,a,s′,r)}\mathcal{D} = \{(s, a, s', r)\} collected under an unknown behavior policy β\beta, the target for the Bellman backup is computed using the current Q-function estimate: r+γmax⁡a′Qk(s′,a′)r + \gamma \max_{a'} Q_k(s', a').

    Let ζk(s,a)=∣Qk(s,a)−Q∗(s,a)∣\zeta_k(s, a) = |Q_k(s, a) - Q^*(s, a)| denote the total error between iterate QkQ_k and the optimal value function Q∗Q^* at iteration kk, and let δk(s,a)=∣Qk(s,a)−TQk−1(s,a)∣\delta_k(s, a) = |Q_k(s, a) - \mathcal{T}Q_{k-1}(s, a)| denote the single-step Bellman projection error. The error bound propagates across iterations as: ζk(s,a)≤δk(s,a)+γmax⁡a′Es′∼P(⋅∣s,a)[ζk−1(s′,a′)]\zeta_k(s, a) \le \delta_k(s, a) + \gamma \max_{a'} \mathbb{E}_{s' \sim P(\cdot|s, a)} [\zeta_{k-1}(s', a')]

    When computing max⁡a′Q(s′,a′)\max_{a'} Q(s', a'), the maximizer can select out-of-distribution (OOD) actions for which β(a′∣s′)≈0\beta(a'|s') \approx 0. Because the function approximator is never trained on these state-action pairs, single-step errors δk(s′,a′)\delta_k(s', a') on OOD inputs are arbitrarily large and cannot be corrected by collecting new data, causing runaway bootstrapping error accumulation and divergence.

  7. Knowl 7 — Robust Offline Policy Performance Across Diverse Dataset Distributions

    empirical result

    BEAR-QL was evaluated across continuous control environments in MuJoCo (HalfCheetah-v2, Walker2d-v2, Hopper-v2, Ant-v2, Humanoid-v2) under three distinct static dataset regimes (1,000,000 transitions each):

    1. Medium-quality (suboptimal demonstration data): BEAR-QL consistently achieves the highest average returns, outperforming Batch-Constrained Q-learning (BCQ), naive actor-critic (TD3), Behavior Cloning (BC), KL-control, and DQfD by large margins. In this regime, BCQ matches the performance of BC, failing to improve significantly beyond the behavior policy return, whereas BEAR-QL leverages off-policy data to find improved policies.
    2. Random exploration data: BEAR-QL and naive RL both successfully learn policies that significantly exceed the average dataset return, whereas BCQ performs poorly because its density constraint forces it to imitate the random policy.
    3. Near-optimal/expert data: BEAR-QL matches the near-optimal return of expert data along with BCQ, whereas naive RL fails to learn due to out-of-distribution action extrapolation when bootstrapping.

    BEAR-QL is the only evaluated offline method that performs robustly across all three data quality regimes.

  8. Knowl 8 — Sampled MMD Support Constraint vs KL-Divergence Density Constraint in Offline RL

    empirical result

    Ablation experiments comparing support matching via Maximum Mean Discrepancy (MMD) to distribution matching via Kullback-Leibler (KL) divergence constraint (Es∼D[DKL(π(⋅∣s)∥β(⋅∣s))]≤ε\mathbb{E}_{s \sim \mathcal{D}}[\mathcal{D}_{\text{KL}}(\pi(\cdot|s) \parallel \beta(\cdot|s))] \le \varepsilon) show that:

    1. KL-divergence constrains the density of the learned policy to be close to the behavior policy. On medium-quality data, this prevents the policy from placing higher probability mass on higher-reward actions within the data support, leading to lower returns or instability (diverging Q-values in Hopper-v2 and Walker2d-v2), even when the Lagrange multiplier is extensively hand-tuned.
    2. MMD support matching provides a more stable training signal that does not penalize non-uniformity within the support of β\beta, achieving significantly higher evaluation returns on suboptimal datasets.
    3. When varying the sample size nn used to estimate MMD, smaller sample sizes (n≈4n \approx 4 or 55) outperform larger sample sizes (n=10n = 10) on environments such as Hopper-v2 and Walker2d-v2, supporting the hypothesis that low-to-intermediate sample sizes relax density matching while enforcing support coverage.
  9. Knowl 9 — Limitations of Action-Space Bootstrapping Error Reduction

    limitation

    The methodology and practical implementation of BEAR-QL have several noted limitations:

    1. Long-run training degradation: Although BEAR-QL stabilizes offline training, policies can still exhibit performance degradation over long training runs. There is currently no standard validation error or early-stopping criterion for model selection in offline reinforcement learning.
    2. Over-conservatism of per-step action constraints: Constraining policy action selection at every single state to the behavior support can be overly conservative compared to directly constraining state-visitation distributions, particularly when datasets are generated by mixtures of multiple distinct policies.
    3. High-dimensional support matching: Evaluating and constraining support overlap in high-dimensional continuous action spaces using kernel-based MMD distances is challenging and sensitive to kernel bandwidth selection.

Coverage note — None was omitted.

References

  1. 1.Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. Striving for simplicity in off-policy deep reinforcement learning. CoRR, abs/1907.04543, 2019. URL http://arxiv.org/abs/1907.04543.
  2. 2.Andr"{a}s Antos, Csaa Szepesvari, and Remi Munos. Value-iteration based fitted policy iteration: Learning with a single trajectory. In 2007 IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning, pages 330–337, April 2007. doi: 10.1109/ADPRL.2007.368207.
  3. 3.Andr'{a}s Antos, Csaba Szepesv'{a}ri, and R'{e}mi Munos. Fitted q-iteration in continuous action-space mdps. In Advances in Neural Information Processing Systems 20, pages 9–16. Curran Associates, Inc., 2008.
  4. 4.James Bennett, Stan Lanning, et al. The netflix prize. 2007.
  5. 5.Dimitri P Bertsekas and John N Tsitsiklis. Neuro-dynamic programming. Athena Scientific, 1996.
  6. 6.Jonathon Byrd and Zachary Lipton. What is the effect of importance weighting in deep learning? In ICML 2019.
  7. 7.Tim de Bruin, Jens Kober, Karl Tuyls, and Robert Babuska. The importance of experience replay database composition in deep reinforcement learning. 01 2015.
  8. 8.Jia Deng, Wei Dong, Richard S. Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09, 2009.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  10. 10.Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al. Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. In Proceedings of the International Conference on Machine Learning (ICML), 2018.
  11. 11.Amir-massoud Farahmand, Csaba Szepesv'{a}ri, and R'{e}mi Munos. Error propagation for approximate policy and value iteration. In Advances in Neural Information Processing Systems, pages 568–576, 2010.
  12. 12.Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine. Diagnosing bottlenecks in deep q-learning algorithms. arXiv preprint arXiv:1902.10250, 2019.
  13. 13.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. arXiv preprint arXiv:1812.02900, 2018.
  14. 14.Scott Fujimoto, Herke van Hoof, and David Meger. Addressing function approximation error in actor-critic methods. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 1587–1596. PMLR, 2018.
  15. 15.Yang Gao, Huazhe Xu, Ji Lin, Fisher Yu, Sergey Levine, and Trevor Darrell. Reinforcement learning from imperfect demonstrations. In ICLR (Workshop). OpenReview.net, 2018.
  16. 16.Carles Gelada and Marc G. Bellemare. Off-policy deep reinforcement learning by bootstrapping the covariate shift. CoRR, abs/1901.09455, 2019.
  17. 17.Jordi Grau-Moya, Felix Leibfried, and Peter Vrancx. Soft q-learning with mutual-information regularization. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyEtjoCqFX.
  18. 18.Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Sch"{o}lkopf, and Alexander Smola. A kernel two-sample test. J. Mach. Learn. Res., 13:723–773, March 2012. ISSN 1532-4435. URL http://dl.acm.org/citation.cfm?id=2188385.2188410.
  19. 19.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv preprint arXiv:1801.01290, 2018.
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016.
  21. 21.Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al. Deep q-learning from demonstrations. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  22. 22.Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, `{A}gata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. CoRR, abs/1907.00456, 2019. URL http://arxiv.org/abs/1907.00456.
  23. 23.Sham Kakade and John Langford. Approximately optimal approximate reinforcement learning. In Proceedings of the Nineteenth International Conference on Machine Learning, pages 267–274. Morgan Kaufmann Publishers Inc., 2002.
  24. 24.Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, and Sergey Levine. Scalable deep reinforcement learning for vision-based robotic manipulation. In Proceedings of The 2nd Conference on Robot Learning, volume 87 of Proceedings of Machine Learning Research, pages 651–673. PMLR, 2018.
  25. 25.Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes. Safe policy improvement with baseline bootstrapping. In International Conference on Machine Learning (ICML), 2019.
  26. 26.Sergey Levine. Reinforcement learning and control as probabilistic inference: Tutorial and review. CoRR, abs/1805.00909, 2018. URL http://arxiv.org/abs/1805.00909.
  27. 27.A Rupam Mahmood, Huizhen Yu, Martha White, and Richard S Sutton. Emphatic temporal-difference learning. arXiv preprint arXiv:1507.01569, 2015.
  28. 28.R'{e}mi Munos. Error bounds for approximate policy iteration. In Proceedings of the Twentieth International Conference on International Conference on Machine Learning, pages 560–567. AAAI Press, 2003.
  29. 29.R'{e}mi Munos. Error bounds for approximate value iteration. In Proceedings of the National Conference on Artificial Intelligence, 2005.
  30. 30.R'{e}mi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare. Safe and efficient off-policy reinforcement learning. In Advances in Neural Information Processing Systems, pages 1054–1062, 2016.
  31. 31.Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta. Off-policy temporal-difference learning with function approximation. In International Conference on Machine Learning (ICML), 2001.
  32. 32.Stefan Schaal. Is imitation learning the route to humanoid robots?, 1999.
  33. 33.Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. Prioritized experience replay. CoRR, abs/1511.05952, 2016.
  34. 34.Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist. Approximate modified policy iteration and its application to the game of tetris. Journal of Machine Learning Research, 16:1629–1676, 2015. URL http://jmlr.org/papers/v16/scherrer15a.html.
  35. 35.John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. Trust region policy optimization. In Francis Bach and David Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1889–1897, Lille, France, 07–09 Jul 2015. PMLR.
  36. 36.Richard S Sutton and Andrew G Barto. Reinforcement learning: An introduction. Second edition, 2018.
  37. 37.Emanuel Todorov, Tom Erez, and Yuval Tassa. MuJoCo: A physics engine for model-based control. In IROS, pages 5026–5033, 2012.
  38. 38.Yifan Wu, Ezra Winston, Divyansh Kaushik, and Zachary Lipton. Domain adaptation with asymmetrically-relaxed distribution alignment. In ICML 2019.
  39. 39.Fisher Yu, Wenqi Xian, Yingying Chen, Fangchen Liu, Mike Liao, Vashisht Madhavan, and Trevor Darrell. BDD100K: A diverse driving video database with scalable annotation tooling. CoRR, abs/1805.04687, 2018. URL http://arxiv.org/abs/1805.04687.

Citation

MLA
Kumar, A., et al. “Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction”. arXiv, 2019, http://arxiv.org/abs/1906.00949v2.
APA
Kumar, A., Fu, J., Tucker, G., & Levine, S. (2019). Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction. arXiv. http://arxiv.org/abs/1906.00949v2
Chicago
Kumar, A., J. Fu, G. Tucker, and S. Levine. 2019. “Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction”. arXiv. http://arxiv.org/abs/1906.00949v2.
Harvard
Kumar, A. et al. (2019) “Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1906.00949v2.
Vancouver
1. Kumar A, Fu J, Tucker G, Levine S (2019) Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction. arXiv

BibTeX

@article{kumar2019stabilizing,
  title = {Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction},
  author = {Kumar, Aviral and Fu, Justin and Tucker, George and Levine, Sergey},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1906.00949v2},
  eprint = {1906.00949}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission