RORL: Robust Offline Reinforcement Learning via Conservative Smoothing

Rui YangChenjia BaiXiaoteng MaZhaoran WangChongjie ZhangLei Han

article2022NeurIPS109 citations

Proposes a conservative smoothing technique for offline reinforcement learning that simultaneously regularizes policy and value functions near the dataset support, achieving state-of-the-art D4RL benchmark performance while defending against adversarial observation perturbations.

Listen

Offline reinforcement learning enables automated systems to learn decision-making strategies strictly from pre-collected historical datasets, eliminating the risks and expenses of live trial-and-error. However, conventional methods heavily prioritize conservatism to avoid unfamiliar actions, resulting in fragile decision models that degrade sharply when exposed to small observation errors, sensor noise, or adversarial inputs. The article introduces and evaluates Robust Offline Reinforcement Learning (RORL), an algorithm designed to balance conservative decision-making with robustness against state and observation disturbances.

To achieve this, RORL introduces a conservative smoothing framework. The method applies smoothness constraints to both the policy and its value functions near observed data while penalizing unfamiliar state-action pairs using uncertainty quantified across an ensemble of value networks. The authors evaluated the approach across standard continuous control benchmark tasks in the D4RL benchmark suite and subjected the system to multiple adversarial observation attack scenarios across different perturbation scales.

Key findings demonstrate that RORL achieves top-tier benchmark performance, yielding an average normalized return of 85.7 across standard continuous control tasks and outperforming strong ensemble-based baselines like EDAC (82.9) and SAC-10 (50.8). Crucially, RORL accomplishes this with only 10 ensemble networks, whereas previous competitive baselines require up to 50 networks in certain complex tasks. Under adversarial observation perturbations, RORL maintains significantly higher stability and performance than baseline methods across all tested attack scales. Theoretical analysis confirms that RORL provides a provably tighter suboptimality bound than earlier pessimistic offline learning methods. Furthermore, ablation experiments show that penalizing out-of-distribution values is the single most critical component for preserving robustness under observation attacks.

These results indicate that offline decision systems can achieve operational robustness without sacrificing baseline performance or requiring prohibitively large computational architectures. By mitigating vulnerability to sensory noise and adversarial shifts, the method reduces operational risk in safety-critical deployments such as robotics and industrial automation. While RORL introduces modest computational overhead due to adversarial state generation—running at roughly 29.6 seconds per training epoch compared to 17.9 seconds for EDAC—it runs substantially faster than previous uncertainty-based methods while maintaining moderate GPU memory requirements.

Organizations evaluating offline learning pipelines for physical or security-sensitive domains should consider incorporating conservative smoothing to protect against observation drift. Future work should focus on accelerating the adversarial state generation step and exploring smoothing directly within compact latent representations rather than raw observation spaces.

arXiv: 2206.02829
  • Paper: Is Value Learning Really the Main Bottleneck in Offline RL?, Seohong Park et al. (2024). It critically examines whether value learning or policy generalization and extraction are the primary failure points in offline reinforcement learning, directly extending the inquiry into value-level vs. policy-level conservatism.
  • Paper: Efficient Diffusion Policies For Offline Reinforcement Learning, Bingyi Kang et al. (2023). It investigates expressive diffusion-based policy representations in offline reinforcement learning to bypass traditional conservative policy parameterization bottlenecks.
  • Paper: On Training in Imagination, Nadav Timor et al. (2026). It analyzes robustness guarantees and Lipschitz regularity under model errors when training policies entirely within learned imaginary rollouts.
Cover for RORL: Robust Offline Reinforcement Learning via Conservative Smoothing

Abstract

Offline reinforcement learning (RL) provides a promising direction to exploit massive amount of offline data for complex decision-making tasks. Due to the distribution shift issue, current offline RL algorithms are generally designed to be conservative in value estimation and action selection. However, such conservatism can impair the robustness of learned policies when encountering observation deviation under realistic conditions, such as sensor errors and adversarial attacks. To trade off robustness and conservatism, we propose Robust Offline Reinforcement Learning (RORL) with a novel conservative smoothing technique. In RORL, we explicitly introduce regularization on the policy and the value function for states near the dataset, as well as additional conservative value estimation on these states. Theoretically, we show RORL enjoys a tighter suboptimality bound than recent theoretical results in linear MDPs. We demonstrate that RORL can achieve state-of-the-art performance on the general offline RL benchmark and is considerably robust to adversarial observation perturbations.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Robustness of Offline RL: A Motivating Example
  • 4 Robust Offline RL via Conservative Smoothing
  • 5 Theoretical Analysis
  • 6 Experiments
  • 6.1 Benchmark Results
  • 6.2 Adversarial Attack
  • 6.3 Ablations
  • 6.4 Computational Cost Comparison
  • 7 Related Works
  • 8 Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — RORL conservative smoothing framework

    model/method

    Robust Offline Reinforcement Learning (RORL) addresses adversarial observation perturbations while retaining the conservatism required by offline reinforcement learning. Let D\mathcal D be a fixed offline dataset of state-action transitions, let πθ\pi_\theta be the learned policy, and let Qϕ1,…,QϕKQ_{\phi_1},\ldots,Q_{\phi_K} be an ensemble of KK critics. For each dataset state-action pair (s,a)(s,a), RORL samples a perturbed state s^\hat s from the ℓ1\ell_1 ball

    Bd(s,ϵ)={s^:d(s,s^)≤ϵ},\mathcal B_d(s,\epsilon)=\{\hat s:d(s,\hat s)\leq \epsilon\},

    where dd is the ℓ1\ell_1 distance and ϵ\epsilon is the perturbation radius. RORL uses three types of training samples: (s,a)(s,a) for ordinary Bellman learning, (s^,a)(\hat s,a) for conservative smoothing of nearby states, and (s^,a^)(\hat s,\hat a) for pessimistic treatment of unfamiliar state-action pairs, where a^∼πθ(⋅∣s^)\hat a\sim\pi_\theta(\cdot\mid\hat s). The policy and critics are both regularized to behave similarly at ss and s^\hat s, while the critic ensemble uncertainty penalizes values at perturbed states with policy-generated actions. This combination is intended to trade off smoothness near the data support against overestimation outside the support.

  2. Knowl 2 — Conservative smoothing of the critic

    equation

    For a dataset transition (s,a,r,s′)(s,a,r,s'), RORL trains critic QϕiQ_{\phi_i} using an ensemble-minimized soft Bellman target

    T^Qϕi(s,a)=r+γ Es′∼P(⋅∣s,a)a′∼πθ(⋅∣s′)[min⁡1≤j≤KQϕˉj(s′,a′)−αlog⁡πθ(a′∣s′)],\widehat{\mathcal T}Q_{\phi_i}(s,a)=r+\gamma\,\mathbb E_{\substack{s'\sim P(\cdot\mid s,a)\\a'\sim\pi_\theta(\cdot\mid s')}}\left[\min_{1\leq j\leq K}Q_{\bar\phi_j}(s',a')-\alpha\log\pi_\theta(a'\mid s')\right],

    where PP is the transition kernel, γ∈[0,1]\gamma\in[0,1] is the discount factor, α≥0\alpha\geq0 is the entropy coefficient, QϕˉjQ_{\bar\phi_j} is the target critic for ensemble member jj, and KK is the number of critics.

    For a perturbed state s^∈Bd(s,ϵ)\hat s\in\mathcal B_d(s,\epsilon) paired with the original action aa, RORL chooses the perturbation that maximizes the critic smoothness loss. Define

    δi(s,s^,a)=Qϕi(s^,a)−Qϕi(s,a),x+=max⁡(x,0),x−=min⁡(x,0).\delta_i(s,\hat s,a)=Q_{\phi_i}(\hat s,a)-Q_{\phi_i}(s,a),\qquad x^+=\max(x,0),\qquad x^- =\min(x,0).

    The asymmetric pointwise loss is

    L ⁣(Qϕi(s^,a),Qϕi(s,a))=(1−τ)(δi(s,s^,a)+)2+τ(δi(s,s^,a)−)2,\mathcal L\!\left(Q_{\phi_i}(\hat s,a),Q_{\phi_i}(s,a)\right)=(1-\tau)\bigl(\delta_i(s,\hat s,a)^+\bigr)^2+\tau\bigl(\delta_i(s,\hat s,a)^-\bigr)^2,

    where τ≤0.5\tau\leq 0.5 weights decreases in value less than increases in value. The adversarial smoothing loss is

    Lsmooth(s,a;ϕi)=max⁡s^∈Bd(s,ϵ)L ⁣(Qϕi(s^,a),Qϕi(s,a)).\mathcal L_{\mathrm{smooth}}(s,a;\phi_i)=\max_{\hat s\in\mathcal B_d(s,\epsilon)}\mathcal L\!\left(Q_{\phi_i}(\hat s,a),Q_{\phi_i}(s,a)\right).

    Thus, overestimated values induced by nearby perturbed states receive stronger smoothing pressure, whereas underestimated perturbed values are penalized less.

  3. Knowl 3 — Uncertainty penalty for perturbed state-action pairs

    model/method

    RORL controls overestimation on perturbed states and policy-generated actions using the standard deviation of an ensemble of KK critics. For a perturbed state s^\hat s and action a^∼πθ(⋅∣s^)\hat a\sim\pi_\theta(\cdot\mid\hat s), define the ensemble mean and uncertainty by

    Qˉ(s^,a^)=1K∑k=1KQϕk(s^,a^),\bar Q(\hat s,\hat a)=\frac{1}{K}\sum_{k=1}^{K}Q_{\phi_k}(\hat s,\hat a), u(s^,a^)=[1K∑k=1K(Qϕk(s^,a^)−Qˉ(s^,a^))2]1/2. u(\hat s,\hat a)=\left[\frac{1}{K}\sum_{k=1}^{K}\left(Q_{\phi_k}(\hat s,\hat a)-\bar Q(\hat s,\hat a)\right)^2\right]^{1/2}.

    For critic ii, RORL forms the detached pessimistic target

    T^oodQϕi(s^,a^)=Qϕi(s^,a^)−u(s^,a^),\widehat{\mathcal T}_{\mathrm{ood}}Q_{\phi_i}(\hat s,\hat a)=Q_{\phi_i}(\hat s,\hat a)-u(\hat s,\hat a),

    and minimizes

    Lood(s;ϕi)=Es^∼Bd(s,ϵ)a^∼πθ(⋅∣s^)[(T^oodQϕi(s^,a^)−Qϕi(s^,a^))2].\mathcal L_{\mathrm{ood}}(s;\phi_i)=\mathbb E_{\substack{\hat s\sim\mathcal B_d(s,\epsilon)\\\hat a\sim\pi_\theta(\cdot\mid\hat s)}}\left[\left(\widehat{\mathcal T}_{\mathrm{ood}}Q_{\phi_i}(\hat s,\hat a)-Q_{\phi_i}(\hat s,\hat a)\right)^2\right].

    The complete critic objective is

    min⁡ϕiE(s,a,r,s′)∼D[(T^Qϕi(s,a)−Qϕi(s,a))2+βQLsmooth(s,a;ϕi)+βoodLood(s;ϕi)],\min_{\phi_i}\mathbb E_{(s,a,r,s')\sim\mathcal D}\left[\left(\widehat{\mathcal T}Q_{\phi_i}(s,a)-Q_{\phi_i}(s,a)\right)^2+\beta_Q\mathcal L_{\mathrm{smooth}}(s,a;\phi_i)+\beta_{\mathrm{ood}}\mathcal L_{\mathrm{ood}}(s;\phi_i)\right],

    where βQ,βood≥0\beta_Q,\beta_{\mathrm{ood}}\geq0 are loss weights. Unlike a method that penalizes only out-of-distribution actions at dataset states, RORL penalizes both unfamiliar states and unfamiliar actions.

  4. Knowl 4 — Robust policy objective

    equation

    RORL trains the policy to maximize a conservative critic value while making its action distribution insensitive to nearby state perturbations. Let DJ(P∥Q)=12[DKL(P∥Q)+DKL(Q∥P)]D_J(P\|Q)=\tfrac12[D_{\mathrm{KL}}(P\|Q)+D_{\mathrm{KL}}(Q\|P)] denote the Jeffrey divergence between action distributions. The policy objective is

    min⁡θ  Es∼Da∼πθ(⋅∣s)[−min⁡1≤j≤KQϕj(s,a)+αlog⁡πθ(a∣s)+βPmax⁡s^∈Bd(s,ϵ)DJ ⁣(πθ(⋅∣s)∥πθ(⋅∣s^))],\min_\theta\;\mathbb E_{\substack{s\sim\mathcal D\\a\sim\pi_\theta(\cdot\mid s)}}\left[-\min_{1\leq j\leq K}Q_{\phi_j}(s,a)+\alpha\log\pi_\theta(a\mid s)+\beta_P\max_{\hat s\in\mathcal B_d(s,\epsilon)}D_J\!\left(\pi_\theta(\cdot\mid s)\middle\|\pi_\theta(\cdot\mid\hat s)\right)\right],

    where βP≥0\beta_P\geq0 weights policy smoothing, α≥0\alpha\geq0 is the entropy coefficient, KK is the number of critic ensemble members, and Bd(s,ϵ)\mathcal B_d(s,\epsilon) is the ℓ1\ell_1 perturbation ball around state ss. The minimum over critics encourages conservative action selection, while the maximization over perturbed states imposes robustness against the most damaging observation perturbation within the allowed ball.

  5. Knowl 5 — Linear-MDP formulation and induced covariance

    model/method

    The theoretical analysis considers a simplified RORL objective in a linear Markov decision process. Let ϕ:S×A→Rd\phi:\mathcal S\times\mathcal A\to\mathbb R^d be a state-action feature map and let Qw(s,a)=ϕ(s,a)⊤wQ_w(s,a)=\phi(s,a)^\top w for parameter w∈Rdw\in\mathbb R^d. At time tt, suppose mm offline samples (sti,ati,yti)(s_t^i,a_t^i,y_t^i) are available, with ytiy_t^i the temporal-difference target. Let Dood\mathcal D_{\mathrm{ood}} contain perturbed states sampled from ℓ1\ell_1 balls around the offline states, policy-generated actions, and corresponding targets (s^,a^,y^)(\hat s,\hat a,\hat y). With discount factor set to 11 in this analysis, the RORL least-squares objective is

    w~t=arg⁡min⁡w∈Rd{∑i=1m(yti−Qw(sti,ati))2+∑i=1m1∣Bd(sti,ϵ)∣∑s^ti∼Bd(sti,ϵ)(Qw(sti,ati)−Qw(s^ti,ati))2+∑(s^,a^,y^)∈Dood(y^−Qw(s^,a^))2}.\widetilde w_t=\arg\min_{w\in\mathbb R^d}\left\{\sum_{i=1}^{m}\left(y_t^i-Q_w(s_t^i,a_t^i)\right)^2+\sum_{i=1}^{m}\frac{1}{|\mathcal B_d(s_t^i,\epsilon)|}\sum_{\hat s_t^i\sim\mathcal B_d(s_t^i,\epsilon)}\left(Q_w(s_t^i,a_t^i)-Q_w(\hat s_t^i,a_t^i)\right)^2+\sum_{(\hat s,\hat a,\hat y)\in\mathcal D_{\mathrm{ood}}}\left(\hat y-Q_w(\hat s,\hat a)\right)^2\right\}.

    The corresponding feature covariance is

    Λ~t=∑i=1mϕ(sti,ati)ϕ(sti,ati)⊤+∑(s^,a^)∈Doodϕ(s^,a^)ϕ(s^,a^)⊤+∑i=1m1∣Bd(sti,ϵ)∣∑s^ti∼Bd(sti,ϵ)Δϕti(s^ti)Δϕti(s^ti)⊤,\widetilde\Lambda_t=\sum_{i=1}^{m}\phi(s_t^i,a_t^i)\phi(s_t^i,a_t^i)^\top+\sum_{(\hat s,\hat a)\in\mathcal D_{\mathrm{ood}}}\phi(\hat s,\hat a)\phi(\hat s,\hat a)^\top+\sum_{i=1}^{m}\frac{1}{|\mathcal B_d(s_t^i,\epsilon)|}\sum_{\hat s_t^i\sim\mathcal B_d(s_t^i,\epsilon)}\Delta\phi_t^i(\hat s_t^i)\Delta\phi_t^i(\hat s_t^i)^\top,

    where Δϕti(s^)=ϕ(s^,ati)−ϕ(sti,ati)\Delta\phi_t^i(\hat s)=\phi(\hat s,a_t^i)-\phi(s_t^i,a_t^i). The first two terms are the usual offline and OOD covariances; the third term is the additional covariance created by conservative state smoothing.

  6. Knowl 6 — Positive-definite covariance guarantee

    theoretical result

    Under the linear-MDP formulation, assume that for every offline index i∈{1,…,m}i\in\{1,\ldots,m\}, the collection of feature-difference vectors

    {ϕ(s^ti,ati)−ϕ(sti,ati):s^ti∼Dood(sti)}\left\{\phi(\hat s_t^i,a_t^i)-\phi(s_t^i,a_t^i):\hat s_t^i\sim\mathcal D_{\mathrm{ood}}(s_t^i)\right\}

    has full rank in Rd\mathbb R^d. Let Λ~tdiff\widetilde\Lambda_t^{\mathrm{diff}} denote the covariance formed by the third, conservative-smoothing term in Λ~t\widetilde\Lambda_t. Then there exists a constant λ>0\lambda>0 such that

    Λ~tdiff⪰λI,\widetilde\Lambda_t^{\mathrm{diff}}\succeq \lambda I,

    where II is the d×dd\times d identity matrix and ⪰\succeq denotes positive-semidefinite ordering. If Λ~tPBRL\widetilde\Lambda_t^{\mathrm{PBRL}} denotes the covariance consisting only of the offline and OOD-sample terms used by PBRL, RORL satisfies

    Λ~t=Λ~tPBRL+Λ~tdiff⪰Λ~tPBRL,\widetilde\Lambda_t=\widetilde\Lambda_t^{\mathrm{PBRL}}+\widetilde\Lambda_t^{\mathrm{diff}}\succeq\widetilde\Lambda_t^{\mathrm{PBRL}},

    and Λ~t⪰λI\widetilde\Lambda_t\succeq\lambda I. Thus, the conservative smoothing term supplies positive curvature even when the offline and ordinary OOD samples do not yield a well-conditioned covariance.

  7. Knowl 7 — Tighter linear-MDP suboptimality bound

    theoretical result

    Under the linear-MDP assumptions and the full-rank condition for perturbed-state feature differences, RORL's bootstrapped ensemble uncertainty provides a valid lower-confidence-bound uncertainty quantifier. For state-action feature vector ϕ(st,at)\phi(s_t,a_t) and RORL covariance Λ~t\widetilde\Lambda_t, the corresponding penalty has the form

    βtLCB(st,at)=βt[ϕ(st,at)⊤Λ~t−1ϕ(st,at)]1/2,\beta_t^{\mathrm{LCB}}(s_t,a_t)=\beta_t\left[\phi(s_t,a_t)^\top\widetilde\Lambda_t^{-1}\phi(s_t,a_t)\right]^{1/2},

    where βt\beta_t is the time-dependent confidence coefficient. Let π∗\pi^* be an optimal policy, let π^RORL\widehat\pi_{\mathrm{RORL}} be the policy learned by RORL, let TT be the episode horizon, and let βtLCB-PBRL\beta_t^{\mathrm{LCB\text{-}PBRL}} be the analogous penalty using PBRL's covariance. The paper establishes

    SubOpt⁡(π∗,π^RORL)≤∑t=1TEπ∗[βtLCB(st,at)]<∑t=1TEπ∗[βtLCB-PBRL(st,at)].\operatorname{SubOpt}(\pi^*,\widehat\pi_{\mathrm{RORL}})\leq\sum_{t=1}^{T}\mathbb E_{\pi^*}\left[\beta_t^{\mathrm{LCB}}(s_t,a_t)\right]<\sum_{t=1}^{T}\mathbb E_{\pi^*}\left[\beta_t^{\mathrm{LCB\text{-}PBRL}}(s_t,a_t)\right].

    Consequently, within this linear-MDP analysis, RORL has a strictly tighter suboptimality guarantee than the corresponding PBRL guarantee.

  8. Knowl 8 — D4RL benchmark performance

    data/table

    RORL was evaluated on the D4RL Gym benchmark comprising HalfCheetah, Hopper, and Walker2d, each with random, medium, medium-replay, medium-expert, and expert datasets. The reported metric is normalized average return over four random seeds; training used small perturbation scales ϵP,ϵQ,ϵood∈{0.001,0.005,0.01}\epsilon_P,\epsilon_Q,\epsilon_{\mathrm{ood}}\in\{0.001,0.005,0.01\} and evaluation used unperturbed observations. RORL used 10 ensemble critics. The comparison demonstrates that RORL is competitive with or better than ensemble-based offline RL baselines, especially on Hopper and Walker2d, while using fewer critics than EDAC's strongest configurations.

    Could not parse LaTeX table
  9. Knowl 9 — Robustness to adversarial observation attacks

    empirical result

    RORL was compared with EDAC and SAC-10 on HalfCheetah-medium-v2, Walker2d-medium-v2, and Hopper-medium-v2 under observation perturbations with evaluation scales in [0,0.3][0,0.3]. Five attack settings were tested: uniformly random perturbations in an ℓ1\ell_1 ball; action-difference attacks maximizing the Jeffrey divergence between πθ(⋅∣s)\pi_\theta(\cdot\mid s) and πθ(⋅∣s^)\pi_\theta(\cdot\mid\hat s); minimum-QQ attacks minimizing Q(s,πθ(s^))Q(s,\pi_\theta(\hat s)); and mixed-order versions of the latter two attacks.

    For the non-mixed attacks, the strongest perturbed state was selected from 50 uniformly sampled candidates. For mixed-order attacks, 20 initial states were sampled and optimized for 10 gradient-descent steps with step size ϵ/10\epsilon/10, clipping every iterate to the ℓ1\ell_1 perturbation ball. RORL was trained with ϵP,ϵQ,ϵood∈{0.005,0.02,0.03,0.05,0.07}\epsilon_P,\epsilon_Q,\epsilon_{\mathrm{ood}}\in\{0.005,0.02,0.03,0.05,0.07\}.

    Across all three tasks and all five attack settings, RORL retained higher performance under increasing perturbation scales than EDAC and SAC-10. Random attacks were comparatively ineffective against the ensemble-based methods, whereas mixed-order attacks caused larger performance degradation than the corresponding direct candidate-selection attacks.

  10. Knowl 10 — Ablation of smoothing and OOD penalties

    empirical result

    An ablation on Walker2d-medium-v2 removed each of RORL's three additional losses: policy smoothing, critic QQ-smoothing, and the perturbed-state/action OOD penalty. Each component improved robustness under adversarial observation attacks, but their contributions differed.

    The OOD penalty was the most important component: removing it produced lower performance than full RORL at nearly every perturbation scale and for every attack type. Policy smoothing became particularly important when the perturbation scale exceeded 0.20.2. Removing the critic smoothing term had the smallest effect, which the paper attributes to the underlying SAC-10-style use of a 10-critic ensemble.

  11. Knowl 11 — Computational overhead and stated limitation

    limitation

    The additional adversarial state sampling and robust training increase computational cost. On Hopper-medium-v2, measured on one Tesla V100 32GB GPU, an epoch consisted of 10310^3 training steps and had the following runtime and memory usage:

    Could not parse LaTeX table

    RORL was faster than CQL and substantially faster than PBRL, but slower than SAC-10 and EDAC. Its memory usage was comparable to PBRL and EDAC, with 2.1 GB versus 1.8 GB, or 16.7% more than the latter value. The paper identifies adversarial state sampling as the main limitation and suggests that future work could smooth or penalize policies and critics in latent spaces rather than in normalized observation space.

Coverage note — The motivating CQL-versus-CQL-smooth visualization and some appendix hyperparameter studies were omitted because they support, rather than materially extend, the core method, theory, benchmark, robustness, and ablation results.

References

  1. 1.Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári. Improved algorithms for linear stochastic bandits. In Advances in neural information processing systems, volume 24, pages 2312–2320, 2011.
  2. 2.Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song. Uncertainty-based offline reinforcement learning with diversified q-ensemble. Advances in Neural Information Processing Systems, 34, 2021.
  3. 3.Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos. Minimax regret bounds for reinforcement learning. In International Conference on Machine Learning, 2017.
  4. 4.Chenjia Bai, Lingxiao Wang, Lei Han, Animesh Garg, Jianye Hao, Peng Liu, and Zhaoran Wang. Dynamic bottleneck for robust self-supervised exploration. Advances in Neural Information Processing Systems, 34:17007–17020, 2021.
  5. 5.Chenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao, Animesh Garg, Peng Liu, and Zhaoran Wang. Principled exploration via optimistic bootstrapping and backward induction. In International Conference on Machine Learning, pages 577–587. PMLR, 2021.
  6. 6.Chenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng, Animesh Garg, Peng Liu, and Zhaoran Wang. Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning. In International Conference on Learning Representations, 2022.
  7. 7.Tamer Ba¸sar and Pierre Bernhard. H-infinity optimal control and related minimax design problems: a dynamic game approach. Springer Science & Business Media, 2008.
  8. 8.Vahid Behzadan and Arslan Munir. Vulnerability of deep reinforcement learning to policy induction attacks. In International Conference on Machine Learning and Data Mining in Pattern Recognition, pages 262–275. Springer, 2017.
  9. 9.Jinglin Chen and Nan Jiang. Information-theoretic considerations in batch reinforcement learning. In International Conference on Machine Learning, pages 1042–1051. PMLR, 2019.
  10. 10.Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal. Adversarially trained actor critic for offline reinforcement learning. arXiv preprint arXiv:2202.02446, 2022.
  11. 11.Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune. First return, then explore. Nature, 590(7847):580–586, 2021.
  12. 12.Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
  13. 13.Scott Fujimoto and Shixiang Shane Gu. A minimalist approach to offline reinforcement learning. Advances in Neural Information Processing Systems, 34, 2021.
  14. 14.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In ICML, 2019.
  15. 15.Seyed Kamyar Seyed Ghasemipour, Shixiang Shane Gu, and Ofir Nachum. Why so pessimistic? estimating uncertainties for offline rl through ensembles, and why their independence matters. arXiv preprint arXiv:2205.13703, 2022.
  16. 16.Adam Gleave, Michael Dennis, Cody Wild, Neel Kant, Sergey Levine, and Stuart Russell. Adversarial policies: Attacking deep reinforcement learning. In International Conference on Learning Representations, 2019.
  17. 17.Florin Gogianu, Tudor Berariu, Mihaela C Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu. Spectral normalisation for deep reinforcement learning: an optimisation perspective. In International Conference on Machine Learning, pages 3734–3744. PMLR, 2021.
  18. 18.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.
  19. 19.Chin Pang Ho, Marek Petrik, and Wolfram Wiesemann. Fast Bellman Updates for Robust MDPs. In Proceedings of the 35th International Conference on Machine Learning, pages 1979–1988. PMLR, 2018.
  20. 20.Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel. Adversarial attacks on neural network policies. arXiv preprint arXiv:1702.02284, 2017.
  21. 21.Garud N Iyengar. Robust dynamic programming. Mathematics of Operations Research, 30(2):257–280, 2005.
  22. 22.Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. Advances in neural information processing systems, 34, 2021.
  23. 23.Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan. Provably efficient reinforcement learning with linear function approximation. In Conference on Learning Theory, pages 2137–2143. PMLR, 2020.
  24. 24.Ying Jin, Zhuoran Yang, and Zhaoran Wang. Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096. PMLR, 2021.
  25. 25.Parameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland, Cheng Shi, and Volkan Cevher. Robust reinforcement learning via adversarial training with langevin dynamics. Advances in Neural Information Processing Systems, 33:8127–8138, 2020.
  26. 26.Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims. Morel: Model-based offline reinforcement learning. In NeurIPS, 2020.
  27. 27.Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations, 2021.
  28. 28.Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine. Stabilizing off-policy q-learning via bootstrapping error reduction. In NeurIPS, 2019.
  29. 29.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In NeurIPS, 2020.
  30. 30.Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020.
  31. 31.Jialian Li, Tongzheng Ren, Dong Yan, Hang Su, and Jun Zhu. Policy learning for robust markov decision process with a mismatched generative model. arXiv preprint arXiv:2203.06587, 2022.
  32. 32.Lanqing Li, Rui Yang, and Dijun Luo. Focal: Efficient fully-offline meta-reinforcement learning via distance metric learning and behavior regularization. In International Conference on Learning Representations, 2021.
  33. 33.Xiaoteng Ma, Yiqin Yang, Hao Hu, Jun Yang, Chongjie Zhang, Qianchuan Zhao, Bin Liang, and Qihan Liu. Offline reinforcement learning with value-based episodic memory. In International Conference on Learning Representations, 2022.
  34. 34.Yuzhe Ma, Xuezhou Zhang, Wen Sun, and Jerry Zhu. Policy poisoning in batch reinforcement learning and control. Advances in Neural Information Processing Systems, 32, 2019.
  35. 35.Bhairav Mehta, Manfred Diaz, Florian Golemo, Christopher J Pal, and Liam Paull. Active domain randomization. In Conference on Robot Learning, pages 1162–1176. PMLR, 2020.
  36. 36.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning. nature, 518(7540):529–533, 2015.
  37. 37.Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine. Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020.
  38. 38.Arnab Nilim and Laurent Ghaoui. Robustness in markov decision problems with uncertain transition matrices. Advances in neural information processing systems, 16, 2003.
  39. 39.Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. Deep exploration via bootstrapped dqn. In NeurIPS, 2016.
  40. 40.Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami. Practical black-box attacks against deep learning systems using adversarial examples. arXiv preprint arXiv:1602.02697, 1(2):3, 2016.
  41. 41.Anay Pattanaik, Zhenyi Tang, Shuijing Liu, Gautham Bommannan, and Girish Chowdhary. Robust deep reinforcement learning with adversarial attacks. arXiv preprint arXiv:1712.03632, 2017.
  42. 42.Anay Pattanaik, Zhenyi Tang, Shuijing Liu, Gautham Bommannan, and Girish Chowdhary. Robust deep reinforcement learning with adversarial attacks. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, pages 2040–2042, 2018.
  43. 43.Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta. Robust adversarial reinforcement learning. In International Conference on Machine Learning, pages 2817–2826. PMLR, 2017.
  44. 44.Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine. Epopt: Learning robust neural network policies using model ensembles. arXiv preprint arXiv:1610.01283, 2016.
  45. 45.Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell. Bridging offline reinforcement learning and imitation learning: A tale of pessimism. Advances in Neural Information Processing Systems, 34, 2021.
  46. 46.Aurko Roy, Huan Xu, and Sebastian Pokutta. Reinforcement learning under model mismatch. Advances in neural information processing systems, 30, 2017.
  47. 47.Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al. Mastering atari, go, chess and shogi by planning with a learned model. Nature, 588(7839):604–609, 2020.
  48. 48.Qianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang, and Tuo Zhao. Deep reinforcement learning with robust and smooth policy. In International Conference on Machine Learning, pages 8707–8718. PMLR, 2020.
  49. 49.David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484–489, 2016.
  50. 50.Samarth Sinha, Ajay Mandlekar, and Animesh Garg. S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics. In Conference on Robot Learning, pages 907–917. PMLR, 2022.
  51. 51.Hao Sun, Lei Han, Rui Yang, Xiaoteng Ma, Jian Guo, and Bolei Zhou. Exploiting reward shifting in value-based deep rl. In Advances in Neural Information Processing Systems, 2022.
  52. 52.Hao Sun, Boris van Breugel, Jonathan Crabbe, Nabeel Seedat, and Mihaela van der Schaar. Daux: a density-based approach for uncertainty explanations. arXiv preprint arXiv:2207.05161, 2022.
  53. 53.Hao Sun, Ziping Xu, Meng Fang, Zhenghao Peng, Jiadong Guo, Bo Dai, and Bolei Zhou. Safe exploration by solving early terminated mdp. arXiv preprint arXiv:2107.04200, 2021.
  54. 54.Aviv Tamar, Huan Xu, and Shie Mannor. Scaling up robust mdps by reinforcement learning. arXiv preprint arXiv:1306.6189, 2013.
  55. 55.Michael E Tipping and Christopher M Bishop. Probabilistic principal component analysis. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 61(3):611–622, 1999.
  56. 56.Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23–30. IEEE, 2017.
  57. 57.Eugene Vinitsky, Yuqing Du, Kanaad Parvate, Kathy Jang, Pieter Abbeel, and Alexandre Bayen. Robust reinforcement learning using adversarial populations. arXiv preprint arXiv:2008.01825, 2020.
  58. 58.Jianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu, Siyuan Li, and Chongjie Zhang. Offline reinforcement learning with reverse model-based imagination. Advances in Neural Information Processing Systems, 34, 2021.
  59. 59.Qing Wang, Jiechao Xiong, Lei Han, Han Liu, Tong Zhang, et al. Exponentially weighted imitation learning for batched historical data. Advances in Neural Information Processing Systems, 31, 2018.
  60. 60.Ruosong Wang, Simon S Du, Lin F Yang, and Ruslan Salakhutdinov. On reward-free reinforcement learning with linear function approximation. In Advances in neural information processing systems, 2020.
  61. 61.Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh S Merel, Jost Tobias Springenberg, Scott E Reed, Bobak Shahriari, Noah Siegel, Caglar Gulcehre, Nicolas Heess, et al. Critic regularized regression. Advances in Neural Information Processing Systems, 33:7768–7778, 2020.
  62. 62.Fan Wu, Linyi Li, Chejian Xu, Huan Zhang, Bhavya Kailkhura, Krishnaram Kenthapadi, Ding Zhao, and Bo Li. Copa: Certifying robust policies for offline reinforcement learning against poisoning attacks. arXiv preprint arXiv:2203.08398, 2022.
  63. 63.Yifan Wu, George Tucker, and Ofir Nachum. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  64. 64.Lihua Xie and Carlos E de Souza. Robust h/sub infinity/control for linear systems with norm-bounded time-varying uncertainty. In 29th IEEE Conference on Decision and Control, pages 1034–1035. IEEE, 1990.
  65. 65.Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal. Bellman-consistent pessimism for offline reinforcement learning. Advances in neural information processing systems, 34, 2021.
  66. 66.Rui Yang, Yiming Lu, Wenzhe Li, Hao Sun, Meng Fang, Yali Du, Xiu Li, Lei Han, and Chongjie Zhang. Rethinking goal-conditioned supervised learning and its connection to offline rl. In International Conference on Learning Representations, 2022.
  67. 67.Tianpei Yang, Hongyao Tang, Chenjia Bai, Jinyi Liu, Jianye Hao, Zhaopeng Meng, and Peng Liu. Exploration in deep reinforcement learning: a comprehensive survey. arXiv preprint arXiv:2109.06668, 2021.
  68. 68.Wenhao Yang, Liangyu Zhang, and Zhihua Zhang. Towards theoretical understandings of robust markov decision processes: Sample complexity and asymptotics. arXiv preprint arXiv:2105.03863, 2021.
  69. 69.Yiqin Yang, Xiaoteng Ma, Li Chenghao, Zewu Zheng, Qiyuan Zhang, Gao Huang, Jun Yang, and Qianchuan Zhao. Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning. Advances in Neural Information Processing Systems, 34, 2021.
  70. 70.Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang. Near-optimal offline reinforcement learning with linear representation: Leveraging variance information with pessimism. arXiv preprint arXiv:2203.05804, 2022.
  71. 71.Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn. Combo: Conservative offline model-based policy optimization. Advances in Neural Information Processing Systems, 34, 2021.
  72. 72.Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma. Mopo: Model-based offline policy optimization. In NeurIPS, 2020.
  73. 73.Huan Zhang, Hongge Chen, Chaowei Xiao, Bo Li, Mingyan Liu, Duane Boning, and Cho-Jui Hsieh. Robust deep reinforcement learning against adversarial perturbations on state observations. Advances in Neural Information Processing Systems, 33:21024–21037, 2020.
  74. 74.Xuezhou Zhang, Yiding Chen, Xiaojin Zhu, and Wen Sun. Robust policy gradient against strong data corruption. In International Conference on Machine Learning, pages 12391–12401. PMLR, 2021.
  75. 75.Zhengqing Zhou, Zhengyuan Zhou, Qinxun Bai, Linhai Qiu, Jose Blanchet, and Peter Glynn. Finite-sample regret bound for distributionally robust offline tabular reinforcement learning. In International Conference on Artificial Intelligence and Statistics, pages 3331–3339. PMLR, 2021.

Citation

MLA
Yang, R., et al. “RORL: Robust Offline Reinforcement Learning via Conservative Smoothing”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 23851–66, https://proceedings.neurips.cc/paper_files/paper/2022/file/96bbdd0ed2a9e7cd2fb7caf2fae15f3d-Paper-Conference.pdf.
APA
Yang, R., Bai, C., Ma, X., Wang, Z., Zhang, C., & Han, L. (2022). RORL: Robust Offline Reinforcement Learning via Conservative Smoothing. Advances in Neural Information Processing Systems, 35, 23851–23866. https://proceedings.neurips.cc/paper_files/paper/2022/file/96bbdd0ed2a9e7cd2fb7caf2fae15f3d-Paper-Conference.pdf
Chicago
Yang, R., C. Bai, X. Ma, Z. Wang, C. Zhang, and L. Han. 2022. “RORL: Robust Offline Reinforcement Learning via Conservative Smoothing”. Advances in Neural Information Processing Systems 35: 23851–66. https://proceedings.neurips.cc/paper_files/paper/2022/file/96bbdd0ed2a9e7cd2fb7caf2fae15f3d-Paper-Conference.pdf.
Harvard
Yang, R. et al. (2022) “RORL: Robust Offline Reinforcement Learning via Conservative Smoothing”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 23851–23866. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/96bbdd0ed2a9e7cd2fb7caf2fae15f3d-Paper-Conference.pdf.
Vancouver
1. Yang R, Bai C, Ma X, Wang Z, Zhang C, Han L (2022) RORL: Robust Offline Reinforcement Learning via Conservative Smoothing. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 23851–23866

BibTeX

@inproceedings{yang2022rorl,
  title = {RORL: Robust Offline Reinforcement Learning via Conservative Smoothing},
  author = {Yang, Rui and Bai, Chenjia and Ma, Xiaoteng and Wang, Zhaoran and Zhang, Chongjie and Han, Lei},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {23851-23866},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/96bbdd0ed2a9e7cd2fb7caf2fae15f3d-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors