Adversarially Trained Actor Critic for Offline Reinforcement Learning

Ching-An ChengTengyang XieNan JiangAlekh Agarwal

article2022ICML160 citationsOutstanding Paper Runner Up Award

Develops a Stackelberg game framework using relative pessimism to guarantee safe policy improvement over data-collection baselines across broad hyperparameter ranges while outperforming state-of-the-art offline reinforcement learning methods on complex continuous control benchmarks.

Listen

Deploying reinforcement learning in high-stakes operational environments such as healthcare, robotics, and customer-facing platforms is often limited because collecting live exploratory data carries severe safety, ethical, and financial risks. Organizations must instead rely on offline reinforcement learning, training automated decision-making models solely on historical logs generated by existing baseline practices. However, existing offline methods struggle with datasets that have limited coverage: standard algorithms frequently produce unstable, overly optimistic policies that fail in deployment, while conservatively regularized models remain overly constrained and fail to extract meaningful performance gains.

The article develops and evaluates Adversarially Trained Actor Critic, a model-free offline reinforcement learning algorithm designed to guarantee robust policy improvement. The central objective is to demonstrate that an autonomous decision agent can reliably perform at least as well as historical baseline behavior across wide hyperparameter settings, while actively competing with the best achievable policy covered within the historical data.

To achieve this, the authors frame offline learning as a two-player game between a learner policy and an adversarial evaluation critic. The critic actively searches for data-consistent scenarios where the candidate policy underperforms the historical baseline, enforcing relative pessimism directly against past actions rather than absolute pessimism over overall returns. The methodology couples this theoretical framework with a practical deep neural network implementation utilizing a two-timescale stochastic optimization schedule, parameter projections, and a double Q residual algorithm loss that stabilizes off-policy evaluation. The evaluation assesses empirical performance across standard continuous-control simulation benchmarks spanning robot locomotion and robotic manipulation tasks, comparing the approach against prevailing offline baselines.

The findings show that the proposed method consistently achieves state-of-the-art performance across standard benchmark suites, delivering substantial margin improvements over existing approaches in complex tasks such as simulated walking and hopping. Second, the method provably and empirically maintains safe policy improvement across hyperparameter ranges spanning up to three orders of magnitude, anchored safely at baseline imitation behavior when pessimism penalties are zeroed. Third, ablation testing confirms that the specialized residual loss formulation prevents numerical divergence, enabling stable deep network training that pure temporal-difference bootstrapping fails to sustain. Fourth, across intermediate iterates and random seeds, more than half of the generated policy checkpoints actively outperformed the data-collection baseline in standard continuous control environments.

These findings indicate that organizations can mitigate deployment risk in sequential decision-making systems by establishing mathematical safety floors anchored to current operational standards. By shifting from absolute return pessimism to relative baseline comparison, the method reduces the danger of hyperparameter misconfiguration causing catastrophic real-world failures. Furthermore, when limited live testing is permitted, operators can safely tune performance parameters upward from zero without dipping below baseline operational performance.

Decision-makers and engineering teams should adopt relative-pessimism frameworks for offline policy training, utilizing the zero-penalty anchor point to initialize safe tuning procedures. Teams should deploy the double Q residual surrogate loss within their off-policy architectures to prevent optimization instability. Where immediate full-scale rollout is risky, practitioners should run low-risk online fine-tuning sweeps starting from conservative hyperparameter settings.

Confidence in these findings is high for standard robotic and algorithmic environments adhering to Markovian conditions. However, decision-makers should exercise caution when training on non-Markovian human demonstration datasets, such as cloned human manipulation logs, where the algorithm did not reliably improve upon human baselines due to modeling mismatches. Additionally, solving the underlying adversarial bilevel optimization introduces higher computational demands than standard dynamic programming, requiring appropriate computing resource allocation.

Cover for Adversarially Trained Actor Critic for Offline Reinforcement Learning

Abstract

We propose Adversarially Trained Actor Critic (ATAC), a new model-free algorithm for offline reinforcement learning (RL) under insufficient data coverage, based on the concept of relative pessimism. ATAC is designed as a two-player Stackelberg game: A policy actor competes against an adversarially trained value critic, who finds data-consistent scenarios where the actor is inferior to the data-collection behavior policy. We prove that, when the actor attains no regret in the two-player game, running ATAC produces a policy that provably 1) outperforms the behavior policy over a wide range of hyperparameters that control the degree of pessimism, and 2) competes with the best policy covered by data with appropriately chosen hyperparameters. Compared with existing works, notably our framework offers both theoretical guarantees for general function approximation and a deep RL implementation scalable to complex environments and large datasets. In the D4RL benchmark, ATAC consistently outperforms state-of-the-art offline RL algorithms on a range of continuous control tasks.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. A Game Theoretic Formulation of Offline RL with Robust Policy Improvement
  • 3.1. A Stackelberg Game Formulation of Offline RL
  • 3.2. Relative Pessimism and Robust Policy Improvement
  • 4. Adversarially Trained Actor Critic
  • 4.1. Theory of ATAC with Optimization Oracles
  • 4.1.1. ALGORITHM
  • 4.1.2. THEORETICAL GUARANTEES
  • 4.2. A Practical Implementation of ATAC
  • 4.2.1. CRITIC UPDATE
  • 4.2.2. ACTOR UPDATE
  • 5. Experiments
  • 6. Discussion and Conclusion
  • Acknowledgment
  • References
  • A. Related Works
  • B. Guarantees of Theoretical Algorithm
  • B.1. Concentration Analysis
  • B.2. Decomposition of Performance Difference
  • B.3. Performance Guarantee of the Theoretical Algorithm
  • C. Experiment Details
  • C.1. Implementation Details
  • C.2. Detailed Experimental Results
  • C.3. Robust Policy Improvement
  • D. Comparison between ATAC and CQL
  • D.1. Conceptual Algorithm
  • D.2. Maximin vs. Minimax
  • D.3. Robust Policy Improvement

Knowls

  1. Knowl 1 — Relative-pessimism Stackelberg formulation

    model/method

    ATAC formulates offline reinforcement learning as a two-player Stackelberg game between a policy actor and an adversarial critic. Consider an MDP with state space S\mathcal S, action space A\mathcal A, transition kernel PP, reward R∈[0,Rmax⁡]R\in[0,R_{\max}], discount factor γ∈[0,1)\gamma\in[0,1), and behavior-policy occupancy distribution μ\mu. Let Π\Pi be the policy class, F⊆(S×A→[0,Vmax⁡])\mathcal F\subseteq(\mathcal S\times\mathcal A\to[0,V_{\max}]) the critic class, and Vmax⁡=Rmax⁡/(1−γ)V_{\max}=R_{\max}/(1-\gamma). For a function ff, define f(s,π)=Ea∼π(⋅∣s)[f(s,a)]f(s,\pi)=\mathbb E_{a\sim\pi(\cdot\mid s)}[f(s,a)] and the policy Bellman operator

    (Tπf)(s,a)=R(s,a)+γ Es′∼P(⋅∣s,a)[f(s′,π)].(T^\pi f)(s,a)=R(s,a)+\gamma\,\mathbb E_{s'\sim P(\cdot\mid s,a)}[f(s',\pi)].

    For the discounted return J(π)J(\pi), ATAC uses the relative-performance objective

    Lμ(π,f)=E(s,a)∼μ[f(s,π)−f(s,a)]L_\mu(\pi,f)=\mathbb E_{(s,a)\sim\mu}[f(s,\pi)-f(s,a)]

    and the population Bellman-residual penalty

    Eμ(π,f)=E(s,a)∼μ[(f−Tπf)(s,a)2].E_\mu(\pi,f)=\mathbb E_{(s,a)\sim\mu}\left[(f-T^\pi f)(s,a)^2\right].

    For a pessimism coefficient β≥0\beta\ge 0, the critic associated with a candidate policy is

    fπ∈arg⁡min⁡f∈F{Lμ(π,f)+βEμ(π,f)},f^\pi\in\arg\min_{f\in\mathcal F}\left\{L_\mu(\pi,f)+\beta E_\mu(\pi,f)\right\},

    and the learned policy is

    πˉ∗∈arg⁡max⁡π∈ΠLμ(π,fπ).\bar\pi^*\in\arg\max_{\pi\in\Pi}L_\mu(\pi,f^\pi).

    The critic therefore searches for a Bellman-consistent scenario in which the candidate policy performs poorly relative to the behavior policy, while the actor selects a policy that remains effective against that adversarial critic. When f=Qπf=Q^\pi, Lμ(π,Qπ)=(1−γ)(J(π)−J(μ))L_\mu(\pi,Q^\pi)=(1-\gamma)(J(\pi)-J(\mu)), which explains why the objective is called relative rather than absolute pessimism.

  2. Knowl 2 — Population robust policy improvement

    theoretical result

    Suppose the critic class exactly realizes the Bellman equations required by the policies, meaning that for every π∈Π\pi\in\Pi, the corresponding QQ-function is available in F\mathcal F, and suppose the behavior policy μ\mu itself belongs to Π\Pi. Then, for every candidate policy π∈Π\pi\in\Pi and every pessimism coefficient β≥0\beta\ge 0,

    Lμ(π,fπ)≤(1−γ)(J(π)−J(μ)).L_\mu(\pi,f^\pi)\le (1-\gamma)\bigl(J(\pi)-J(\mu)\bigr).

    Because Lμ(μ,fμ)=0L_\mu(\mu,f^\mu)=0, maximizing the lower bound gives

    J(πˉ∗)≥J(μ).J(\bar\pi^*)\ge J(\mu).

    Thus relative pessimism provides safe, or robust, policy improvement for every nonnegative value of β\beta, including β=0\beta=0. This differs from absolute-pessimism objectives, whose safe-improvement guarantee generally requires carefully tuning the pessimism coefficient. At β=0\beta=0, ATAC ignores reward and transition labels and reduces to an integral-probability-metric-style imitation objective; under exact realizability, matching the behavior occupancy is the robust solution.

  3. Knowl 3 — Function-approximation assumptions for ATAC guarantees

    assumption

    The theoretical guarantees use a value-function class F\mathcal F and policy class Π\Pi together with two approximate assumptions. For a policy π\pi, let dπd^\pi denote its normalized discounted state-action occupancy, let ∥h∥2,ν2=E(s,a)∼ν[h(s,a)2]\|h\|_{2,\nu}^2=\mathbb E_{(s,a)\sim\nu}[h(s,a)^2], and call distributions in {dπ′:π′∈Π}\{d^{\pi'}:\pi'\in\Pi\} admissible.

    Approximate realizability requires that every policy has a critic with small Bellman error simultaneously over all admissible occupancies:

    ∀π∈Π,min⁡f∈Fmax⁡ν admissible∥f−Tπf∥2,ν2≤εF.\forall\pi\in\Pi,\qquad \min_{f\in\mathcal F}\max_{\nu\ \mathrm{admissible}}\|f-T^\pi f\|_{2,\nu}^2\le \varepsilon_F.

    Approximate completeness requires that applying a policy Bellman operator to any critic can be approximated on the data distribution:

    ∀π∈Π, ∀f∈F,min⁡g∈F∥g−Tπf∥2,μ2≤εF,F.\forall\pi\in\Pi,\ \forall f\in\mathcal F,\qquad \min_{g\in\mathcal F}\|g-T^\pi f\|_{2,\mu}^2\le \varepsilon_{\mathcal F,\mathcal F}.

    The exact-realizability version of the main results sets both approximation errors to zero. These assumptions are weaker than requiring exact Bellman closure or uniformly small ℓ∞\ell_\infty error; completeness is required only under the behavior distribution μ\mu.

  4. Knowl 4 — Oracle-based theoretical ATAC algorithm

    algorithm

    The theoretical ATAC procedure alternates adversarial policy evaluation and no-regret policy optimization. Given an offline dataset D={(si,ai,ri,si′)}i=1ND=\{(s_i,a_i,r_i,s_i')\}_{i=1}^N, define

    LD(π,f)=1N∑i=1N[f(si,π)−f(si,ai)],L_D(\pi,f)=\frac1N\sum_{i=1}^N[f(s_i,\pi)-f(s_i,a_i)], ED(π,f)=1N∑i=1N(f(si,ai)−ri−γf(si′,π))2−min⁡g∈F1N∑i=1N(g(si,ai)−ri−γf(si′,π))2.E_D(\pi,f)=\frac1N\sum_{i=1}^N\bigl(f(s_i,a_i)-r_i-\gamma f(s_i',\pi)\bigr)^2-\min_{g\in\mathcal F}\frac1N\sum_{i=1}^N\bigl(g(s_i,a_i)-r_i-\gamma f(s_i',\pi)\bigr)^2.

    The subtraction makes EDE_D an estimated excess Bellman error. The policy-optimization oracle must satisfy, for any adaptively generated critic sequence f1,…,fKf_1,\ldots,f_K and any comparator π∈Π\pi\in\Pi,

    1K(1−γ)∑k=1KEs∼dπ[fk(s,π)−fk(s,πk)]=o(1).\frac{1}{K(1-\gamma)}\sum_{k=1}^K\mathbb E_{s\sim d^\pi}\left[f_k(s,\pi)-f_k(s,\pi_k)\right]=o(1).
    Input: Offline dataset DD, pessimism coefficient β≥0\beta\ge 0, number of iterations KK
    Initialize π1\pi_1 as the uniform policy
    for k=1,2,…,Kk=1,2,\ldots,K do
        Compute fkf_k in arg⁡min⁡f∈F{LD(πk,f)+βED(πk,f)}\arg\min_{f\in\mathcal F}\{L_D(\pi_k,f)+\beta E_D(\pi_k,f)\}
        Give fkf_k and DD to the no-regret policy-optimization oracle
        Set πk+1\pi_{k+1} to the policy returned by the oracle
    end for
    Output the trajectory-level uniform mixture πˉ\bar\pi of π1,…,πK\pi_1,\ldots,\pi_K

    The actor is updated against the sequence of adversarial critics rather than against only the last critic. A finite-action implementation of the oracle uses multiplicative-weights updates πk+1(a∣s)∝πk(a∣s)exp⁡(ηfk(s,a))\pi_{k+1}(a\mid s)\propto\pi_k(a\mid s)\exp(\eta f_k(s,a)) with η=log⁡∣A∣/(2Vmax⁡2K)\eta=\sqrt{\log|\mathcal A|/(2V_{\max}^2K)} and cumulative regret of order O ⁣(Vmax⁡Klog⁡∣A∣/(1−γ))O\!\left(V_{\max}\sqrt{K\log|\mathcal A|}/(1-\gamma)\right).

  5. Knowl 5 — Consistency under partial data coverage

    theoretical result

    Let πˉ\bar\pi be the mixture returned by theoretical ATAC after KK iterations, let N=∣D∣N=|D|, and assume exact realizability and completeness. Define the coverage coefficient for a distribution ν\nu by

    C(ν;μ,F,π)=max⁡f∈F∥f−Tπf∥2,ν2∥f−Tπf∥2,μ2,\mathcal C(\nu;\mu,\mathcal F,\pi)=\max_{f\in\mathcal F}\frac{\|f-T^\pi f\|_{2,\nu}^2}{\|f-T^\pi f\|_{2,\mu}^2},

    and let dF,Πd_{\mathcal F,\Pi} denote the joint statistical complexity of the critic and policy classes. If a distribution ν\nu satisfies max⁡k≤KC(ν;μ,F,πk)≤C\max_{k\le K}\mathcal C(\nu;\mu,\mathcal F,\pi_k)\le C for some C≥1C\ge1, then choosing

    β=Θ ⁣((Vmax⁡N2dF,Π2)1/3)\beta=\Theta\!\left(\left(\frac{V_{\max}N^2}{d_{\mathcal F,\Pi}^2}\right)^{1/3}\right)

    ensures, with high probability, that every comparator policy π∈Π\pi\in\Pi satisfies

    J(π)−J(πˉ)≤εoptπ+O ⁣(Vmax⁡C dF,Π1/3(1−γ)N1/3)+1K(1−γ)∑k=1K⟨dπ∖ν,fk−Tπkfk⟩.J(\pi)-J(\bar\pi)\le \varepsilon_{\mathrm{opt}}^\pi+O\!\left(\frac{V_{\max}\sqrt C\,d_{\mathcal F,\Pi}^{1/3}}{(1-\gamma)N^{1/3}}\right)+\frac{1}{K(1-\gamma)}\sum_{k=1}^K\left\langle d^\pi\setminus\nu, f_k-T^{\pi_k}f_k\right\rangle.

    Here εoptπ\varepsilon_{\mathrm{opt}}^\pi is the comparator's average no-regret error, dπ∖νd^\pi\setminus\nu is the pointwise positive part max⁡{dπ(s,a)−ν(s,a),0}\max\{d^\pi(s,a)-\nu(s,a),0\}, and ⟨d,h⟩=∑s,ad(s,a)h(s,a)\langle d,h\rangle=\sum_{s,a}d(s,a)h(s,a). Taking ν=dπ\nu=d^\pi removes the off-support term, yielding a guarantee for any policy whose occupancy has finite coverage relative to the offline data. Consequently, when the data covers an optimal policy, ATAC competes with that policy up to optimization and statistical error. The N−1/3N^{-1/3} rate is the cost of the regularized, computationally tractable formulation.

  6. Knowl 6 — Finite-sample robust improvement

    theoretical result

    Under exact realizability, with the behavior policy μ\mu included in the policy class, the finite-sample ATAC output satisfies, with high probability,

    J(μ)−J(πˉ)≤O ⁣(Vmax⁡1−γdF,ΠN+βVmax⁡2dF,Π(1−γ)N)+εoptμ.J(\mu)-J(\bar\pi)\le O\!\left(\frac{V_{\max}}{1-\gamma}\sqrt{\frac{d_{\mathcal F,\Pi}}{N}}+\frac{\beta V_{\max}^2d_{\mathcal F,\Pi}}{(1-\gamma)N}\right)+\varepsilon_{\mathrm{opt}}^\mu.

    Thus ATAC is asymptotically no worse than the behavior policy for any pessimism coefficient satisfying β=o(N)\beta=o(N), provided the optimization error also vanishes. The result is weaker in its assumptions than the general consistency theorem: it does not require coverage of arbitrary comparator policies. For β=0\beta=0, the method still has this robust-improvement guarantee even though the critic objective contains no Bellman-error term, because the relative objective reduces to imitation learning and is anchored at the behavior policy. Choosing β≲N1/2\beta\lesssim N^{1/2} gives the sharper statistical regime discussed by the paper.

  7. Knowl 7 — Double-Q residual algorithm loss

    model/method

    To stabilize neural-network optimization of the Bellman penalty, ATAC replaces the empirical Bellman error by the double-Q residual algorithm loss. For critic networks f1,f2f_1,f_2 and slowly updated target networks fˉ1,fˉ2\bar f_1,\bar f_2, define

    EDtd(f,f′,π)=1∣D∣∑(s,a,r,s′)∈D(f(s,a)−r−γf′(s′,π))2,E_D^{\mathrm{td}}(f,f',\pi)=\frac1{|D|}\sum_{(s,a,r,s')\in D}\left(f(s,a)-r-\gamma f'(s',\pi)\right)^2, fˉmin⁡(s,a)=min⁡i∈{1,2}fˉi(s,a),\bar f_{\min}(s,a)=\min_{i\in\{1,2\}}\bar f_i(s,a),

    and, for w∈[0,1]w\in[0,1],

    EDw(f,π)=(1−w)EDtd(f,f,π)+wEDtd(f,fˉmin⁡,π).E_D^w(f,\pi)=(1-w)E_D^{\mathrm{td}}(f,f,\pi)+wE_D^{\mathrm{td}}(f,\bar f_{\min},\pi).

    The first term is a fixed residual-gradient objective, while the second uses delayed double-Q targets. The convex combination is called DQRA. In the paper's experiments, w=0.5w=0.5 provided a stable compromise: w=1w=1 caused conventional bootstrapped double-Q training to become unstable, whereas w=0w=0 was numerically stable but often produced poor performance or slow convergence. Intermediate values substantially improved both temporal-difference stability and policy performance on hopper-medium-replay and related tasks.

  8. Knowl 8 — Scalable deep ATAC implementation

    algorithm

    The practical algorithm uses two critics, a stochastic actor, projected or clipped neural-network parameters, and two-timescale optimization. For a minibatch DbD_b, the critic loss is LDb(π,f)+βEDbw(π,f)L_{D_b}(\pi,f)+\beta E_{D_b}^w(\pi,f), where LDbL_{D_b} is the empirical relative objective and EDbwE_{D_b}^w is the DQRA loss. The actor loss is −LDb(π,f1)-L_{D_b}(\pi,f_1) using only the first critic; the second critic is reserved for stable critic training.

    Input: Offline dataset DD, actor π\pi, critics f1,f2f_1,f_2, β≥0\beta\ge0, w∈[0,1]w\in[0,1], target-update rate τ\tau
    Initialize target critics fˉ1←f1\bar f_1\leftarrow f_1 and fˉ2←f2\bar f_2\leftarrow f_2
    Initialize the actor entropy multiplier α←1\alpha\leftarrow1
    for each training iteration do
        Sample minibatch DbD_b from DD
        for each critic fif_i, i∈{1,2}i\in\{1,2\}, do
            Compute ℓi=LDb(π,fi)+βEDbw(π,fi)\ell_i=L_{D_b}(\pi,f_i)+\beta E_{D_b}^w(\pi,f_i)
            Update fif_i with an ADAM step at the fast rate ηfast\eta_{fast}
            Clip or project the critic weights to the prescribed ℓ2\ell_2 bound
        end for
        Compute the actor loss ℓπ=−LDb(π,f1)\ell_\pi=-L_{D_b}(\pi,f_1)
        Add a Lagrange penalty so the actor entropy stays above EntropyminEntropy_{min}
        Update the actor with an ADAM step at the slow rate ηslow\eta_{slow}
        Update α\alpha at the fast rate and project α\alpha onto [0,∞)[0,\infty)
        For each critic, update its target by fˉi←(1−τ)fˉi+τfi\bar f_i\leftarrow(1-\tau)\bar f_i+\tau f_i
    end for
    Return the trained actor or a selected training checkpoint

    The implementation used three-layer fully connected networks with 256 ReLU units per hidden layer, a Gaussian actor, critic weight norm bound 100 excluding biases, minibatches of 256, ηfast=0.0005\eta_{fast}=0.0005, ηslow=10−3ηfast\eta_{slow}=10^{-3}\eta_{fast}, τ=0.005\tau=0.005, w=0.5w=0.5, and discount factor γ=0.99\gamma=0.99. The faster critic timescale approximates the adversarial evaluation of the current actor, while the slower actor update approximates no-regret optimization.

  9. Knowl 9 — D4RL benchmark performance

    empirical result

    ATAC was evaluated on D4RL MuJoCo and Adroit offline datasets. Each run used 100 behavior-cloning warm-start epochs followed by 900 ATAC epochs, with 2,000 gradient updates per epoch. The coefficient β\beta was selected from {0,4−4,4−3,4−2,4−1,1,4,42,43,44}\{0,4^{-4},4^{-3},4^{-2},4^{-1},1,4,4^2,4^3,4^4\}, results were aggregated over 10 random seeds, and the reported score was the median normalized return. ATAC denotes the last iterate; ATAC∗^* denotes the best of nine checkpoints; ATAC0_0 and ATAC0∗_0^* are the corresponding absolute-pessimism variants.

    Representative results reported on page 8 are:

    Could not parse LaTeX table

    Across the full benchmark, ATAC and ATAC∗^* generally exceeded model-free baselines and often exceeded COMBO, with especially large gains on walker2d-medium, walker2d-medium-replay, hopper-medium-replay, and pen-exp. The main clear weakness was halfcheetah-random, where insufficient convergence left ATAC below CQL and COMBO.

  10. Knowl 10 — Empirical robustness across pessimism coefficients

    empirical result

    The practical ATAC implementation empirically reproduced the theoretical robust-improvement pattern. On the hopper tasks shown on page 2, relative-pessimism ATAC improved over the behavior return across a broad range of β\beta values; on hopper-medium-expert, improvement persisted across approximately three orders of magnitude, from β=0.01\beta=0.01 to β=10\beta=10. The corresponding absolute-pessimism variant required a much narrower, well-tuned interval to remain above the behavior policy.

    The complete MuJoCo plots on page 24 show the same broad relative-pessimism trend across random, medium, medium-replay, and medium-expert datasets. The Adroit plots on page 25 show robust improvement mainly for expert datasets. Human and cloned datasets frequently failed to improve over behavior even with favorable β\beta values; the paper attributes this to a likely violation of the realizability assumption because human demonstrations may be non-Markovian and are not well represented by the Markovian Gaussian policy class.

    Across all tested β\beta values, seeds, and training iterates from epochs 100 through 900, the robust-improvement analysis reported on page 23 found that more than half of the iterates were better than behavior on most non-human, non-cloned datasets. On the remaining datasets, more than 60% of iterates stayed within 80% of behavior performance. This supports tuning in practice by starting at β=0\beta=0 and increasing β\beta without immediately deploying a policy substantially worse than the data-collection behavior.

Coverage note — Detailed CQL maximin-versus-minimax derivations and auxiliary concentration/proof-only lemmas were omitted because they support the main guarantees without constituting separate load-bearing contributions.

References

  1. 1.Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. On the theory of policy gradient methods: Optimality, approximation, and distribution shift. Journal of Machine Learning Research, 22(98):1–76, 2021.
  2. 2.Antos, A., Szepesvari, C., and Munos, R. Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path. Machine Learning, 71(1):89–129, 2008.
  3. 3.Arjovsky, M., Chintala, S., and Bottou, L. Wasserstein generative adversarial networks. In International conference on machine learning, pp. 214–223. PMLR, 2017.
  4. 4.Baird, L. Residual algorithms: Reinforcement learning with function approximation. In Machine Learning Proceedings 1995, pp. 30–37. Elsevier, 1995.
  5. 5.Bertsekas, D. P. and Tsitsiklis, J. N. Neuro-dynamic programming: an overview. In Proceedings of 1995 34th IEEE conference on decision and control, volume 1, pp. 560–564. IEEE, 1995.
  6. 6.Borkar, V. S. Stochastic approximation with two time scales. Systems & Control Letters, 29(5):291–294, 1997.
  7. 7.Cesa-Bianchi, N. and Lugosi, G. Prediction, learning, and games. Cambridge university press, 2006.
  8. 8.Chen, J. and Jiang, N. Information-theoretic considerations in batch reinforcement learning. In International Conference on Machine Learning, pp. 1042–1051, 2019.
  9. 9.Cheng, C.-A., Yan, X., Ratliff, N., and Boots, B. Predictor-corrector policy optimization. In International Conference on Machine Learning, pp. 1151–1161. PMLR, 2019.
  10. 10.Cheng, C.-A., Kolobov, A., and Agarwal, A. Policy improvement via imitation of multiple oracles. Advances in Neural Information Processing Systems, 33, 2020.
  11. 11.Cheng, C.-A., Kolobov, A., and Swaminathan, A. Heuristic-guided reinforcement learning. Advances in Neural Information Processing Systems, 34:13550–13563, 2021.
  12. 12.Even-Dar, E., Kakade, S. M., and Mansour, Y. Online markov decision processes. Mathematics of Operations Research, 34(3):726–736, 2009.
  13. 13.Farahmand, A. M., Munos, R., and Szepesvari, C. Error propagation for approximate policy and value iteration. In Advances in Neural Information Processing Systems, 2010.
  14. 14.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
  15. 15.Fujimoto, S. and Gu, S. S. A minimalist approach to offline reinforcement learning. Advances in neural information processing systems, 34:20132–20145, 2021.
  16. 16.Fujimoto, S., Hoof, H., and Meger, D. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning, pp. 1587–1596. PMLR, 2018.
  17. 17.Fujimoto, S., Meger, D., and Precup, D. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pp. 2052–2062, 2019.
  18. 18.Geist, M., Scherrer, B., and Pietquin, O. A theory of regularized markov decision processes. In International Conference on Machine Learning, pp. 2160–2169. PMLR, 2019.
  19. 19.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, pp. 1861–1870. PMLR, 2018.
  20. 20.Jin, Y., Yang, Z., and Wang, Z. Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp. 5084–5096. PMLR, 2021.
  21. 21.Kakade, S. and Langford, J. Approximately optimal approximate reinforcement learning. In ICML, volume 2, pp. 267–274, 2002.
  22. 22.Kakade, S. M. A natural policy gradient. Advances in neural information processing systems, 14, 2001.
  23. 23.Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. Morel: Model-based offline reinforcement learning. In NeurIPS, 2020.
  24. 24.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, 2015.
  25. 25.Konda, V. R. and Tsitsiklis, J. N. Actor-critic algorithms. In Advances in neural information processing systems, pp. 1008–1014. Citeseer, 2000.
  26. 26.Kostrikov, I., Nair, A., and Levine, S. Offline reinforcement learning with implicit q-learning. arXiv preprint arXiv:2110.06169, 2021.
  27. 27.Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. Stabilizing off-policy q-learning via bootstrapping error reduction. Advances in Neural Information Processing Systems, 32:11784–11794, 2019.
  28. 28.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. Advances in Neural Information Processing Systems, 33:1179–1191, 2020.
  29. 29.Laroche, R., Trichelair, P., and Des Combes, R. T. Safe policy improvement with baseline bootstrapping. In International Conference on Machine Learning, pp. 3652–3661. PMLR, 2019.
  30. 30.Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. Provably good batch reinforcement learning without great exploration. Advances in Neural Information Processing Systems, 33, 2020.
  31. 31.Maei, H. R., Szepesvari, C., Bhatnagar, S., Precup, D., Silver, D., and Sutton, R. Convergent temporal-difference learning with arbitrary smooth function approximation. In NIPS, pp. 1204–1212, 2009.
  32. 32.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fijdeland, A. K., Ostrovski, G., et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
  33. 33.Muller, A. Integral probability metrics and their generating classes of functions. Advances in Applied Probability, 1997.
  34. 34.Munos, R. Error bounds for approximate policy iteration. In Proceedings of the Twentieth International Conference on International Conference on Machine Learning, pp. 560–567, 2003.
  35. 35.Munos, R. and Szepesvari, C. Finite-time bounds for fitted value iteration. Journal of Machine Learning Research, 9(5), 2008.
  36. 36.Neu, G., Jonsson, A., and Gomez, V. A unified view of entropy-regularized markov decision processes. arXiv preprint arXiv:1705.07798, 2017.
  37. 37.Paine, T. L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N. Hyperparameter selection for offline reinforcement learning. arXiv preprint arXiv:2007.09055, 2020.
  38. 38.Rajeswaran, A., Mordatch, I., and Kumar, V. A game theoretic framework for model based reinforcement learning. In International Conference on Machine Learning, pp. 7953–7963. PMLR, 2020.
  39. 39.Schoknecht, R. and Merke, A. Td (0) converges provably faster than the residual gradient algorithm. In Proceedings of the 20th International Conference on Machine Learning (ICML-03), pp. 680–687, 2003.
  40. 40.Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489, 2016.
  41. 41.Sun, W., Venkatraman, A., Gordon, G. J., Boots, B., and Bagnell, J. A. Deeply aggrevated: Differentiable imitation learning for sequential prediction. In International Conference on Machine Learning, pp. 3309–3318. PMLR, 2017.
  42. 42.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT press, 2018.
  43. 43.Swamy, G., Choudhury, S., Bagnell, J. A., and Wu, S. Of moments and matching: A game-theoretic framework for closing the imitation gap. In International Conference on Machine Learning, pp. 10022–10032. PMLR, 2021.
  44. 44.Uehara, M., Zhang, X., and Sun, W. Representation learning for online and offline rl in low-rank mdps. arXiv preprint arXiv:2110.04652, 2021.
  45. 45.Von Stackelberg, H. Market structure and equilibrium. Springer Science & Business Media, 2010.
  46. 46.Wainwright, M. J. High-dimensional statistics: A non-asymptotic viewpoint, volume 48. Cambridge University Press, 2019.
  47. 47.Wang, Z. T. and Ueda, M. A convergent and efficient deep q network algorithm. arXiv preprint arXiv:2106.15419, 2021.
  48. 48.Wu, Y., Tucker, G., and Nachum, O. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  49. 49.Xie, T. and Jiang, N. Q* approximation schemes for batch reinforcement learning: A theoretical comparison. In Conference on Uncertainty in Artificial Intelligence, pp. 550–559. PMLR, 2020.
  50. 50.Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A. Bellman-consistent pessimism for offline reinforcement learning. Advances in neural information processing systems, 34:6683–6694, 2021.
  51. 51.Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T. Mopo: Model-based offline policy optimization. Advances in Neural Information Processing Systems, 33:14129–14142, 2020.
  52. 52.Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C. Combo: Conservative offline model-based policy optimization. Advances in neural information processing systems, 34:28954–28967, 2021.
  53. 53.Zanette, A., Wainwright, M. J., and Brunskill, E. Provable benefits of actor-critic methods for offline reinforcement learning. Advances in neural information processing systems, 34, 2021.
  54. 54.Zhang, S. and Jiang, N. Towards hyperparameter-free policy selection for offline reinforcement learning. Advances in Neural Information Processing Systems, 34, 2021.
  55. 55.Zheng, L., Fiez, T., Alumbaugh, Z., Chasnov, B., and Ratliff, L. J. Stackelberg actor-critic: Game-theoretic reinforcement learning algorithms. arXiv preprint arXiv:2109.12286, 2021.

Citation

MLA
Cheng, C.-A., et al. “Adversarially Trained Actor Critic for Offline Reinforcement Learning”. International Conference on Machine Learning, vol. 162, 2022, pp. 3852–78, https://proceedings.mlr.press/v162/cheng22b.html.
APA
Cheng, C.-A., Xie, T., Jiang, N., & Agarwal, A. (2022). Adversarially Trained Actor Critic for Offline Reinforcement Learning. International Conference on Machine Learning, 162, 3852–3878. https://proceedings.mlr.press/v162/cheng22b.html
Chicago
Cheng, C.-A., T. Xie, N. Jiang, and A. Agarwal. 2022. “Adversarially Trained Actor Critic for Offline Reinforcement Learning”. International Conference on Machine Learning 162: 3852–78. https://proceedings.mlr.press/v162/cheng22b.html.
Harvard
Cheng, C.-A. et al. (2022) “Adversarially Trained Actor Critic for Offline Reinforcement Learning”, International Conference on Machine Learning. PMLR, pp. 3852–3878. Available at: https://proceedings.mlr.press/v162/cheng22b.html.
Vancouver
1. Cheng C-A, Xie T, Jiang N, Agarwal A (2022) Adversarially Trained Actor Critic for Offline Reinforcement Learning. In: International Conference on Machine Learning. PMLR, pp 3852–3878

BibTeX

@InProceedings{pmlr-v162-cheng22b,
  title = 	 {Adversarially Trained Actor Critic for Offline Reinforcement Learning},
  author =       {Cheng, Ching-An and Xie, Tengyang and Jiang, Nan and Agarwal, Alekh},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {3852--3878},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/cheng22b/cheng22b.pdf},
  url = 	 {https://proceedings.mlr.press/v162/cheng22b.html},
  abstract = 	 {We propose Adversarially Trained Actor Critic (ATAC), a new model-free algorithm for offline reinforcement learning (RL) under insufficient data coverage, based on the concept of relative pessimism. ATAC is designed as a two-player Stackelberg game framing of offline RL: A policy actor competes against an adversarially trained value critic, who finds data-consistent scenarios where the actor is inferior to the data-collection behavior policy. We prove that, when the actor attains no regret in the two-player game, running ATAC produces a policy that provably 1) outperforms the behavior policy over a wide range of hyperparameters that control the degree of pessimism, and 2) competes with the best policy covered by data with appropriately chosen hyperparameters. Compared with existing works, notably our framework offers both theoretical guarantees for general function approximation and a deep RL implementation scalable to complex environments and large datasets. In the D4RL benchmark, ATAC consistently outperforms state-of-the-art offline RL algorithms on a range of continuous control tasks.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/