Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification

Ling PanLongbo HuangTengyu MaHuazhe Xu

article2022ICML97 citations

Proposes Offline Multi-Agent RL with Actor Rectification (OMAR), a framework that integrates first-order policy gradients with zeroth-order optimization to prevent multi-agent policies from getting trapped in suboptimal local optima when learning from conservative value functions.

Listen

Deploying artificial intelligence in complex, real-world multi-agent systems—such as autonomous vehicle fleets, warehouse robotics, and industrial automation—often requires training models solely from pre-collected, historical datasets because active online experimentation is costly, hazardous, or impractical. While conservative offline reinforcement learning methods have successfully prevented errors in single-agent environments, directly applying these techniques to multi-agent settings causes severe performance degradation as the number of agents grows. The article investigates the root cause of this coordination breakdown and develops a solution that enables multiple autonomous agents to reliably learn cooperative policies from static data.

The main objective of the article is to demonstrate why conventional conservatism-based offline algorithms fail in multi-agent environments and to evaluate a proposed hybrid optimization method, named Offline Multi-Agent Reinforcement Learning with Actor Rectification (OMAR). The study aims to provide theoretical guarantees and empirical proof that this framework improves multi-agent coordination across varying data qualities and task complexities.

To evaluate this approach, the researchers conducted extensive simulations across multiple continuous and discrete control benchmarks. These testbeds included standard multi-agent particle cooperative tasks, a complex multi-agent locomotion domain (HalfCheetah), the StarCraft II micromanagement challenge, and single-agent navigation mazes. The evaluation tested policies trained on datasets spanning random exploration, intermediate training replays, and expert demonstrations. The proposed method integrates standard first-order policy gradient updates—which compute directional adjustments—with a zeroth-order sampling mechanism based on iteratively refined Gaussian distributions. This hybrid mechanism samples candidate actions and regularizes the learning policy toward high-value regions, bypassing deceptive local optima without requiring expensive online trial and error.

The findings demonstrate substantial performance gains over existing baseline methods. First, OMAR consistently achieved state-of-the-art results across all evaluated benchmarks, notably outperforming standard multi-agent implementations of Conservative Q-Learning (CQL) and behavior-regularized methods. In the StarCraft II discrete control benchmarks, OMAR improved average test win rates by 76.7% over multi-agent CQL. Second, the framework proved robust across diverse data distributions, excelling even when training on low-quality or mixed replay data where traditional baselines failed. Third, the study established that standard first-order policy updates get trapped in poor local optima because conservative value landscapes are non-concave; in cooperative settings, one agent settling for a suboptimal action triggers systemic coordination failure. Finally, OMAR achieved these improvements with negligible computational overhead, requiring only about 4.7% more runtime than standard conservative baseline algorithms.

These results indicate that organizations can safely train cooperative multi-agent systems using historical operational data without incurring the safety risks or financial burdens of active online testing. The findings challenge the assumption that standard single-agent offline algorithms transfer seamlessly to multi-agent tasks, showing that multi-agent deployment requires explicit mechanisms to avoid uncoordinated local failures. In addition, the article shows that decentralized value functions generally provide more stable and superior offline performance than centralized value functions, which suffer from higher dimensionality.

Decision-makers and engineering teams developing multi-agent systems should adopt hybrid optimization frameworks that combine gradient updates with sampling-based actor rectification when training on offline logs. Implementation teams should tune the policy regularization coefficient based on data quality—using lower values for diverse datasets and higher values for narrower expert datasets—while keeping sampling hyperparameters fixed. Before large-scale deployment in production, teams should conduct simulation pilots on domain-specific logs to confirm coordination dynamics. Further research should extend the theoretical analysis of deep neural network landscape dynamics in offline multi-agent games and investigate more expressive sampling distributions to scale to even larger multi-agent fleets.

arXiv: 2111.11188
Cover for Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification

Abstract

Conservatism has led to significant progress in offline reinforcement learning (RL) where an agent learns from pre-collected datasets. However, as many real-world scenarios involve interaction among multiple agents, it is important to resolve offline RL in the multi-agent setting. Given the recent success of transferring online RL algorithms to the multi-agent setting, one may expect that offline RL algorithms will also transfer to multi-agent settings directly. Surprisingly, we empirically observe that conservative offline RL algorithms do not work well in the multi-agent setting—the performance degrades significantly with an increasing number of agents. Towards mitigating the degradation, we identify a key issue that non-concavity of the value function makes the policy gradient improvements prone to local optima. Multiple agents exacerbate the problem severely, since the suboptimal policy by any agent can lead to uncoordinated global failure. Following this intuition, we propose a simple yet effective method, Offline Multi-Agent RL with Actor Rectification (OMAR), which combines the first-order policy gradients and zeroth-order optimization methods to better optimize the conservative value functions over the actor parameters. Despite the simplicity, OMAR achieves state-of-the-art results in a variety of multi-agent control tasks.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Multi-Agent Actor Critic
  • 2.2. Conservative Q-Learning
  • 3. Proposed Method
  • 3.1. The Motivating Example
  • 3.2. Offline MARL with Actor Rectification
  • 3.2.1. THE EFFECT OF OMAR IN THE SPREAD TASK
  • 4. Experiments
  • 4.1. Multi-Agent Particle Environments
  • 4.1.1. PERFORMANCE COMPARISON
  • 4.1.2. ABLATION STUDY
  • 4.2. Applicability to Other Algorithms
  • 4.3. Multi-Agent MuJoCo
  • 4.4. StarCraft II Micromanagement Benchmark
  • 4.5. Compatibility in Single-Agent D4RL Environments
  • 5. Related Work
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Additional Results for Spread
  • A.1. Results with Larger Learning Rates and Number of Updates of Actors in MA-CQL
  • A.2. The Performance of MA-CQL in a Non-Cooperative Version of the Multi-Agent Spread Task
  • B. Proof of Theorem 3.1
  • C. Experimental Details
  • C.1. Experimental Setup
  • C.1.1. TASKS.
  • C.1.2. BASELINES.
  • C.2. Learning Curves
  • C.3. Additional Ablation Study on the Effect of the Size of the Dataset
  • C.4. Applicability on Centralized Training with Decentralized Execution
  • C.4.1. RESULTS BASED ON MATD3
  • C.4.2. DISCUSSION ABOUT THE CENTRALIZED AND DECENTRALIZED CRITICS IN OFFLINE MULTI-AGENT RL

Knowls

  1. Knowl 1 — Offline Multi-Agent Reinforcement Learning with Actor Rectification Algorithm

    algorithm

    Offline Multi-Agent Reinforcement Learning with Actor Rectification (OMAR) trains decentralized actors to optimize conservative critic functions in offline multi-agent settings by combining policy gradient updates with zeroth-order action searches.

    Input: Offline dataset D={Di}i=1N\mathcal{D} = \{\mathcal{D}_i\}_{i=1}^N with transitions (oi,ai,ri,oi′)(o_i, a_i, r_i, o'_i), number of agents NN, total training steps TT, batch size SS, discount factor γ\gamma, soft update rate ρ\rho, actor regularization coefficient τ∈[0,1]\tau \in [0, 1], zeroth-order iterations JJ, sample population size KK, temperature parameter β\beta.
    Output: Decentralized policy networks πi\pi_i parameterized by ϕi\phi_i for i∈{1,…,N}i \in \{1, \dots, N\}.
    Initialize critic networks Qi1,Qi2Q_i^1, Q_i^2 with parameters θi1,θi2\theta_i^1, \theta_i^2 and actor πi\pi_i with parameters ϕi\phi_i for each agent i∈{1,…,N}i \in \{1, \dots, N\}
    Initialize target networks θˉi1←θi1\bar{\theta}_i^1 \leftarrow \theta_i^1, θˉi2←θi2\bar{\theta}_i^2 \leftarrow \theta_i^2, and ϕˉi←ϕi\bar{\phi}_i \leftarrow \phi_i
    for step t=1t = 1 to TT do
        for each agent i=1i = 1 to NN do
            Sample a random minibatch of SS transitions (oi,ai,ri,oi′)(o_i, a_i, r_i, o'_i) from Di\mathcal{D}_i
            Compute target value yi=ri+γmin⁡j=1,2Qˉij(oi′,πˉi(oi′+ϵ))y_i = r_i + \gamma \min_{j=1,2} \bar{Q}_i^j(o'_i, \bar{\pi}_i(o'_i + \epsilon)) with target noise ϵ\epsilon
            Update critic parameters θi\theta_i by minimizing the Conservative Q-Learning (CQL) loss:
                L(θi)=EDi[(Qi(oi,ai)−yi)2]+αEDi[log⁡∑aexp⁡(Qi(oi,a))−Ea∼π^βi[Qi(oi,a)]]\mathcal{L}(\theta_i) = \mathbb{E}_{\mathcal{D}_i} [(Q_i(o_i, a_i) - y_i)^2] + \alpha \mathbb{E}_{\mathcal{D}_i} [\log \sum_{a} \exp(Q_i(o_i, a)) - \mathbb{E}_{a \sim \hat{\pi}_{\beta_i}}[Q_i(o_i, a)]]
            Initialize zeroth-order Gaussian distribution N(μi0,σi0)\mathcal{N}(\mu_i^0, \sigma_i^0)
            for iteration j=1j = 1 to JJ do
                Draw candidate population A^i={a^ik∼N(μij−1,σij−1)}k=1K\hat{\mathcal{A}}_i = \{\hat{a}_i^k \sim \mathcal{N}(\mu_i^{j-1}, \sigma_i^{j-1})\}_{k=1}^K
                Evaluate critic values {Qi1(oi,a^ik)}k=1K\{Q_i^1(o_i, \hat{a}_i^k)\}_{k=1}^K
                Update distribution parameters:
                    μij=∑k=1Kexp⁡(βQi1(oi,a^ik))a^ik∑m=1Kexp⁡(βQi1(oi,a^im))\mu_i^j = \frac{\sum_{k=1}^K \exp(\beta Q_i^1(o_i, \hat{a}_i^k)) \hat{a}_i^k}{\sum_{m=1}^K \exp(\beta Q_i^1(o_i, \hat{a}_i^m))}
                    σij=∑k=1K(a^ik−μij−1)2\sigma_i^j = \sqrt{\sum_{k=1}^K (\hat{a}_i^k - \mu_i^{j-1})^2}
            end for
            Select candidate action a^i=arg⁡max⁡a∈A^i∪{πi(oi)}Qi1(oi,a)\hat{a}_i = \arg\max_{a \in \hat{\mathcal{A}}_i \cup \{\pi_i(o_i)\}} Q_i^1(o_i, a)
            Update actor parameters ϕi\phi_i to minimize the regularized loss:
                L(ϕi)=EDi[−(1−τ)Qi1(oi,πi(oi;ϕi))+τ∥πi(oi;ϕi)−a^i∥22]\mathcal{L}(\phi_i) = \mathbb{E}_{\mathcal{D}_i} [ - (1 - \tau) Q_i^1(o_i, \pi_i(o_i; \phi_i)) + \tau \|\pi_i(o_i; \phi_i) - \hat{a}_i\|_2^2 ]
            Update target network parameters:
                θˉij←ρθij+(1−ρ)θˉij\bar{\theta}_i^j \leftarrow \rho \theta_i^j + (1 - \rho) \bar{\theta}_i^j for j∈{1,2}j \in \{1, 2\}
                ϕˉi←ρϕi+(1−ρ)ϕˉi\bar{\phi}_i \leftarrow \rho \phi_i + (1 - \rho) \bar{\phi}_i
        end for
    end for
  2. Knowl 2 — Actor Loss with Zeroth-Order Rectification Regularizer

    equation

    In Offline Multi-Agent Reinforcement Learning with Actor Rectification (OMAR), each agent ii updates its policy parameters ϕi\phi_i by optimizing an objective that balances first-order policy gradient exploitation of the conservative critic with proximity to candidate actions discovered via zeroth-order optimization:

    max⁡ϕiEoi∼Di[(1−τ)Qi(oi,πi(oi;ϕi))−τ∥πi(oi;ϕi)−a^i∥22]\max_{\phi_i} \mathbb{E}_{o_i \sim \mathcal{D}_i} \left[ (1 - \tau) Q_i(o_i, \pi_i(o_i; \phi_i)) - \tau \|\pi_i(o_i; \phi_i) - \hat{a}_i\|_2^2 \right]

    where:

    • oi∈Oio_i \in \mathcal{O}_i is the private local observation of agent ii sampled from offline dataset Di\mathcal{D}_i.
    • πi(⋅;ϕi)\pi_i(\cdot; \phi_i) is the deterministic policy parameterized by ϕi\phi_i.
    • Qi(oi,ai)Q_i(o_i, a_i) is the conservative critic value function for agent ii.
    • a^i\hat{a}_i is the high-value candidate action discovered by the zeroth-order optimizer for observation oio_i.
    • τ∈[0,1]\tau \in [0, 1] is a trade-off regularization coefficient controlling the relative importance of policy gradient ascent versus action rectification towards a^i\hat{a}_i.
  3. Knowl 3 — Zeroth-Order Exponential Q-Weighted Action Sampling Mechanism

    model/method

    To identify actions that escape local optima of conservative critic functions without relying solely on first-order gradients, OMAR maintains and iteratively refines an action-sampling Gaussian distribution N(μi,σi)\mathcal{N}(\mu_i, \sigma_i) over JJ iterations for each agent ii.

    At iteration jj, KK candidate actions are sampled from the current proposal: a^ik∼N(μij,σij)\hat{a}_i^k \sim \mathcal{N}(\mu_i^j, \sigma_i^j) for k∈{1,…,K}k \in \{1, \dots, K\}. Each candidate is evaluated using the critic Qi1(oi,a^ik)Q_i^1(o_i, \hat{a}_i^k). The distribution parameters are then updated softly via exponential Q-value weighting:

    μij+1=∑k=1Kexp⁡(βQi1(oi,a^ik))a^ik∑m=1Kexp⁡(βQi1(oi,a^im))\mu_i^{j+1} = \frac{\sum_{k=1}^K \exp\left(\beta Q_i^1(o_i, \hat{a}_i^k)\right) \hat{a}_i^k}{\sum_{m=1}^K \exp\left(\beta Q_i^1(o_i, \hat{a}_i^m)\right)}

    σij+1=∑k=1K(a^ik−μij)2\sigma_i^{j+1} = \sqrt{\sum_{k=1}^K \left(\hat{a}_i^k - \mu_i^j\right)^2}

    where β\beta is an inverse temperature hyperparameter. After JJ refinement steps, the final rectified target action a^i\hat{a}_i is selected greedily from the union of all generated samples and the current policy action:

    a^i=arg⁡max⁡a∈A^i∪{πi(oi)}Qi1(oi,a)\hat{a}_i = \arg\max_{a \in \hat{\mathcal{A}}_i \cup \{\pi_i(o_i)\}} Q_i^1(o_i, a)

    This provides a robust target that prevents actors from becoming trapped in suboptimal flat regions or local maxima of the critic landscape.

  4. Knowl 4 — Safe Policy Improvement Bound for OMAR

    theoretical result

    Let M^i={(oi,ai,ri,oi′)∈Di}\hat{M}_i = \{(o_i, a_i, r_i, o'_i) \in \mathcal{D}_i\} be the empirical Markov Decision Process induced by agent ii's dataset Di\mathcal{D}_i, and let J(πi)J(\pi_i) denote the discounted expected return of policy πi\pi_i in M^i\hat{M}_i. Let πi∗\pi^*_i denote the policy optimizing the OMAR objective with regularization parameter τ∈[0,1)\tau \in [0, 1), and let π^βi\hat{\pi}_{\beta_i} denote the empirical behavior policy.

    Define the observation distribution ratio penalty:

    D(πi,π^βi)(oi)=1−π^βi(πi(oi)∣oi)π^βi(πi(oi)∣oi)D(\pi_i, \hat{\pi}_{\beta_i})(o_i) = \frac{1 - \hat{\pi}_{\beta_i}(\pi_i(o_i) \mid o_i)}{\hat{\pi}_{\beta_i}(\pi_i(o_i) \mid o_i)}

    Let dπi(oi)d^{\pi_i}(o_i) be the discounted marginal observation distribution of πi\pi_i, γ\gamma the discount factor, α\alpha the conservative critic regularization weight, and a^i\hat{a}_i the action found by the zeroth-order optimizer. Then, the performance difference satisfies the lower bound:

    J(πi∗)−J(π^βi)≥α1−γEoi∼dπi∗(oi)[D(πi∗,π^βi)(oi)]+τ1−τEoi∼dπi∗(oi)[(πi∗(oi)−a^i)2]−τ1−τEoi∼dπ^βi(oi),ai∼π^βi[(ai−a^i)2]J(\pi^*_i) - J(\hat{\pi}_{\beta_i}) \ge \frac{\alpha}{1-\gamma} \mathbb{E}_{o_i \sim d^{\pi^*_i}(o_i)} [D(\pi^*_i, \hat{\pi}_{\beta_i})(o_i)] + \frac{\tau}{1-\tau} \mathbb{E}_{o_i \sim d^{\pi^*_i}(o_i)} \left[ (\pi^*_i(o_i) - \hat{a}_i)^2 \right] - \frac{\tau}{1-\tau} \mathbb{E}_{o_i \sim d^{\hat{\pi}_{\beta_i}}(o_i), a_i \sim \hat{\pi}_{\beta_i}} \left[ (a_i - \hat{a}_i)^2 \right]

    Because the first term is non-negative and the difference between the second and third terms is bounded by the distance from the behavior policy and optimal policy to the zeroth-order optimizer's action a^i\hat{a}_i, OMAR provides a safe policy improvement guarantee over the empirical behavior policy.

  5. Knowl 5 — Degradation of Conservative Policy Gradient Methods in Multi-Agent RL

    empirical result

    Directly transferring single-agent conservative offline RL methods (such as Conservative Q-Learning, CQL) to multi-agent cooperative tasks leads to severe performance degradation as the number of agents increases.

    In cooperative environments such as Spread (where NN agents must coordinate without colliding to cover NN landmarks):

    1. In non-cooperative multi-agent variants where each agent has an independent goal, CQL performance improvement over the behavior policy remains stable across 1 to 5 agents.
    2. In cooperative multi-agent variants, CQL's performance improvement over the behavior policy drops sharply as the number of agents increases from 1 to 5.
    3. The failure is driven by non-concavity and flat regions in the conservative value function landscape, which trap standard first-order policy gradient updates in suboptimal local optima. In cooperative MARL, a single agent stuck in a local optimum leads to catastrophic failure of joint coordination.
  6. Knowl 6 — Benchmark Results in Multi-Agent Particle Environments

    data/table

    Normalized score comparisons across three cooperative continuous control environments in the Multi-Agent Particle Environment (MPE) benchmark: Cooperative Navigation (3 agents, 3 landmarks), Predator-Prey (3 predators, 1 prey), and World (4 agents, 2 adversaries). Four dataset quality levels (Random, Medium-Replay, Medium, Expert) of 1M transitions each were evaluated using decentralized critics across 5 random seeds (normalized as 100×(S−Srandom)/(Sexpert−Srandom)100 \times (S - S_{\text{random}})/(S_{\text{expert}} - S_{\text{random}})):

    Dataset Environment MA-ICQ MA-TD3+BC MA-CQL OMAR
    Random Cooperative navigation 6.3 ±\pm 3.5 9.8 ±\pm 4.9 24.0 ±\pm 9.8 34.4 ±\pm 5.3
    Predator-prey 2.2 ±\pm 2.6 5.7 ±\pm 3.5 5.0 ±\pm 8.2 11.1 ±\pm 2.8
    World 1.0 ±\pm 3.2 2.8 ±\pm 5.5 0.6 ±\pm 2.0 5.9 ±\pm 5.2
    Medium-replay Cooperative navigation 13.6 ±\pm 5.7 15.4 ±\pm 5.6 20.0 ±\pm 8.4 37.9 ±\pm 12.3
    Predator-prey 34.5 ±\pm 27.8 28.7 ±\pm 20.9 24.8 ±\pm 17.3 47.1 ±\pm 15.3
    World 12.0 ±\pm 9.1 17.4 ±\pm 8.1 29.6 ±\pm 13.8 42.9 ±\pm 19.5
    Medium Cooperative navigation 29.3 ±\pm 5.5 29.3 ±\pm 4.8 34.1 ±\pm 7.2 47.9 ±\pm 18.9
    Predator-prey 63.3 ±\pm 20.0 65.1 ±\pm 29.5 61.7 ±\pm 23.1 66.7 ±\pm 23.2
    World 71.9 ±\pm 20.0 73.4 ±\pm 9.3 58.6 ±\pm 11.2 74.6 ±\pm 11.5
    Expert Cooperative navigation 104.0 ±\pm 3.4 108.3 ±\pm 3.3 98.2 ±\pm 5.2 114.9 ±\pm 2.6
    Predator-prey 113.0 ±\pm 14.4 115.2 ±\pm 12.5 93.9 ±\pm 14.0 116.2 ±\pm 19.8
    World 109.5 ±\pm 22.8 110.3 ±\pm 21.3 71.9 ±\pm 28.1 110.4 ±\pm 25.7

    OMAR consistently outperforms MA-ICQ, MA-TD3+BC, and MA-CQL across all settings while adding only 4.7% runtime overhead relative to MA-CQL.

  7. Knowl 7 — Ablation of Zeroth-Order Action Sampling Schemes in OMAR

    data/table

    Comparison of different candidate action generation mechanisms integrated into OMAR on the Cooperative Navigation task across four offline dataset types (mean ±\pm standard deviation normalized scores over 5 seeds):

    Dataset OMAR (Random) OMAR (CEM) OMAR (Exponential Q-Weighted)
    Random 24.3 ±\pm 7.0 25.8 ±\pm 7.3 34.4 ±\pm 5.3
    Medium-replay 23.5 ±\pm 5.3 32.6 ±\pm 5.1 37.9 ±\pm 5.3
    Medium 41.2 ±\pm 11.1 45.0 ±\pm 13.3 47.9 ±\pm 18.9
    Expert 101.0 ±\pm 5.2 106.4 ±\pm 13.8 114.9 ±\pm 2.6

    The soft exponential Q-weighted Gaussian update rule outperforms both standard Cross-Entropy Method (CEM) selection and uniform random sampling across all dataset qualities, with the largest margins occurring on lower-quality and more diverse datasets.

  8. Knowl 8 — Performance on Continuous Multi-Agent MuJoCo Locomotion

    data/table

    Normalized score evaluation on the multi-agent HalfCheetah locomotion benchmark (2 agents controlling distinct joint sets of a single robot) across datasets of varying quality levels (1M samples each):

    Dataset MA-ICQ MA-TD3+BC MA-CQL OMAR
    Random 7.4 ±\pm 0.0 7.4 ±\pm 0.0 7.4 ±\pm 0.0 13.5 ±\pm 7.0
    Medium-replay 35.6 ±\pm 2.7 27.1 ±\pm 5.5 41.2 ±\pm 10.1 57.7 ±\pm 5.1
    Medium 73.6 ±\pm 5.0 75.5 ±\pm 3.7 50.4 ±\pm 10.8 80.4 ±\pm 10.2
    Expert 110.6 ±\pm 3.3 114.4 ±\pm 3.8 64.2 ±\pm 24.9 113.5 ±\pm 4.3

    OMAR significantly outperforms baseline methods on Random, Medium-replay, and Medium datasets, and matches MA-TD3+BC on Expert data while achieving higher stability than MA-CQL.

  9. Knowl 9 — Performance on StarCraft II Micromanagement Discrete Benchmarks

    empirical result

    In large-scale discrete multi-agent control on the StarCraft Multi-Agent Challenge (SMAC), OMAR was evaluated across four scenarios with increasing agent numbers and task difficulties: 2s3z (5 vs 5 units), 3s5z (8 vs 8 units), 1c3s5z (9 vs 9 units), and 2c_vs_64zg (2 vs 64 units).

    Discrete actions are made differentiable for MATD3-based training via the Gumbel-Softmax reparameterization trick. Across all four maps:

    • OMAR significantly outperforms the primary baseline MA-CQL in learning sample efficiency and asymptotic test win rate.
    • OMAR achieves an average test win rate performance improvement of 76.7% over MA-CQL across the tested map suite.
  10. Knowl 10 — Decentralized vs. Centralized Critics in Offline Multi-Agent RL

    empirical result

    In offline multi-agent reinforcement learning, decentralized critic architectures (Independent TD3, ITD3) consistently outperform centralized critic architectures (Multi-Agent TD3, MATD3) across dataset qualities.

    Normalized performance on Cooperative Navigation:

    • Random dataset: ITD3 achieves 18.7 ±\pm 8.0 vs. MATD3 16.1 ±\pm 5.6.
    • Medium-replay dataset: ITD3 achieves 19.9 ±\pm 4.7 vs. MATD3 12.7 ±\pm 6.1.
    • Medium dataset: ITD3 achieves 18.6 ±\pm 4.4 vs. MATD3 12.1 ±\pm 14.2.
    • Expert dataset: ITD3 achieves 75.5 ±\pm 7.9 vs. MATD3 1.6 ±\pm 2.7.

    Centralized value functions condition on joint actions and global state representations, creating an exponentially larger input space that is difficult to approximate accurately from offline data without exploratory interactions.

  11. Knowl 11 — Generalization of Actor Rectification to Ensemble-Based Offline RL (EDAC)

    data/table

    The actor rectification framework of OMAR is modular and applies to other conservative offline RL algorithms, such as Ensemble-Diversified Actor-Critic (EDAC). Normalized score performance in Multi-Agent Particle Environments demonstrates that incorporating OMAR's zeroth-order action rectification into MA-EDAC consistently improves over baseline MA-EDAC:

    Dataset Task MA-EDAC OMAR (on MA-EDAC)
    Random Cooperative navigation 29.9 ±\pm 13.3 35.7 ±\pm 12.2
    Predator-prey 10.1 ±\pm 6.5 11.5 ±\pm 1.7
    World 13.5 ±\pm 5.3 16.5 ±\pm 4.0
    Med-replay Cooperative navigation 34.1 ±\pm 8.2 38.5 ±\pm 6.7
    Predator-prey 31.4 ±\pm 8.9 58.0 ±\pm 14.9
    World 34.5 ±\pm 15.0 43.5 ±\pm 11.3
    Medium Cooperative navigation 39.9 ±\pm 14.2 51.6 ±\pm 13.8
    Predator-prey 64.4 ±\pm 14.6 70.8 ±\pm 11.1
    World 72.1 ±\pm 7.0 76.3 ±\pm 13.0
    Expert Cooperative navigation 99.6 ±\pm 7.7 114.7 ±\pm 3.7
    Predator-prey 93.7 ±\pm 11.0 114.5 ±\pm 8.1
    World 93.5 ±\pm 20.3 103.7 ±\pm 21.2

Coverage note — Single-agent Maze2D tabular results (Table 5) and hyperparameter sensitivity curves across dataset size/sampling iterations were omitted in favor of the primary multi-agent algorithmic, theoretical, and benchmark contributions.

References

  1. 1.Ackermann, J., Gabler, V., Osa, T., and Sugiyama, M. Reducing overestimation bias in multi-agent domains using double centralized critics. arXiv preprint arXiv:1910.01465, 2019.
  2. 2.Agarwal, R., Schuurmans, D., and Norouzi, M. An optimistic perspective on offline reinforcement learning. In International Conference on Machine Learning, pp. 104–114. PMLR, 2020.
  3. 3.Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D. Understanding the impact of entropy on policy optimization. In International Conference on Machine Learning, pp. 151–160. PMLR, 2019.
  4. 4.Amato, C. Decision-making under uncertainty in multiagent and multi-robot systems: Planning and learning. In IJCAI, pp. 5662–5666, 2018.
  5. 5.An, G., Moon, S., Kim, J.-H., and Song, H. O. Uncertainty-based offline reinforcement learning with diversified q-ensemble. Advances in Neural Information Processing Systems, 34, 2021.
  6. 6.Argenson, A. and Dulac-Arnold, G. Model-based offline planning. arXiv preprint arXiv:2008.05556, 2020.
  7. 7.Conti, E., Madhavan, V., Such, F. P., Lehman, J., Stanley, K. O., and Clune, J. Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents. arXiv preprint arXiv:1712.06560, 2017.
  8. 8.Daskalakis, C. On the complexity of approximating a nash equilibrium. ACM Transactions on Algorithms (TALG), 9(3):1–35, 2013.
  9. 9.Daskalakis, C., Skoulakis, S., and Zampetakis, M. The complexity of constrained min-max optimization. In Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing, pp. 1466–1478, 2021.
  10. 10.Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In Advances in Neural Information Processing Systems, pp. 2933–2941, 2014.
  11. 11.De Boer, P.-T., Kroese, D. P., Mannor, S., and Rubinstein, R. Y. A tutorial on the cross-entropy method. Annals of operations research, 134(1):19–67, 2005.
  12. 12.Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S. Counterfactual multi-agent policy gradients. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  13. 13.Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. D4rl: Datasets for deep data-driven reinforcement learning. arXiv preprint arXiv:2004.07219, 2020.
  14. 14.Fujimoto, S. and Gu, S. S. A minimalist approach to offline reinforcement learning. arXiv preprint arXiv:2106.06860, 2021.
  15. 15.Fujimoto, S., Hoof, H., and Meger, D. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning, pp. 1587–1596. PMLR, 2018.
  16. 16.Fujimoto, S., Meger, D., and Precup, D. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pp. 2052–2062. PMLR, 2019.
  17. 17.Ge, R., Lee, J. D., and Ma, T. Learning one-hidden-layer neural networks with landscape design. arXiv preprint arXiv:1711.00501, 2017.
  18. 18.Gulcehre, C., Wang, Z., Novikov, A., Le Paine, T., Gomez Colmenarejo, S., Zolna, K., Agarwal, R., Merel, J., Mankowitz, D., Paduraru, C., et al. Rl unplugged: Benchmarks for offline reinforcement learning. arXiv e-prints, pp. arXiv–2006, 2020.
  19. 19.Hu, J., Wellman, M. P., et al. Multiagent reinforcement learning: theoretical framework and an algorithm. In ICML, volume 98, pp. 242–250. Citeseer, 1998.
  20. 20.Iqbal, S. and Sha, F. Actor-attention-critic for multi-agent reinforcement learning. In International Conference on Machine Learning, pp. 2961–2970. PMLR, 2019.
  21. 21.Jang, E., Gu, S., and Poole, B. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016.
  22. 22.Jiang, J. and Lu, Z. Offline decentralized multi-agent reinforcement learning. arXiv preprint arXiv:2108.01832, 2021.
  23. 23.Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al. Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation. arXiv preprint arXiv:1806.10293, 2018.
  24. 24.Khadka, S. and Tumer, K. Evolution-guided policy gradient in reinforcement learning. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 1196–1208, 2018.
  25. 25.Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. Morel: Model-based offline reinforcement learning. arXiv preprint arXiv:2005.05951, 2020.
  26. 26.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  27. 27.Kostrikov, I., Fergus, R., Tompson, J., and Nachum, O. Offline reinforcement learning with fisher divergence critic regularization. In International Conference on Machine Learning, pp. 5774–5783. PMLR, 2021a.
  28. 28.Kostrikov, I., Nair, A., and Levine, S. Offline reinforcement learning with implicit q-learning. arXiv preprint arXiv:2110.06169, 2021b.
  29. 29.Kumar, A., Fu, J., Tucker, G., and Levine, S. Stabilizing off-policy q-learning via bootstrapping error reduction. arXiv preprint arXiv:1906.00949, 2019.
  30. 30.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Systems, volume 33, pp. 1179–1191, 2020.
  31. 31.Lee, S., Seo, Y., Lee, K., Abbeel, P., and Shin, J. Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble. arXiv preprint arXiv:2107.00591, 2021.
  32. 32.Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. Continuous control with deep reinforcement learning. In ICLR (Poster), 2016.
  33. 33.Lim, S., Joseph, A., Le, L., Pan, Y., and White, M. Actor-expert: A framework for using q-learning in continuous action spaces. arXiv preprint arXiv:1810.09103, 2018.
  34. 34.Littman, M. L. Markov games as a framework for multi-agent reinforcement learning. In Machine learning proceedings 1994, pp. 157–163. Elsevier, 1994.
  35. 35.Lowe, R., Wu, Y., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I. Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in Neural Information Processing Systems, 30:6379–6390, 2017.
  36. 36.Lowrey, K., Rajeswaran, A., Kakade, S., Todorov, E., and Mordatch, I. Plan online, learn offline: Efficient learning and exploration via model-based control. arXiv preprint arXiv:1811.01848, 2018.
  37. 37.Lyu, X., Xiao, Y., Daley, B., and Amato, C. Contrasting centralized and decentralized critics in multi-agent reinforcement learning. In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems, pp. 844–852, 2021.
  38. 38.Mania, H., Guy, A., and Recht, B. Simple random search provides a competitive approach to reinforcement learning. arXiv preprint arXiv:1803.07055, 2018.
  39. 39.Mathieu, M., Ozair, S., Srinivasan, S., Gulcehre, C., Zhang, S., Jiang, R., Le Paine, T., Zolna, K., Powell, R., Schrittwieser, J., et al. Starcraft ii unplugged: Large scale offline reinforcement learning. In Deep RL Workshop NeurIPS 2021, 2021.
  40. 40.Nachum, O., Norouzi, M., and Schuurmans, D. Improving policy gradient by exploring under-appreciated rewards. arXiv preprint arXiv:1611.09321, 2016.
  41. 41.Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D. Algaedice: Policy gradient from arbitrary experience. arXiv preprint arXiv:1912.02074, 2019.
  42. 42.Nagabandi, A., Konolige, K., Levine, S., and Kumar, V. Deep dynamics models for learning dexterous manipulation. In Conference on Robot Learning, pp. 1101–1112. PMLR, 2020.
  43. 43.Pan, L., Rashid, T., Peng, B., Huang, L., and Whiteson, S. Regularized softmax deep multi-agent q-learning. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  44. 44.Peng, B., Rashid, T., de Witt, C. A. S., Kamienny, P.-A., Torr, P. H., Bohmer, W., and Whiteson, S. Facmac: Factored multi-agent centralised policy gradients. arXiv preprint arXiv:2003.06709, 2020.
  45. 45.Pomerleau, D. A. Alvinn: An autonomous land vehicle in a neural network. Technical report, CARNEGIE-MELLON UNIV PITTSBURGH PA ARTIFICIAL INTELLIGENCE AND PSYCHOLOGY . . . , 1989.
  46. 46.Pourchot and Sigaud. CEM-RL: Combining evolutionary and gradient-based methods for policy search. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=BkeU5j0ctQ.
  47. 47.Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S. Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International Conference on Machine Learning, pp. 4295–4304. PMLR, 2018.
  48. 48.Rubinstein, R. Y. and Kroese, D. P. The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning. Springer Science & Business Media, 2013.
  49. 49.Sadigh, D., Sastry, S., Seshia, S. A., and Dragan, A. D. Planning for autonomous cars that leverage effects on human actions. In Robotics: Science and Systems, volume 2, pp. 1–9. Ann Arbor, MI, USA, 2016.
  50. 50.Safran, I. and Shamir, O. Spurious local minima are common in two-layer relu neural networks. arXiv preprint arXiv:1712.08968, 2017.
  51. 51.Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864, 2017.
  52. 52.Samvelyan, M., Rashid, T., Schroeder de Witt, C., Farquhar, G., Nardelli, N., Rudner, T. G., Hung, C.-M., Torr, P. H., Foerster, J., and Whiteson, S. The starcraft multi-agent challenge. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, pp. 2186–2188, 2019.
  53. 53.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  54. 54.Simmons-Edler, R., Eisner, B., Mitchell, E., Seung, S., and Lee, D. Q-learning for continuous actions with cross-entropy guided policies. arXiv preprint arXiv:1903.10605, 2019.
  55. 55.Such, F. P., Madhavan, V., Conti, E., Lehman, J., Stanley, K. O., and Clune, J. Deep neuroevolution: Genetic algorithms are a competitive alternative for training deep neural networks for reinforcement learning. arXiv preprint arXiv:1712.06567, 2017.
  56. 56.Sun, H., Xu, Z., Song, Y., Fang, M., Xiong, J., Dai, B., Zhang, Z., and Zhou, B. Zeroth-order supervised policy improvement. arXiv preprint arXiv:2006.06600, 2020.
  57. 57.Thomas, P. S. Safe reinforcement learning. PhD thesis, University of Massachusetts Libraries, 2015.
  58. 58.Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., et al. Critic regularized regression. Advances in Neural Information Processing Systems, 33: 7768–7778, 2020.
  59. 59.Williams, G., Aldrich, A., and Theodorou, E. Model predictive path integral control using covariance variable importance sampling. arXiv preprint arXiv:1509.01149, 2015.
  60. 60.Witt, C. S., Gupta, T., Makoviichuk, D., Makoviychuk, V., Torr, P. H., Sun, M., and Whiteson, S. Is independent learning all you need in the starcraft multi-agent challenge? arXiv preprint arXiv:2011.09533, 2020.
  61. 61.Wu, Y., Tucker, G., and Nachum, O. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361, 2019.
  62. 62.Wu, Y., Zhai, S., Srivastava, N., Susskind, J., Zhang, J., Salakhutdinov, R., and Goh, H. Uncertainty weighted actor-critic for offline reinforcement learning. arXiv preprint arXiv:2105.08140, 2021.
  63. 63.Yang, Y., Ma, X., Li, C., Zheng, Z., Zhang, Q., Huang, G., Yang, J., and Zhao, Q. Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning. arXiv preprint arXiv:2106.03400, 2021.
  64. 64.Yu, C., Velu, A., Vinitsky, E., Wang, Y., Bayen, A., and Wu, Y. The surprising effectiveness of mappo in cooperative, multi-agent games. arXiv preprint arXiv:2103.01955, 2021.
  65. 65.Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T. Mopo: Model-based offline policy optimization. Advances in Neural Information Processing Systems, 33:14129–14142, 2020.

Citation

MLA
Pan, L., et al. “Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification”. International Conference on Machine Learning, vol. 162, 2022, pp. 17221–37, https://proceedings.mlr.press/v162/pan22a.html.
APA
Pan, L., Huang, L., Ma, T., & Xu, H. (2022). Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification. International Conference on Machine Learning, 162, 17221–17237. https://proceedings.mlr.press/v162/pan22a.html
Chicago
Pan, L., L. Huang, T. Ma, and H. Xu. 2022. “Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification”. International Conference on Machine Learning 162: 17221–37. https://proceedings.mlr.press/v162/pan22a.html.
Harvard
Pan, L. et al. (2022) “Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification”, International Conference on Machine Learning. PMLR, pp. 17221–17237. Available at: https://proceedings.mlr.press/v162/pan22a.html.
Vancouver
1. Pan L, Huang L, Ma T, Xu H (2022) Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification. In: International Conference on Machine Learning. PMLR, pp 17221–17237

BibTeX

@InProceedings{pmlr-v162-pan22a,
  title = 	 {Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification},
  author =       {Pan, Ling and Huang, Longbo and Ma, Tengyu and Xu, Huazhe},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {17221--17237},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/pan22a/pan22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/pan22a.html},
  abstract = 	 {Conservatism has led to significant progress in offline reinforcement learning (RL) where an agent learns from pre-collected datasets. However, as many real-world scenarios involve interaction among multiple agents, it is important to resolve offline RL in the multi-agent setting. Given the recent success of transferring online RL algorithms to the multi-agent setting, one may expect that offline RL algorithms will also transfer to multi-agent settings directly. Surprisingly, we empirically observe that conservative offline RL algorithms do not work well in the multi-agent setting—the performance degrades significantly with an increasing number of agents. Towards mitigating the degradation, we identify a key issue that non-concavity of the value function makes the policy gradient improvements prone to local optima. Multiple agents exacerbate the problem severely, since the suboptimal policy by any agent can lead to uncoordinated global failure. Following this intuition, we propose a simple yet effective method, Offline Multi-Agent RL with Actor Rectification (OMAR), which combines the first-order policy gradients and zeroth-order optimization methods to better optimize the conservative value functions over the actor parameters. Despite the simplicity, OMAR achieves state-of-the-art results in a variety of multi-agent control tasks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/