Offline Multi-Agent Reinforcement Learning with Knowledge Distillation

Wei-Cheng TsengTsun-Hsuan Johnson WangYen-Chen LinPhillip Isola

article2022NeurIPS56 citations

Proposes a sequence modeling framework for offline multi-agent reinforcement learning that trains a centralized Decision Transformer teacher and distills both its representations and relational feature structures into decentralized student policies to achieve state-of-the-art coordination without online environment interaction.

Listen

Deploying multi-agent artificial intelligence systems in physical environments—such as fleets of autonomous vehicles or robotic swarms—often carries severe safety risks and high operational costs when agents must learn through real-world trial and error. Offline multi-agent reinforcement learning addresses this bottleneck by training models exclusively on previously recorded datasets. However, traditional approaches rely on complex value-estimation techniques that struggle to determine individual credit during collaborative tasks. The article demonstrates a novel framework that reformulates offline multi-agent training into a sequence modeling problem, combining centralized sequence models with structural knowledge transfer to enable effective, decentralized execution.

To overcome the limitations of existing methods, the proposed approach employs a centralized "teacher" model that processes the combined observations, actions, and rewards of all agents simultaneously. This unified perspective allows the teacher to identify effective cooperative behaviors across the dataset. The framework then distills the teacher's knowledge into independent "student" models designed for decentralized deployment. Crucially, the transfer objective preserves the geometric, structural relationships between agents' internal representations using specialized mapping networks and momentum updates, rather than merely copying individual actions.

Empirical evaluations across diverse benchmarks—including StarCraft combat scenarios, particle environments, grid navigation, and highway traffic simulations—reveal that the proposed method consistently achieves state-of-the-art results. The approach outperforms established offline reinforcement learning and imitation learning baselines across all tested environments, showing particular strength in complex coordination tasks such as the Equal Space and Highway benchmarks. Furthermore, the framework exhibits superior resilience when learning from suboptimal or poor-quality datasets, demonstrates high data efficiency with limited demonstrations, and adds minimal computational overhead, requiring less than 10% additional training time compared to standard sequence models.

These findings suggest that structural distillation effectively bridges the gap between centralized coordination and safe, decentralized execution without requiring live environment interaction. Organizations can leverage existing operational logs to train coordinated systems safely and cost-effectively before deployment. Practitioners looking to adopt this framework should utilize centralized-to-decentralized relational distillation pipelines and consider online fine-tuning on pre-trained models to further accelerate deployment readiness. However, decision-makers should note that centralized sequence modeling may encounter scalability constraints as team sizes grow significantly, and further research is needed to validate performance in very large-scale agent networks.

Tseng et al (2022).pdf
Cover for Offline Multi-Agent Reinforcement Learning with Knowledge Distillation

Abstract

We introduce an offline multi-agent reinforcement learning (offline MARL) framework that utilizes previously collected data without additional online data collection. Our method reformulates offline MARL as a sequence modeling problem and thus builds on top of the simplicity and scalability of the Transformer architecture. In the fashion of centralized training and decentralized execution, we propose to first train a teacher policy who has the privilege to access every agent’s observations, actions, and rewards. After the teacher policy has identified and recombined the “good” behavior in the dataset, we create separate student policies and distill not only the teacher policy’s features but also its structural relations among different agents’ features to student policies. We show that our framework significantly improves performances on a range of tasks and outperforms state-of-the-art offline MARL baselines. Furthermore, we demonstrate that the proposed method has a better convergence rate, is more sample efficient, and is more robust to various demonstration qualities compared with baselines.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Independent Decision Transformer
  • 3.2 Teacher-Student Policy Distillation
  • 4 Experiments
  • 4.1 Quantitative Results
  • 4.2 Ablation Study
  • 4.3 Finetuning
  • 4.4 Data Efficiency
  • 4.5 Convergence Rate
  • 5 Limitation and Broader Impact
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Centralized-to-Decentralized Sequence Modeling Policy Distillation Framework

    model/method

    Offline multi-agent reinforcement learning (MARL) is formulated as a sequence modeling problem under the centralized training and decentralized execution (CTDE) paradigm.

    First, a centralized teacher policy πteacher\pi_{\text{teacher}}, instantiated as a Decision Transformer parameterized by θ\theta, is trained on static multi-agent offline trajectories. For an environment with nn agents and past context length BB, the teacher receives the joint sequence of concatenated observations, actions, and rewards for all agents: (o(t−B):t1:n,a(t−B):t1:n,r(t−B):t1:n)(o_{(t-B):t}^{1:n}, a_{(t-B):t}^{1:n}, r_{(t-B):t}^{1:n}) and predicts the joint next actions a^t+11:n\hat{a}_{t+1}^{1:n}. Through self-attention over the joint trajectories of all agents, the centralized teacher performs cross-agent credit assignment and extracts coordinated behaviors.

    Second, to enable decentralized execution without privileged communication at inference time, nn independent student Decision Transformers π1,…,πn\pi_1, \dots, \pi_n parameterized by ϕ1,…,ϕn\phi_1, \dots, \phi_n are initialized. Each student agent ii receives only its local trajectory (o(t−B):ti,a(t−B):ti,r(t−B):ti)(o_{(t-B):t}^i, a_{(t-B):t}^i, r_{(t-B):t}^i) to predict at+1ia_{t+1}^i. Knowledge from the frozen teacher policy πteacher\pi_{\text{teacher}} is transferred to the decentralized student policies via a multi-agent policy distillation objective.

  2. Knowl 2 — Structural Relation Distillation Objective

    equation

    To transfer cooperative relational structures across agents from a centralized teacher policy to decentralized student policies, relational policy distillation transfers the angular geometric relationships between agents' latent feature representations.

    Let f^i\hat{f}^i and f^j\hat{f}^j denote the feature vectors generated by the teacher Decision Transformer for agents ii and jj, and let fif^i and fjf^j denote the corresponding feature vectors generated by the student Decision Transformers. Let MM and NN denote multi-layer perceptron mapping networks that project teacher and student features into a relational space. The structural relation distillation loss for agent ii is defined as:

    Lreli=∑j≠iH(cos⁡−1(M(f^i),M(f^j)),cos⁡−1(N(fi),N(fj)))L_{\text{rel}}^i = \sum_{j \neq i} H\left(\cos^{-1}\left(M(\hat{f}^i), M(\hat{f}^j)\right), \cos^{-1}\left(N(f^i), N(f^j)\right)\right)

    where cos⁡−1(u,v)=arccos⁡(u⋅v∥u∥2∥v∥2)\cos^{-1}(u, v) = \arccos\left(\frac{u \cdot v}{\|u\|_2 \|v\|_2}\right) computes the angle between feature vectors uu and vv, and H(y,y^)H(y, \hat{y}) denotes the Huber loss between the scalar teacher angle and student angle:

    H(y,y^)={12(y−y^)2if ∣y−y^∣≤δδ∣y−y^∣−12δ2otherwiseH(y, \hat{y}) = \begin{cases} \frac{1}{2}(y - \hat{y})^2 & \text{if } |y - \hat{y}| \le \delta \\ \delta |y - \hat{y}| - \frac{1}{2}\delta^2 & \text{otherwise} \end{cases}

    This loss preserves the relative feature configuration among agents (such as agents of the same unit type or with similar observation contexts maintaining similar representations) across the student policies without requiring identical absolute feature values.

  3. Knowl 3 — Asymmetric Mapping Networks with Momentum Update for Relational Distillation

    model/method

    To prevent relational policy distillation from discarding feature information useful for downstream action prediction, multi-layer perceptron (MLP) mapping networks MM and NN project teacher features f^i\hat{f}^i and student features fif^i before computing pairwise relational angles.

    Because the centralized teacher operates on privileged joint information across all agents while student policies only process local observations, the representations are structurally asymmetric. Therefore, the mapping networks have non-shared weights (M≠NM \neq N).

    To prevent the learnable projection head MM from creating a destabilizing moving target during distillation, a momentum-like update frequency schedule is used: the student mapping network NN is updated at every optimization step, whereas the teacher mapping network MM is updated only once every ee optimization steps of NN (with e=4e = 4).

  4. Knowl 4 — Total Student Policy Distillation Objective

    equation

    The overall learning objective for training each decentralized student Decision Transformer policy πi\pi_i with parameters ϕi\phi_i combines sequence modeling action prediction with relational and point-wise knowledge distillation:

    Ltotali=Lactioni+αLreli+βLKLiL_{\text{total}}^i = L_{\text{action}}^i + \alpha L_{\text{rel}}^i + \beta L_{\text{KL}}^i

    where:

    • Lactioni=∥at+1i−a^t+1i∥22L_{\text{action}}^i = \|a_{t+1}^i - \hat{a}_{t+1}^i\|_2^2 is the mean squared error (or action prediction loss) between the student's predicted action and the offline dataset action at+1ia_{t+1}^i.
    • Lreli=∑j≠iH(cos⁡−1(M(f^i),M(f^j)),cos⁡−1(N(fi),N(fj)))L_{\text{rel}}^i = \sum_{j \neq i} H\left(\cos^{-1}(M(\hat{f}^i), M(\hat{f}^j)), \cos^{-1}(N(f^i), N(f^j))\right) is the structural relational distillation loss with Huber loss HH and mapping networks MM (teacher) and NN (student).
    • LKLi=DKL(a^i∥ai)L_{\text{KL}}^i = D_{\text{KL}}(\hat{a}^i \parallel a^i) is the Kullback-Leibler divergence between the action probability distribution a^i\hat{a}^i predicted by the teacher and the action distribution aia^i predicted by student ii.
    • α≥0\alpha \ge 0 and β≥0\beta \ge 0 are hyperparameter weights balancing relational structure preservation and action distribution matching.
  5. Knowl 5 — Offline MARL with Policy Distillation Training Algorithm

    algorithm

    The training pipeline consists of two stages: centralized sequence modeling of the offline multi-agent dataset, followed by relational knowledge distillation into decentralized student policies.

    Input: Offline dataset D={τ}\mathcal{D} = \{\tau\} where τ=(ot1:n,at1:n,rt1:n)t=1T\tau = (o_t^{1:n}, a_t^{1:n}, r_t^{1:n})_{t=1}^T for nn agents, learning rate αlr\alpha_{\text{lr}}, hyperparameters α,β\alpha, \beta, momentum frequency e=4e=4
    Initialize: Centralized teacher parameters θ\theta, student parameters ϕ1,…,ϕn\phi_1, \dots, \phi_n, teacher mapping network MM, student mapping network NN
    // Stage 1: Centralized Decision Transformer Training
    for each trajectory τ∈D\tau \in \mathcal{D} do
        Predict joint action a^t1:n=arg⁡max⁡πteacher(at1:n∣τ<t,θ)\hat{a}_t^{1:n} = \arg\max \pi_{\text{teacher}}(a_t^{1:n} \mid \tau_{<t}, \theta)
        Compute centralized prediction loss Lcentralized=∥at1:n−a^t1:n∥22\mathcal{L}_{\text{centralized}} = \|a_t^{1:n} - \hat{a}_t^{1:n}\|_2^2
        Update teacher weights: θ←θ−αlr∇θLcentralized\theta \leftarrow \theta - \alpha_{\text{lr}} \nabla_\theta \mathcal{L}_{\text{centralized}}
    end for
    // Stage 2: Decentralized Student Distillation
    Freeze teacher parameters θ\theta
    for each trajectory τ∈D\tau \in \mathcal{D} do
        Extract teacher actions a^t1:n\hat{a}_t^{1:n} and features f^1:n\hat{f}^{1:n} using frozen πteacher\pi_{\text{teacher}}
        for i=1i = 1 to nn do
            Predict student action ati=arg⁡max⁡πi(ati∣τ<ti,ϕi)a_t^i = \arg\max \pi_i(a_t^i \mid \tau_{<t}^i, \phi_i) and feature fif^i
            Compute Lactioni=∥ati−a^ti∥22L_{\text{action}}^i = \|a_t^i - \hat{a}_t^i\|_2^2
            Compute Lreli=∑j≠iH(cos⁡−1(M(f^i),M(f^j)),cos⁡−1(N(fi),N(fj)))L_{\text{rel}}^i = \sum_{j \neq i} H\left(\cos^{-1}(M(\hat{f}^i), M(\hat{f}^j)), \cos^{-1}(N(f^i), N(f^j))\right)
            Compute LKLi=DKL(a^ti∥ati)L_{\text{KL}}^i = D_{\text{KL}}(\hat{a}_t^i \parallel a_t^i)
            Compute Ltotali=Lactioni+αLreli+βLKLiL_{\text{total}}^i = L_{\text{action}}^i + \alpha L_{\text{rel}}^i + \beta L_{\text{KL}}^i
            Update student weights: ϕi←ϕi−αlr∇ϕiLtotali\phi_i \leftarrow \phi_i - \alpha_{\text{lr}} \nabla_{\phi_i} L_{\text{total}}^i
        end for
        Update student mapping network NN via gradient descent
        Every ee iterations of updating NN, update teacher mapping network MM
    end for
  6. Knowl 6 — Independent Decision Transformer Baseline Formulation

    model/method

    Independent Decision Transformer (IDT) is a direct adaptation of Decision Transformer to multi-agent environments without parameter sharing or cross-agent communication.

    For each agent i∈{1,…,n}i \in \{1, \dots, n\}, an isolated Decision Transformer πi\pi_i is trained exclusively on agent ii's local history of observations, actions, and rewards (o(t−B):ti,a(t−B):ti,r(t−B):ti)(o_{(t-B):t}^i, a_{(t-B):t}^i, r_{(t-B):t}^i) over a context window of length BB. The training objective minimizes the action prediction error along the agent's individual trajectory:

    Lactioni=∥at+1i−a^t+1i∥22L_{\text{action}}^i = \|a_{t+1}^i - \hat{a}_{t+1}^i\|_2^2

    Although each agent uses self-attention to perform temporal credit assignment along its own past sequence, IDT cannot attend across agents. As a result, it fails to perform cross-agent credit assignment necessary for discovering coordinated cooperative strategies from suboptimal multi-agent datasets.

  7. Knowl 7 — Benchmark Performance Across Multi-Agent Environments

    data/table

    Evaluation of offline MARL algorithms across continuous navigation and coordination benchmarks (Fill-In, Equal Space, Grid-World, Highway) and StarCraft Multi-Agent Challenge (SMAC) scenarios (2s3z, 3s5z, 8m9m, 3s5z vs 3s6z). Returns report mean and standard deviation of per-agent return across 10 random seeds.

    Method Fill-In Equal Space Grid-World Highway
    BC −12.43±0.21-12.43 \pm 0.21 −9.35±0.64-9.35 \pm 0.64 1.49±0.161.49 \pm 0.16 13.38±1.1413.38 \pm 1.14
    IDT −7.83±0.42-7.83 \pm 0.42 −7.99±0.42-7.99 \pm 0.42 1.52±0.271.52 \pm 0.27 18.71±1.5318.71 \pm 1.53
    MADT −6.51±0.21-6.51 \pm 0.21 −6.91±0.92-6.91 \pm 0.92 1.57±0.341.57 \pm 0.34 18.78±1.2718.78 \pm 1.27
    MA-CQL −9.41±1.72-9.41 \pm 1.72 −6.99±0.38-6.99 \pm 0.38 1.44±0.301.44 \pm 0.30 17.69±1.3217.69 \pm 1.32
    MA-ICQ −9.72±0.39-9.72 \pm 0.39 −7.12±0.29-7.12 \pm 0.29 1.62±0.291.62 \pm 0.29 18.01±1.2718.01 \pm 1.27
    MA-BCQ −8.11±0.20-8.11 \pm 0.20 −7.06±0.59-7.06 \pm 0.59 1.49±0.441.49 \pm 0.44 17.92±1.4817.92 \pm 1.48
    MA-GAIL −3.41±0.12-3.41 \pm 0.12 −8.43±0.42-8.43 \pm 0.42 1.51±0.321.51 \pm 0.32 16.48±1.8016.48 \pm 1.80
    MA-AIRL −11.41±0.07-11.41 \pm 0.07 −8.43±0.63-8.43 \pm 0.63 1.52±0.371.52 \pm 0.37 16.76±1.9516.76 \pm 1.95
    Ours −3.41±0.12\mathbf{-3.41 \pm 0.12} −2.43±0.72\mathbf{-2.43 \pm 0.72} 2.09±0.22\mathbf{2.09 \pm 0.22} 23.35±0.91\mathbf{23.35 \pm 0.91}
    Method SMAC 2s3z SMAC 3s5z SMAC 8m9m SMAC 3s5z vs 3s6z
    BC 14.77±1.0114.77 \pm 1.01 11.32±0.7911.32 \pm 0.79 11.45±1.1411.45 \pm 1.14 10.86±0.9910.86 \pm 0.99
    IDT 17.63±1.8017.63 \pm 1.80 15.99±1.1115.99 \pm 1.11 15.93±0.8615.93 \pm 0.86 16.33±1.9316.33 \pm 1.93
    MADT 18.09±1.2618.09 \pm 1.26 16.18±1.0516.18 \pm 1.05 17.11±1.8317.11 \pm 1.83 16.91±2.1016.91 \pm 2.10
    MA-CQL 17.04±1.3817.04 \pm 1.38 15.02±1.9315.02 \pm 1.93 14.92±1.8714.92 \pm 1.87 15.32±2.4215.32 \pm 2.42
    MA-ICQ 17.42±1.5217.42 \pm 1.52 15.36±2.0115.36 \pm 2.01 14.72±1.2214.72 \pm 1.22 14.99±2.2114.99 \pm 2.21
    MA-BCQ 17.08±1.1217.08 \pm 1.12 15.09±0.8415.09 \pm 0.84 14.32±1.0214.32 \pm 1.02 15.78±1.6415.78 \pm 1.64
    MA-GAIL 15.01±1.1215.01 \pm 1.12 13.99±0.8413.99 \pm 0.84 13.99±0.6913.99 \pm 0.69 14.98±2.0414.98 \pm 2.04
    MA-AIRL 15.11±1.1215.11 \pm 1.12 14.02±0.8414.02 \pm 0.84 14.01±0.7914.01 \pm 0.79 14.95±2.1814.95 \pm 2.18
    Ours 18.12±1.31\mathbf{18.12 \pm 1.31} 16.98±1.19\mathbf{16.98 \pm 1.19} 18.33±0.99\mathbf{18.33 \pm 0.99} 18.78±2.01\mathbf{18.78 \pm 2.01}

    The relational knowledge distillation approach achieves the highest return per agent across all eight benchmarks. Sequence modeling approaches (IDT, MADT, Ours) outperform TD-learning offline RL methods (MA-CQL, MA-ICQ, MA-BCQ) and inverse MARL/imitation learning methods (BC, MA-GAIL, MA-AIRL), while structural distillation yields substantial improvements over independent sequence modeling (IDT).

  8. Knowl 8 — Robustness to Demonstration Quality Across Dataset Tiers

    data/table

    Comparison of offline MARL algorithms across three tiers of trajectory demonstration quality (good, normal, poor) on Grid-World and Highway environments. Results report mean return and standard deviation over 10 random seeds.

    Grid-World Highway
    Method good normal poor good normal poor
    BC 1.49±0.161.49 \pm 0.16 1.29±0.051.29 \pm 0.05 1.01±0.121.01 \pm 0.12 13.38±1.1413.38 \pm 1.14 10.23±0.9110.23 \pm 0.91 8.71±0.728.71 \pm 0.72
    IDT 1.52±0.271.52 \pm 0.27 1.45±0.121.45 \pm 0.12 1.43±0.091.43 \pm 0.09 18.71±1.5318.71 \pm 1.53 18.01±1.3918.01 \pm 1.39 17.58±1.0117.58 \pm 1.01
    MADT 1.57±0.341.57 \pm 0.34 1.50±0.171.50 \pm 0.17 1.44±0.111.44 \pm 0.11 18.78±1.2718.78 \pm 1.27 18.03±1.0218.03 \pm 1.02 17.84±0.8417.84 \pm 0.84
    MA-CQL 1.44±0.301.44 \pm 0.30 1.40±0.211.40 \pm 0.21 1.31±0.111.31 \pm 0.11 17.69±1.3217.69 \pm 1.32 16.88±0.9316.88 \pm 0.93 15.98±0.6715.98 \pm 0.67
    MA-ICQ 1.62±0.291.62 \pm 0.29 1.37±0.101.37 \pm 0.10 1.32±0.091.32 \pm 0.09 18.01±1.2718.01 \pm 1.27 16.54±0.7416.54 \pm 0.74 16.03±0.8316.03 \pm 0.83
    MA-BCQ 1.49±0.441.49 \pm 0.44 1.39±0.141.39 \pm 0.14 1.37±0.051.37 \pm 0.05 17.92±1.4817.92 \pm 1.48 16.89±0.7816.89 \pm 0.78 15.73±0.5815.73 \pm 0.58
    MA-GAIL 1.51±0.321.51 \pm 0.32 1.19±0.311.19 \pm 0.31 1.16±0.281.16 \pm 0.28 14.48±1.8014.48 \pm 1.80 11.27±1.5811.27 \pm 1.58 9.12±1.789.12 \pm 1.78
    MA-AIRL 1.52±0.371.52 \pm 0.37 1.21±0.271.21 \pm 0.27 1.11±0.251.11 \pm 0.25 16.76±1.9516.76 \pm 1.95 15.82±1.9315.82 \pm 1.93 10.22±1.8710.22 \pm 1.87
    Ours 2.09±0.22\mathbf{2.09 \pm 0.22} 1.89±0.12\mathbf{1.89 \pm 0.12} 1.81±0.19\mathbf{1.81 \pm 0.19} 23.35±0.91\mathbf{23.35 \pm 0.91} 20.01±1.01\mathbf{20.01 \pm 1.01} 17.13±1.1217.13 \pm 1.12

    The relational distillation method outperforms all baselines across all quality levels on Grid-World and across good and normal quality on Highway. On Highway-poor, it performs on par with sequence modeling baselines because constructing an effective centralized Decision Transformer teacher is constrained by low demonstration quality, which in turn limits the supervision signal available for student distillation.

  9. Knowl 9 — Ablation Analysis of Distillation and Architectural Components

    data/table

    Ablation study evaluating the contribution of individual distillation components, mapping networks, weight sharing, and update frequency across eight tasks. Values represent mean and standard deviation of return per agent over 10 random seeds.

    Variant Fill-In Equal Space Grid-World Highway
    Ours (conventional distillation) −4.11±0.22-4.11 \pm 0.22 −5.43±0.51-5.43 \pm 0.51 1.62±0.291.62 \pm 0.29 18.33±1.2718.33 \pm 1.27
    Ours (M=NM = N) −3.11±0.22-3.11 \pm 0.22 −2.61±0.33-2.61 \pm 0.33 1.42±0.191.42 \pm 0.19 14.92±1.8714.92 \pm 1.87
    Ours - Momentum Update −3.92±0.31-3.92 \pm 0.31 −2.44±0.53-2.44 \pm 0.53 1.79±0.191.79 \pm 0.19 18.62±1.9818.62 \pm 1.98
    Ours - Mapping Network −3.32±0.42-3.32 \pm 0.42 −2.87±0.19-2.87 \pm 0.19 1.88±0.111.88 \pm 0.11 18.21±1.9918.21 \pm 1.99
    Ours - KL Distillation −3.89±0.33-3.89 \pm 0.33 −2.99±0.27-2.99 \pm 0.27 1.99±0.511.99 \pm 0.51 20.11±1.8820.11 \pm 1.88
    Ours - KL Distillation - Mapping Network −3.30±0.21-3.30 \pm 0.21 −2.71±0.42-2.71 \pm 0.42 1.59±0.441.59 \pm 0.44 18.40±1.0918.40 \pm 1.09
    Ours Full method −3.41±0.12\mathbf{-3.41 \pm 0.12} −2.43±0.72\mathbf{-2.43 \pm 0.72} 2.09±0.22\mathbf{2.09 \pm 0.22} 23.35±0.91\mathbf{23.35 \pm 0.91}
    Variant SMAC 2s3z SMAC 3s5z SMAC 8m9m SMAC 3s5z vs 3s6z
    Ours (conventional distillation) 16.11±1.8716.11 \pm 1.87 15.78±1.7815.78 \pm 1.78 15.99±1.1115.99 \pm 1.11 17.93±1.3217.93 \pm 1.32
    Ours (M=NM = N) 14.14±1.3114.14 \pm 1.31 12.78±1.0812.78 \pm 1.08 15.40±2.1115.40 \pm 2.11 12.27±1.1112.27 \pm 1.11
    Ours - Momentum Update 15.69±1.9115.69 \pm 1.91 15.20±0.8115.20 \pm 0.81 16.39±2.7016.39 \pm 2.70 17.51±1.4917.51 \pm 1.49
    Ours - Mapping Network 15.23±1.8115.23 \pm 1.81 15.01±0.3115.01 \pm 0.31 16.87±0.9116.87 \pm 0.91 17.01±0.3217.01 \pm 0.32
    Ours - KL Distillation 16.32±0.9116.32 \pm 0.91 16.21±0.5516.21 \pm 0.55 17.01±1.2617.01 \pm 1.26 17.80±1.0817.80 \pm 1.08
    Ours - KL Distillation - Mapping Network 15.01±1.2315.01 \pm 1.23 14.98±1.0614.98 \pm 1.06 16.42±1.2216.42 \pm 1.22 17.71±1.8817.71 \pm 1.88
    Ours Full method 18.12±1.31\mathbf{18.12 \pm 1.31} 16.98±1.19\mathbf{16.98 \pm 1.19} 18.33±0.99\mathbf{18.33 \pm 0.99} 18.78±2.01\mathbf{18.78 \pm 2.01}

    Key takeaways:

    1. Conventional point-wise KL distillation performs poorly relative to relational distillation across tasks.
    2. Non-shared mapping networks (M≠NM \neq N) substantially outperform shared mapping networks (M=NM = N), reflecting the asymmetry in information access between teacher and student.
    3. Removing mapping networks degrades performance on 7 of 8 tasks, confirming their role in decoupling relational alignment from raw feature representations.
    4. Momentum updates to MM stabilize training and reduce performance variance.
  10. Knowl 10 — Scalability Overhead and Centralized Teacher Limitations

    limitation

    The framework relies on a centralized teacher Decision Transformer that processes the concatenated trajectories, joint actions, and returns of all nn agents simultaneously during training. As the number of agents nn increases, the input dimensionality and sequence modeling complexity scale up, presenting a potential computational bottleneck in environments with very large agent populations.

    Empirically, training the centralized teacher adds less than 10% training time overhead relative to decentralized training alone on the tested benchmarks (up to 9 agents in SMAC 8m9m), and student policies achieve comparable convergence rates to parameter-sharing approaches like MADT. However, scaling the centralized teacher architecture to hundreds or thousands of agents remains an open challenge.

Coverage note — No substantial contributed material was omitted. All architectural components, distillation objectives, baseline formulations, benchmark evaluations, demonstration quality analyses, and ablations are covered.

References

  1. 1.Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. An optimistic perspective on offline reinforcement learning, 2020.
  2. 2.Dian Chen, Brady Zhou, Vladlen Koltun, and Philipp Krähenbühl. Learning by cheating. In CoRL, 2019.
  3. 3.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. In NeurIPS, 2021.
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  5. 5.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning, 2020.
  6. 6.Wojciech M. Czarnecki, Razvan Pascanu, Simon Osindero, Siddhant Jayakumar, Grzegorz Swirszcz, and Max Jaderberg. Distilling policy distillation. In Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, 2019.
  7. 7.Wojciech Marian Czarnecki, Siddhant M. Jayakumar, Max Jaderberg, Leonard Hasenclever, Yee Whye Teh, Simon Osindero, Nicolas Heess, and Razvan Pascanu. Mix match - agent curricula for reinforcement learning, 2018.
  8. 8.Abhishek Das, Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Mike Rabbat, and Joelle Pineau. TarMAC: Targeted multi-agent communication. In ICML, 2019.
  9. 9.Jakob Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch. Learning with opponent-learning awareness. In AAMAS, 2018.
  10. 10.Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In ICML, 2019.
  11. 11.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  12. 12.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network, 2015.
  13. 13.Siyi Hu, Fengda Zhu, Xiaojun Chang, and Xiaodan Liang. {UPD}et: Universal multi-agent {rl} via policy decoupling with transformers. In ICLR, 2021.
  14. 14.Wenlong Huang, Igor Mordatch, and Deepak Pathak. One policy to control them all: Shared modular policies for agent-agnostic control. In ICML, 2020.
  15. 15.Shariq Iqbal, Christian A Schroeder de Witt, Bei Peng, Wendelin Böhmer, Shimon Whiteson, and Fei Sha. Randomized entity-wise factorization for multi-agent reinforcement learning. In ICML, 2021.
  16. 16.Shariq Iqbal and Fei Sha. Actor-attention-critic for multi-agent reinforcement learning. In ICML, 2019.
  17. 17.Jiechuan Jiang and Zongqing Lu. Learning attentional communication for multi-agent cooperation. In NeurIPS, page 7265–7275, Red Hook, NY, USA, 2018. Curran Associates Inc.
  18. 18.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. 2020.
  19. 19.Kwei-Herng Lai, Daochen Zha, Yuening Li, and Xia Hu. Dual policy distillation, 2020.
  20. 20.Sheng Li, Jayesh K. Gupta, Peter Morales, Ross Allen, and Mykel J. Kochenderfer. Deep implicit coordination graphs for multi-agent reinforcement learning. In AAMAS, 2021.
  21. 21.Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. In ICLR, 2016.
  22. 22.Michael L. Littman. Markov games as a framework for multi-agent reinforcement learning. In ICML, 1994.
  23. 23.Bo Liu, Qiang Liu, Peter Stone, Animesh Garg, Yuke Zhu, and Animashree Anandkumar. Coach-player multi-agent reinforcement learning for dynamic team composition, 2021.
  24. 24.Yong Liu, Weixun Wang, Yujing Hu, Jianye Hao, Xingguo Chen, and Yang Gao. Multi-agent game abstraction via graph attention neural network. In AAAI, 2020.
  25. 25.Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. Neural Information Processing Systems (NIPS), 2017.
  26. 26.Linghui Meng, Muning Wen, Yaodong Yang, Chenyang Le, Xiyun Li, Weinan Zhang, Ying Wen, Haifeng Zhang, Jun Wang, and Bo Xu. Offline pre-trained multi-agent decision transformer: One big sequence model tackles all SMAC tasks. 2021.
  27. 27.Ofir Nachum and Bo Dai. Reinforcement learning via fenchel-rockafellar duality, 2020.
  28. 28.Yaru Niu, Rohan Paleja, and Matthew Gombolay. Multi-agent graph-attention communication and teaming. In AAMAS, 2021.
  29. 29.Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. Relational knowledge distillation. In CVPR, 2019.
  30. 30.Deepak Pathak, Chris Lu, Trevor Darrell, Phillip Isola, and Alexei A. Efros. Learning to control self- assembling morphologies: A study of generalization via modularity. In NeurIPS, 2019.
  31. 31.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  32. 32.Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning. In ICML, 2018.
  33. 33.Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2014.
  34. 34.Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell. Policy distillation. arXiv preprint arXiv:1511.06295, 2015.
  35. 35.Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson. The StarCraft Multi-Agent Challenge. CoRR, abs/1902.04043, 2019.
  36. 36.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms, 2017.
  37. 37.Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar. Individualized controlled continuous communication model for multiagent cooperative and competitive tasks. In ICLR, 2019.
  38. 38.Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning. In ICML, 2019.
  39. 39.Jiaming Song, Hongyu Ren, Dorsa Sadigh, and Stefano Ermon. Multi-agent generative adversarial imitation learning, 2018.
  40. 40.Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. Value-decomposition networks for cooperative multi-agent learning based on team reward. In AAMAS, 2018.
  41. 41.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive representation distillation, 2020.
  42. 42.Frederick Tung and Greg Mori. Similarity-preserving knowledge distillation, 2019.
  43. 43.Ikechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu, Mengyuan Yan, Joséphine Simon, Matthew Bennice, Chuyuan Fu, Cong Ma, Jiantao Jiao, Sergey Levine, and Karol Hausman. Jump-start reinforcement learning. https://arxiv.org/abs/2204.02372, 2022.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  45. 45.Yiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng, Qiyuan Zhang, Gao Huang, Jun Yang, and Qianchuan Zhao. Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning. In NeurIPS, 2021.
  46. 46.Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu. The surprising effectiveness of ppo in cooperative, multi-agent games, 2021.
  47. 47.Lantao Yu, Jiaming Song, and Stefano Ermon. Multi-agent adversarial inverse reinforcement learning, 2019.
  48. 48.Chi Zhang, Sanmukh Kuppannagari, and Prasanna Viktor. Brac+: Improved behavior regularized actor critic for offline reinforcement learning. In Proceedings of The 13th Asian Conference on Machine Learning, 2021.

Citation

MLA
Tseng, W.-C., et al. “Offline Multi-Agent Reinforcement Learning with Knowledge Distillation”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 226–37, https://proceedings.neurips.cc/paper_files/paper/2022/file/01d78b294d80491fecddea897cf03642-Paper-Conference.pdf.
APA
Tseng, W.-C., Wang, T.-H. J., Lin, Y.-C., & Isola, P. (2022). Offline Multi-Agent Reinforcement Learning with Knowledge Distillation. Advances in Neural Information Processing Systems, 35, 226–237. https://proceedings.neurips.cc/paper_files/paper/2022/file/01d78b294d80491fecddea897cf03642-Paper-Conference.pdf
Chicago
Tseng, W.-C., T.-H. J. Wang, Y.-C. Lin, and P. Isola. 2022. “Offline Multi-Agent Reinforcement Learning with Knowledge Distillation”. Advances in Neural Information Processing Systems 35: 226–37. https://proceedings.neurips.cc/paper_files/paper/2022/file/01d78b294d80491fecddea897cf03642-Paper-Conference.pdf.
Harvard
Tseng, W.-C. et al. (2022) “Offline Multi-Agent Reinforcement Learning with Knowledge Distillation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 226–237. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/01d78b294d80491fecddea897cf03642-Paper-Conference.pdf.
Vancouver
1. Tseng W-C, Wang T-HJ, Lin Y-C, Isola P (2022) Offline Multi-Agent Reinforcement Learning with Knowledge Distillation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 226–237

BibTeX

@inproceedings{tseng2022offline,
  title = {Offline Multi-Agent Reinforcement Learning with Knowledge Distillation},
  author = {Tseng, Wei-Cheng and Wang, Tsun-Hsuan Johnson and Lin, Yen-Chen and Isola, Phillip},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {226-237},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/01d78b294d80491fecddea897cf03642-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors