Offline Multi-Agent Reinforcement Learning with Knowledge Distillation
Wei-Cheng TsengTsun-Hsuan Johnson WangYen-Chen LinPhillip Isola
Proposes a sequence modeling framework for offline multi-agent reinforcement learning that trains a centralized Decision Transformer teacher and distills both its representations and relational feature structures into decentralized student policies to achieve state-of-the-art coordination without online environment interaction.
Deploying multi-agent artificial intelligence systems in physical environments—such as fleets of autonomous vehicles or robotic swarms—often carries severe safety risks and high operational costs when agents must learn through real-world trial and error. Offline multi-agent reinforcement learning addresses this bottleneck by training models exclusively on previously recorded datasets. However, traditional approaches rely on complex value-estimation techniques that struggle to determine individual credit during collaborative tasks. The article demonstrates a novel framework that reformulates offline multi-agent training into a sequence modeling problem, combining centralized sequence models with structural knowledge transfer to enable effective, decentralized execution.
To overcome the limitations of existing methods, the proposed approach employs a centralized "teacher" model that processes the combined observations, actions, and rewards of all agents simultaneously. This unified perspective allows the teacher to identify effective cooperative behaviors across the dataset. The framework then distills the teacher's knowledge into independent "student" models designed for decentralized deployment. Crucially, the transfer objective preserves the geometric, structural relationships between agents' internal representations using specialized mapping networks and momentum updates, rather than merely copying individual actions.
Empirical evaluations across diverse benchmarks—including StarCraft combat scenarios, particle environments, grid navigation, and highway traffic simulations—reveal that the proposed method consistently achieves state-of-the-art results. The approach outperforms established offline reinforcement learning and imitation learning baselines across all tested environments, showing particular strength in complex coordination tasks such as the Equal Space and Highway benchmarks. Furthermore, the framework exhibits superior resilience when learning from suboptimal or poor-quality datasets, demonstrates high data efficiency with limited demonstrations, and adds minimal computational overhead, requiring less than 10% additional training time compared to standard sequence models.
These findings suggest that structural distillation effectively bridges the gap between centralized coordination and safe, decentralized execution without requiring live environment interaction. Organizations can leverage existing operational logs to train coordinated systems safely and cost-effectively before deployment. Practitioners looking to adopt this framework should utilize centralized-to-decentralized relational distillation pipelines and consider online fine-tuning on pre-trained models to further accelerate deployment readiness. However, decision-makers should note that centralized sequence modeling may encounter scalability constraints as team sizes grow significantly, and further research is needed to validate performance in very large-scale agent networks.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). Introduces the Decision Transformer framework that reformulates offline reinforcement learning as conditional sequence modeling, establishing the foundational architecture used by the source.
- Paper: Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments, Ryan Lowe et al. (2017). Establishes the centralized training with decentralized execution (CTDE) paradigm in multi-agent reinforcement learning that underlies the source's teacher-student architecture.
- Paper: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning, Tabish Rashid et al. (2020). Provides fundamental principles of cooperative multi-agent reinforcement learning and decentralized policy extraction under centralized training regimes.
- Paper: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, Sergey Levine et al. (2020). Presents a comprehensive review of the challenges, formulations, and core methodologies in offline reinforcement learning necessary for understanding static data training.
- Paper: Counterfactual Multi-Agent Policy Gradients, Jakob N. Foerster et al. (2017). Develops centralized training for decentralized policies in cooperative multi-agent tasks, offering essential context on managing multi-agent credit assignment.
- Paper: Conservative Q-Learning for Offline Reinforcement Learning, Aviral Kumar et al. (2020). Formulates conservative value estimation to prevent distributional shift in offline RL datasets, providing key baseline mechanics for offline policy learning.
- Paper: D4RL: Datasets for Deep Data-Driven Reinforcement Learning, Justin Fu et al. (2020). Defines standard offline reinforcement learning datasets and evaluation benchmarks utilized throughout data-driven decision-making research.
- Paper: Plan Better Amid Conservatism: Offline Multi-Agent Reinforcement Learning with Actor Rectification, Ling Pan et al. (2022). Explores failure modes of standard conservatism in offline multi-agent RL and develops actor rectification, directly extending offline MARL methodologies.
- Paper: Multi-Game Decision Transformers, Kuang-Huei Lee et al. (2022). Scales transformer-based sequence modeling for offline decision making to generalist multi-task domains across extensive offline datasets.
- Paper: Is Value Learning Really the Main Bottleneck in Offline RL?, Seohong Park et al. (2024). Provides an in-depth empirical investigation of bottlenecks across value learning, policy extraction, and policy generalization in offline RL.
