Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms
Kaiqing ZhangZhuoran YangTamer Başar
Synthesizes the theoretical foundations of multi-agent reinforcement learning across stochastic and extensive-form games, categorizing algorithmic guarantees for cooperative, competitive, decentralized, and mean-field settings.
Modern autonomous systems increasingly rely on reinforcement learning to make sequential decisions, but critical real-world deployments—including autonomous driving, drone swarms, robotics, and distributed energy systems—involve multiple interacting participants rather than a single actor. Multi-agent reinforcement learning addresses this setting, where agents optimize their individual or shared goals in a shared environment. Despite high-profile empirical successes in complex board games and real-time strategy environments, theoretical understanding has lagged behind practical implementations, creating uncertainty regarding system stability, scalability, and safety in high-stakes environments.
The article provides a systematic overview of multi-agent reinforcement learning theories and algorithms across cooperative, competitive, and mixed-sum environments. It examines theoretical convergence and sample complexity guarantees across two primary frameworks: Markov games and extensive-form games.
The authors conducted a structured literature review spanning dynamic programming, computational game theory, decentralized control, and optimization theory. They analyzed core mathematical foundations across fully cooperative teams, two-player zero-sum competitions, and general-sum mixed games. The review also examined key operating structures, comparing centralized coordination with fully decentralized execution over communication networks.
The article establishes several major findings regarding the theoretical guarantees of multi-agent systems. First, the core difficulty in multi-agent learning stems from environment non-stationarity, because concurrent learning by multiple agents invalidates the static environment assumption foundational to single-agent reinforcement learning. Second, theoretical guarantees vary substantially by game setting: two-player zero-sum competitive games and cooperative team problems have established convergence proofs, whereas general-sum mixed settings remain largely intractable without restrictive assumptions due to inherent computational complexity barriers. Third, decentralized networked cooperative algorithms can successfully reach global optima using local neighbor-to-neighbor communication, eliminating the need for single-point centralized controllers. Fourth, extensive-form games combined with regret minimization techniques offer provably convergent, polynomial-time frameworks for managing imperfect information and partial observability.
These findings indicate that deploying multi-agent learning in safety-critical systems requires carefully matching the problem setting to known theoretical boundaries. While cooperative systems and two-player competitive systems offer predictable behavior, mixed multi-agent environments carry significant risks of cycling or non-convergence when standard policy gradient methods are used. Practitioners cannot assume that methods performing well in single-agent environments will remain stable in multi-agent deployments.
Organizations developing multi-agent systems should select algorithmic architectures supported by proven convergence properties, such as consensus-based decentralized algorithms for collaborative fleets or regret-minimization methods for imperfect-information settings. Before adopting multi-agent reinforcement learning in high-risk operational domains, decision-makers should invest in rigorous simulation testing, establish formal safety constraints, and investigate model-based approaches that offer better sample efficiency and clearer stability profiles.
The conclusions are limited by the scarcity of non-asymptotic, finite-sample guarantees, as well as the absence of unified theoretical foundations for deep neural network function approximation. Additionally, general partially observable settings remain computationally hard in the worst case. Consequently, while stakeholders can place high confidence in the theoretical foundations of cooperative and two-player zero-sum tabular frameworks, they should exercise caution when scaling deep multi-agent learning to complex, unconstrained environments.
- Paper: Markov Games as a Framework for Multi-Agent Reinforcement Learning, Michael L. Littman (1994). This seminal paper introduces the framework of Markov games to multi-agent reinforcement learning along with the Minimax-Q algorithm, providing the foundational theoretical problem formulation surveyed in the source.
- Paper: Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments, Ryan Lowe et al. (2017). This paper establishes the centralized training with decentralized execution actor-critic paradigm (MADDPG) for mixed cooperative-competitive environments, a central algorithmic architecture analyzed in the survey.
- Paper: QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning, Tabish Rashid et al. (2018). This work introduces monotonic value function factorisation for cooperative multi-agent RL, forming a key benchmark and theoretical mechanism discussed in cooperative game overviews.
- Paper: Counterfactual Multi-Agent Policy Gradients, Jakob N. Foerster et al. (2017). This paper formulates counterfactual multi-agent policy gradients to solve the multi-agent credit assignment problem via centralized critics, providing crucial background for policy-based MARL theory.
- Paper: Learning to Communicate with Deep Multi-Agent Reinforcement Learning, Jakob N. Foerster et al. (2016). This study introduces foundational centralized training with decentralized execution algorithms (DIAL/RIAL) for learning emergent inter-agent communication under partial observability.
- Paper: Reinforcement Learning: A Survey, Leslie Pack Kaelbling et al. (1996). This survey provides the classical single-agent reinforcement learning and Markov decision process foundations that multi-agent theoretical frameworks directly generalize.
- Paper: Technical Note: Q-Learning, CHRISTOPHER J.C.H. WATKINS et al. (2004). This foundational note proves the convergence of single-agent Q-learning, which serves as the core theoretical baseline from which multi-agent value-based algorithms are derived and analyzed.
- Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). This chapter establishes the Policy Gradient Theorem under function approximation, which is fundamental to understanding the convergence and non-convergence of policy-based methods in multi-agent games.
- Paper: The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games, Chao Yu et al. (2022). This paper empirically investigates and extends multi-agent policy optimization by systematically evaluating on-policy centralized and decentralized PPO variants across major cooperative MARL benchmarks.
- Paper: Emergent Tool Use From Multi-Agent Autocurricula, Bowen Baker et al. (2020). This study demonstrates how multi-agent competitive games and autocurricula drive the spontaneous emergence of complex behaviors and tool use in physical multi-agent environments.
- Paper: Deep Reinforcement Learning for Autonomous Driving: A Survey, Bangalore Ravi Kiran et al. (2020). This survey explores the application and operational scaling of deep reinforcement learning algorithms in autonomous driving, a key multi-agent sequential decision-making domain highlighted in the source.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). This survey explores the modern extension of agent frameworks to large language model-based controllers, analyzing multi-agent collaboration and game-theoretic interaction architectures.
- Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). This work establishes scaling laws and architectural principles for multi-agent systems, providing empirical analysis of coordination costs versus single-agent performance.
