Value-Decomposition Networks For Cooperative Multi-Agent Learning

Peter SunehagGuy LeverAudrunas GruslysWojciech Marian CzarneckiVinicius ZambaldiMax JaderbergMarc LanctotNicolas SonneratJoel Z. LeiboKarl Tuyls

article2017arXiv1,380 citations

Introduces Value-Decomposition Networks to solve the cooperative multi-agent credit assignment problem by decomposing joint team value functions into individual agent utilities to overcome spurious rewards and lazy-agent behavior under partial observability.

Listen

Coordinating autonomous systems—such as self-driving vehicles, traffic management systems, and automated factory components—is a critical challenge in modern artificial intelligence. In these cooperative multi-agent environments, individual agents must act based on limited local observations while working toward a single shared team goal. Standard approaches generally perform poorly: centralized control suffers from an exponentially exploding action space and produces "lazy agents" where one unit remains inactive to avoid disrupting another, while fully decentralized independent learning fails because agents misinterpret teammates' actions as environmental randomness and struggle with spurious rewards.

The article introduces a novel Value-Decomposition Network architecture designed to solve cooperative multi-agent reinforcement learning problems with a single joint reward signal. The primary objective is to evaluate whether learning to decompose a shared team value function into individual, agent-specific value components enables decentralized execution while maintaining high team performance.

To evaluate this framework, the authors conducted simulations across seven challenging two-player grid-world tasks (Switch, Fetch, and Checkers) featuring severe partial observability. They benchmarked nine distinct system configurations, comparing value-decomposition networks against standard centralized and independent baselines. The study also examined enhancements including weight sharing across agents, role identification tags, and differentiable communication channels. Each architectural configuration was evaluated across ten independent experimental runs over 50,000 training episodes.

The findings show that value-decomposition architectures consistently and significantly outperformed both centralized approaches and independent learners across all benchmark environments. First, learning an additive decomposition successfully eliminated spurious reward interference and enabled agents to act independently during execution based only on local views. Second, pairing value decomposition with shared network weights prevented the lazy agent problem on symmetrical coordination tasks. Third, adding role identification tags resolved bottlenecks when agents needed distinct responsibilities or asymmetric behaviors. Finally, adding low-level communication channels between agent networks accelerated learning speed compared to higher-level communication in environments requiring tight coordination.

These results demonstrate that complex collective tasks can be autonomously broken down into simpler, learnable local responsibilities without manually engineering individual reward functions. In practice, this approach substantially reduces computational and operational risk: teams can be trained centrally using global feedback and then safely deployed as autonomous, decentralized units. This architecture resolves a long-standing trade-off between the scalability of decentralized systems and the coordination stability of centralized controllers.

Decision-makers should consider value-decomposition methods for autonomous multi-agent coordination, incorporating shared network weights for symmetric tasks and explicit role identifiers when agents have asymmetric reward scales. Future work should pilot this framework on systems with larger team sizes, validate more complex non-linear value aggregation techniques, and test performance in physical hardware environments.

Confidence in these findings is high for two-agent cooperative settings with joint reward structures, supported by consistent performance across multiple seeds and task designs. However, because the study is limited to simulated two-player grid environments with linear value summation, stakeholders should exercise caution when extrapolating these results directly to large-scale teams or systems with highly non-linear reward dependencies without further testing.

arXiv: 1706.05296
Cover for Value-Decomposition Networks For Cooperative Multi-Agent Learning

Abstract

We study the problem of cooperative multi-agent reinforcement learning with a single joint reward signal. This class of learning problems is difficult because of the often large combined action and observation spaces. In the fully centralized and decentralized approaches, we find the problem of spurious rewards and a phenomenon we call the "lazy agent" problem, which arises due to partial observability. We address these problems by training individual agents with a novel value decomposition network architecture, which learns to decompose the team value function into agent-wise value functions. We perform an experimental evaluation across a range of partially-observable multi-agent domains and show that learning such value-decompositions leads to superior results, in particular when combined with weight sharing, role information and information channels.

Table of Contents

  • 1 Introduction
  • 1.1 Other Related Work
  • 2 Background
  • 2.1 Reinforcement Learning
  • 2.2 Deep QQ-Learning
  • 2.3 Multi-Agent Reinforcement Learning
  • 3 A Deep-RL Architecture for Coop-MARL
  • 4 Experiments
  • 4.1 Agents
  • 4.2 Environments
  • 4.3 Results
  • 4.4 The Learned QQ-Decomposition
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Additive Value-Decomposition Network (VDN)

    model/method

    Value-Decomposition Networks (VDN) address cooperative multi-agent reinforcement learning in decentralized partially observable Markov decision processes (Dec-POMDPs) where dd agents observe local histories hih^i and execute individual actions ai∈Aa^i \in \mathcal{A}, receiving only a single scalar team reward r(s,a)r(s, \mathbf{a}).

    VDN assumes and enforces that the joint action-value function Q(h,a)Q(\mathbf{h}, \mathbf{a}) additively factorizes into individual agent value functions Q~i(hi,ai)\tilde{Q}_i(h^i, a^i):

    Q((h1,h2,…,hd),(a1,a2,…,ad))≈∑i=1dQ~i(hi,ai)Q((h^1, h^2, \dots, h^d), (a^1, a^2, \dots, a^d)) \approx \sum_{i=1}^d \tilde{Q}_i(h^i, a^i)

    During centralized training, the loss is computed using deep Q-learning targets evaluated on the joint reward signal rtr_t, and gradients from the TD error are backpropagated through the sum to train the individual network components Q~i\tilde{Q}_i. The functions Q~i\tilde{Q}_i are learned implicitly without needing individual agent-specific reward signals.

    During decentralized execution, each agent selects its action greedily with respect to its local value function:

    a∗,i=arg⁡max⁡ai∈AQ~i(hi,ai)a^{*, i} = \arg\max_{a^i \in \mathcal{A}} \tilde{Q}_i(h^i, a^i)

    Because the sum of maximums equals the maximum of the sum under additive factorization, local greedy choices are guaranteed to maximize the joint action-value function Q(h,a)Q(\mathbf{h}, \mathbf{a}).

  2. Knowl 2 — Lazy Agent Problem

    definition

    The lazy agent problem is a failure mode in cooperative multi-agent reinforcement learning with shared team rewards and partial observability. It occurs when one agent successfully learns a partially effective policy that generates positive reward, while a teammate is discouraged from exploring or learning because the exploratory actions of the second agent interfere with the first agent's routine and temporarily degrade the joint team reward. Consequently, the second agent learns an inactive or passive policy ("laziness"), leading to suboptimal system-level performance.

  3. Knowl 3 — Agent Invariance and Role Information Conditioning

    definition

    A joint multi-agent policy π:Hd→P(A)d\pi : \mathcal{H}^d \to \mathcal{P}(\mathcal{A})^d mapping joint observation histories hˉ=(h1,…,hd)\bar{h} = (h^1, \dots, h^d) to joint action distributions is defined as agent invariant if for any permutation (bijection) p:{1,…,d}→{1,…,d}p : \{1, \dots, d\} \to \{1, \dots, d\}:

    π(p(hˉ))=p(π(hˉ))\pi(p(\bar{h})) = p(\pi(\bar{h}))

    Agent invariance is enforced by sharing all neural network weights across agents, which drastically reduces parameter count and prevents asymmetric lazy-agent failure modes.

    When a task requires specialized agent behaviors (asymmetric roles), agents are conditioned on role information by concatenating a 1-hot encoding of each agent's identity to its local observation at the input layer. Networks with shared weights conditioned on role identifiers satisfy conditional agent invariance, yielding identical policies only when conditioned on identical roles.

  4. Knowl 4 — Recurrent Dueling Deep Q-Network Architecture for VDN

    model/method

    The individual utility function Q~i(hi,ai)\tilde{Q}_i(h^i, a^i) for each agent ii is parameterized by a deep neural network structured as follows:

    1. Input & Representation Layer: Local observation of size 3×5×53 \times 5 \times 5 (RGB grid) passes through a fully connected linear layer with 32 hidden units followed by a ReLU non-linearity.
    2. Recurrent Memory Layer: The 32-dimensional feature vector is passed into an LSTM recurrent layer with 32 units, followed by a ReLU non-linearity to handle partial observability over history hih^i.
    3. Dueling Output Layer: A dueling architecture with 32 units produces a state-value stream V(hi)V(h^i) and an advantage stream A(hi,ai)A(h^i, a^i), combined to compute individual utilities:

    Q~i(hi,ai)=V(hi)+(A(hi,ai)−1∣A∣∑a′∈AA(hi,a′))\tilde{Q}_i(h^i, a^i) = V(h^i) + \left( A(h^i, a^i) - \frac{1}{|\mathcal{A}|} \sum_{a' \in \mathcal{A}} A(h^i, a') \right)

    Training Hyperparameters & Procedures:

    • Optimization: Adam optimizer with an initial learning rate of 0.00010.0001.
    • Sequence Learning: Multi-step updates using forward-view eligibility traces with parameter λ=0.9\lambda = 0.9 over trajectories of length 8.
    • Recurrence Truncation: Truncated backpropagation through time (BPTT) unrolled over 8 steps.
    • Stabilization: Experience replay buffer and target networks.
  5. Knowl 5 — Low-Level and High-Level Inter-Agent Communication Channels

    model/method

    VDN architectures can be augmented with differentiable inter-agent communication channels to share representations across agents:

    • Low-Level Communication: Each agent processes its local observation through its first fully connected layer. The resulting 32-dimensional feature vectors from all dd agents are concatenated into a 32d32d-dimensional vector (ordered with the receiving agent's own features first to preserve agent invariance) before being passed through a ReLU into the receiving agent's LSTM.
    • High-Level Communication: Each agent processes its history independently through its LSTM. The resulting hidden states from all dd agents are concatenated before being passed to the linear dueling layer.

    Because gradients from the joint team reward backpropagate through the communication channels into other agents' feature extractors, communication channels induce joint optimization of representations across teammates.

  6. Knowl 6 — Partially Observable Multi-Agent Gridworld Benchmark Suite

    experimental setup

    The experimental evaluation consists of 2-player cooperative 2D grid worlds with 8 discrete actions: step forward, step backward, step left, step right, rotate left, rotate right, use beam, and stand still. Observations are local 3×5×53 \times 5 \times 5 RGB windows (extending 4 squares ahead and 2 squares laterally relative to agent orientation). Episodes last 5,000 steps during training and 2,000 steps during testing across three primary domains:

    1. Switch: Two agents spawn at opposite ends of a map and must navigate through narrow corridor bottlenecks to reach target goals at the opposing ends. If agents meet in a corridor, one must step aside or reverse to let the other pass. Scoring occurs when both reach their goals (+1 per player reaching goal).
    2. Fetch: Both agents start at base, travel across maps (open, 1 corridor, or 2 corridors) to pickup items (+3 team points on pickup), and return them to base (+5 team points on drop-off). Optimal play requires cycle synchronization where one agent returns as the other departs.
    3. Checkers: Agents must clear obstructing lemons (penalty) to collect apples (+ reward). Agent 1 is sensitive (+10+10 for apple, −10-10 for lemon), while Agent 2 is less sensitive (+1+1 for apple, −1-1 for lemon). Optimal coordination requires Agent 2 to clear lemons and leave apples for Agent 1.
  7. Knowl 7 — Performance Advantage of VDN Over Centralized and Independent Learners

    empirical result

    Across 10 independent runs over 50,000 training episodes on the Switch, Fetch, and Checkers domains:

    • VDN vs. Independent Learners: Independent Q-learners (IL) learning directly from team rewards fail across most tasks due to non-stationarity and spurious reward signals caused by unobserved teammate actions.
    • VDN vs. Centralized Baselines: Combinatorially centralized joint action learners fail on multi-corridor and synchronization tasks due to exponential action-space scaling and susceptibility to the lazy agent problem.
    • Role of Weight Sharing: Weight sharing prevents lazy agent behaviors in symmetric tasks such as Fetch with 1 corridor, where unshared VDN suffers from lazy agent collapse.
    • Role of Role Identifiers and Channels: Combining weight sharing with 1-hot role identifiers enables optimal performance in asymmetric tasks (Checkers and single-corridor Switch). Low-level communication channels achieve faster convergence than high-level channels by sharing raw feature representations prior to recurrent temporal integration.
  8. Knowl 8 — Autonomous Credit Assignment in Learned Individual Utilities

    empirical result

    In the Fetch domain, tracking the individual learned component functions Q~1(h1,a1)\tilde{Q}_1(h^1, a^1) and Q~2(h2,a2)\tilde{Q}_2(h^2, a^2) demonstrates autonomous credit assignment from a single scalar team reward:

    • When Agent 1 executes a pickup (+3 team points) or drop-off (+5 team points), its individual utility Q~1\tilde{Q}_1 spikes sharply in anticipation of the reward event, while Q~2\tilde{Q}_2 for Agent 2 remains flat.
    • Conversely, imminent drop-off events by Agent 2 trigger spikes exclusively in Q~2\tilde{Q}_2 while Q~1\tilde{Q}_1 remains baseline.

    This shows that additive backpropagation through the joint QQ-value naturally decomposes the scalar team reward into agent-attributable sub-utilities based on local action-observation histories.

Coverage note — None was omitted; all key theoretical models, architecture definitions, experimental environments, and empirical findings were extracted.

References

  1. 1.A. K. Agogino and K. Tumer. Analyzing and visualizing multiagent rewards in dynamic and stochastic environments. Journal of Autonomous Agents and Multi-Agent Systems, 17(2):320–338, 2008.
  2. 2.M. Babes, E. M. de Cote, and M. L. Littman. Social reward shaping in the prisoner’s dilemma. In 7th International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS 2008), Estoril, Portugal, May 12-16, 2008, Volume 3, pages 1389–1392, 2008.
  3. 3.D. S. Bernstein, S. Zilberstein, and N. Immerman. The complexity of decentralized control of Markov Decision Processes. In UAI ’00: Proceedings of the 16th Conference in Uncertainty in Artificial Intelligence, Stanford University, Stanford, California, USA, June 30 - July 3, 2000, pages 32–37, 2000.
  4. 4.L. Busoniu, R. Babuska, and B. D. Schutter. A comprehensive survey of multiagent reinforcement learning. IEEE Transactions of Systems, Man, and Cybernetics Part C: Applications and Reviews, 38(2), 2008.
  5. 5.C. Claus and C. Boutilier. The dynamics of reinforcement learning in cooperative multiagent systems. In Proceedings of the Fifteenth National Conference on Artificial Intelligence and Tenth Innovative Applications of Artificial Intelligence Conference, AAAI 98, IAAI 98, July 26-30, 1998, Madison, Wisconsin, USA., pages 746–752, 1998.
  6. 6.M. Colby, T. Duchow-Pressley, J. J. Chung, and K. Tumer. Local approximation of difference evaluation functions. In Proceedings of the Fifteenth International Joint Conference on Autonomous Agents and Multiagent Systems, Singapore, May 2016.
  7. 7.S. Devlin, L. Yliniemi, D. Kudenko, and K. Tumer. Potential-based difference rewards for multiagent reinforcement learning. In Proceedings of the Thirteenth International Joint Conference on Autonomous Agents and Multiagent Systems, May 2014.
  8. 8.A. Eck, L. Soh, S. Devlin, and D. Kudenko. Potential-based reward shaping for finite horizon online POMDP planning. Autonomous Agents and Multi-Agent Systems, 30(3):403–445, 2016.
  9. 9.J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson. Learning to communicate with deep multi-agent reinforcement learning. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pages 2137–2145, 2016.
  10. 10.C. Guestrin, M. G. Lagoudakis, and R. Parr. Coordinated reinforcement learning. In Proceedings of the Nineteenth International Conference on Machine Learning, ICML ’02, pages 227–234, San Francisco, CA, USA, 2002. Morgan Kaufmann Publishers Inc. ISBN 1-55860-873-7. URL http://dl.acm.org/citation.cfm?id=645531.757784.
  11. 11.J. Harb and D. Precup. Investigating recurrence and eligibility traces in deep Q-networks. In Deep Reinforcement Learning Workshop, NIPS 2016, Barcelona, Spain, 2016.
  12. 12.M. J. Hausknecht. Cooperation and Communication in Multiagent Deep Reinforcement Learning. PhD thesis, The University of Texas at Austin, 2016.
  13. 13.M. J. Hausknecht and P. Stone. Deep recurrent Q-learning for partially observable MDPs. CoRR, abs/1507.06527, 2015.
  14. 14.C. HolmesParker, A. Agogino, and K. Tumer. Combining reward shaping and hierarchies for scaling to large multiagent systems. Knowledge Engineering Review, 2016. to appear.
  15. 15.J. Hu and M. P. Wellman. Nash q-learning for general-sum stochastic games. Journal of Machine Learning Research, 4:1039–1069, 2003.
  16. 16.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. URL http://arxiv.org/abs/1412.6980.
  17. 17.L. Kuyer, S. Whiteson, B. Bakker, and N. A. Vlassis. Multiagent reinforcement learning for urban traffic control using coordination graphs. In Machine Learning and Knowledge Discovery in Databases, European Conference, ECML/PKDD 2008, Antwerp, Belgium, September 15-19, 2008, Proceedings, Part I, pages 656–671, 2008.
  18. 18.G. J. Laurent, L. Matignon, and N. L. Fort-Piat. The world of independent learners is not Markovian. Int. J. Know.-Based Intell. Eng. Syst., 15(1):55–64, 2011.
  19. 19.J. Z. Leibo, V. Zambaldi, M. Lanctot, J. Marecki, and T. Graepel. Multi-agent Reinforcement Learning in Sequential Social Dilemmas. In Proceedings of the 16th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2017), Sao Paulo, Brazil, 2017.
  20. 20.M. L. Littman. Markov games as a framework for multi-agent reinforcement learning. In Machine Learning, Proceedings of the Eleventh International Conference, Rutgers University, New Brunswick, NJ, USA, July 10-13, 1994, pages 157–163, 1994.
  21. 21.M. L. Littman. Friend-or-foe q-learning in general-sum games. In Proceedings of the Eighteenth International Conference on Machine Learning (ICML 2001), Williams College, Williamstown, MA, USA, June 28 - July 1, 2001, pages 322–328, 2001.
  22. 22.V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 02 2015.
  23. 23.V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, pages 1928–1937, 2016.
  24. 24.A. Y. Ng, D. Harada, and S. J. Russell. Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the Sixteenth International Conference on Machine Learning (ICML 1999), Bled, Slovenia, June 27 - 30, 1999, pages 278–287, 1999.
  25. 25.F. A. Oliehoek and C. Amato. A Concise Introduction to Decentralized POMDPs. SpringerBriefs in Intelligent Systems. Springer, 2016.
  26. 26.F. A. Oliehoek, M. T. J. Spaan, and N. A. Vlassis. Optimal and approximate q-value functions for decentralized pomdps. J. Artif. Intell. Res. (JAIR), 32:289–353, 2008.
  27. 27.L. Panait and S. Luke. Cooperative multi-agent learning: The state of the art. Autonomous Agents and Multi-Agent Systems, 11(3):387–434, 2005.
  28. 28.S. Proper and K. Tumer. Modeling difference rewards for multiagent learning (extended abstract). In Proceedings of the Eleventh International Joint Conference on Autonomous Agents and Multiagent Systems, Valencia, Spain, June 2012.
  29. 29.M. Puterman. Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley, New York, 1994.
  30. 30.S. J. Russell and P. Norvig. Artificial Intelligence: A Modern Approach. Prentice Hall, Englewood Cliffs, NJ, 3^{nd} edition, 2010.
  31. 31.S. J. Russell and A. Zimdars. Q-decomposition for reinforcement learning agents. In Machine Learning, Proceedings of the Twentieth International Conference (ICML 2003), August 21-24, 2003, Washington, DC, USA, pages 656–663, 2003.
  32. 32.J. G. Schneider, W. Wong, A. W. Moore, and M. A. Riedmiller. Distributed value functions. In Proceedings of the Sixteenth International Conference on Machine Learning (ICML 1999), Bled, Slovenia, June 27 - 30, 1999, pages 371–378, 1999.
  33. 33.S. Sukhbaatar, A. Szlam, and R. Fergus. Learning multiagent communication with backpropagation. CoRR, abs/1605.07736, 2016. URL http://arxiv.org/abs/1605.07736.
  34. 34.R. Sutton and A. Barto. Reinforcement Learning. The MIT Press, 1998.
  35. 35.C. Szepesvári. Algorithms for Reinforcement Learning. Synthesis Lectures on Artificial Intelligence and Machine Learning. Morgan & Claypool Publishers, 2010.
  36. 36.K. Tumer and D. Wolpert. A survey of collectives. In K. Tumer and D. Wolpert, editors, Collectives and the Design of Complex Systems, pages 1–42. Springer, 2004.
  37. 37.K. Tuyls and G. Weiss. Multiagent learning: Basics, challenges, and prospects. AI Magazine, 33(3): 41–52, 2012.
  38. 38.E. van der Pol and F. A. Oliehoek. Coordinated deep reinforcement learners for traffic light control. NIPS Workshop on Learning, Inference and Control of Multi-Agent Systems, 2016.
  39. 39.Video. Video for the q-decomposition plot. 2017. URL https://youtu.be/aAH1eyUQsRo.
  40. 40.Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas. Dueling network architectures for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning (ICML), pages 1995–2003, 2016.

Citation

MLA
Sunehag, P., et al. “Value-Decomposition Networks For Cooperative Multi-Agent Learning”. arXiv, 2017, http://arxiv.org/abs/1706.05296v1.
APA
Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J. Z., Tuyls, K., & Graepel, T. (2017). Value-Decomposition Networks For Cooperative Multi-Agent Learning. arXiv. http://arxiv.org/abs/1706.05296v1
Chicago
Sunehag, P., G. Lever, A. Gruslys, et al. 2017. “Value-Decomposition Networks For Cooperative Multi-Agent Learning”. arXiv. http://arxiv.org/abs/1706.05296v1.
Harvard
Sunehag, P. et al. (2017) “Value-Decomposition Networks For Cooperative Multi-Agent Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1706.05296v1.
Vancouver
1. Sunehag P, Lever G, Gruslys A, et al (2017) Value-Decomposition Networks For Cooperative Multi-Agent Learning. arXiv

BibTeX

@article{sunehag2017value,
  title = {Value-Decomposition Networks For Cooperative Multi-Agent Learning},
  author = {Sunehag, Peter and Lever, Guy and Gruslys, Audrunas and Czarnecki, Wojciech Marian and Zambaldi, Vinicius and Jaderberg, Max and Lanctot, Marc and Sonnerat, Nicolas and Leibo, Joel Z. and Tuyls, Karl and Graepel, Thore},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1706.05296v1},
  eprint = {1706.05296}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission