Learning Multiagent Communication with Backpropagation

Sainbayar SukhbaatarArthur SzlamRob Fergus

article2016NeurIPS1,416 citations

Introduces CommNet, a neural model that enables cooperative agents to learn continuous, interpretable communication protocols end-to-end using backpropagation to solve collaborative tasks.

Listen

Coordinating multiple autonomous agents in complex, partially observable environments is a core challenge across modern engineering domains, including autonomous vehicle networks, sensor arrays, and robotic systems. Most existing systems rely either on independent controllers that cannot coordinate effectively or on rigid, hand-crafted communication rules that cannot adapt when operational conditions change. The objective of the article is to demonstrate and evaluate a neural network controller, termed CommNet, that enables fully cooperative multi-agent teams to learn their own continuous communication protocols simultaneously with their operational policies using standard backpropagation.

To evaluate this architecture, the authors conducted extensive simulated experiments across four diverse multi-agent benchmarks: a coordinated lever-pulling coordination task, a multi-car traffic junction management task, a team combat simulation against rule-based adversaries, and a natural-language question answering dataset (bAbI). The architecture processes incoming agent states, facilitates continuous communication exchanges through dynamically sized broadcast or local channels, and outputs coordinated action distributions. Performance was benchmarked against independent non-communicating controllers, fixed fully connected networks, and reinforcement learning baselines with discrete communication protocols.

The findings demonstrate substantial performance gains from learned continuous communication. In the lever-pulling task, the proposed model achieved 94% to 99% coordination success compared to only 59% for independent agents. In traffic junction simulations, the model reduced collision failure rates to 1.6% to 2.2%—drastically outperforming independent baselines (9.4% to 20.6% failure) and remaining 90% successful even when agent vision was completely obstructed. In team combat, the model increased win rates across all configurations, reaching up to 49.5% against hard-coded bots that held vision advantages. Across language reasoning tasks, it halved the error rate of independent models from 15.2% to 7.1%. Furthermore, internal analysis revealed that agents autonomously developed sparse, interpretable communication strategies, sending signals primarily when critical coordination—such as braking to avoid collisions—was necessary.

These results demonstrate that multi-agent systems can autonomously generate efficient, lightweight communication without needing expensive manual protocol engineering. By handling dynamic team sizes and fluctuating local neighborhood topologies, the framework reduces operational risk and collision hazards in shared environments while maintaining high throughput. For organizations deploying cooperative automated fleets or multi-unit robotic systems, this approach offers a flexible path to scalable coordination.

Teams considering multi-agent automation should pilot end-to-end differentiable communication architectures over rigid rule-based messaging protocols, especially in environments with limited or changing sensor visibility. Future development should focus on extending the architecture to heterogeneous agent types and testing performance on larger swarms with specialized local connectivity. While the current evidence is highly compelling within controlled simulation environments, confidence in physical-world deployment should be qualified until further validation is conducted under real-world noise, hardware latency, and communication packet drops.

  • Paper: Cooperative Multi-Agent Learning: The State of the Art, Liviu Panait et al. (2005). Provides a comprehensive foundational survey of cooperative multi-agent learning architectures, team versus concurrent learning dynamics, and credit-assignment trade-offs that motivate neural multiagent communication models like CommNet.
Cover for Learning Multiagent Communication with Backpropagation

Abstract

Many tasks in AI require the collaboration of multiple agents. Typically, the communication protocol between agents is manually specified and not altered during training. In this paper we explore a simple neural model, called CommNet, that uses continuous communication for fully cooperative tasks. The model consists of multiple agents and the communication between them is learned alongside their policy. We apply this model to a diverse set of tasks, demonstrating the ability of the agents to learn to communicate amongst themselves, yielding improved performance over non-communicative agents and baselines. In some cases, it is possible to interpret the language devised by the agents, revealing simple but effective strategies for solving the task at hand.

Table of Contents

  • 1 Introduction
  • 2 Communication Model
  • 2.1 Controller Structure
  • 2.2 Model Extensions
  • 3 Related Work
  • 4 Experiments
  • 4.1 Baselines
  • 4.2 Simple Demonstration with a Lever Pulling Task
  • 4.3 Multi-turn Games
  • 4.3.1 Traffic Junction
  • 4.3.2 Analysis of Communication
  • 4.3.3 Combat Task
  • 4.4 bAbI Tasks
  • 5 Discussion and Future Work
  • References
  • A Reinforcement Training
  • B Lever Pulling Task Analysis
  • C Details of Traffic Junction
  • D Traffic Junction Analysis
  • E bAbI Tasks Details

Knowls

  1. Knowl 1 — CommNet Neural Architecture and Continuous Message Passing

    model/method

    The Communication Neural Net (CommNet) is a neural network architecture designed for cooperative multi-agent tasks that allows multiple agents to learn a continuous communication protocol end-to-end via standard backpropagation.

    Let JJ be the number of cooperating agents, and let sjs_j denote the local observation or state-view of agent j∈{1,…,J}j \in \{1, \dots, J\}. The state-view of each agent is first mapped into an initial hidden feature representation hj0∈Rd0h_j^0 \in \mathbb{R}^{d_0} using a task-dependent encoder function rr: hj0=r(sj)h_j^0 = r(s_j) Unless otherwise specified, the initial communication vector for every agent is initialized to zero: cj0=0c_j^0 = 0.

    The model executes KK communication steps (or layers) indexed by i∈{0,…,K−1}i \in \{0, \dots, K-1\}. At communication step ii, each agent jj is updated using a module fif^i parameterized by neural network weights shared across all agents. The module takes as input the agent's current hidden state hjih_j^i and an incoming continuous communication vector cjic_j^i: hji+1=fi(hji,cji)h_j^{i+1} = f^i(h_j^i, c_j^i)

    In the global broadcast setting, the incoming communication vector cji+1c_j^{i+1} sent to agent jj for the next step is computed as the average of the hidden states produced by all other agents: cji+1=1J−1∑j′≠jhj′i+1c_j^{i+1} = \frac{1}{J - 1} \sum_{j' \neq j} h_{j'}^{i+1}

    After KK communication steps, a decoder function q(hjK)q(h_j^K)—consisting of a linear transformation followed by a softmax function—outputs a categorical probability distribution over the discrete action space of agent jj. The discrete action aja_j for agent jj is then sampled from this distribution: aj∼q(hjK)a_j \sim q(h_j^K)

    Because the message aggregation and module updates consist entirely of differentiable operations, the communication protocol is learned jointly with the policy via backpropagation.

  2. Knowl 2 — Block Matrix Linear Formulation and Permutation Invariance of CommNet

    equation

    When the module fif^i at step ii consists of a single linear layer followed by an element-wise non-linearity σ\sigma, the update for agent j∈{1,…,J}j \in \{1, \dots, J\} takes the form: hji+1=σ(Hihji+Cicji)h_j^{i+1} = \sigma(H^i h_j^i + C^i c_j^i) where HiH^i and CiC^i are weight matrices applied to the hidden state and incoming communication vector, respectively.

    Letting hi=[h1i,h2i,…,hJi]Th^i = [h_1^i, h_2^i, \dots, h_J^i]^T be the concatenated vector of all agent hidden states, the entire communication step can be expressed as a linear transformation followed by σ\sigma: hi+1=σ(Tihi)h^{i+1} = \sigma(T^i h^i) where TiT^i is a block matrix defined by: Ti=(HiCˉiCˉi…CˉiCˉiHiCˉi…CˉiCˉiCˉiHi…Cˉi⋮⋮⋮⋱⋮CˉiCˉiCˉi…Hi)T^i = \begin{pmatrix} H^i & \bar{C}^i & \bar{C}^i & \dots & \bar{C}^i \\ \bar{C}^i & H^i & \bar{C}^i & \dots & \bar{C}^i \\ \bar{C}^i & \bar{C}^i & H^i & \dots & \bar{C}^i \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ \bar{C}^i & \bar{C}^i & \bar{C}^i & \dots & H^i \end{pmatrix} with Cˉi=CiJ−1\bar{C}^i = \frac{C^i}{J - 1}.

    The normalization factor 1J−1\frac{1}{J - 1} rescales the aggregated communication by the number of other agents present, allowing the transformation matrix TiT^i to scale dynamically when the number of agents JJ varies at run time. Because the diagonal blocks are identical and all off-diagonal blocks are identical, TiT^i is invariant under permutations of agent ordering.

  3. Knowl 3 — CommNet Architectural Extensions: Local Connectivity, Skip Connections, and Recurrence

    model/method

    CommNet supports several architectural variations:

    1. Local Connectivity: When agents only communicate with neighbors within a spatial or topological range, broadcast aggregation is replaced by local neighborhood averaging. Let N(j)N(j) be the set of agents within communication range of agent jj. The incoming communication vector becomes: cji+1=1∣N(j)∣∑j′∈N(j)hj′i+1c_j^{i+1} = \frac{1}{|N(j)|} \sum_{j' \in N(j)} h_{j'}^{i+1} As agents enter, move, or leave, N(j)N(j) forms a dynamic graph where edges represent communication channels and communication corresponds to message/belief propagation.

    2. Skip Connections: To retain original observation features across multiple communication layers, the initial encoding hj0h_j^0 is provided as an additional direct input to module fif^i at subsequent steps ii: hji+1=fi(hji,cji,hj0)h_j^{i+1} = f^i(h_j^i, c_j^i, h_j^0)

    3. Temporal Recurrence: Instead of using KK distinct feedforward layers per environment step, the communication step index ii can be mapped directly to environment time steps tt. In this recurrent setup, module ftf^t shares its parameters across all time steps tt, taking the form of an RNN cell or an LSTM cell, with action distributions sampled from q(hjt)q(h_j^t) at every time step.

  4. Knowl 4 — Policy Gradient Objective with State-Specific Baseline for CommNet

    equation

    In partially observed multi-agent reinforcement learning settings with sporadic cooperative reward, CommNet parameters θ\theta are trained using a policy gradient algorithm with a state-specific baseline.

    Let an episode consist of TT time steps with states s(1),…,s(T)s(1), \dots, s(T), actions a(1),…,a(T)a(1), \dots, a(T), and scalar rewards r(1),…,r(T)r(1), \dots, r(T) shared among all agents. A state-specific scalar baseline b(s(t),θ)b(s(t), \theta) is predicted by an auxiliary head of the model from the joint state. The parameter update Δθ\Delta \theta minimizes baseline estimation error while maximizing expected future return: Δθ=∑t=1T[∂log⁡p(a(t)∣s(t),θ)∂θ(∑i=tTr(i)−b(s(t),θ))−α∂∂θ(∑i=tTr(i)−b(s(t),θ))2]\Delta \theta = \sum_{t=1}^T \left[ \frac{\partial \log p(a(t) \mid s(t), \theta)}{\partial \theta} \left( \sum_{i=t}^T r(i) - b(s(t), \theta) \right) - \alpha \frac{\partial}{\partial \theta} \left( \sum_{i=t}^T r(i) - b(s(t), \theta) \right)^2 \right] where α\alpha is a hyperparameter balancing the policy gradient and baseline regression objectives, set to α=0.03\alpha = 0.03.

  5. Knowl 5 — Discrete Communication Multi-Agent Baseline

    model/method

    As a baseline to continuous communication, discrete communication requires agents to broadcast discrete categorical symbols whose meanings are learned during training. Because discrete sampling is non-differentiable, communication symbols are treated as internal actions trained via reinforcement learning policy gradients.

    At internal communication step ii, agent jj samples a discrete symbol index wjiw_j^i from a learned distribution over the vocabulary: wji∼Softmax(Dhji)w_j^i \sim \text{Softmax}(D h_j^i) where DD is a learned projection parameter matrix.

    Let w^ji\hat{w}_j^i denote the one-hot binary vector representation of symbol wjiw_j^i. The broadcast communication vector received by agent jj at step i+1i+1 is the element-wise boolean OR (denoted by ∨\vee) of the one-hot symbol vectors emitted by all other agents: cji+1=⋁j′≠jw^j′ic_j^{i+1} = \bigvee_{j' \neq j} \hat{w}_{j'}^i Credit assignment through discrete communication steps is handled by treating symbol transmissions as intermediate actions within an internal temporal policy gradient loop.

  6. Knowl 6 — Lever Pulling Task Performance Under Supervised and RL Training

    data/table

    The lever pulling task tests communication-dependent coordination. From a total pool of N=500N = 500 agents, m=5m = 5 agents are sampled at random in each round. The round requires the mm agents to simultaneously pull one of mm available levers. Each agent only observes its own identity (sj=js_j = j). The cooperative objective is for all mm agents to choose pairwise distinct levers. Performance is evaluated by the ratio of the number of distinct levers pulled to the total number of levers (m=5m = 5), averaged over 500 test trials after 50,000 training batches of size 64.

    Model Φ\Phi Supervised Reinforcement
    Independent 0.59 0.59
    CommNet 0.99 0.94

    The independent controller fails to coordinate beyond the chance level of random choice (0.59), whereas CommNet achieves near-perfect coordination under both supervised sorting targets (0.99) and reinforcement learning (0.94) by communicating agent IDs across continuous channels.

  7. Knowl 7 — Traffic Junction Benchmark Failure Rates Across Architectures and Visibility

    data/table

    The traffic junction task simulates grid intersections where cars enter with probability parrivep_{\text{arrive}}, follow one of three routes, and choose between gas and brake at each step. A collision gives reward rcoll=−10r_{\text{coll}} = -10, and each step spent in the grid incurs a delay cost of τrtime=−0.01τ\tau r_{\text{time}} = -0.01\tau, where τ\tau is elapsed steps since arrival. Episodes run for 40 steps, and an episode is classified as a failure if one or more collisions occur. Agents have limited vision (3×33 \times 3 grid around the vehicle).

    Default Junction Failure Rate (%) Game Variants
    Model Φ\Phi MLP RNN LSTM Easy (MLP) Hard (RNN)
    Independent 20.6±14.120.6 \pm 14.1 19.5±4.519.5 \pm 4.5 9.4±5.69.4 \pm 5.6 15.8±12.515.8 \pm 12.5 26.9±6.026.9 \pm 6.0
    Fully-connected 12.5±4.412.5 \pm 4.4 34.8±19.734.8 \pm 19.7 4.8±2.44.8 \pm 2.4 – –
    Discrete comm. 15.8±9.315.8 \pm 9.3 15.2±2.115.2 \pm 2.1 8.4±3.48.4 \pm 3.4 1.1±2.41.1 \pm 2.4 28.2±5.728.2 \pm 5.7
    CommNet 2.2±0.62.2 \pm 0.6 7.6±1.47.6 \pm 1.4 1.6±1.01.6 \pm 1.0 0.3±0.10.3 \pm 0.1 22.5±6.122.5 \pm 6.1
    CommNet local – – – – 21.1±3.421.1 \pm 3.4

    CommNet reduces junction failure rates across all module types, with LSTM CommNet reaching 1.6% failure compared to 9.4% for the independent LSTM. On the harder four-connected junction grid, CommNet with local connectivity achieves the best performance (21.1±3.421.1 \pm 3.4% failure rate).

  8. Knowl 8 — Combat Task Win Rates Across Team Sizes and Visibility Levels

    data/table

    In the combat gridworld task, a team of mm learning agents battles a team of mm heuristic enemy bots on a 15×1515 \times 15 grid. Each agent has 3 health points, a 3×33 \times 3 visual field, and a 3×33 \times 3 firing range. An attack deals 1 damage and requires a 1-step cooldown. The heuristic bot team has an asymmetric advantage due to pooled vision (if one bot sees an agent, all bots see it). Models receive −1-1 reward for a loss or draw, and −0.1-0.1 times the total remaining health of enemy bots. Results report win rate percentages (mean ±\pm std over 5 runs):

    Default (m=5m=5, 3×33\times 3 vision) Game Variations (MLP)
    Model Φ\Phi MLP RNN LSTM m=3m=3 m=10m=10 5×55\times 5 vision
    Independent 34.2±1.334.2 \pm 1.3 37.3±4.637.3 \pm 4.6 44.3±0.444.3 \pm 0.4 29.2±5.929.2 \pm 5.9 30.5±8.730.5 \pm 8.7 60.5±2.160.5 \pm 2.1
    Fully-connected 17.7±7.117.7 \pm 7.1 2.9±1.82.9 \pm 1.8 19.6±4.219.6 \pm 4.2 – – –
    Discrete comm. 29.1±6.729.1 \pm 6.7 33.4±9.433.4 \pm 9.4 46.4±0.746.4 \pm 0.7 – – –
    CommNet 44.5±13.444.5 \pm 13.4 44.4±11.944.4 \pm 11.9 49.5±12.649.5 \pm 12.6 51.0±14.151.0 \pm 14.1 45.4±12.445.4 \pm 12.4 73.0±0.773.0 \pm 0.7

    Continuous CommNet outperforms independent and discrete communication baselines across all module variants. Fully connected baselines underperform independent models due to overfitting and lack of structural permutation invariance.

  9. Knowl 9 — bAbI Toy Question Answering Multi-Agent Formulation and Results

    data/table

    The bAbI question answering benchmark (20 tasks) is formulated as a multi-agent problem where each story sentence sjs_j is assigned to an individual agent jj, and question sentence qq is broadcast initially as cj0=r(q,θq)c_j^0 = r(q, \theta_q). Each sentence includes a relative temporal position tag t=J−jt = J - j. Agents communicate for K=2K=2 steps using a 2-layer MLP module with skip connections: hji+1=σ(Wiσ(Hihji+Cicji+hj0))h_j^{i+1} = \sigma(W^i \sigma(H^i h_j^i + C^i c_j^i + h_j^0)) After K=2K=2 steps, final hidden states are summed and decoded into an answer word prediction y=Softmax(D∑j=1JhjK)y = \text{Softmax}\left(D \sum_{j=1}^J h_j^K\right), trained via supervised cross-entropy.

    Model Mean Error (%) Failed Tasks (Error >5> 5%)
    LSTM baseline 36.4 16
    MemN2N 4.2 3
    DMN+ 2.8 1
    Independent (MLP module) 15.2 9
    CommNet (MLP module) 7.1 3

    CommNet achieves a mean error of 7.1% and fails only 3 tasks (tasks 3, 16, and 17), whereas the communication-free independent MLP controller fails 9 tasks with a 15.2% mean error.

  10. Knowl 10 — Emergence of Sparse Communication and Coordination Protocols

    empirical result

    Analysis of learned continuous communication vectors c~ji+1=Ci+1hji\tilde{c}_j^{i+1} = C^{i+1} h_j^i in the traffic junction task reveals that CommNet learns a sparse communication strategy.

    Two-dimensional PCA projections of the communication contributions c~ji+1\tilde{c}_j^{i+1} demonstrate that the majority of transmissions have norms close to zero ('silent' agents), even when the underlying hidden state vectors hjih_j^i show high variance across agents. Non-zero communication vectors form tight, distinct spatial clusters emitted when vehicles reach specific critical approach locations near the intersection.

    Tracking two-car interactions shows that the emission of these clustered non-zero messages correlates directly with the other car executing a braking action at potential collision points. In addition, spatial visualizations of braking frequency demonstrate that vehicles approaching from the left never brake for cross traffic, indicating that agents learn an asymmetrical priority rule (a right-of-way convention) to resolve intersection conflicts without collisions.

Coverage note — None was omitted; all primary architectural components, RL training objectives, baseline definitions, experimental tasks (Lever Pulling, Traffic Junction, Combat, bAbI), and empirical communication analyses from the paper are covered.

References

  1. 1.Y. Bengio, J. Louradour, R. Collobert, and J. Weston. Curriculum learning. In ICML, 2009.
  2. 2.L. Busoniu, R. Babuska, and B. De Schutter. A comprehensive survey of multiagent reinforcement learning. Systems, Man, and Cybernetics, IEEE Transactions on, 38(2):156–172, 2008.
  3. 3.Y. Cao, W. Yu, W. Ren, and G. Chen. An overview of recent progress in the study of distributed multi-agent coordination. IEEE Transactions on Industrial Informatics, 1(9):427–438, 2013.
  4. 4.R. H. Crites and A. G. Barto. Elevator group control using multiple reinforcement learning agents. Machine Learning, 33(2):235–262, 1998.
  5. 5.J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson. Learning to communicate to solve riddles with deep distributed recurrent Q-networks. arXiv, abs/1602.02672, 2016.
  6. 6.D. Fox, W. Burgard, H. Kruppa, and S. Thrun. Probabilistic approach to collaborative multi-robot localization. Autonomous Robots, 8(3):325––344, 2000.
  7. 7.C. L. Giles and K. C. Jim. Learning communication for multi-agent systems. In Innovative Concepts for Agent Based Systems, pages 377—-390. Springer, 2002.
  8. 8.C. Guestrin, D. Koller, and R. Parr. Multiagent planning with factored MDPs. In NIPS, 2001.
  9. 9.X. Guo, S. Singh, H. Lee, R. L. Lewis, and X. Wang. Deep learning for real-time atari game play using offline monte-carlo tree search planning. In NIPS, 2014.
  10. 10.L. Kaiser and I. Sutskever. Neural gpus learn algorithms. In ICLR, 2016.
  11. 11.T. Kasai, H. Tenmoto, and A. Kamiya. Learning of communication codes in multi-agent reinforcement learning problem. IEEE Conference on Soft Computing in Industrial Applications, pages 1–6, 2008.
  12. 12.D. Kingma and J. Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  13. 13.M. Lauer and M. A. Riedmiller. An algorithm for distributed reinforcement learning in cooperative multi-agent systems. In ICML, 2000.
  14. 14.S. Levine, C. Finn, T. Darrell, and P. Abbeel. End-to-end training of deep visuomotor policies. Journal of Machine Learning Research, 17(39):1–40, 2016.
  15. 15.Y. Li, D. Tarlow, M. Brockschmidt, and R. Zemel. Gated graph sequence neural networks. In ICLR, 2015.
  16. 16.M. L. Littman. Value-function reinforcement learning in markov games. Cognitive Systems Research, 2(1):55–66, 2001.
  17. 17.C. J. Maddison, A. Huang, I. Sutskever, and D. Silver. Move evaluation in go using deep convolutional neural networks. In ICLR, 2015.
  18. 18.D. Maravall, J. De Lope, and R. Domnguez. Coordination of communication in robot teams by reinforcement learning. Robotics and Autonomous Systems, 61(7):661–666, 2013.
  19. 19.M. Matari. Reinforcement learning in the multi-robot domain. Autonomous Robots, 4(1):73–83, 1997.
  20. 20.F. S. Melo, M. Spaan, and S. J. Witwicki. Querypomdp: Pomdp-based communication in multiagent systems. In Multi-Agent Systems, pages 189–204, 2011.
  21. 21.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, D. Wierstra, S. Legg, and D. Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
  22. 22.R. Olfati-Saber, J. Fax, and R. Murray. Consensus and cooperation in networked multi-agent systems. Proceedings of the IEEE, 95(1):215–233, 2007.
  23. 23.J. Pearl. Reverend bayes on inference engines: A distributed hierarchical approach. In AAAI, 1982.
  24. 24.B. Peng, Z. Lu, H. Li, and K. Wong. Towards Neural Network-based Reasoning. ArXiv preprint: 1508.05508, 2015.
  25. 25.F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini. The graph neural network model. IEEE Trans. Neural Networks, 20(1):61–80, 2009.
  26. 26.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489, 2016.
  27. 27.P. Stone and M. Veloso. Towards collaborative and adversarial learning: A case study in robotic soccer. International Journal of Human Computer Studies, (48), 1998.
  28. 28.S. Sukhbaatar, A. Szlam, G. Synnaeve, S. Chintala, and R. Fergus. Mazebase: A sandbox for learning from games. CoRR, abs/1511.07401, 2015.
  29. 29.S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus. End-to-end memory networks. NIPS, 2015.
  30. 30.R. S. Sutton and A. G. Barto. Introduction to Reinforcement Learning. MIT Press, 1998.
  31. 31.A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, and R. Vicente. Multiagent cooperation and competition with deep reinforcement learning. arXiv:1511.08779, 2015.
  32. 32.M. Tan. Multi-agent reinforcement learning: Independent vs. cooperative agents. In ICML, 1993.
  33. 33.T. Tieleman and G. Hinton. Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural Networks for Machine Learning, 2012.
  34. 34.P. Varshavskaya, L. P. Kaelbling, and D. Rus. Distributed Autonomous Robotic Systems 8, chapter Efficient Distributed Reinforcement Learning through Agreement, pages 367–378. 2009.
  35. 35.X. Wang and T. Sandholm. Reinforcement learning to play an optimal nash equilibrium in team markov games. In NIPS, pages 1571–1578, 2002.
  36. 36.J. Weston, A. Bordes, S. Chopra, and T. Mikolov. Towards ai-complete question answering: A set of prerequisite toy tasks. In ICLR, 2016.
  37. 37.R. J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. In Machine Learning, pages 229–256, 1992.
  38. 38.C. Xiong, S. Merity, and R. Socher. Dynamic memory networks for visual and textual question answering. ICML, 2016.
  39. 39.C. Zhang and V. Lesser. Coordinating multi-agent reinforcement learning with limited communication. In Proc. AAMAS, pages 1101–1108, 2013.

Citation

MLA
Sukhbaatar, S., et al. “Learning Multiagent Communication with Backpropagation”. arXiv, 2016, http://arxiv.org/abs/1605.07736v2.
APA
Sukhbaatar, S., Szlam, A., & Fergus, R. (2016). Learning Multiagent Communication with Backpropagation. arXiv. http://arxiv.org/abs/1605.07736v2
Chicago
Sukhbaatar, S., A. Szlam, and R. Fergus. 2016. “Learning Multiagent Communication with Backpropagation”. arXiv. http://arxiv.org/abs/1605.07736v2.
Harvard
Sukhbaatar, S., Szlam, A. and Fergus, R. (2016) “Learning Multiagent Communication with Backpropagation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1605.07736v2.
Vancouver
1. Sukhbaatar S, Szlam A, Fergus R (2016) Learning Multiagent Communication with Backpropagation. arXiv

BibTeX

@article{sukhbaatar2016learning,
  title = {Learning Multiagent Communication with Backpropagation},
  author = {Sukhbaatar, Sainbayar and Szlam, Arthur and Fergus, Rob},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1605.07736v2},
  eprint = {1605.07736}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors