Cooperative Multi-Agent Learning: The State of the Art

Liviu PanaitSean Luke

article2005AAMAS1,459 citations

Synthesizes cooperative multi-agent learning across robotics, evolutionary computation, and reinforcement learning by categorizing approaches into team and concurrent learning while identifying key challenges in communication, scalability, and task decomposition.

Listen

Modern computational challenges in logistics, robotics, telecommunications, and automated management increasingly rely on decentralized systems where multiple autonomous entities must collaborate to achieve common goals. Manually programming optimal strategies for these interacting entities is often prohibitively complex due to massive decision spaces and unpredictable group dynamics. The article provides a comprehensive evaluation of machine learning methods applied to cooperative multi-agent systems, synthesizing literature across reinforcement learning, evolutionary computation, game theory, and robotics to establish a unifying structural framework for the discipline.

The analysis identifies a fundamental structural division in the field between team learning, where a single central learner optimizes behavior for the entire group, and concurrent learning, where multiple independent learners adapt simultaneously. Team learning simplifies the optimization process by avoiding complex inter-agent credit assignment, but it suffers from severe exponential state-space growth as the number of agents increases. In contrast, concurrent learning projects the problem into smaller, individual search spaces per agent, but it violates standard stationary environment assumptions because agents continuously alter one another's operational context during adaptation.

Key findings emphasize critical trade-offs in system design and learning dynamics. First, agent homogeneity drastically reduces search complexity and performs well in general tasks, whereas heterogeneous specialization yields superior performance only in inherently decomposable domains that demand distinct roles. Second, concurrent learning frequently leads agents to converge on suboptimal equilibria, as individual rational optimization can prevent the discovery of globally optimal team behaviors. Third, reward structuring fundamentally dictates system outcomes: global rewards encourage complete collaboration but scale poorly, whereas local rewards accelerate individual learning rates at the risk of creating counterproductive, uncooperative competition. Finally, while unconstrained communication can simplify coordination, real-world latency and bandwidth limits require selective, structured, or indirect communication methods such as digital pheromones.

These findings indicate that deploying multi-agent learning in operational environments carries notable risks if algorithms are selected without considering problem structure. Organizations attempting to build automated multi-agent systems must carefully balance the trade-offs between centralized coordination overhead and the instability of decentralized adaptation. To advance system capabilities, future development should prioritize techniques that decompose large-scale tasks into manageable subtasks, scale beyond simplistic two- or three-agent models, and evaluate algorithms based on global team efficiency rather than purely individual equilibrium metrics.

Decision-makers should approach these conclusions with measured caution regarding large-scale real-world deployments. Much of the underlying empirical literature relies on low-dimensional simulations, stateless game scenarios, or small teams of identical agents. Consequently, while the article provides high confidence in its theoretical taxonomy and structural trade-offs, applying these techniques to complex, safety-critical systems with hundreds of heterogeneous agents will require targeted validation and pilot testing in high-fidelity environments.

Cover for Cooperative Multi-Agent Learning: The State of the Art

Abstract

Cooperative multi-agent systems are ones in which several agents attempt, through their interaction, to jointly solve tasks or to maximize utility. Due to the interactions among the agents, multi-agent problem complexity can rise rapidly with the number of agents or their behavioral sophistication. The challenge this presents to the task of programming solutions to multi-agent systems problems has spawned increasing interest in machine learning techniques to automate the search and optimization process.

We provide a broad survey of the cooperative multi-agent learning literature. Previous surveys of this area have largely focused on issues common to specific subareas (for example, reinforcement learning or robotics). In this survey we attempt to draw from multi-agent learning work in a spectrum of areas, including reinforcement learning, evolutionary computation, game theory, complex systems, agent modeling, and robotics.

We find that this broad view leads to a division of the work into two categories, each with its own special issues: applying a single learner to discover joint solutions to multi-agent problems (team learning), or using multiple simultaneous learners, often one per agent (concurrent learning). Additionally, we discuss direct and indirect communication in connection with learning, plus open issues in task decomposition, scalability, and adaptive dynamics. We conclude with a presentation of multi-agent learning problem domains, and a list of multi-agent learning resources.

Table of Contents

  • 1 Introduction
  • 1.1 Multi-Agent Systems
  • 1.2 Multi-Agent Learning
  • 1.3 Machine Learning Methods
  • 1.4 Survey Structure and Taxonomy
  • 2 Team Learning
  • 2.1 Homogeneous Team Learning
  • 2.2 Heterogeneous Team Learning
  • 2.3 Hybrid Team Learning
  • 3 Concurrent Learning
  • 3.1 Credit Assignment
  • 3.2 The Dynamics of Learning
  • 3.2.1 Fully Cooperative Scenarios
  • 3.2.2 General Sum Games
  • 3.2.3 Related Pathologies from Competitive Learning
  • 3.3 Teammate Modeling
  • 4 Learning and Communication
  • 4.1 Direct Communication
  • 4.2 Indirect Communication
  • 5 Major Open Topics
  • 5.1 Scalability
  • 5.2 Adaptive Dynamics and Nash Equilibria
  • 5.3 Problem Decomposition
  • 6 Problem Domains and Applications
  • 6.1 Embodied Agents
  • 6.2 Game-Theoretic Environments
  • 6.3 Real-World Applications
  • 7 Resources
  • 8 Conclusions
  • 9 Acknowledgments
  • References

Knowls

  1. Knowl 1 — Taxonomy of Cooperative Multi-Agent Learning: Team Learning vs. Concurrent Learning

    model/method

    Cooperative multi-agent learning (CMAL) approaches divide fundamentally into two structural categories based on whether learning is centralized or decentralized:

    1. Team Learning: A single centralized machine learning process searches for a global joint behavior for the entire team of agents. Because a single learner evaluates the performance of the full collective, team learning avoids game-theoretic co-adaptation dynamics and bypasses the inter-agent credit assignment problem. However, the state-action search space grows exponentially with the number of agents NN (for NN agents with individual state space size ∣S∣|S|, the joint state space size is ∣S∣N|S|^N), and the central controller requires centralized data aggregation.

    2. Concurrent Learning: Multiple learning processes execute simultaneously, typically with one independent learner assigned per agent. This projects the exponential joint team search space into NN smaller, tractable individual search spaces. However, because agents adapt their behaviors concurrently, each learner faces a non-stationary environment where the actions and adaptations of other agents invalidate the assumptions underlying standard single-agent learning algorithms.

  2. Knowl 2 — Credit Assignment Mechanisms in Concurrent Multi-Agent Learning

    model/method

    In concurrent multi-agent learning, apportioning credit from joint team outcomes to individual agent learners requires choosing among several reward distribution strategies:

    • Global Reward: The team reward is shared equally or uniformly across all learners. While this guarantees alignment between individual incentives and team success, it creates a low signal-to-noise ratio and encourages laziness, as an individual agent's specific contribution is diluted across the group.
    • Local Reward: Each agent receives feedback based strictly on its own individual actions and task completions. This provides immediate, low-variance feedback and faster convergence, but fails to incentivize cooperative assistance, often producing selfish or greedy behaviors.
    • Social Reinforcement: Combines local reward with observational reinforcement (rewarding agents for imitating successful teammate behaviors) and vicarious reinforcement (broadcasting small portions of an agent's individual reward to nearby peers) to strike a balance between individual task execution and collective coordination.
    • Wonderful Life Utility (Difference Reward): Rewards each agent based on its marginal contribution to the overall collective performance. The utility UiU_i for agent ii is the difference between the actual global reward G(z)G(z) for the joint system trajectory zz and the counterfactual reward G(z−i)G(z_{-i}) achieved by the system without agent ii:

    Ui(z)=G(z)−G(z−i)U_i(z) = G(z) - G(z_{-i})

    • Reward Filtering: Models the observed reward as the sum of the agent's direct contribution and an external Markovian noise process representing teammates, applying a Kalman filter to isolate the agent's true contribution signal.
  3. Knowl 3 — Non-Stationarity and Co-adaptation Dynamics in Concurrent Learning

    theoretical result

    Concurrent learning violates the core stationarity assumption of traditional single-agent machine learning and Markov Decision Processes (MDPs). In a system with NN concurrent learners, the transition probability distribution P(s' ar s, a_i) perceived by a single agent ii depends dynamically on the joint policy π−i\boldsymbol{\pi}_{-i} of all other agents:

    P(s′∣s,ai)=∑a−iP(s′∣s,ai,a−i)∏j≠iπj(aj∣s)P(s' \mid s, a_i) = \sum_{\boldsymbol{a}_{-i}} P(s' \mid s, a_i, \boldsymbol{a}_{-i}) \prod_{j \neq i} \pi_j(a_j \mid s)

    As the other agents update their respective policies πj\pi_j over time, the state transition dynamics and reward landscapes for agent ii continuously change. Consequently, policies learned under earlier teammate configurations become obsolete, causing co-adaptation instability, divergence, or oscillatory behavior unless algorithms explicitly account for the adaptive dynamics of peers.

  4. Knowl 4 — Limitations of Nash Equilibrium and Strict Rationality in Cooperative Multi-Agent Learning

    theoretical result

    Applying game-theoretic Nash equilibrium concepts and strict individual rationality to cooperative multi-agent learning introduces significant failure modes:

    1. Suboptimal Equilibrium Traps: In cooperative matrix games with multiple equilibria (such as the climb or penalty games), standard rational learners frequently converge to risk-dominant, suboptimal Nash equilibria because any unilateral exploratory deviation towards the globally optimal equilibrium incurs severe immediate penalties when teammates have not yet coordinated.
    2. Equilibrium Selection Failures: When multiple optimal Nash equilibria exist, decentralized rational agents operating independently lack a coordination mechanism to select the same equilibrium, leading to cross-equilibrium action collisions and poor team outcomes.
    3. Optimism vs. Rationality: In cooperative domains where agents share a mutual objective, strict individual rationality (assuming worst-case or uncoordinated responses from peers) degrades performance. Learning rules that maintain optimistic heuristics about teammate collaboration—such as updating action values based on the maximum possible payoff across joint actions—consistently outperform standard rational best-response learners in converging to globally optimal team solutions.
  5. Knowl 5 — Agent Specialization Spectrum in Team Learning: Homogeneous, Heterogeneous, and Hybrid Teams

    model/method

    Candidate solution representations in centralized team learning fall along a spectrum defined by the degree of behavioral specialization:

    • Homogeneous Teams: A single shared behavior or policy is learned and executed identically by every agent in the collective. This dramatically restricts the dimensionality of the search space, making it scalable to large swarms. Homogeneous agents can still exhibit differentiated physical roles if the shared policy conditions actions on relative spatial positions or local sensory contexts.
    • Heterogeneous Teams: The centralized learner develops a distinct, unique policy for every agent on the team. This allows specialized division of labor (e.g., distinct player positions in robotic soccer) but increases the search space exponentially with team size and requires constrained variation operators (such as restricted breeding between corresponding agent genotypes) to maintain co-adapted sub-behaviors.
    • Hybrid Teams: The collective is partitioned into squads, where agents within each squad share an identical policy while distinct squads execute different behaviors. Squad boundaries and policy assignments can be pre-configured or discovered automatically using techniques such as Automatically Defined Groups (ADG).
  6. Knowl 6 — Problem Decomposition Techniques: Layered Learning, Shaping, and Coordination Graphs

    model/method

    To reduce the complexity of learning high-dimensional joint policies, multi-agent systems employ structured decomposition methods:

    • Layered Learning: A bottom-up, hierarchical decomposition where complex tasks are factored into sequential layers of sub-behaviors. Lower-level skills (such as ball acquisition or kicking in robot soccer) are learned first, fixed, and subsequently used as foundational primitives to train higher-level coordination behaviors (such as passing or offensive formations).
    • Reward Shaping and Fitness Switching: Reward shaping modifies the learning feedback by providing intermediate rewards for completing partial milestones toward a complex objective, reducing the time required to achieve coordinated team behaviors. Fitness switching dynamically shifts the reward function during training to prioritize sub-tasks on which the team is making the least progress.
    • Coordination Graphs: Exploits conditional independence among agents by decomposing the global joint action-value function into a sum of local payoff components over interacting subsets of agents:

    Q(s,a)=∑(i,j)∈EQij(sij,ai,aj)+∑i∈VQi(si,ai)Q(\boldsymbol{s}, \boldsymbol{a}) = \sum_{(i,j) \in \mathcal{E}} Q_{ij}(s_{ij}, a_i, a_j) + \sum_{i \in \mathcal{V}} Q_i(s_i, a_i)

    where V\mathcal{V} is the set of agents and E\mathcal{E} is the set of edges defining required agent-to-agent interactions, enabling exact joint action selection via local variable elimination.

  7. Knowl 7 — Hierarchy of Recursive Teammate Modeling

    model/method

    Teammate modeling enables an agent to predict the policies, beliefs, or action distributions of other agents to improve collaboration. Vidal and Durfee formalize this via a recursive depth hierarchy:

    • 0-Level Agent: Assumes that other agents are static, non-adaptive entities in the environment and does not attempt to model their learning or decision-making processes.
    • 1-Level Agent: Models its teammates as 0-level agents, assuming they follow fixed action probability distributions or simple reactive policies.
    • NN-Level Agent: Models other agents as (N−1)(N-1)-level agents, explicitly incorporating the belief that teammates are modeling others up to depth N−1N-1.

    This finite hierarchy terminates the infinite recursive reasoning loop (AA predicts that BB predicts that AA predicts...). While higher modeling levels can enhance coordination, agent performance is sensitive to initial prior beliefs; inaccurate models of teammate capabilities can prevent convergence to optimal cooperative behaviors.

  8. Knowl 8 — Direct and Indirect (Stigmergic) Communication in Multi-Agent Learning

    model/method

    Communication mechanisms in multi-agent learning operate across two distinct modalities:

    • Direct Communication: Explicit information exchange via discrete message passing, shared blackboards, or shared state-utility tables. Direct communication allows agents to share immediate sensor data, past trajectory episodes ⟨s,a,r,s′⟩\langle s, a, r, s' \rangle, or learned policy parameters. Alternatively, agents can learn emergent communication lexicons by evolving or reinforcing symbolic signals. Unconstrained direct communication expands the joint search space and can trivialize the multi-agent system into a centralized controller, requiring selective, bandwidth-constrained communication protocols.
    • Indirect Communication (Stigmergy): Implicit communication mediated entirely through environmental modifications. Common techniques include synthetic digital pheromone trails that diffuse and evaporate to guide foraging, routing, and exploration. Indirect communication also encompasses spatial embodiment, where agents use physical body positioning, alignment, and movement trajectories to signal leadership or role allocation without dedicated communication channels.
  9. Knowl 9 — Pathologies of Multi-Agent Co-adaptation: Loss of Gradient, Red Queen Effect, and Non-transitive Cycles

    definition

    Concurrent and co-evolutionary multi-agent learning algorithms are vulnerable to three key dynamical pathologies:

    • Loss of Gradient: Occurs when asymmetric learning rates cause one agent or sub-population to adapt substantially faster than its peers, dominating the interaction. The dominated learners receive invariant negative feedback with no learning gradient, while the dominant learner receives uniform success feedback and ceases improving, stalling the entire learning process.
    • Red Queen Effect: A condition in co-adapting systems where an agent's objective capability against an absolute benchmark remains static or degrades over time, despite continuous adaptation, because its peers are simultaneously adapting to counteract its changes (running continuously just to maintain relative fitness).
    • Non-transitive Cycles: Arises in environments containing cyclic dominance relations (analogous to Rock-Paper-Scissors dynamics where policy AA defeats BB, BB defeats CC, and CC defeats AA). Instead of progressing toward a globally optimal arms race or stable policy, co-adapting learners cycle indefinitely through the strategy space.
  10. Knowl 10 — Taxonomy of Benchmark Environments for Cooperative Multi-Agent Learning

    model/method

    Cooperative multi-agent learning benchmarks are structured into three main classes:

    1. Embodied Agent and Robotics Domains: Spatial environments requiring coordinated movement, physical interaction, and sensor processing:

      • Predator-Prey Pursuit: Multiple predator agents coordinate movement to encircle and capture an evasive prey.
      • Foraging and Clustering: Agents discover, collect, and transport items to central repositories, or cluster items by type.
      • Box Pushing & Cooperative Navigation: Multiple robots jointly exert force to move heavy obstacles or maintain geometric formations across obstacle fields.
      • Robotic Soccer (RoboCup) & Keep-Away: Highly dynamic domains requiring role specialization (offense/defense), ball passing, and real-time adversary resistance.
      • Cooperative Target Observation & Patrolling: Multi-robot teams coordinate coverage to keep moving targets within sensor ranges or continuously patrol regions.
    2. Game-Theoretic Matrix Domains: Mathematical abstractions isolating coordination and payoff structures:

      • Coordination Games: Matrix games with multiple equilibria and steep miscoordination penalties (e.g., climb and penalty games).
      • Social Dilemmas: Games modeling tensions between individual payoffs and collective welfare (e.g., Iterated Prisoner's Dilemma, Tragedy of the Commons, Braess's Paradox, Santa Fe Bar problem).
    3. Real-World Distributed Problem Solving: Large-scale distributed resource management including distributed traffic light control, air traffic sector management, packet routing in ad-hoc networks, power grid distribution, and manufacturing supply chain scheduling.

  11. Knowl 11 — Critical Dimensions for Scaling Up Multi-Agent Learning Systems

    limitation

    Scaling multi-agent learning from simplified benchmark scenarios to complex real-world applications requires addressing four major structural dimensions:

    1. Large Agent Populations: Expanding learning algorithms from two- or three-agent toy models to teams comprising dozens, hundreds, or thousands of agents, where centralized representations are computationally intractable.
    2. Structured Heterogeneity: Moving beyond the dichotomy of fully homogeneous swarms or completely heterogeneous agents to teams organized into structured classes of heterogeneous capabilities and roles.
    3. Internal and Partially Observable State: Scaling beyond stateless or fully observable MDPs to agents with rich internal hidden states and recurrent memory, overcoming the PSPACE/NEXP complexity inherent in decentralized partially observable Markov decision processes (Dec-POMDPs).
    4. Dynamic Teams and Non-Stationary Environments: Developing algorithms robust to open multi-agent environments where agents dynamically enter, fail, leave, or change task objectives during operation.

Coverage note — None was omitted; all major contributions including the core taxonomy (team vs. concurrent learning), credit assignment models, game-theoretic and co-adaptation dynamics, communication mechanisms, open research challenges, and domain taxonomies are fully covered.

References

  1. 1.D. H. Ackley and M. Littman. Altruism in the evolution of communication. In Artificial Life IV: Proceedings of the International Workshop on the Synthesis and Simulation of Living Systems, third edition. MIT Press, 1994.
  2. 2.D. Andre and A. Teller. Evolving team Darwin United. In M. Asada and H. Kitano, editors, RoboCup-98: Robot Soccer World Cup II. Springer Verlag, 1999.
  3. 3.D. Andre, F. Bennett III, and J. Koza. Discovery by genetic programming of a cellular automata rule that is better than any known rule for the majority classification problem. In Genetic Programming 1996: Proceedings of the First Annual Conference. MIT Press, 1996.
  4. 4.P. Angeline and J. Pollack. Competitive environments evolve better solutions for complex tasks. In S. Forrest, editor, Proceedings of the Fifth International Conference on Genetic Algorithms (ICGA), pages 264–270, San Mateo, CA, 1993. Morgan Kaufmann.
  5. 5.W. Arthur. Inductive reasoning and bounded rationality. Complexity in Economic Theory, 84(2):406–411, 1994.
  6. 6.T. Bäck. Evolutionary Algorithms in Theory and Practice: Evolutionary Straegies, Evolutionary Programming, and Genetic Algorithms. Oxford Press, 1996.
  7. 7.T. Balch. Learning roles: Behavioral diversity in robot teams. Technical Report GIT-CC-97-12, Georgia Institute of Technology, 1997.
  8. 8.T. Balch. Behavioral Diversity in Learning Robot Teams. PhD thesis, College of Computing, Georgia Institute of Technology, 1998.
  9. 9.T. Balch. Reward and diversity in multirobot foraging. In IJCAI-99 Workshop on Agents Learning About, From and With other Agents, 1999.
  10. 10.B. Banerjee, R. Mukherjee, and S. Sen. Learning mutual trust. In Working Notes of AGENTS-00 Workshop on Deception, Fraud and Trust in Agent Societies, pages 9–14, 2000.
  11. 11.A. Barto, R. Sutton, and C. Watkins. Learning and sequential decision making. In M. Gabriel and J. Moore, editors, Learning and computational neuroscience : foundations of adaptive networks. M.I.T. Press, Cambridge, Mass, 1990.
  12. 12.J. Bassett and K. De Jong. Evolving behaviors for cooperating agents. In Z. Ras, editor, Proceedings from the Twelfth International Symposium on Methodologies for Intelligent Systems, pages 157–165, Charlotte, NC., 2000. Springer-Verlag.
  13. 13.J. K. Bassett. A study of generalization techniques in evolutionary rule learning. Master’s thesis, George Mason University, Fairfax VA, USA, 2002.
  14. 14.R. Beckers, O. E. Holland, and J.-L. Deneubourg. From local actions to global tasks: Stigmergy and collective robotics. In Artificial Life IV: Proceedings of the International Workshop on the Synthesis and Simulation of Living Systems, third edition. MIT Press, 1994.
  15. 15.M. Benda, V. Jagannathan, and R. Dodhiawala. On optimal cooperation of knowledge sources — an empirical investigation. Technical Report BCS-G2010-28, Boeing Advanced Technology Center, Boeing Computer Services, 1986.
  16. 16.H. Berenji and D. Vengerov. Advantages of cooperation between reinforcement learning agents in difficult stochastic problems. In Proceedings of 9th IEEE International Conference on Fuzzy Systems, 2000.
  17. 17.H. Berenji and D. Vengerov. Learning, cooperation, and coordination in multi-agent systems. Technical Report IIS-00-10, Intelligent Inference Systems Corp., 333 W. Maude Avennue, Suite 107, Sunnyvale, CA 94085-4367, 2000.
  18. 18.D. Bernstein, S. Zilberstein, and N. Immerman. The complexity of decentralized control of MDPs. In Proceedings of UAI-2000: The Sixteenth Conference on Uncertainty in Artificial Intelligence, 2000.
  19. 19.H. J. Blumenthal and G. Parker. Co-evolving team capture strategies for dissimilar robots. In Proceedings of Artificial Multiagent Learning. Papers from the 2004 AAAI Fall Symposium. Technical Report FS-04-02, 2004.
  20. 20.E. Bonabeau, M. Dorigo, and G. Theraulaz. Swarm Intelligence: From Natural to Artificial Systems. SFI Studies in the Sciences of Complexity. Oxford University Press, 1999.
  21. 21.J. C. Bongard. The legion system: A novel approach to evolving heterogeneity for collective problem solving. In R. Poli, W. Banzhaf, W. B. Langdon, J. F. Miller, P. Nordin, and T. C. Fogarty, editors, Genetic Programming: Proceedings of EuroGP-2000, volume 1802, pages 16–28, Edinburgh, 15-16 2000. Springer-Verlag. ISBN 3-540-67339-3.
  22. 22.C. Boutilier. Learning conventions in multiagent stochastic domains using likelihood estimates. In Uncertainty in Artificial Intelligence, pages 106–114, 1996.
  23. 23.C. Boutilier. Planning, learning and coordination in multiagent decision processes. In Proceedings of the Sixth Conference on Theoretical Aspects of Rationality and Knowledge (TARK96), pages 195–210, 1996.
  24. 24.M. Bowling. Convergence problems of general-sum multiagent reinforcement learning. In Proceedings of the Seventeenth International Conference on Machine Learning, pages 89–94. Morgan Kaufmann, San Francisco, CA, 2000.
  25. 25.M. Bowling. Multiagent learning in the presence of agents with limitations. PhD thesis, Computer Science Department, Carnegie Mellon University, 2003.
  26. 26.M. Bowling and M. Veloso. An analysis of stochastic game theory for multiagent reinforcement learning. Technical Report CMU-CS-00-165, Computer Science Department, Carnegie Mellon University, 2000.
  27. 27.M. Bowling and M. Veloso. Rational and convergent learning in stochastic games. In Proceedings of Seventeenth International Joint Conference on Artificial Intelligence (IJCAI-01), pages 1021–1026, 2001.
  28. 28.M. Bowling and M. Veloso. Existence of multiagent equilibria with limited agents. Technical Report CMU-CS-02-104, Computer Science Department, Carnegie Mellon University, 2002.
  29. 29.M. Bowling and M. Veloso. Multiagent learning using a variable learning rate. Artificial Intelligence, 136(2):215–250, 2002.
  30. 30.J. A. Boyan and M. Littman. Packet routing in dynamically changing networks: A reinforcement learning approach. In J. D. Cowan, G. Tesauro, and J. Alspector, editors, Advances in Neural Information Processing Systems, volume 6, pages 671–678. Morgan Kaufmann Publishers, Inc., 1994.
  31. 31.R. Brafman and M. Tennenholtz. Efficient learning equilibrium. In Advances in Neural Information Processing Systems (NIPS-2002), 2002.
  32. 32.W. Brauer and G. Weiß. Multi-machine scheduling — a multi-agent learning approach. In Proceedings of the Third International Conference on Multi-Agent Systems, pages 42–48, 1998.
  33. 33.P. Brazdil, M. Gams, S. Sian, L. Torgo, and W. van de Velde. Learning in distributed systems and multi-agent environments. In Y. Kodratoff, editor, Lecture Notes in Artificial Intelligence, Volume 482, pages 412–423. Springer-Verlag, 1991.
  34. 34.O. Buffet, A. Dutech, and F. Charpillet. Incremental reinforcement learning for designing multi-agent systems. In J. P. Müller, E. André, S. Sen, and C. Frasson, editors, Proceedings of the Fifth International Conference on Autonomous Agents, pages 31–32, Montreal, Canada, 2001. ACM Press.
  35. 35.O. Buffet, A. Dutech, and F. Charpillet. Learning to weigh basic behaviors in scalable agents. In Proceedings of the 1st International Joint Conference on Autonomous Agents and MultiAgent Systems (AAMAS’02), 2002.
  36. 36.H. Bui, S. Venkatesh, and D. Kieronska. A framework for coordination and learning among team of agents. In W. Wobcke, M. Pagnucco, and C. Zhang, editors, Agents and Multi-Agent Systems: Formalisms, Methodologies and Applications, Lecture Notes in Artificial Intelligence, Volume 1441, pages 164–178. Springer-Verlag, 1998.
  37. 37.H. Bui, S. Venkatesh, and D. Kieronska. Learning other agents’ preferences in multi-agent negotiation using the Bayesian classifier. International Journal of Cooperative Information Systems, 8(4):275–294, 1999.
  38. 38.L. Bull. Evolutionary computing in multi-agent environments: Partners. In T. Bäck, editor, Proceedings of the Seventh International Conference on Genetic Algorithms, pages 370–377. Morgan Kaufmann, 1997.
  39. 39.L. Bull. Evolutionary computing in multi-agent environments: Operators. In D. W. V W Porto, N Saravanan and A. E. Eiben, editors, Proceedings of the Seventh Annual Conference on Evolutionary Programming, pages 43–52. Springer Verlag, 1998.
  40. 40.L. Bull and T. C. Fogarty. Evolving cooperative communicating classifier systems. In A. V. Sebald and L. J. Fogel, editors, Proceedings of the Fourth Annual Conference on Evolutionary Programming (EP94), pages 308–315, 1994.
  41. 41.L. Bull and O. Holland. Evolutionary computing in multiagent environments: Eusociality. In Proceedings of Seventh Annual Conference on Genetic Algorithms, 1997.
  42. 42.A. Cangelosi. Evolution of communication and language using signals, symbols, and words. IEEE Transactions on Evolutionary Computation, 5(2):93–101, 2001.
  43. 43.Y. U. Cao, A. S. Fukunaga, and A. B. Kahng. Cooperative mobile robotics: Antecedents and directions. Autonomous Robots, 4(1):7–23, 1997.
  44. 44.D. Carmel. Model-based Learning of Interaction Strategies in Multi-agent systems. PhD thesis, Technion — Israel Institute of Technology, 1997.
  45. 45.D. Carmel and S. Markovitch. The M* algorithm: Incorporating opponent models into adversary search. Technical Report 9402, Technion - Israel Institute of Technology, March 1994.
  46. 46.L.-E. Cederman. Emergent Actors in World Politics: How States and Nations Develop and Dissolve. Princeton University Press, 1997.
  47. 47.G. Chalkiadakis and C. Boutilier. Coordination in multiagent reinforcement learning: A Bayesian approach. In Proceedings of The Second International Joint Conference on Autonomous Agents & Multiagent Systems (AAMAS 2003). ACM, 2003. ISBN 1-58113-683-8.
  48. 48.H. Chalupsky, Y. Gil, C. A. Knoblock, K. Lerman, J. Oh, D. Pynadath, T. Russ, and M. Tambe. Electric elves: agent technology for supporting human organizations. In AI Magazine - Summer 2002. AAAI Press, 2002.
  49. 49.Y.-H. Chang, T. Ho, and L. Kaelbling. All learning is local: Multi-agent learning in global reward games. In Proceedings of Neural Information Processing Systems (NIPS-03), 2003.
  50. 50.Y.-H. Chang, T. Ho, and L. Kaelbling. Multi-agent learning in mobilized ad-hoc networks. In Proceedings of Artificial Multiagent Learning. Papers from the 2004 AAAI Fall Symposium. Technical Report FS-04-02, 2004.
  51. 51.C. Claus and C. Boutilier. The dynamics of reinforcement learning in cooperative multiagent systems. In Proceedings of National Conference on Artificial IntelligenceAAAI/IAAI, pages 746–752, 1998.
  52. 52.D. Cliff and G. F. Miller. Tracking the red queen: Measurements of adaptive progress in co-evolutionary simulations. In Proceedings of the Third European Conference on Artificial Life, pages 200–218. Springer-Verlag, 1995.
  53. 53.R. Collins and D. Jefferson. An artificial neural network representation for artificial organisms. In H.-P. Schwefel and R. Männer, editors, Parallel Problem Solving from Nature: 1st Workshop (PPSN I), pages 259–263, Berlin, 1991. Springer-Verlag.
  54. 54.R. Collins and D. Jefferson. AntFarm : Towards simulated evolution. In C. Langton, C. Taylor, J. D. Farmer, and S. Rasmussen, editors, Artificial Life II, pages 579–601. Addison-Wesley, Redwood City, CA, 1992.
  55. 55.E. Crawford and M. Veloso. Opportunities for learning in multi-agent meeting scheduling. In Proceedings of Artificial Multiagent Learning. Papers from the 2004 AAAI Fall Symposium. Technical Report FS-04-02, 2004.
  56. 56.V. Crespi, G. Cybenko, M. Santini, and D. Rus. Decentralized control for coordinated flow of multi-agent systems. Technical Report TR2002-414, Dartmouth College, Computer Science, Hanover, NH, January 2002.
  57. 57.R. H. Crites. Large-Scale Dynamic Optimization Using Teams of Reinforcement Learning Agents. PhD thesis, University of Massachusetts Amherst, 1996.
  58. 58.M. R. Cutkosky, R. S. Englemore, R. E. Fikes, M. R. Genesereth, T. R. Gruber, W. S. Mark, J. M. Tenenbaum, and J. C. Weber. PACT: An experiment in integrating concurrent engineering systems. In M. N. Huhns and M. P. Singh, editors, Readings in Agents, pages 46–55. Morgan Kaufmann, San Francisco, CA, USA, 1997.
  59. 59.T. Dahl, M. Mataric, and G. Sukhatme. Adaptive spatio-temporal organization in groups of robots. In Proceedings of the 2002 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS-02), 2002.
  60. 60.R. Das, M. Mitchell, and J. Crutchfield. A genetic algorithm discovers particle-based computation in cellular automata. In Parallel Problem Solving from Nature III, LNCS 866, pages 344–353. Springer-Verlag, 1994.
  61. 61.J. Davis and G. Kendall. An investigation, using co-evolution, to evolve an awari player. In Proceedings of 2002 Congress on Evolutionary Computation (CEC2002), 2002.
  62. 62.B. de Boer. Generating vowel systems in a population of agents. In Proceedings of the Fourth European Conference Artificial Life. MIT Press, 1997.
  63. 63.K. De Jong. Evolutionary Computation: A unified approach. MIT Press, 2005.
  64. 64.K. De Jong. An Analysis of the Behavior of a Class of Genetic Adaptive Systems. PhD thesis, University of Michigan, Ann Arbor, MI, 1975.
  65. 65.K. Decker, E. Durfee, and V. Lesser. Evaluating research in cooperative distributed problem solving. In L. Gasser and M. Huhns, editors, Distributed Artificial Intelligence Volume II, pages 487–519. Pitman Publishing and Morgan Kaufmann, 1989.
  66. 66.K. Decker, M. Fisher, M. Luck, M. Tennenholtz, and UKMAS’98 Contributors. Continuing research in multi-agent systems. Knowledge Engineering Review, 14(3):279–283, 1999.
  67. 67.J. L. Deneubourg, S. Goss, N. Franks, A. Sendova-Franks, C. Detrain, and L. Chretien. The dynamics of collective sorting: Robot-like ants and ant-like robots. In From Animals to Animats: Proceedings of the First International Conference on Simulation of Adaptive Behavior, pages 356–363. MIT Press, 1991.
  68. 68.J. Denzinger and M. Fuchs. Experiments in learning prototypical situations for variants of the pursuit game. In Proceedings on the International Conference on Multi-Agent Systems (ICMAS-1996), pages 48–55, 1996.
  69. 69.M. Dowell. Learning in Multiagent Systems. PhD thesis, University of South Carolina, 1995.
  70. 70.K. Dresner and P. Stone. Multiagent traffic management: A reservation-based intersection control mechanism. In AAMAS-2004 — Proceedings of the Third International Joint Conference on Autonomous Agents and Multi Agent Systems, 2004.
  71. 71.G. Dudek, M. Jenkin, R. Milios, and D. Wilkes. A taxonomy for swarm robots. In Proceedings of IEEE/RSJ Conference on Intelligent Robots and Systems, 1993.
  72. 72.E. Durfee. What your computer really needs to know, you learned in kindergarten. In National Conference on Artificial Intelligence, pages 858–864, 1992.
  73. 73.E. Durfee, V. Lesser, and D. Corkill. Coherent cooperation among communicating problem solvers. IEEE Transactions on Computers, C-36 (11):1275–1291, 1987.
  74. 74.E. Durfee, V. Lesser, and D. Corkill. Trends in cooperative distributed problem solving. IEEE Transactions on Knowledge and Data Engineering, KDE-1(1):63–83, March 1989.
  75. 75.A. Dutech, O. Buffet, and F. Charpillet. Multi-agent systems by incremental gradient reinforcement learning. In Proceedings of Seventeenth International Joint Conference on Artificial Intelligence (IJCAI-01), pages 833–838, 2001.
  76. 76.F. Fernandez and L. Parker. Learning in large cooperative multi-robot domains. International Journal of Robotics and Automation, 16(4): 217–226, 2001.
  77. 77.S. Ficici and J. Pollack. A game-theoretic approach to the simple coevolutionary algorithm. In Proceedings of the Sixth International Conference on Parallel Problem Solving from Nature (PPSN VI). Springer Verlag, 2000.
  78. 78.S. Ficici and J. Pollack. Challenges in coevolutionary learning: Arms-race dynamics, open-endedness, and mediocre stable states. In C. Adami et al, editor, Proceedings of the Sixth International Conference on Artificial Life, pages 238–247, Cambridge, MA, 1998. MIT Press.
  79. 79.K. Fischer, N. Kuhn, H. J. Müller, J. P. Müller, and M. Pischel. Sophisticated and distributed: The transportation domain. In Proceedings of the Fifth European Workshop on Modelling Autonomous Agents in a Multi-Agent World (MAAMAW’93), 1993.
  80. 80.D. Fogel. Blondie24: Playing at the Edge of Artificial Intelligence. Morgan Kaufmann, 2001. ISBN 1-55860-783-8.
  81. 81.L. Fogel. Intelligence Through Simulated Evolution: Forty Years of Evolutionary Programming. Wiley Series on Intelligent Systems, 1999.
  82. 82.D. Fudenberg and D. Levine. The Theory of Learning in Games. MIT Press, 1998.
  83. 83.A. Garland and R. Alterman. Autonomous agents that learn to better coordinate. Autonomous Agents and Multi-Agent Systems, 8:267–301, 2004.
  84. 84.M. Ghavamzadeh and S. Mahadevan. Learning to communicate and act using hierarchical reinforcement learning. In AAMAS-2004 — Proceedings of the Third International Joint Conference on Autonomous Agents and Multi Agent Systems, 2004.
  85. 85.N. Glance and B. Huberman. The dynamics of social dilemmas. Scientific American, 270(3):76–81, March 1994.
  86. 86.P. Gmytrasiewicz. A Decision-Theoretic Model of Coordination and Communication in Autonomous Systems (Reasoning Systems). PhD thesis, University of Michigan, 1992.
  87. 87.D. E. Goldberg. Genetic Algorithms in Search, Optimization, and Machine Learning. Addison Wesley, Reading, MA, 1989.
  88. 88.C. Goldman and J. Rosenschein. Mutually supervised learning in multiagent systems. In G. Weiß and S. Sen, editors, Adaptation and Learning in Multi-Agent Systems, pages 85–96. Springer-Verlag: Heidelberg, Germany, Berlin, 1996.
  89. 89.B. M. Good. Evolving multi-agent systems: Comparing existing approaches and suggesting new directions. Master’s thesis, University of Sussex, 2000.
  90. 90.M. Gordin, S. Sen, and N. Puppala. Evolving cooperative groups: Preliminary results. In Working Papers of the AAAI-97 Workshop on Multiagent Learning, pages 31–35, 1997.
  91. 91.S. Grand and D. Cliff. Creatures: Entertainment software agents with artificial life. Autonomous Agents and Multi-Agent Systems, 1(1): 39–57, 1998.
  92. 92.S. Grand, D. Cliff, and A. Malhotra. Creatures : Artificial life autonomous software agents for home entertainment. In Proceedings of the First International Conference on Autonomous Agents (Agents-97), pages 22–29, 1997.
  93. 93.D. L. Grecu. Using Learning to Improve Multi-Agent Systems for Design. PhD thesis, Worcester Polytechnic Institute, 1997.
  94. 94.A. Greenwald and K. Hall. Correlated Q-learning. In Proceedings of the Twentieth International Conference on Machine Learning, 2003.
  95. 95.A. Greenwald, J. Farago, and K. Hall. Fair and efficient solutions to the Santa Fe bar problem. In Proceedings of the Grace Hopper Celebration of Women in Computing 2002, 2002.
  96. 96.J. Grefenstette. Lamarckian learning in multi-agent environments. In R. Belew and L. Booker, editors, Proceedings of the Fourth International Conference on Genetic Algorithms, pages 303–310, San Mateo, CA, 1991. Morgan Kaufman.
  97. 97.J. Grefenstette, C. L. Ramsey, and A. Schultz. Learning sequential decision rules using simulation models and competition. Machine Learning, 5:355–381, 1990.
  98. 98.C. Guestrin, M. Lagoudakis, and R. Parr. Coordinated reinforcement learning. In Proceedings of the 2002 AAAI Symposium Series: Collaborative Learning Agents, 2002.
  99. 99.S. M. Gustafson. Layered learning in genetic programming for a co-operative robot soccer problem. Master’s thesis, Kansas State University, Manhattan, KS, USA, 2000.
  100. 100.S. M. Gustafson and W. H. Hsu. Layered learning in genetic programming for a co-operative robot soccer problem. In J. F. Miller, M. Tomassini, P. L. Lanzi, C. Ryan, A. G. B. Tettamanzi, and W. B. Langdon, editors, Genetic Programming: Proceedings of EuroGP-2001, volume 2038, pages 291–301, Lake Como, Italy, 18-20 2001. Springer-Verlag. ISBN 3-540-41899-7.
  101. 101.A. Hara and T. Nagao. Emergence of cooperative behavior using ADG; Automatically Defined Groups. In Proceedings of the 1999 Genetic and Evolutionary Computation Conference (GECCO-99), pages 1038–1046, 1999.
  102. 102.I. Harvey, P. Husbands, D. Cliff, A. Thompson, and N. Jakobi. Evolutionary robotics : the Sussex approach. Robotics and Autonomous Systems, 20(2–4):205–224, June 1997.
  103. 103.T. Haynes and S. Sen. Evolving behavioral strategies in predators and prey. In G. Weiß and S. Sen, editors, Adaptation and Learning in Multiagent Systems, Lecture Notes in Artificial Intelligence. Springer Verlag, Berlin, Germany, 1995.
  104. 104.T. Haynes and S. Sen. Adaptation using cases in cooperative groups. In I. Imam, editor, Working Notes of the AAAI-96 Workshop on Intelligent Adaptive Agents, Portland, OR, 1996.
  105. 105.T. Haynes and S. Sen. Cooperation of the fittest. Technical Report UTULSA-MCS-96-09, The University of Tulsa, Apr. 12, 1996.
  106. 106.T. Haynes and S. Sen. Learning cases to resolve conflicts and improve group behavior. In M. Tambe and P. Gmytrasiewicz, editors, Working Notes of the AAAI-96 Workshop on Agent Modeling, pages 46–52, Portland, OR, 1996.
  107. 107.T. Haynes and S. Sen. Crossover operators for evolving a team. In J. R. Koza, K. Deb, M. Dorigo, D. B. Fogel, M. Garzon, H. Iba, and R. L. Riolo, editors, Genetic Programming 1997: Proceedings of the Second Annual Conference, pages 162–167, Stanford University, CA, USA, 13-16 July 1997. Morgan Kaufmann.
  108. 108.T. Haynes, S. Sen, D. Schoenefeld, and R. Wainwright. Evolving multiagent coordination strategies with genetic programming. Technical Report UTULSA-MCS-95-04, The University of Tulsa, May 31, 1995.
  109. 109.T. Haynes, S. Sen, D. Schoenefeld, and R. Wainwright. Evolving a team. In E. V. Siegel and J. R. Koza, editors, Working Notes for the AAAI Symposium on Genetic Programming, pages 23–30, MIT, Cambridge, MA, USA, 10–12 Nov. 1995. AAAI.
  110. 110.T. Haynes, R. Wainwright, S. Sen, and D. Schoenefeld. Strongly typed genetic programming in evolving cooperation strategies. In L. Eshelman, editor, Genetic Algorithms: Proceedings of the Sixth International Conference (ICGA95), pages 271–278, Pittsburgh, PA, USA, 15-19 July 1995. Morgan Kaufmann. ISBN 1-55860-370-0.
  111. 111.T. Haynes, K. Lau, and S. Sen. Learning cases to compliment rules for conflict resolution in multiagent systems. In S. Sen, editor, AAAI Spring Symposium on Adaptation, Coevolution, and Learning in Multiagent Systems, pages 51–56, 1996.
  112. 112.T. D. Haynes and S. Sen. Co-adaptation in a team. International Journal of Computational Intelligence and Organizations (IJCIO), 1(4), 1997.
  113. 113.D. Hillis. Co-evolving parasites improve simulated evolution as an optimization procedure. Artificial Life II, SFI Studies in the Sciences of Complexity, 10:313–324, 1991.
  114. 114.J. Holland. Adaptation in Natural and Artificial Systems. The MIT Press, Cambridge, MA, 1975.
  115. 115.J. Holland. Properties of the bucket brigade. In Proceedings of an International Conference on Genetic Algorithms, 1985.
  116. 116.B. Hölldobler and E. O. Wilson. The Ants. Harvard University Press, 1990.
  117. 117.W. H. Hsu and S. M. Gustafson. Genetic programming and multi-agent layered learning by reinforcements. In W. B. Langdon, E. Cantú-Paz, K. Mathias, R. Roy, D. Davis, R. Poli, K. Balakrishnan, V. Honavar, G. Rudolph, J. Wegener, L. Bull, M. Potter, A. C. Schultz, J. F. Miller, E. Burke, and N. Jonoska, editors, GECCO 2002: Proceedings of the Genetic and Evolutionary Computation Conference, pages 764–771, New York, 9-13 July 2002. Morgan Kaufmann Publishers. ISBN 1-55860-878-8.
  118. 118.J. Hu and M. Wellman. Nash Q-learning for general-sum stochastic games. Journal of Machine Learning Research, 4:1039–1069, 2003.
  119. 119.J. Hu and M. Wellman. Self-fulfilling bias in multiagent learning. In Proceedings of the Second International Conference on Multi-Agent Systems, 1996.
  120. 120.J. Hu and M. Wellman. Multiagent reinforcement learning: theoretical framework and an algorithm. In Proceedings of the Fifteenth International Conference on Machine Learning, pages 242–250. Morgan Kaufmann, San Francisco, CA, 1998.
  121. 121.J. Hu and M. Wellman. Online learning about other agents in a dynamic multiagent system. In K. P. Sycara and M. Wooldridge, editors, Proceedings of the Second International Conference on Autonomous Agents (Agents’98), pages 239–246, New York, 1998. ACM Press. ISBN 0-89791-983-1.
  122. 122.J. Huang, N. R. Jennings, and J. Fox. An agent architecture for distributed medical care. In M. Wooldridge and N. R. Jennings, editors, Intelligent Agents: Theories, Architectures, and Languages (LNAI Volume 890), pages 219–232. Springer-Verlag: Heidelberg, Germany, 1995.
  123. 123.M. Huhns and M. Singh. Agents and multiagent systems: themes, approaches and challenges. In M. Huhns and M. Singh, editors, Readings in Agents, pages 1–23. Morgan Kaufmann, 1998.
  124. 124.M. Huhns and G. Weiß. Special issue on multiagent learning. Machine Learning Journal, 33(2-3), 1998.
  125. 125.H. Iba. Emergent cooperation for multiple agents using genetic programming. In H.-M. Voigt, W. Ebeling, I. Rechenberg, and H.-P. Schwefel, editors, Parallel Problem Solving from Nature IV: Proceedings of the International Conference on Evolutionary Computation, volume 1141 of LNCS, pages 32–41, Berlin, Germany, 1996. Springer Verlag. ISBN 3-540-61723-X.
  126. 126.H. Iba. Evolutionary learning of communicating agents. Information Sciences, 108(1–4):181–205, 1998.
  127. 127.H. Iba. Evolving multiple agents by genetic programming. In L. Spector, W. Langdon, U.-M. O’Reilly, and P. Angeline, editors, Advances in Genetic Programming 3, pages 447–466. The MIT Press, Cambridge, MA, 1999.
  128. 128.I. Imam, editor. Intelligent Adaptive Agents. Papers from the 1996 AAAI Workshop. Technical Report WS-96-04. AAAI Press, 1996.
  129. 129.A. Ito. How do selfish agents learn to cooperate? In Artificial Life V: Proceedings of the Fifth International Workshop on the Synthesis and Simulation of Living Systems, pages 185–192. MIT Press, 1997.
  130. 130.T. Jansen and R. P. Wiegand. Exploring the explorative advantage of the cooperative coevolutionary (1+1) EA. In E. Cantu-Paz et al, editor, Prooceedings of the Genetic and Evolutionary Computation Conference (GECCO). Springer-Verlag, 2003.
  131. 131.N. Jennings, L. Varga, R. Aarnts, J. Fuchs, and P. Skarek. Transforming standalone expert systems into a community of cooperating agents. International Journal of Engineering Applications of Artificial Intelligence, 6(4):317–331, 1993.
  132. 132.N. Jennings, K. Sycara, and M. Wooldridge. A roadmap of agents research and development. Autonomous Agents and Multi-Agent Systems, 1:7–38, 1998.
  133. 133.K.-C. Jim and C. L. Giles. Talking helps: Evolving communicating agents for the predator-prey pursuit problem. Artificial Life, 6(3): 237–254, 2000.
  134. 134.H. Juille and J. Pollack. Coevolving the “ideal” trainer: Application to the discovery of cellular automata rules. In Proceedings of the Third Annual Genetic Programming Conference (GP-98), 1998.
  135. 135.L. Kaelbling, M. Littman, and A. Moore. Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4:237–285, 1996.
  136. 136.S. Kapetanakis and D. Kudenko. Improving on the reinforcement learning of coordination in cooperative multi-agent systems. In Proceedings of the Second Symposium on Adaptive Agents and Multi-agent Systems (AISB02), 2002.
  137. 137.S. Kapetanakis and D. Kudenko. Reinforcement learning of coordination in cooperative multi-agent systems. In Proceedings of the Nineteenth National Conference on Artificial Intelligence (AAAI02), 2002.
  138. 138.G. Kendall and G. Whitwell. An evolutionary approach for the tuning of a chess evaluation function using population dynamics. In Proceedings of the 2001 Congress on Evolutionary Computation (CEC-2001), pages 995–1002. IEEE Press, 27-30 2001.
  139. 139.G. Kendall and M. Willdig. An investigation of an adaptive poker player. In Proceedings of the 14th Australian Joint Conference on Artificial Intelligence (AI’01), 2001.
  140. 140.H. Kitano, M. Asada, Y. Kuniyoshi, I. Noda, and E. Osawa. RoboCup: The robot world cup initiative. In W. L. Johnson and B. Hayes-Roth, editors, Proceedings of the First International Conference on Autonomous Agents (Agents’97), pages 340–347, New York, 5–8, 1997. ACM Press. ISBN 0-89791-877-0.
  141. 141.J. Koza. Genetic Programming: on the Programming of Computers by Means of Natural Selection. MIT Press, 1992.
  142. 142.M. Lauer and M. Riedmiller. An algorithm for distributed reinforcement learning in cooperative multi-agent systems. In Proceedings of the Seventeenth International Conference on Machine Learning, pages 535–542. Morgan Kaufmann, San Francisco, CA, 2000.
  143. 143.L. R. Leerink, S. R. Schultz, and M. A. Jabri. A reinforcement learning exploration strategy based on ant foraging mechanisms. In Proceedings of the Sixth Australian Conference on Neural Networks, Sydney, Australia, 1995.
  144. 144.V. Lesser. Cooperative multiagent systems : A personal view of the state of the art. IEEE Trans. on Knowledge and Data Engineering, 11 (1):133–142, 1999.
  145. 145.V. Lesser, D. Corkill, and E. Durfee. An update on the distributed vehicle monitoring testbed. Technical Report UM-CS-1987-111, University of Massachessets Amherst, 1987.
  146. 146.M. I. Lichbach. The cooperator’s dilemma. University of Michigan Press, 1996. ISBN 0472105728.
  147. 147.M. Littman. Friend-or-foe Q-learning in general-sum games. In Proceedings of the Eighteenth International Conference on Machine Learning, pages 322–328. Morgan Kaufmann Publishers Inc., 2001.
  148. 148.M. Littman. Markov games as a framework for multi-agent reinforcement learning. In Proceedings of the 11th International Conference on Machine Learning (ML-94), pages 157–163, New Brunswick, NJ, 1994. Morgan Kaufmann.
  149. 149.A. Lubberts and R. Miikkulainen. Co-evolving a go-playing neural network. In Coevolution: Turning Adaptive Algorithms upon Themselves, (Birds-on-a-Feather Workshop, Genetic and Evolutionary Computation Conference), 2001.
  150. 150.M. Luck, M. d’Inverno, M. Fisher, and FoMAS’97 Contributors. Foundations of multi-agent systems: Techniques, tools and theory. Knowledge Engineering Review, 13(3):297–302, 1998.
  151. 151.S. Luke. Genetic programming produced competitive soccer softbot teams for RoboCup97. In J. R. Koza et al, editor, Genetic Programming 1998: Proceedings of the Third Annual Conference, pages 214–222. Morgan Kaufmann, 1998.
  152. 152.S. Luke and L. Spector. Evolving teamwork and coordination with genetic programming. In J. R. Koza, D. E. Goldberg, D. B. Fogel, and R. L. Riolo, editors, Genetic Programming 1996: Proceedings of the First Annual Conference, pages 150–156, Stanford University, CA, USA, 28–31 1996. MIT Press.
  153. 153.S. Luke and R. P. Wiegand. Guaranteeing coevolutionary objective measures. In Poli et al. [201], pages 237–251.
  154. 154.S. Luke, C. Hohn, J. Farris, G. Jackson, and J. Hendler. Co-evolving soccer softbot team coordination with genetic programming. In Proceedings of the First International Workshop on RoboCup, at the International Joint Conference on Artificial Intelligence, Nagoya, Japan, 1997.
  155. 155.S. Luke, K. Sullivan, G. C. Balan, and L. Panait. Tunably decentralized algorithms for cooperative target observation. Technical Report GMU-CS-TR-2004-1, Department of Computer Science, George Mason University, 2004.
  156. 156.S. Mahadevan and J. Connell. Automatic programming of behavior-based robots using reinforcement learning. In National Conference on Artificial Intelligence, pages 768–773, 1991.
  157. 157.R. Makar, S. Mahadevan, and M. Ghavamzadeh. Hierarchical multi-agent reinforcement learning. In J. P. Müller, E. André, S. Sen, and C. Frasson, editors, Proceedings of the Fifth International Conference on Autonomous Agents, pages 246–253, Montreal, Canada, 2001. ACM Press.
  158. 158.M. Mataric. Interaction and Intelligent Behavior. PhD thesis, Massachusetts Institute of Technology, Cambridge, MA, 1994. Also Technical Report AITR-1495.
  159. 159.M. Mataric. Learning to behave socially. In Third International Conference on Simulation of Adaptive Behavior, 1994.
  160. 160.M. Mataric. Reward functions for accelerated learning. In International Conference on Machine Learning, pages 181–189, 1994.
  161. 161.M. Mataric. Reinforcement learning in the multi-robot domain. Autonomous Robots, 4(1):73–83, 1997.
  162. 162.M. Mataric. Using communication to reduce locality in distributed multi-agent learning. Joint Special Issue on Learning in Autonomous Robots, Machine Learning, 31(1-3), 141-167, and Autonomous Robots, 5(3-4), Jul/Aug 1998, 335-354, 1998.
  163. 163.M. Mataric, M. Nilsson, and K. Simsarian. Cooperative multi-robot box-pushing. In Proceedings of IEEE/RSJ Conference on Intelligent Robots and Systems, pages 556–561, 1995.
  164. 164.Z. Michalewicz. Genetic Algorithms + Data Structures = Evolution Programs. Springer-Verlag, Berlin, 3rd edition, 1996.
  165. 165.T. Miconi. A collective genetic algorithm. In E. Cantu-Paz et al, editor, Proceedings of the Genetic and Evolutionary Computation Conference (GECCO), pages 876–883, 2001.
  166. 166.T. Miconi. When evolving populations is better than coevolving individuals: The blind mice problem. In Proceedings of the Eighteenth International Joint Conference on Artificial Intelligence (IJCAI-03), 2003.
  167. 167.M. Mitchell, J. Crutchfield, and R. Das. Evolving cellular automata with genetic algorithms: A review of recent work. In Proceedings of the First International Conference on Evolutionary Computation and its Applications (EvCA’96),, 1996.
  168. 168.N. Monekosso, P. Remagnino, and A. Szarowicz. An improved Q-learning algorithm using synthetic pheromones. In E. N. B. Dunin-Keplicz, editor, From Theory to Practice in Multi-Agent Systems, Second International Workshop of Central and Eastern Europe on Multi-Agent Systems, CEEMAS 2001 Cracow, Poland, September 26-29, 2001. Revised Papers, Lecture Notes in Artificial Intelligence LNAI-2296. Springer-Verlag, 2002.
  169. 169.N. D. Monekosso and P. Remagnino. Phe-Q: A pheromone based Q-learning. In Australian Joint Conference on Artificial Intelligence, pages 345–355, 2001.
  170. 170.N. D. Monekosso and P. Remagnino. An analysis of the pheromone Q-learning algorithm. In Proceedings of the VIII Iberoamerican Conference on Artificial Intelligence IBERAMIA-02, pages 224–232, 2002.
  171. 171.N. D. Monekosso, P. Remagnino, and A. Szarowicz. An improved Q-learning algorithm using synthetic pheromones. In Proceedings of the Second Workshop of Central and Eastern Europe on Multi-Agent Systems CEEMAS-01, pages 197–206, 2001.
  172. 172.J. Moody, Y. Liu, M. Saffell, and K. Youn. Stochastic direct reinforcement: Application to simple games with recurrence. In Proceedings of Artificial Multiagent Learning. Papers from the 2004 AAAI Fall Symposium. Technical Report FS-04-02, 2004.
  173. 173.R. Mukherjee and S. Sen. Towards a pareto-optimal solution in general-sum games. In Agents-2001 Workshop on Learning Agents, 2001.
  174. 174.U. Mukhopadjyay, L. Stephens, and M. Huhns. An intelligent system for document retrieval in distributed office environments. Journal of the American Society for Information Science, 37(3):123–135, May 1986.
  175. 175.J. Muller and M. Pischel. An architecture for dynamically interacting agents. Journal of Intelligent and Cooperative Information Systems, 3 (1):25–45, 1994.
  176. 176.M. Mundhe and S. Sen. Evaluating concurrent reinforcement learners. In Proceedings of the International Conference on Multiagent System, 2000.
  177. 177.M. Mundhe and S. Sen. Evolving agent societies that avoid social dilemmas. In D. Whitley, D. Goldberg, E. Cantu-Paz, L. Spector, I. Parmee, and H.-G. Beyer, editors, Proceedings of the Genetic and Evolutionary Computation Conference (GECCO-2000), pages 809–816, Las Vegas, Nevada, USA, 10-12 2000. Morgan Kaufmann. ISBN 1-55860-708-0.
  178. 178.Y. Nagayuki, S. Ishii, and K. Doya. Multi-agent reinforcement learning: An approach based on the other agent’s internal model. In Proceedings of the International Conference on Multi-Agent Systems (ICMAS-00), 2000.
  179. 179.M. V. Nagendra-Prasad. Learning Situation-Specific Control in Multi-Agent Systems. PhD thesis, University of Massachusetts Amherst, 1997.
  180. 180.R. Nair, D. Pynadath, M. Yokoo, M. Tambe, and S. Marsella. Taming decentralized POMDPs: Towards efficient policy computation for multiagent settings. In Proceedings of the Eighteenth International Joint Conference on Artificial Intelligence (IJCAI-03), 2003.
  181. 181.M. Nowak and K. Sigmund. Evolution of indirect reciprocity by image scoring / the dynamics of indirect reciprocity. Nature, 393:573–577, 1998.
  182. 182.A. Nowe, K. Verbeeck, and T. Lenaerts. Learning agents in a homo egualis society. Technical report, Computational Modeling Lab - VUB, March 2001.
  183. 183.L. Nunes and E. Oliveira. Learning from multiple sources. In AAMAS-2004 — Proceedings of the Third International Joint Conference on Autonomous Agents and Multi Agent Systems, 2004.
  184. 184.T. Ohko, K. Hiraki, and Y. Arzai. Addressee learning and message interception for communication load reduction in multiple robots environments. In G. Weiß, editor, Distributed Artificial Intelligence Meets Machine Learning : Learning in Multi-Agent Environments, Lecture Notes in Artificial Intelligence 1221. Springer-Verlag, 1997.
  185. 185.E. Ostergaard, G. Sukhatme, and M. Mataric. Emergent bucket brigading - a simple mechanism for improving performance in multi-robot constrainedspace foraging tasks. In Proceedings of the Fifth International Conference on Autonomous Agents,, 2001.
  186. 186.L. Pagie and M. Mitchell. A comparison of evolutionary and coevolutionary search. In R. K. Belew and H. Juille, editors, Coevolution: Turning Adaptive Algorithms upon Themselves, pages 20–25, San Francisco, California, USA, 7 2001.
  187. 187.L. Panait and S. Luke. A pheromone-based utility model for collaborative foraging. In AAMAS-2004 — Proceedings of the Third International Joint Conference on Autonomous Agents and Multi Agent Systems, 2004.
  188. 188.L. Panait and S. Luke. Ant foraging revisited. In Proceedings of the Ninth International Conference on the Simulation and Synthesis of Living Systems (ALIFE9), 2004.
  189. 189.L. Panait and S. Luke. Learning ant foraging behaviors. In Proceedings of the Ninth International Conference on the Simulation and Synthesis of Living Systems (ALIFE9), 2004.
  190. 190.L. Panait, R. P. Wiegand, and S. Luke. A sensitivity analysis of a cooperative coevolutionary algorithm biased for optimization. In Genetic and Evolutionary Computation Conference — GECCO-2004. Springer, 2004.
  191. 191.L. Panait, R. P. Wiegand, and S. Luke. A visual demonstration of convergence properties of cooperative coevolution. In Parallel Problem Solving from Nature — PPSN-2004. Springer, 2004.
  192. 192.L. A. Panait, R. P. Wiegand, and S. Luke. Improving coevolutionary search for optimal multiagent behaviors. In Proceedings of the Eighteenth International Joint Conference on Artificial Intelligence (IJCAI-03), 2003.
  193. 193.C. Papadimitriou and J. Tsitsiklis. Complexity of markov decision processes. Mathematics of Operations Research, 12(3):441–450, 1987.
  194. 194.L. Parker. Current state of the art in distributed autonomous mobile robotics. In L. Parker, G. Bekey, and J. Barhen, editors, Distributed Autonomous Robotic Systems 4,, pages 3–12. Springer-Verlag, 2000.
  195. 195.L. Parker. Multi-robot learning in a cooperative observation task. In Proceedings of Fifth International Symposium on Distributed Autonomous Robotic Systems (DARS 2000), 2000.
  196. 196.L. Parker. Distributed algorithms for multi-robot observation of multiple moving targets. Autonomous Robots, 12(3), 2002.
  197. 197.L. Parker, C. Touzet, and F. Fernandez. Techniques for learning in multi-robot teams. In T. Balch and L. Parker, editors, Robot Teams: From Diversity to Polymorphism. AK Peters, 2001.
  198. 198.M. Peceny, G. Weiß, and W. Brauer. Verteiltes maschinelles lernen in fertigungsumgebungen. Technical Report FKI-218-96, Institut fur Informatik, Technische Universitat Munchen, 1996.
  199. 199.M. Peeters, K. Verbeeck, and A. Nowe. Multi-agent learning in conflicting multi-level games with incomplete information. In Proceedings of Artificial Multiagent Learning. Papers from the 2004 AAAI Fall Symposium. Technical Report FS-04-02, 2004.
  200. 200.L. Peshkin, K.-E. Kim, N. Meuleau, and L. Kaelbling. Learning to cooperate via policy search. In Sixteenth Conference on Uncertainty in Artificial Intelligence, pages 307–314. Morgan Kaufmann, 2000.
  201. 201.R. Poli, J. Rowe, and K. D. Jong, editors. Foundations of Genetic Algorithms (FOGA) VII, 2002. Morgan Kaufmann.
  202. 202.J. Pollack and A. Blair. Coevolution in the successful learning of backgammon strategy. Machine Learning, 32(3):225–240, 1998.
  203. 203.J. Pollack, A. Blair, and M. Land. Coevolution of a backgammon player. In C. G. Langton and K. Shimohara, editors, Artificial Life V: Proc. of the Fifth Int. Workshop on the Synthesis and Simulation of Living Systems, pages 92–98, Cambridge, MA, 1997. The MIT Press.
  204. 204.E. Popovici and K. DeJong. Understanding competitive co-evolutionary dynamics via fitness landscapes. In Artificial Multiagent Symposium. Part of the 2004 AAAI Fall Symposium on Artificial Intelligence, 2004.
  205. 205.M. Potter. The Design and Analysis of a Computational Model of Cooperative Coevolution. PhD thesis, George Mason University, Fairfax, Virginia, 1997.
  206. 206.M. Potter and K. De Jong. Cooperative coevolution: An architecture for evolving coadapted subcomponents. Evolutionary Computation, 8 (1):1–29, 2000.
  207. 207.M. Potter and K. De Jong. A cooperative coevolutionary approach to function optimization. In Y. Davidor and H.-P. Schwefel, editors, Proceedings of the Third International Conference on Parallel Problem Solving from Nature (PPSN III), pages 249–257. Springer-Verlag, 1994.
  208. 208.M. Potter, K. De Jong, and J. J. Grefenstette. A coevolutionary approach to learning sequential decision rules. In Proceedings from the Sixth International Conference on Genetic Algorithms, pages 366–372. Morgan Kaufmann Publishers, Inc., 1995.
  209. 209.M. Potter, L. Meeden, and A. Schultz. Heterogeneity in the coevolved behaviors of mobile robots: The emergence of specialists. In Proceedings of The Seventeenth International Conference on Artificial Intelligence (IJCAI-2001), 2001.
  210. 210.N. Puppala, S. Sen, and M. Gordin. Shared memory based cooperative coevolution. In Proceedings of the 1998 IEEE World Congress on Computational Intelligence, pages 570–574, Anchorage, Alaska, USA, 1998. IEEE Press.
  211. 211.M. Quinn. A comparison of approaches to the evolution of homogeneous multi-robot teams. In Proceedings of the 2001 Congress on Evolutionary Computation (CEC2001), pages 128–135, COEX, World Trade Center, 159 Samseong-dong, Gangnam-gu, Seoul, Korea, 27-30 2001. IEEE Press. ISBN 0-7803-6658-1.
  212. 212.M. Quinn. Evolving communication without dedicated communication channels. In Advances in Artificial Life: Sixth European Conference on Artificial Life (ECAL01), 2001.
  213. 213.M. Quinn, L. Smith, G. Mayley, and P. Husbands. Evolving formation movement for a homogeneous multi-robot system: Teamwork and role-allocation with real robots. Cognitive Science Research Paper 515. School of Cognitive and Computing Sciences, University of Sussex, Brighton, BN1 9QG. ISSN 1350-3162, 2002.
  214. 214.C. Reynolds. Competition, coevolution and the game of tag. In R. A. Brooks and P. Maes, editors, Artificial Life IV, Proceedings of the Fourth International Workshop on the Synthesis and Simulation of Living Systems., pages 59–69. MIT Press, 1994.
  215. 215.C. Reynolds. An evolved, vision-based behavioral model of coordinated group motion. In From Animals to Animats 2: Proceedings of the Second International Conference on Simulation of Adaptive Behavior (SAB92), pages 384–392, 1993.
  216. 216.C. W. Reynolds. Flocks, herds, and schools: a distributed behavioral model. Computer Graphics, 21(4):25–34, 1987.
  217. 217.P. Riley and M. Veloso. On behavior classification in adversarial environments. In L. Parker, G. Bekey, and J. Barhen, editors, Distributed Autonomous Robotic Systems 4, pages 371–380. Springer-Verlag, 2000.
  218. 218.A. Robinson and L. Spector. Using genetic programming with multiple data types and automatic modularization to evolve decentralized and coordinated navigation in multi-agent systems. In In Late-Breaking Papers of the Genetic and Evolutionary Computation Conference (GECCO-2002). The International Society for Genetic and Evolutionary Computation, 2002.
  219. 219.C. Rosin and R. Belew. New methods for competitive coevolution. Evolutionary Computation, 5(1):1–29, 1997.
  220. 220.R. Salustowicz, M. Wiering, and J. Schmidhuber. Learning team strategies with multiple policy-sharing agents: A soccer case study. Technical report, ISDIA, Corso Elvezia 36, 6900 Lugano, Switzerland, 1997.
  221. 221.R. Salustowicz, M. Wiering, and J. Schmidhuber. Learning team strategies: Soccer case studies. Machine Learning, 33(2/3):263–282, 1998.
  222. 222.A. Samuel. Some studies in machine learning using the game of checkers. IBM Journal of Research and Development, 3(3):210–229, 1959.
  223. 223.T. Sandholm and R. H. Crites. On multiagent Q-learning in a semi-competitive domain. In Adaption and Learning in Multi-Agent Systems, pages 191–205, 1995.
  224. 224.H. Santana, G. Ramalho, V. Corruble, and B. Ratitch. Multi-agent patrolling with reinforcement learning. In AAMAS-2004 — Proceedings of the Third International Joint Conference on Autonomous Agents and Multi Agent Systems, 2004.
  225. 225.G. Saunders and J. Pollack. The evolution of communication schemes over continuous channels. In From Animals to Animats 4 - Proceedings of the Fourth International Conference on Adaptive Behaviour,, 1996.
  226. 226.J. Sauter, H. Van Dyke Parunak, S. Brueckner, and R. Matthews. Tuning synthetic pheromones with evolutionary computing. In R. E. Smith, C. Bonacina, C. Hoile, and P. Marrow, editors, Evolutionary Computation and Multi-Agent Systems (ECOMAS), pages 321–324, San Francisco, California, USA, 7 2001.
  227. 227.J. Sauter, R. S. Matthews, H. Van Dyke Parunak, and S. Brueckner. Evolving adaptive pheromone path planning mechanisms. In Proceedings of First International Joint Conference on Autonomous Agents and Multi-Agent Systems (AAMAS-02), pages 434–440, 2002.
  228. 228.J. Schmidhuber. Realistic multi-agent reinforcement learning. In Learning in Distributed Artificial Intelligence Systems, Working Notes of the 1996 ECAI Workshop, 1996.
  229. 229.J. Schmidhuber and J. Zhao. Multi-agent learning with the success-story algorithm. In ECAI Workshop LDAIS / ICMAS Workshop LIOME, pages 82–93, 1996.
  230. 230.J. Schneider, W.-K. Wong, A. Moore, and M. Riedmiller. Distributed value functions. In Proceedings of the Sixteenth International Conference on Machine Learning, pages 371–378, 1999.
  231. 231.A. Schultz, J.Grefenstette, and W. Adams. Robo-shepherd: Learning complex robotic behaviors. In Robotics and Manufacturing: Recent Trends in Research and Applications, Volume 6, pages 763–768. ASME Press, 1996.
  232. 232.U. M. Schwuttke and A. G. Quan. Enhancing performance of cooperating agents in realtime diagnostic systems. In Proceedings of the Thirteenth International Joint Conference on Artificial Intelligence (IJCAI-93), 1993.
  233. 233.M. Sekaran and S. Sen. To help or not to help. In Proceedings of the Seventeenth Annual Conference of the Cognitive Science Society, pages 736–741, Pittsburgh, PA, 1995.
  234. 234.S. Sen. Multiagent systems: Milestones and new horizons. Trends in Cognitive Science, 1(9):334–339, 1997.
  235. 235.S. Sen. Special issue on evolution and learning in multiagent systems. International Journal of Human-Computer Studies, 48(1), 1998.
  236. 236.S. Sen and M. Sekaran. Using reciprocity to adapt to others. In G. Weiß and S. Sen, editors, International Joint Conference on Artificial Intelligence Workshop on Adaptation and Learning in Multiagent Sytems, Lecture Notes in Artificial Intelligence, pages 206–217. Springer-Verlag, 1995.
  237. 237.S. Sen and M. Sekaran. Multiagent coordination with learning classifier systems. In G. Weiß and S. Sen, editors, Proceedings of the IJCAI Workshop on Adaption and Learning in Multi-Agent Systems, volume 1042, pages 218–233. Springer Verlag, 1996. ISBN 3-540-60923-7.
  238. 238.S. Sen and M. Sekaran. Individual learning of coordination knowledge. Journal of Experimental and Theoretical Artificial Intelligence, 10 (3):333–356, 1998.
  239. 239.S. Sen, M. Sekaran, and J. Hale. Learning to coordinate without sharing information. In Proceedings of the Twelfth National Conference on Artificial Intelligence, pages 426–431, 1994.
  240. 240.Y. Shoham, R. Powers, and T. Grenager. On the agenda(s) of research on multi-agent learning. In Proceedings of Artificial Multiagent Learning. Papers from the 2004 AAAI Fall Symposium. Technical Report FS-04-02, 2004.
  241. 241.R. Smith and B. Gray. Co-adaptive genetic algorithms: An example in othello strategy. Technical Report TCGA 94002, University of Alabama, Department of Engineering Science and Mechanics, 1993.
  242. 242.L. Spector and J. Klein. Evolutionary dynamics discovered via visualization in the breve simulation environment. In Workshop Proceedings of the 8th International Conference on the Simulation and Synthesis of Living Systems, pages 163–170, 2002.
  243. 243.L. Spector, J. Klein, C. Perry, and M. Feinstein. Emergence of collective behavior in evolving populations of flying agents. In E. Cantu-Paz et al, editor, Prooceedings of the Genetic and Evolutionary Computation Conference (GECCO). Springer-Verlag, 2003.
  244. 244.R. Steeb, S. Cammarata, F. Hayes-Roth, P. Thorndyke, and R. Wesson. Distributed intelligence for air fleet control. In A. Bond and L. Gasser, editors, Readings in Distributed Artificial Intelligence, pages 90–101. Morgan Kaufmann Publishers, 1988.
  245. 245.L. Steels. The puzzle of language evolution. Kognitionswissenschaft, 8(4):143–150, 2000.
  246. 246.L. Steels. A self-organizing spatial vocabulary. Artificial Life, 2(3):319–332, 1995.
  247. 247.L. Steels. Emergent adaptive lexicons. In P. Maes, editor, Proceedings of the Simulation of Adaptive Behavior Conference. MIT Press, 1996.
  248. 248.L. Steels. Self-organising vocabularies. In Proceedings of Artificial Life V, 1996.
  249. 249.L. Steels. The spontaneous self-organization of an adaptive language. In S. Muggleton, editor, Machine Intelligence 15. Oxford University Press, Oxford, UK, 1996.
  250. 250.L. Steels. Synthesising the origins of language and meaning using co-evolution, self-organisation and level formation. In J. Hurford, C. Knight, and M. Studdert-Kennedy, editors, Approaches to the Evolution of Language: Social and Cognitive bases. Edinburgh University Press, 1997.
  251. 251.L. Steels and F. Kaplan. Collective learning and semiotic dynamics. In Proceedings of the European Conference on Artificial Life, pages 679–688, 1999.
  252. 252.P. Stone. Layered learning in multiagent systems. In Proceedings of National Conference on Artificial IntelligenceAAAI/IAAI, 1997.
  253. 253.P. Stone. Layered Learning in Multi-Agent Systems. PhD thesis, Carnegie Mellon University, 1998.
  254. 254.P. Stone and R. Sutton. Keepaway soccer: A machine learning testbed. In A. Birk, S. Coradeschi, and S. Tadokoro, editors, RoboCup 2001: Robot Soccer World Cup V, volume 2377 of Lecture Notes in Computer Science, pages 214–223. Springer, 2002. ISBN 3-540-43912-9.
  255. 255.P. Stone and M. M. Veloso. Multiagent systems: A survey from a machine learning perspective. Autonomous Robots, 8(3):345–383, 2000.
  256. 256.N. Sturtevant and R. Korf. On pruning techniques for multi-player games. In Proceedings of National Conference on Artificial Intelligence (AAAI), pages 201–207, 2000.
  257. 257.D. Subramanian, P. Druschel, and J. Chen. Ants and reinforcement learning: A case study in routing in dynamic networks. In Proceedings of Fifteenth International Joint Conference on Artificial Intelligence (IJCAI-97), pages 832–839, 1997.
  258. 258.N. Suematsu and A. Hayashi. A multiagent reinforcement learning algorithm using extended optimal response. In Proceedings of First International Joint Conference on Autonomous Agents and Multi-Agent Systems (AAMAS-02), pages 370–377, 2002.
  259. 259.D. Suryadi and P. J. Gmytrasiewicz. Learning models of other agents using influence diagrams. In Preceedings of the 1999 International Conference on User Modeling, pages 223–232, 1999.
  260. 260.R. Sutton. Learning to predict by the methods of temporal differences. Machine Learning, 3:9–44, 1988.
  261. 261.R. Sutton and A. Barto. Reinforcement Learning: An Introduction. MIT Press, 1998.
  262. 262.J. Svennebring and S. Koenig. Trail-laying robots for robust terrain coverage. In Proceedings of the International Conference on Robotics and Automation (ICRA-03), 2003.
  263. 263.P. ’t Hoen and K. Tuyls. Analyzing multi-agent reinforcement learning using evolutionary dynamics. In Proceedings of the 15th European Conference on Machine Learning (ECML), 2004.
  264. 264.M. Tambe. Recursive agent and agent-group tracking in a real-time dynamic environment. In V. Lesser and L. Gasser, editors, Proceedings of the First International Conference on Multiagent Systems (ICMAS-95). AAAI Press, 1995.
  265. 265.M. Tan. Multi-agent reinforcement learning: Independent vs. cooperative learning. In M. N. Huhns and M. P. Singh, editors, Readings in Agents, pages 487–494. Morgan Kaufmann, San Francisco, CA, USA, 1993.
  266. 266.P. Tangamchit, J. Dolan, and P. Khosla. The necessity of average rewards in cooperative multirobot learning. In Proceedings of IEEE Conference on Robotics and Automation, 2002.
  267. 267.G. Tesauro. Temporal difference learning and TD-gammon. Communications of the ACM, 38(3):58–68, 1995.
  268. 268.G. Tesauro and J. O. Kephart. Pricing in agent economies using multi-agent Q-learning. Autonomous Agents and Multi-Agent Systems, 8: 289–304, 2002.
  269. 269.S. Thrun. Learning to play the game of chess. In G. Tesauro, D. Touretzky, and T. Leen, editors, Advances in Neural Information Processing Systems 7, pages 1069–1076. The MIT Press, Cambridge, MA, 1995.
  270. 270.K. Tumer, A. K. Agogino, and D. H. Wolpert. Learning sequences of actions in collectives of autonomous agents. In Proceedings of First International Joint Conference on Autonomous Agents and Multi-Agent Systems (AAMAS-02), pages 378–385, 2002.
  271. 271.K. Tuyls, K. Verbeeck, and T. Lenaerts. A selection-mutation model for Q-learning in multi-agent systems. In AAMAS-2003 — Proceedings of the Second International Joint Conference on Autonomous Agents and Multi Agent Systems, 2003.
  272. 272.W. Uther and M. Veloso. Adversarial reinforcement learning. Technical Report CMU-CS-03-107, School of Computer Science, Carnegie Mellon University, 2003.
  273. 273.H. Van Dyke Parunak. Applications of distributed artificial intelligence in industry. In G. M. P. O’Hare and N. R. Jennings, editors, Foundations of Distributed AI. John Wiley & Sons, 1996.
  274. 274.L. Z. Varga, N. R. Jennings, and D. Cockburn. Integrating intelligent systems into a cooperating community for electricity distribution management. International Journal of Expert Systems with Applications, 7(4):563–579, 1994.
  275. 275.J. Vidal and E. Durfee. Predicting the expected behavior of agents that learn about agents: the CLRI framework. Autonomous Agents and Multi-Agent Systems, January 2003.
  276. 276.J. Vidal and E. Durfee. Agents learning about agents: A framework and analysis. In Working Notes of AAAI-97 Workshop on Multiagent Learning, 1997.
  277. 277.J. Vidal and E. Durfee. The moving target function problem in multiagent learning. In Proceedings of the Third Annual Conference on Multi-Agent Systems, 1998.
  278. 278.K. Wagner. Cooperative strategies and the evolution of communication. Artificial Life, 6(2):149–179, Spring 2000.
  279. 279.X. Wang and T. Sandholm. Reinforcement learning to play an optimal Nash equilibrium in team Markov games. In Advances in Neural Information Processing Systems (NIPS-2002), 2002.
  280. 280.R. Watson and J. Pollack. Coevolutionary dynamics in a minimal substrate. In E. Cantu-Paz et al, editor, Proceedings of the Genetic and Evolutionary Computation Conference (GECCO), 2001.
  281. 281.R. Weihmayer and H. Velthuijsen. Application of distributed AI and cooperative problem solving to telecommunications. In J. Liebowitz and D. Prereau, editors, AI Approaches to Telecommunications and Network Management. IOS Press, 1994.
  282. 282.M. Weinberg and J. Rosenschein. Best-response multiagent learning in non-stationary environments. In AAMAS-2004 — Proceedings of the Third International Joint Conference on Autonomous Agents and Multi Agent Systems, 2004.
  283. 283.G. Weiß. Some studies in distributed machine learning and organizational design. Technical Report FKI-189-94, Institut für Informatik, TU München, 1994.
  284. 284.G. Weiß. Distributed Machine Learning. Sankt Augustin: Infix Verlag, 1995.
  285. 285.G. Weiß, editor. Distributed Artificial Intelligence Meets Machine Learning : Learning in Multi-Agent Environments. Number 1221 in Lecture Notes in Artificial Intelligence. Springer-Verlag, 1997.
  286. 286.G. Weiß. Special issue on learning in distributed artificial intelligence systems. Journal of Experimental and Theoretical Artificial Intelligence, 10(3), 1998.
  287. 287.G. Weiß, editor. Multiagent Systems: A Modern Approach to Distributed Artificial Intelligence. MIT Press, 1999.
  288. 288.G. Weiß and P. Dillenbourg. What is ‘multi’ in multi-agent learning? In P. Dillenbourg, editor, Collaborative learning. Cognitive and computational approaches, pages 64–80. Pergamon Press, 1999.
  289. 289.G. Weiß and S. Sen, editors. Adaptation and Learning in Multiagent Systems. Lecture Notes in Artificial Intelligence, Volume 1042. Springer-Verlag, 1996.
  290. 290.M. Wellman and J. Hu. Conjectural equilibrium in multiagent learning. Machine Learning, 33(2-3):179–200, 1998.
  291. 291.J. Werfel, M. Mitchell, and J. P. Crutchfield. Resource sharing and coevolution in evolving cellular automata. IEEE Transactions on Evolutionary Computation, 4(4):388, November 2000.
  292. 292.B. B. Werger and M. Mataric. Exploiting embodiment in multi-robot teams. Technical Report IRIS-99-378, University of Southern California, Institute for Robotics and Intelligent Systems, 1999.
  293. 293.G. M. Werner and M. G. Dyer. Evolution of herding behavior in artificial animals. In From Animals to Animats 2: Proceedings of the Second International Conference on Simulation of Adaptive Behavior (SAB92), 1993.
  294. 294.T. White, B. Pagurek, and F. Oppacher. ASGA: Improving the ant system by integration with genetic algorithms. In J. R. Koza, W. Banzhaf, K. Chellapilla, K. Deb, M. Dorigo, D. B. Fogel, M. H. Garzon, D. E. Goldberg, H. Iba, and R. Riolo, editors, Genetic Programming 1998: Proceedings of the Third Annual Conference, pages 610–617, University of Wisconsin, Madison, Wisconsin, USA, 22-25 1998. Morgan Kaufmann.
  295. 295.S. Whiteson and P. Stone. Concurrent layered learning. In AAMAS-2003 — Proceedings of the Second International Joint Conference on Autonomous Agents and Multi Agent Systems, 2003.
  296. 296.R. P. Wiegand. Analysis of Cooperative Coevolutionary Algorithms. PhD thesis, Department of Computer Science, George Mason University, 2003.
  297. 297.R. P. Wiegand and J. Sarma. Spatial embedding and loss of gradient in cooperative coevolutionary algorithms. In Parallel Problem Solving from Nature — PPSN-2004. Springer, 2004.
  298. 298.R. P. Wiegand, W. Liles, and K. De Jong. An empirical analysis of collaboration methods in cooperative coevolutionary algorithms. In E. Cantu-Paz et al, editor, Proceedings of the Genetic and Evolutionary Computation Conference (GECCO), pages 1235–1242, 2001.
  299. 299.R. P. Wiegand, W. Liles, and K. De Jong. Analyzing cooperative coevolution with evolutionary game theory. In D. Fogel, editor, Proceedings of Congress on Evolutionary Computation (CEC-02), pages 1600–1605. IEEE Press, 2002.
  300. 300.R. P. Wiegand, W. Liles, and K. De Jong. Modeling variation in cooperative coevolution using evolutionary game theory. In Poli et al. [201], pages 231–248.
  301. 301.M. Wiering, R. Salustowicz, and J. Schmidhuber. Reinforcement learning soccer teams with incomplete world models. Journal of Autonomous Robots, 7(1):77–88, 1999.
  302. 302.A. Williams. Learning to share meaning in a multi-agent system. Autonomous Agents and Multi-Agent Systems, 8:165–193, 2004.
  303. 303.E. Wilson. Sociobiology: The New Synthesis. Belknap Press, 1975.
  304. 304.D. H. Wolpert and K. Tumer. Optimal payoff functions for members of collectives. Advances in Complex Systems, 4(2/3):265–279, 2001.
  305. 305.D. H. Wolpert, K. Tumer, and J. Frank. Using collective intelligence to route internet traffic. In Advances in Neural Information Processing Systems-11, pages 952–958, Denver, 1998.
  306. 306.D. H. Wolpert, K. R. Wheller, and K. Tumer. General principles of learning-based multi-agent systems. In O. Etzioni, J. P. Müller, and J. M. Bradshaw, editors, Proceedings of the Third International Conference on Autonomous Agents (Agents’99), pages 77–83, Seattle, WA, USA, 1999. ACM Press.
  307. 307.M. Wooldridge, S. Bussmann, and M. Klosterberg. Production sequencing as negotiation. In Proceedings of the First International Conference on the Practical Application of Intelligent Agents and Multi-Agent Technology (PAAM-96),, 1996.
  308. 308.A. Wu, A. Schultz, and A. Agah. Evolving control for distributed micro air vehicles. In IEEE Computational Intelligence in Robotics and Automation Engineers Conference, 1999.
  309. 309.H. Yanco and L. Stein. An adaptive communication protocol for cooperating mobile robots. In From Animals to Animats: International Conference on Simulation of Adaptive Behavior, pages 478–485, 1993.
  310. 310.N. Zaera, D. Cliff, and J. Bruten. (not) evolving collective behaviours in synthetic fish. Technical Report HPL-96-04, Hewlett-Packard Laboratories, 1996.
  311. 311.B. Zhang and D. Cho. Coevolutionary fitness switching: Learning complex collective behaviors using genetic programming. In Advances in Genetic Programming III, pages 425–445. MIT Press, 1998.
  312. 312.J. Zhao and J. Schmidhuber. Incremental self-improvement for life-time multi-agent reinforcement learning. In P. Maes, M. Mataric, J.-A. Meyer, J. Pollack, and S. W. Wilson, editors, Proceedings of the Fourth International Conference on Simulation of Adaptive Behavior: From Animals to Animats 4, pages 516–525, Cape Code, USA, 9-13 1996. MIT Press. ISBN 0-262-63178-4.

Citation

MLA
Panait, L., and S. Luke. “Cooperative Multi-Agent Learning: The State of the Art”. Autonomous Agents and Multi-Agent Systems, vol. 11, no. 3, 2005, pp. 387–434, https://doi.org/10.1007/s10458-005-2631-2.
APA
Panait, L., & Luke, S. (2005). Cooperative Multi-Agent Learning: The State of the Art. Autonomous Agents and Multi-Agent Systems, 11(3), 387–434. https://doi.org/10.1007/s10458-005-2631-2
Chicago
Panait, L., and S. Luke. 2005. “Cooperative Multi-Agent Learning: The State of the Art”. Autonomous Agents and Multi-Agent Systems 11 (3): 387–434. https://doi.org/10.1007/s10458-005-2631-2.
Harvard
Panait, L. and Luke, S. (2005) “Cooperative Multi-Agent Learning: The State of the Art”, Autonomous Agents and Multi-Agent Systems, 11(3), pp. 387–434. Available at: https://doi.org/10.1007/s10458-005-2631-2.
Vancouver
1. Panait L, Luke S (2005) Cooperative Multi-Agent Learning: The State of the Art. Autonomous Agents and Multi-Agent Systems 11:387–434

BibTeX

@article{Panait_2005, title={Cooperative Multi-Agent Learning: The State of the Art}, volume={11}, ISSN={1573-7454}, url={http://dx.doi.org/10.1007/s10458-005-2631-2}, DOI={10.1007/s10458-005-2631-2}, number={3}, journal={Autonomous Agents and Multi-Agent Systems}, publisher={Springer Science and Business Media LLC}, author={Panait, Liviu and Luke, Sean}, year={2005}, month=Nov, pages={387–434} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF