Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination

Rui ZhaoJinming SongYufeng YuanHaifeng HuYang GaoYi WuZhongqian SunWei Yang

article2023AAAI114 citations

Introduces Maximum Entropy Population-based training, a framework that generates a diverse partner population using a computationally efficient population entropy objective and trains a coordinator agent via prioritized sampling to enable zero-shot human-AI collaboration without requiring human data.

Listen

Building artificial intelligence systems that can seamlessly collaborate with humans remains a major challenge. Standard self-play reinforcement learning trains agents solely against copies of themselves, causing them to develop overly specialized strategies. When paired with unfamiliar partners such as real people, these agents frequently fail due to behavioral mismatches, or distributional shift. Addressing this issue without relying on expensive, time-consuming human training data is essential for deploying collaborative AI in applications like autonomous vehicles, assistive robotics, and digital assistants.

The article demonstrates and evaluates a framework called Maximum Entropy Population-based training to train collaborative AI agents without any human demonstration data. Its primary objective is to verify whether cultivating a diverse population of synthetic partners and training against them via prioritized sampling enables zero-shot human-AI collaboration.

To achieve this, the authors formulated a population diversity metric that merges individual agent exploration with pairwise behavioral differences, then derived an efficient surrogate objective termed Population Entropy. The overall framework operates in two distinct phases: first, a diverse pool of synthetic partner agents is trained using this entropy bonus; second, a primary AI agent is trained against this pool using learning-progress-based prioritized sampling, which emphasizes partners that are harder to coordinate with. The approach was evaluated using simulations in a collaborative matrix game and across five layouts of the cooperative cooking game Overcooked, testing AI performance alongside human proxy models and real participants recruited via Amazon Mechanical Turk.

The evaluation produced several critical findings. First, the proposed method consistently outperformed existing baselines—including standard self-play, basic population training, and recent diversity-driven methods like Trajectory Diversity and Fictitious Co-Play—across all five simulated environments when paired with human proxy models. Second, in live trials with real humans, the proposed model achieved the highest average coordination scores among all evaluated AI methods, performing on par with human-human pairs. Third, ablation studies confirmed that both the entropy bonus and prioritized sampling are vital; removing either component degraded coordination returns. Finally, the method achieved superior results using only half the population size required by alternative frameworks such as Fictitious Co-Play, while converging faster in matrix benchmarks.

These findings indicate that zero-shot human-AI coordination can be achieved efficiently purely through simulated diversity, significantly lowering the cost, time, and data collection overhead associated with human-in-the-loop training. By intentionally exposing AI to challenging and varied non-human partners during training, the system learns robust, adaptable policies that naturally accommodate human unpredictability rather than freezing or failing when conventions are broken.

Organizations aiming to build collaborative AI should consider adopting entropy-regularized population frameworks to minimize data costs and mitigate deployment failure risks. For practical implementations, prioritized sampling should be incorporated to prevent the AI from over-relying on easy partners. Further work should explore integrating this population-based approach with broader multi-agent learning algorithms and evaluating performance in more complex physical or operational domains beyond simulated games.

While the empirical findings provide strong confidence in the method's effectiveness for grid-based cooperative tasks, several limitations remain. The evaluations were conducted within stylized game settings and focused on two-player scenarios. Stakeholders should maintain cautious optimism when extrapolating these performance levels to high-stakes, open-ended, or safety-critical domains where coordination mistakes carry significant operational risks.

Cover for Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination

Abstract

We study the problem of training a Reinforcement Learning (RL) agent that is collaborative with humans without using human data. Although such agents can be obtained through self-play training, they can suffer significantly from the distributional shift when paired with unencountered partners, such as humans. In this paper, we propose Maximum Entropy Population-based training (MEP) to mitigate such distributional shift. In MEP, agents in the population are trained with our derived Population Entropy bonus to promote the pairwise diversity between agents and the individual diversity of agents themselves. After obtaining this diversified population, a common best agent is trained by paring with agents in this population via prioritized sampling, where the prioritization is dynamically adjusted based on the training progress. We demonstrate the effectiveness of our method MEP, with comparison to Self-Play PPO (SP), Population-Based Training (PBT), Trajectory Diversity (TrajeDi), and Fictitious Co-Play (FCP) in both matrix game and Overcooked game environments, with partners being human proxy models and real humans. A supplementary video showing experimental results is available at https://youtu.be/Xh-FKD0AAKE.

Table of Contents

  • Introduction
  • Preliminaries
  • Method
  • Population Diversity
  • Population Entropy
  • Training a Maximum Entropy Population
  • Algorithm 1: Maximum Entropy Population
  • Training a Robust Agent via Prioritized Sampling
  • Experiments
  • Question 3. What does an MEP population look like?
  • Question 4. How does MEP compare to other methods?
  • Question 5. How does MEP perform with real humans?
  • Performance with human proxy model
  • Question 6. What does AI do when paired with humans?
  • Related Work
  • Conclusion
  • References

Knowls

  1. Knowl 1 — Population Diversity Objective

    definition

    In population-based reinforcement learning with nn agents whose policies are parameterized as {π(1),π(2),…,π(n)}\{\pi^{(1)}, \pi^{(2)}, \dots, \pi^{(n)}\}, the Population Diversity (PD) objective at state sts_t measures both the individual exploration capability of each policy and the pairwise distinctness across all pairs in the population. It is defined as:

    PD({π(1),π(2),…,π(n)},st):=1n∑i=1nH(π(i)(⋅∣st))+1n2∑i=1n∑j=1nDKL(π(i)(⋅∣st),π(j)(⋅∣st)),\text{PD}(\{\pi^{(1)}, \pi^{(2)}, \dots, \pi^{(n)}\}, s_t) := \frac{1}{n} \sum_{i=1}^n \mathcal{H}(\pi^{(i)}(\cdot \mid s_t)) + \frac{1}{n^2} \sum_{i=1}^n \sum_{j=1}^n D_{\text{KL}}(\pi^{(i)}(\cdot \mid s_t), \pi^{(j)}(\cdot \mid s_t)),

    where A\mathcal{A} is the discrete action space, the Shannon entropy of policy π(i)\pi^{(i)} is:

    H(π(i)(⋅∣st))=−∑a∈Aπ(i)(a∣st)log⁡π(i)(a∣st),\mathcal{H}(\pi^{(i)}(\cdot \mid s_t)) = -\sum_{a \in \mathcal{A}} \pi^{(i)}(a \mid s_t) \log \pi^{(i)}(a \mid s_t),

    and the Kullback–Leibler divergence between policies π(i)\pi^{(i)} and π(j)\pi^{(j)} is:

    DKL(π(i)(⋅∣st),π(j)(⋅∣st))=∑a∈Aπ(i)(a∣st)log⁡π(i)(a∣st)π(j)(a∣st).D_{\text{KL}}(\pi^{(i)}(\cdot \mid s_t), \pi^{(j)}(\cdot \mid s_t)) = \sum_{a \in \mathcal{A}} \pi^{(i)}(a \mid s_t) \log \frac{\pi^{(i)}(a \mid s_t)}{\pi^{(j)}(a \mid s_t)}.

    Evaluating the full PD objective requires quadratic computational complexity O(n2)\mathcal{O}(n^2) with respect to population size nn. Furthermore, because KL divergence is unbounded, directly optimizing PD within reinforcement learning reward functions can induce training instability.

  2. Knowl 2 — Population Entropy Surrogate Objective

    definition

    The Population Entropy (PE) objective is a computationally efficient, bounded surrogate for the Population Diversity (PD) objective. It is defined as the Shannon entropy of the average action distribution across the population {π(1),π(2),…,π(n)}\{\pi^{(1)}, \pi^{(2)}, \dots, \pi^{(n)}\} at state sts_t:

    PE({π(1),π(2),…,π(n)},st):=H(πˉ(⋅∣st)),\text{PE}(\{\pi^{(1)}, \pi^{(2)}, \dots, \pi^{(n)}\}, s_t) := \mathcal{H}(\bar{\pi}(\cdot \mid s_t)),

    where πˉ\bar{\pi} is the mean policy of the population:

    πˉ(at∣st):=1n∑i=1nπ(i)(at∣st).\bar{\pi}(a_t \mid s_t) := \frac{1}{n} \sum_{i=1}^n \pi^{(i)}(a_t \mid s_t).

    Evaluating PE requires linear time complexity O(n)\mathcal{O}(n) with respect to the population size nn. For categorical action distributions, PE is bounded by [0,log⁡∣A∣][0, \log |\mathcal{A}|], making it numerically stable when incorporated directly into reinforcement learning reward bonuses.

  3. Knowl 3 — Population Entropy Lower Bound on Population Diversity

    theoretical result

    For any population of nn stochastic policies {π(1),π(2),…,π(n)}\{\pi^{(1)}, \pi^{(2)}, \dots, \pi^{(n)}\} operating over a shared discrete action space A\mathcal{A} at state sts_t, the Population Entropy (PE) is a mathematical lower bound of the Population Diversity (PD) objective:

    PD({π(1),π(2),…,π(n)},st)≥PE({π(1),π(2),…,π(n)},st),\text{PD}(\{\pi^{(1)}, \pi^{(2)}, \dots, \pi^{(n)}\}, s_t) \ge \text{PE}(\{\pi^{(1)}, \pi^{(2)}, \dots, \pi^{(n)}\}, s_t),

    where PD\text{PD} denotes the sum of the average individual policy entropies and average pairwise KL divergences:

    PD({π(1),…,π(n)},st)=1n∑i=1nH(π(i)(⋅∣st))+1n2∑i=1n∑j=1nDKL(π(i)(⋅∣st),π(j)(⋅∣st)),\text{PD}(\{\pi^{(1)}, \dots, \pi^{(n)}\}, s_t) = \frac{1}{n} \sum_{i=1}^n \mathcal{H}(\pi^{(i)}(\cdot \mid s_t)) + \frac{1}{n^2} \sum_{i=1}^n \sum_{j=1}^n D_{\text{KL}}(\pi^{(i)}(\cdot \mid s_t), \pi^{(j)}(\cdot \mid s_t)),

    and PE\text{PE} denotes the entropy of the population mean policy πˉ(at∣st)=1n∑i=1nπ(i)(at∣st)\bar{\pi}(a_t \mid s_t) = \frac{1}{n} \sum_{i=1}^n \pi^{(i)}(a_t \mid s_t):

    PE({π(1),…,π(n)},st)=H(πˉ(⋅∣st)).\text{PE}(\{\pi^{(1)}, \dots, \pi^{(n)}\}, s_t) = \mathcal{H}(\bar{\pi}(\cdot \mid s_t)).

    Maximizing the Population Entropy objective directly maximizes a lower bound of the total Population Diversity.

  4. Knowl 4 — Maximum Entropy Population Training

    algorithm

    Maximum Entropy Population (MEP) training learns a diverse set of nn cooperative policies by augmenting the shared environment task reward with a centralized population entropy bonus derived from the ensemble mean policy.

    Input: Population of policies {\pi^{(1)}, \pi^{(2)}, ..., \pi^{(n)}}, entropy weight \alpha, steps per episode T
    Output: Trained maximum entropy population {\pi^{(1)}, \pi^{(2)}, ..., \pi^{(n)}}
    while training not converged do
        Uniformly sample an agent index i ~ {1, ..., n}
        for t = 1 to T do
            Sample action a_t ~ \pi^{(i)}(a_t | s_t)
            Execute joint action and step environment s_{t+1} ~ p(s_{t+1} | s_t, a_t)
            Compute mean policy distribution: \bar{\pi}(a_t | s_t) = (1 / n) * \sum_{j=1}^n \pi^{(j)}(a_t | s_t)
            Compute combined reward: r_t = R(s_t, a_t) - \alpha * log(\bar{\pi}(a_t | s_t))
        end for
        Update policy \pi^{(i)} using policy gradients to maximize expected return E_\tau [\sum_{t=1}^T r_t]
    end while

    Each agent interacts in self-play or with its own duplicate while the entropy penalty is computed against the centralized mean distribution πˉ(at∣st)\bar{\pi}(a_t \mid s_t) of all nn agents, incentivizing individual policies to diversify across complementary behaviors.

  5. Knowl 5 — Robust Agent Training via Prioritized Partner Sampling

    model/method

    Once a maximum entropy population {π(1),…,π(n)}\{\pi^{(1)}, \dots, \pi^{(n)}\} is trained, a robust best-response agent π(A)\pi^{(A)} is trained by pairing it with partners sampled from the population. To prevent the agent from solely exploiting easy-to-collaborate partners (which occurs under uniform partner sampling), MEP employs rank-based prioritized sampling that favors partners with lower expected collaboration returns, approximating a minimax objective:

    π(A)=arg⁡max⁡πmin⁡i∈{1,…,n}J(π,π(i)),\pi^{(A)} = \arg\max_{\pi} \min_{i \in \{1, \dots, n\}} J(\pi, \pi^{(i)}),

    where J(π,π(i))J(\pi, \pi^{(i)}) is the expected return of pairing π\pi with partner π(i)\pi^{(i)}.

    The probability of sampling partner policy π(i)\pi^{(i)} during training is defined by:

    p(π(i))=[rank(1Eτ[∑tR(st,at(A),at(i))])]β∑j=1n[rank(1Eτ[∑tR(st,at(A),at(j))])]β,p(\pi^{(i)}) = \frac{\left[ \text{rank}\left( \frac{1}{\mathbb{E}_\tau \left[ \sum_t R(s_t, a_t^{(A)}, a_t^{(i)}) \right]} \right) \right]^\beta}{\sum_{j=1}^n \left[ \text{rank}\left( \frac{1}{\mathbb{E}_\tau \left[ \sum_t R(s_t, a_t^{(A)}, a_t^{(j)}) \right]} \right) \right]^\beta},

    where rank(⋅)∈{1,…,n}\text{rank}(\cdot) \in \{1, \dots, n\} assigns rank 1 to the lowest inverse reward (easiest partner) and rank nn to the highest inverse reward (hardest partner), and β≥0\beta \ge 0 is a hyperparameter controlling prioritization strength. Setting β=0\beta = 0 corresponds to uniform sampling.

  6. Knowl 6 — Effect of PE Bonus Weight on Population Entropy in Overcooked

    data/table

    The Population Entropy (PE) of an agent population was evaluated across five Overcooked game layouts under varying entropy reward weights α∈[0.000,0.050]\alpha \in [0.000, 0.050]. The table reports the population entropy corresponding to the best task reward achieved during training across five layouts: Cramped Room (Cramped Rm.), Asymmetric Advantages (Asymm. Adv.), Coordination Ring (Coord. Ring), Forced Coordination (Forced Coord.), and Counter Circuit (Counter Circ.).

    α\alpha Cramped Rm. Asymm. Adv. Coord. Ring Forced Coord. Counter Circ.
    0.000 0.971 1.120 0.878 0.970 0.988
    0.001 1.031 1.051 0.907 0.858 1.152
    0.005 0.949 1.075 0.901 0.889 1.038
    0.010 1.057 1.139 0.840 1.079 1.151
    0.020 1.029 1.074 0.947 1.093 1.171
    0.030 1.134 1.203 1.028 0.957 1.715
    0.040 1.194 1.353 1.122 1.460 1.791
    0.050 1.127 1.364 0.996 1.703 1.791

    Across all layouts, setting α>0\alpha > 0 yields higher population entropy than the baseline α=0.000\alpha = 0.000, confirming that the PE bonus actively increases behavioral diversity among agents in the population.

  7. Knowl 7 — Zero-Shot Coordination with Behavior-Cloned Human Proxy Models

    empirical result

    AI agents trained via Maximum Entropy Population-based training (MEP) were evaluated in a zero-shot coordination setup by pairing them with a fixed human proxy model (HProxyH_{\text{Proxy}}), trained via behavior cloning on human gameplay trajectories in the Overcooked environment over 400-timestep episodes. MEP was compared against Self-Play (SP), Population-Based Training (PBT), Trajectory Diversity (TrajeDi), Fictitious Co-Play (FCP), and Maximum Population Diversity (MPD) using a default population size n=5n=5 (and n=10n=10 for FCP) across five layouts:

    • Cramped Room: MEP achieves the highest average reward per episode (approx120 approx 120), outperforming SP (approx103 approx 103), PBT (approx98 approx 98), TrajeDi (approx115 approx 115), FCP (approx110 approx 110), and MPD (approx85 approx 85).
    • Asymmetric Advantages: MEP achieves approx112 approx 112, outperforming SP (approx70 approx 70), PBT (approx75 approx 75), TrajeDi (approx92 approx 92), FCP (approx83 approx 83), and MPD (approx87 approx 87).
    • Coordination Ring: MEP achieves approx102 approx 102, outperforming SP (approx70 approx 70), PBT (approx75 approx 75), TrajeDi (approx98 approx 98), FCP (approx95 approx 95), and MPD (approx87 approx 87).
    • Forced Coordination: MEP achieves approx40 approx 40, outperforming SP (approx30 approx 30), PBT (approx25 approx 25), TrajeDi (approx25 approx 25), FCP (approx31 approx 31), and MPD (approx17 approx 17).
    • Counter Circuit: MEP achieves approx55 approx 55, outperforming SP (approx31 approx 31), PBT (approx35 approx 35), TrajeDi (approx46 approx 46), FCP (approx32 approx 32), and MPD (approx53 approx 53).

    Ablation studies show that removing the population entropy reward (MEPα=0\text{MEP}_{\alpha=0}) or removing prioritized partner sampling (MEPβ=0\text{MEP}_{\beta=0}) leads to degraded performance across all five layouts.

  8. Knowl 8 — Zero-Shot Coordination Performance with Real Human Players

    empirical result

    In a user study conducted via Amazon Mechanical Turk (AMT) across five Overcooked layouts, AI agents trained with MEP, TrajeDi, FCP, SP, and PBT were paired with real human players (HTrueH_{\text{True}}) in a between-subjects experimental design.

    Evaluating the average reward per episode across all five layouts:

    • Human-Human baseline (HTrue+HTrueH_{\text{True}} + H_{\text{True}}): approx92 approx 92
    • MEP+HTrue\text{MEP} + H_{\text{True}}: approx88 approx 88
    • TrajeDi+HTrue\text{TrajeDi} + H_{\text{True}}: approx78 approx 78
    • FCP+HTrue\text{FCP} + H_{\text{True}}: approx74 approx 74
    • SP+HTrue\text{SP} + H_{\text{True}}: approx58 approx 58
    • PBT+HTrue\text{PBT} + H_{\text{True}}: approx52 approx 52

    MEP-trained agents achieve the highest average coordination score with real humans among all evaluated algorithmic approaches, performing on par with human-human team coordination.

  9. Knowl 9 — Convergence Speed in Single-Step Collaborative Matrix Game

    empirical result

    In the single-step collaborative matrix game (a 10×1010 \times 10 payoff matrix coordination task), best-response (BR) agents trained against an MEP population demonstrate faster convergence and higher asymptotic returns than agents trained against TrajeDi populations, baseline populations, or individual agents.

    Specifically, the BR agent trained on an MEP population converges to an average return of approx0.95 approx 0.95 in both self-play and cross-play within 75 training steps, whereas the BR agent trained on TrajeDi requires over 150 training steps to reach comparable return levels, and baseline populations/individual agents fail to surpass an average return of 0.400.40 within 200 steps.

Coverage note — Appendix proofs (Theorem 1 proof in Appendix A, performance lower bound derivations in Appendices B and C, and hyperparameter sensitivity figures in Appendix E) were omitted in accordance with the rule excluding intermediate proofs and derivations.

References

  1. 1.Akkaya, I.; Andrychowicz, M.; Chociej, M.; Litwin, M.; McGrew, B.; Petron, A.; Paino, A.; Plappert, M.; Powell, G.; Ribas, R.; Schneider, J.; Tezak, N.; Tworek, J.; Welinder, P.; Weng, L.; Yuan, Q.; Zaremba, W.; and Zhang, L. 2019. Solving Rubik's Cube with a Robot Hand. arXiv preprint.
  2. 2.Bain, M.; and Sammut, C. 1999. A Framework for Behavioural Cloning. In Machine Intelligence 15, Intelligent Agents [St. Catherine's College, Oxford, July 1995], 103–129. Oxford, UK, UK: Oxford University. ISBN 0-19-853867-7.
  3. 3.Balduzzi, D.; Garnelo, M.; Bachrach, Y.; Czarnecki, W.; Perolat, J.; Jaderberg, M.; and Graepel, T. 2019. Open-ended learning in symmetric zero-sum games. In International Conference on Machine Learning, 434–443. PMLR.
  4. 4.Boutilier, C. 1996. Planning, learning and coordination in multiagent decision processes. In Proceedings of the 6th conference on Theoretical aspects of rationality and knowledge, 195–210. Morgan Kaufmann Publishers Inc.
  5. 5.Carroll, M.; Shah, R.; Ho, M. K.; Griffiths, T.; Seshia, S.; Abbeel, P.; and Dragan, A. 2019. On the Utility of Learning about Humans for Human-AI Coordination. In Advances in Neural Information Processing Systems, 5175–5186.
  6. 6.Carter, S.; and Nielsen, M. 2017. Using Artificial Intelligence to Augment Human Intelligence. Distill. https://distill.pub/2017/aia.
  7. 7.Engelbart, D. C. 1962. Augmenting human intellect: A conceptual framework. Menlo Park, CA.
  8. 8.Eysenbach, B.; Gupta, A.; Ibarz, J.; and Levine, S. 2019. Diversity is All You Need: Learning Skills without a Reward Function. In International Conference on Learning Representations.
  9. 9.Foerster, J.; Assael, I. A.; De Freitas, N.; and Whiteson, S. 2016. Learning to communicate with deep multi-agent reinforcement learning. In Advances in neural information processing systems, 2137–2145.
  10. 10.Foerster, J.; Farquhar, G.; Afouras, T.; Nardelli, N.; and Whiteson, S. 2018. Counterfactual multi-agent policy gradients. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  11. 11.Fox, R.; Pakman, A.; and Tishby, N. 2015. Taming the noise in reinforcement learning via soft updates. arXiv preprint arXiv:1512.08562.
  12. 12.Ghost Town Games. 2016. Overcooked. https://store.steampowered.com/app/448510/Overcooked/. Accessed: 2016-08-03.
  13. 13.Haarnoja, T.; Ha, S.; Zhou, A.; Tan, J.; Tucker, G.; and Levine, S. 2019. Learning to Walk Via Deep Reinforcement Learning. In Robotics: Science and Systems.
  14. 14.Haarnoja, T.; Hartikainen, K.; Abbeel, P.; and Levine, S. 2018a. Latent Space Policies for Hierarchical Reinforcement Learning. arXiv preprint arXiv:1804.02808.
  15. 15.Haarnoja, T.; Tang, H.; Abbeel, P.; and Levine, S. 2017. Reinforcement learning with deep energy-based policies. In International Conference on Machine Learning, 1352–1361. PMLR.
  16. 16.Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018b. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning, 1861–1870. PMLR.
  17. 17.Han, L.; Xiong, J.; Sun, P.; Sun, X.; Fang, M.; Guo, Q.; Chen, Q.; Shi, T.; Yu, H.; and Zhang, Z. 2020. Tstarbot-x: An open-sourced and comprehensive study for efficient league training in starcraft ii full game. arXiv preprint arXiv:2011.13729.
  18. 18.Hu, H.; Lerer, A.; Peysakhovich, A.; and Foerster, J. 2020. “Other-Play” for Zero-Shot Coordination. In International Conference on Machine Learning, 4399–4410. PMLR.
  19. 19.Jaderberg, M.; Czarnecki, W. M.; Dunning, I.; Marris, L.; Lever, G.; Castaneda, A. G.; Beattie, C.; Rabinowitz, N. C.; Morcos, A. S.; Ruderman, A.; Sonnerat, N.; Green, T.; Deason, L.; Leibo, J. Z.; Silver, D.; Hassabis, D.; Kavukcuoglu, K.; and Graepel, T. 2019. Human-level performance in 3D multiplayer games with population-based reinforcement learning. Science, 364(6443): 859–865.
  20. 20.Jaderberg, M.; Dalibard, V.; Osindero, S.; Czarnecki, W. M.; Donahue, J.; Razavi, A.; Vinyals, O.; Green, T.; Dunning, I.; Simonyan, K.; et al. 2017. Population based training of neural networks. arXiv preprint arXiv:1711.09846.
  21. 21.Kleiman-Weiner, M.; Ho, M. K.; Austerweil, J. L.; Littman, M. L.; and Tenenbaum, J. B. 2016. Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction. In CogSci.
  22. 22.Knott, P.; Carroll, M.; Devlin, S.; Ciosek, K.; Hofmann, K.; Dragan, A.; and Shah, R. 2021. Evaluating the Robustness of Collaborative Agents. In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems, 1560–1562.
  23. 23.Lerer, A.; and Peysakhovich, A. 2017. Maintaining cooperation in complex social dilemmas using deep reinforcement learning. arXiv preprint arXiv:1707.01068.
  24. 24.Lerer, A.; and Peysakhovich, A. 2018. Learning social conventions in markov games. arXiv preprint arXiv:1806.10071.
  25. 25.Liu, X.; Jia, H.; Wen, Y.; Yang, Y.; Hu, Y.; Chen, Y.; Fan, C.; and Hu, Z. 2021. Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games. arXiv preprint arXiv:2106.04958.
  26. 26.Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; and Mordatch, I. 2017. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems, 6379–6390.
  27. 27.Lupu, A.; Cui, B.; Hu, H.; and Foerster, J. 2021. Trajectory Diversity for Zero-Shot Coordination. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, 7204–7213. PMLR.
  28. 28.Masood, M. A.; and Doshi-Velez, F. 2019. Diversity-inducing policy gradient: Using maximum mean discrepancy to find a set of diverse policies. Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI).
  29. 29.Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016. Asynchronous methods for deep reinforcement learning. In International conference on machine learning, 1928–1937. PMLR.
  30. 30.Murphy, K. P. 2012. Machine Learning: A Probabilistic Perspective. Adaptive Computation and Machine Learning. MIT press.
  31. 31.OpenAI. 2019. OpenAI Five Finals. https://openai.com/blog/openai-five-finals/. Accessed: 2019-03-26.
  32. 32.Pan, S. J.; and Yang, Q. 2009. A survey on transfer learning. IEEE Transactions on knowledge and data engineering, 22(10): 1345–1359.
  33. 33.Parker-Holder, J.; Pacchiano, A.; Choromanski, K. M.; and Roberts, S. J. 2020. Effective Diversity in Population Based Reinforcement Learning. Advances in Neural Information Processing Systems, 33.
  34. 34.Peng, X. B.; Andrychowicz, M.; Zaremba, W.; and Abbeel, P. 2018. Sim-to-real transfer of robotic control with dynamics randomization. In 2018 IEEE international conference on robotics and automation (ICRA), 1–8. IEEE.
  35. 35.Perez-Nieves, N.; Yang, Y.; Slumbers, O.; Mguni, D. H.; Wen, Y.; and Wang, J. 2021. Modelling Behavioural Diversity for Learning in Open-Ended Games. In International Conference on Machine Learning, 8514–8524. PMLR.
  36. 36.Rawlik, K.; Toussaint, M.; and Vijayakumar, S. 2013. On stochastic optimal control and reinforcement learning by approximate inference. In Twenty-third international joint conference on artificial intelligence.
  37. 37.Resnick, C.; Kulikov, I.; Cho, K.; and Weston, J. 2018. Vehicle community strategies. arXiv preprint arXiv:1804.07178.
  38. 38.Schaul, T.; Quan, J.; Antonoglou, I.; and Silver, D. 2016. Prioritized experience replay. In International Conference on Learning Representations.
  39. 39.Schulman, J.; Chen, X.; and Abbeel, P. 2017. Equivalence between policy gradients and soft q-learning. arXiv preprint arXiv:1704.06440.
  40. 40.Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
  41. 41.Shum, M.; Kleiman-Weiner, M.; Littman, M. L.; and Tenenbaum, J. B. 2019. Theory of minds: Understanding behavior in groups through inverse planning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, 6163–6170.
  42. 42.Silver, D.; Hubert, T.; Schrittwieser, J.; Antonoglou, I.; Lai, M.; Guez, A.; Lanctot, M.; Sifre, L.; Kumaran, D.; Graepel, T.; et al. 2017. Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv preprint arXiv:1712.01815.
  43. 43.Strouse, D.; McKee, K.; Botvinick, M.; Hughes, E.; and Everett, R. 2021. Collaborating with Humans without Human Data. Advances in Neural Information Processing Systems, 34.
  44. 44.Tan, J.; Zhang, T.; Coumans, E.; Iscen, A.; Bai, Y.; Hafner, D.; Bohez, S.; and Vanhoucke, V. 2018. Sim-to-real: Learning agile locomotion for quadruped robots. arXiv preprint arXiv:1804.10332.
  45. 45.Tang, Z.; Yu, C.; Chen, B.; Xu, H.; Wang, X.; Fang, F.; Du, S. S.; Wang, Y.; and Wu, Y. 2020. Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization. In International Conference on Learning Representations.
  46. 46.Tesauro, G. 1994. TD-Gammon, a self-teaching backgammon program, achieves master-level play. Neural computation, 6(2): 215–219.
  47. 47.Tobin, J.; Fong, R.; Ray, A.; Schneider, J.; Zaremba, W.; and Abbeel, P. 2017. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), 23–30. IEEE.
  48. 48.Toussaint, M. 2009. Robot trajectory optimization using approximate inference. In Proceedings of the 26th annual international conference on machine learning, 1049–1056.
  49. 49.Tucker, M.; Zhou, Y.; and Shah, J. 2020. Adversarially Guided Self-Play for Adopting Social Conventions. arXiv preprint arXiv:2001.05994.
  50. 50.Vinyals, O.; Babuschkin, I.; Czarnecki, W. M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D. H.; Powell, R.; Ewalds, T.; Georgiev, P.; et al. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature, 575(7782): 350–354.
  51. 51.Zhao, R.; Gao, Y.; Abbeel, P.; Tresp, V.; and Xu, W. 2021. Mutual Information State Intrinsic Control. In International Conference on Learning Representations.
  52. 52.Zhao, R.; Sun, X.; and Tresp, V. 2019. Maximum Entropy-Regularized Multi-Goal Reinforcement Learning. In Proceedings of the 36th International Conference on Machine Learning, 7553–7562. PMLR.
  53. 53.Ziebart, B. D. 2010. Modeling purposeful adaptive behavior with the principle of maximum causal entropy. Carnegie Mellon University.
  54. 54.Ziebart, B. D.; Maas, A. L.; Bagnell, J. A.; Dey, A. K.; et al. 2008. Maximum entropy inverse reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 8, 1433–1438.

Citation

MLA
Zhao, R., et al. “Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination”. arXiv, 2021, http://arxiv.org/abs/2112.11701v3.
APA
Zhao, R., Song, J., Yuan, Y., Haifeng, H., Gao, Y., Wu, Y., Sun, Z., & Wei, Y. (2021). Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination. arXiv. http://arxiv.org/abs/2112.11701v3
Chicago
Zhao, R., J. Song, Y. Yuan, et al. 2021. “Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination”. arXiv. http://arxiv.org/abs/2112.11701v3.
Harvard
Zhao, R. et al. (2021) “Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.11701v3.
Vancouver
1. Zhao R, Song J, Yuan Y, Haifeng H, Gao Y, Wu Y, Sun Z, Wei Y (2021) Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination. arXiv

BibTeX

@article{zhao2021maximum,
  title = {Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination},
  author = {Zhao, Rui and Song, Jinming and Yuan, Yufeng and Haifeng, Hu and Gao, Yang and Wu, Yi and Sun, Zhongqian and Wei, Yang},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.11701v3},
  eprint = {2112.11701}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF