ProAgent: Building Proactive Cooperative Agents with Large Language Models

Ceyao ZhangKaijie YangSiyi HuZihao WangGuanghe LiYihang SunCheng ZhangZhaowei ZhangAnji LiuSong-Chun Zhu

article2024AAAI136 citations

Proposes an LLM-based cooperative agent framework that dynamically infers teammate intentions and updates beliefs without prior training, outperforming standard self-play and population-based methods in zero-shot coordination on Overcooked-AI.

Listen

Building artificial intelligence systems capable of collaborating smoothly with unfamiliar partners remains a critical challenge across multi-agent environments. Current systems depend heavily on learning-based approaches that require extensive training with diverse teammates. When deployed alongside novel artificial agents or human partners in zero-shot coordination scenarios—where agents must cooperate without prior joint experience—these traditional methods often fail to adapt their strategies, demand excessive computational resources for training, and operate as uninterpretable black boxes.

The article introduces and evaluates ProAgent, a modular framework that leverages large language models to create proactive, adaptive agents for cooperative multi-agent tasks without requiring task-specific training or fine-tuning. The framework aims to demonstrate that large language models can interpret environmental states, infer teammate intentions, and dynamically plan and revise cooperative actions in zero-shot settings.

To evaluate this approach, the researchers conducted extensive simulations in Overcooked-AI, a standard multi-agent benchmark requiring two partners to coordinate task allocation under spatial constraints. The framework was structured into four modular components: a Planner that reasons through intermediate steps to infer partner intentions, a Verificator that validates planned skills and provides failure feedback, a Controller that translates high-level skills into executable low-level movements, and a Memory module supported by a Belief Revision mechanism to update predictions against actual partner actions. ProAgent was evaluated across five classical kitchen layouts against five established zero-shot learning methods—including self-play and population-based training baselines—when paired with diverse AI partners and human proxy models.

The findings show that ProAgent systematically outperforms traditional baselines in multi-agent cooperation. When playing as Player 0 across all five test layouts, ProAgent consistently achieved the highest average scores among all evaluated methods, demonstrating strong zero-shot adaptability without prior joint training. When collaborating with human proxy models, ProAgent outperformed the baselines in four out of five environments, demonstrating an average performance improvement exceeding 10% over current state-of-the-art approaches. Furthermore, ablation experiments confirmed that inferring partner intentions significantly enhances coordination efficiency: direct planning without intention modeling reduced reward scores by over 50% in testing, while removing the failure verification module caused execution success rates to plummet from reliable levels down to 20%.

These results demonstrate that large language models can bridge the strategic coordination gap in multi-agent systems using common-sense reasoning and step-by-step intention inference. By replacing computationally intensive, brittle reinforcement learning pipelines with modular, language-grounded planning, organizations can deploy cooperative agents that adapt to unfamiliar partners in real time. Because the decision-making pipeline operates in natural language, it also provides clear operational interpretability, reducing systemic risks and simplifying compliance monitoring in safety-critical human-machine collaboration.

Organizations developing cooperative autonomous systems should consider modular, language-model-based coordination architectures rather than relying exclusively on population-based training. Before full-scale operational deployment, practitioners should conduct pilot studies paired with live human operators to confirm interaction dynamics beyond synthetic proxy models. Engineering teams should also integrate robust internal verification mechanisms, as the findings show that automated failure analysis is indispensable for preventing execution errors.

While the empirical results are robust across standardized benchmarks, the evaluation is subject to notable boundary conditions. Human collaboration was evaluated using behavior cloning proxy models rather than live, varied human subjects, and the low-level movement controller relied partly on simplified search rules. Confidence in the framework's high-level reasoning and zero-shot adaptability is high, but stakeholders should exercise caution when extrapolating these simulated findings to physical systems with high latency or complex sensory inputs.

  • Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This survey provides a comprehensive framework generalizing agentic reasoning, self-evolving mechanisms, and collective multi-agent collaboration pioneered by systems like ProAgent.
  • Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). It systematically investigates the scaling properties and trade-offs of multi-agent coordination architectures compared to single-agent baselines in complex agentic tasks.
  • Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). It extends hand-designed LLM agent frameworks like ProAgent by using meta-agents to automate the programmatic discovery and optimization of entire agentic systems.
  • Paper: Intelligent AI Delegation, Nenad Tomašev et al. (2026). It builds upon multi-agent coordination protocols to formalize dynamic authority transfer, intent clarity, and accountability in complex AI-to-AI and human-AI delegations.
  • Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). It advances dynamic agent adaptation by systematically evolving context playbooks and strategies from interaction trajectories without parameter updates.
  • Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). It investigates learning automated multi-agent orchestration directly in natural language using reinforcement learning over heterogeneous worker models.
Cover for ProAgent: Building Proactive Cooperative Agents with Large Language Models

Abstract

Building agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they interact with during the training phase. Such reliance, however, constrains the agents’ capacity for strategic adaptation when cooperating with unfamiliar teammates, which becomes a significant challenge in zero-shot coordination scenarios. To address this challenge, we propose ProAgent, a novel framework that harnesses large language models (LLMs) to create proactive agents capable of dynamically adapting their behavior to enhance cooperation with teammates. ProAgent can analyze the present state, and infer the intentions of teammates from observations. It then updates its beliefs in alignment with the teammates’ subsequent actual behaviors. Moreover, ProAgent exhibits a high degree of modularity and interpretability, making it easily integrated into various of coordination scenarios. Experimental evaluations conducted within the Overcooked-AI environment unveil the remarkable performance superiority of ProAgent, outperforming five methods based on self-play and population-based training when cooperating with AI agents. Furthermore, in partnered with human proxy models, its performance exhibits an average improvement exceeding 10% compared to the current state-of-the-art method. For more information about our project, please visit https://pku-proagent.github.io.

Table of Contents

  • Introduction
  • Related Works
  • Method
  • Prompt Construction
  • Cooperative Reasoning and Planning
  • Experiments
  • Experimental Settings
  • Collaborating with AI Agents
  • Collaborating with Humans
  • Discussion
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — ProAgent Modular Coordination Framework

    model/method

    ProAgent is a decentralized, training-free multi-agent framework that uses Large Language Models (LLMs) to achieve zero-shot coordination in cooperative environments. The architecture operates through a multi-stage pipeline without requiring fine-tuning or prior co-training with specific teammates:

    1. Knowledge Library and State Grounding: Translates the raw symbolic environment state into a structured, natural language representation containing kitchen layouts, agent inventories, and environmental objects, constrained by task instructions and skill definitions.
    2. LLM-Based Planner: Employs Chain-of-Thought (CoT) reasoning over the grounded state and historical trajectory memory to interpret the scene, infer the teammate's current intention, and select a high-level skill from a predefined candidate skill pool.
    3. Belief Revision: Compares previously inferred teammate intentions with their observed ground-truth behaviors to iteratively refine belief states and prevent repeated mispredictions.
    4. Verificator: Performs precondition checking on the generated skill before execution. If the skill violates environmental constraints, the Verificator produces a failure diagnosis and prompts the Planner to enter a re-planning loop.
    5. Controller: Decomposes validated high-level skills into atomic low-level actions (e.g., directional movement, object interaction) using search-based path planning with collision avoidance and deadlock resolution.
    6. Memory: Maintains persistent task specifications (Knowledge Library) alongside a temporal trajectory list storing recent grounded states, reasoning analyses, intention beliefs, and executed skills.
  2. Knowl 2 — Belief Revision Mechanism for Teammate Intention Tracking

    model/method

    The Belief Revision mechanism in ProAgent provides online adaptation to partner policies without parameter updates. During cooperative planning at time step tt, the Planner infers the intention (anticipated skill or goal) of the teammate agent and records it in the Memory trajectory. At subsequent decision steps, ProAgent observes the teammate's actual behavior and compares it against the recorded prediction.

    When a discrepancy is detected between assumed intentions and actual actions, the mechanism updates the agent's historical belief records by pairing the inferred intention with the ground-truth behavior. Under intention-augmented prompting (Level 3 prompts), this history of past predictions and verified ground-truth actions is appended to the context prompt for subsequent queries. This iterative feedback forces the LLM to learn from its mispredictions in-context, aligning its internal model of the partner with the partner's actual policy.

  3. Knowl 3 — Hierarchical Prompt Abstraction Levels in the Planner Module

    model/method

    The Planner module formats LLM decision-making using structured prompt templates composed of Instructions (task rules and objectives), Skills (legal skill definitions and arguments), and Examples (in-context demonstrations). Reasoning within the Planner is categorized into three hierarchical prompt levels:

    • Level 1 (L1) Prompting (Direct Planning): The LLM directly outputs a high-level skill from the skill pool given the current state description, without generating intermediate textual reasoning or partner modeling.
    • Level 2 (L2) Prompting (Analysis-Augmented Planning): The LLM is prompted via Chain-of-Thought (CoT) to first generate an explicit natural language analysis of the current state (e.g., status of cooking pots, agent inventories, pending orders) before selecting an action skill.
    • Level 3 (L3) Prompting (Intention-Augmented Planning): The LLM generates a state analysis, explicitly infers and outputs the predicted intention/sub-goal of the teammate, and selects its own skill conditioned on this inferred intention and historical belief revisions.
  4. Knowl 4 — Verificator Module for Closed-Loop Precondition Checking and Reflection

    model/method

    The Verificator module acts as an internal validation and reflection layer that scrutinizes planned high-level skills before environment execution. Its operation follows a multi-round interaction pipeline:

    1. Preconditions Check: Evaluates the planned skill against the preconditions of the current grounded state (expressed in natural language or pseudocode). If a precondition is violated (e.g., attempting to place an ingredient into a pot that is already full or cooking), the Verificator flags the plan as invalid.
    2. Failure Analysis: A trigger prompt (e.g., "Analysis of why I cannot execute this skill in the current scene step by step") queries the LLM to identify the specific violated precondition.
    3. Double-Check and Error Conclusion: Synthesizes the failure cause into an error explanation.
    4. Re-planning Loop: The error explanation is appended to the prompt context, prompting the Planner to generate an alternative, valid skill.
  5. Knowl 5 — Controller Module for Skill Grounding and Deadlock Avoidance

    model/method

    The Controller module translates high-level skills (e.g., pickup(onion), fill_dish(soup)) generated by the Planner into primitive action sequences executable by the environment simulator. In gridworld domains such as Overcooked-AI, the Controller uses search-based path planning augmented with dynamic coordination handling:

    • Dynamic Rerouting: When an agent's path toward a target object is blocked by the teammate, the Controller detects the obstruction and dynamically computes an alternative interconnected path around the obstacle.
    • Deadlock Resolution: When the Planner has no clear goal or when agents encounter symmetrical corridor deadlocks (often caused by rigid conventions learned by reinforcement learning partners), the Controller injects stochastic or random exploratory moves to unblock pathways and restore cooperative flow.
  6. Knowl 6 — Overcooked-AI Multi-Agent Coordination Experimental Setup

    experimental setup

    ProAgent is evaluated on the Overcooked-AI benchmark, where two agents must cooperate to prepare, cook, and deliver soups to maximize cumulative team reward. The evaluation spans five standard layouts:

    1. Cramped Room: A compact layout requiring mutual space sharing and turn-taking.
    2. Asymmetric Advantages: A layout featuring distinct sub-areas where agents have asymmetric efficiencies for gathering ingredients versus plating soups.
    3. Coordination Ring: A loop-shaped kitchen where agents must coordinate movement directions to avoid collisions in narrow corridors.
    4. Forced Coordination: A partitioned layout where ingredients, cooking pots, and serving counters are strictly segregated across a divider, requiring explicit item passing via a shared counter.
    5. Counter Circuit: A ring layout with a central island allowing multiple routing and item placement strategies.

    Evaluation involves 36 cross-play algorithm pairs constructed from ProAgent and five baseline methods: Self-Play (SP), Population-Based Training (PBT), Fictitious Co-Play (FCP), Maximum Entropy Population-Based Training (MEP), and Cooperative Open-ended Learning (COLE). Each pair is evaluated over 5 episodes per role assignment (Player 0 and Player 1) using Level 2 prompts and a recent-1 trajectory retrieval strategy. Human-AI coordination is tested using 5 distinct Behavior Cloning (BC) models trained on human gameplay data over 400 timesteps per episode.

  7. Knowl 7 — Cross-Play Performance Across Diverse AI Partner Algorithms

    data/table

    The performance of ProAgent in cross-play evaluations with heterogeneous AI partners across the five Overcooked-AI layouts is measured in average episode return (mean ±\pm standard error over 5 episodes). In each cell, the first row indicates the algorithm playing as Player 0 (paired across all other algorithms as Player 1), and the second row indicates the algorithm playing as Player 1 (paired across all others as Player 0):

    Layout SP PBT FCP MEP COLE ProAgent (ours)
    Cramped Room 168.5±15.2168.5 \pm 15.2 178.8±16.5178.8 \pm 16.5 196.3±16.8196.3 \pm 16.8 185.0±15.0185.0 \pm 15.0 163.8±24.1163.8 \pm 24.1 197.3±6.1\mathbf{197.3 \pm 6.1}
    172.8±16.1172.8 \pm 16.1 179.8±26.8179.8 \pm 26.8 196.0±11.9\mathbf{196.0 \pm 11.9} 178.2±15.6178.2 \pm 15.6 169.2±16.8169.2 \pm 16.8 194.2±10.5194.2 \pm 10.5
    Asymmetric Advantages 183.3±27.5183.3 \pm 27.5 182.2±27.9182.2 \pm 27.9 185.7±22.7185.7 \pm 22.7 155.7±63.9155.7 \pm 63.9 201.3±34.5201.3 \pm 34.5 228.7±23.0\mathbf{228.7 \pm 23.0}
    177.8±24.6177.8 \pm 24.6 152.3±64.5152.3 \pm 64.5 167.8±21.3167.8 \pm 21.3 184.0±41.8184.0 \pm 41.8 165.5±33.3165.5 \pm 33.3 229.8±21.9\mathbf{229.8 \pm 21.9}
    Coordination Ring 122.0±17.2122.0 \pm 17.2 141.3±28.0141.3 \pm 28.0 148.8±19.4148.8 \pm 19.4 167.2±22.4167.2 \pm 22.4 168.8±26.1168.8 \pm 26.1 175.3±29.0\mathbf{175.3 \pm 29.0}
    133.3±23.7133.3 \pm 23.7 141.3±27.5141.3 \pm 27.5 145.7±17.1145.7 \pm 17.1 159.3±25.3159.3 \pm 25.3 158.3±27.1158.3 \pm 27.1 183.0±31.7\mathbf{183.0 \pm 31.7}
    Forced Coordination 6.7±6.76.7 \pm 6.7 15.3±17.115.3 \pm 17.1 44.7±36.444.7 \pm 36.4 23.3±19.823.3 \pm 19.8 24.0±21.824.0 \pm 21.8 49.7±33.1\mathbf{49.7 \pm 33.1}
    30.2±21.930.2 \pm 21.9 61.7±46.0\mathbf{61.7 \pm 46.0} 32.2±30.232.2 \pm 30.2 39.3±16.939.3 \pm 16.9 57.3±36.457.3 \pm 36.4 31.0±33.931.0 \pm 33.9
    Counter Circuit 64.7±45.864.7 \pm 45.8 64.7±45.964.7 \pm 45.9 58.3±37.558.3 \pm 37.5 74.3±39.174.3 \pm 39.1 95.5±25.295.5 \pm 25.2 126.3±32.3\mathbf{126.3 \pm 32.3}
    60.7±40.860.7 \pm 40.8 54.3±49.154.3 \pm 49.1 60.0±38.360.0 \pm 38.3 81.5±27.581.5 \pm 27.5 100.8±31.1100.8 \pm 31.1 128.5±28.1\mathbf{128.5 \pm 28.1}

    These cross-play results establish that ProAgent achieves the highest average cooperative return across all five layouts when acting as Player 0, and achieves the highest return in three of five layouts as Player 1, outperforming population-trained and self-play reinforcement learning baselines without receiving layout-specific or partner-specific training.

  8. Knowl 8 — Zero-Shot Coordination with Behavior Cloning Human Proxies

    empirical result

    When evaluated in partnership with five distinct Behavior Cloning (BC) human proxy models over 400 timesteps across five Overcooked-AI layouts, ProAgent exhibits the following performance characteristics:

    • ProAgent outperforms all baselines (Self-Play, Population-Based Training, Fictitious Co-Play, Maximum Entropy PBT, and Cooperative Open-ended Learning) in four out of five kitchen layouts, achieving an average performance improvement exceeding 10% compared to the strongest baseline (COLE).
    • In the Forced Coordination layout, ProAgent achieves substantial gains over all baselines when operating in the Player 0 position.
    • ProAgent displays positional invariance: switching between left and right starting positions yields negligible differences in ProAgent's cumulative rewards. In contrast, all learning-based baseline algorithms suffer substantial performance drops when initialized in the right position compared to the left position on asymmetric layouts.
  9. Knowl 9 — Ablation of Reasoning and Intention Levels in Planning

    empirical result

    An ablation study evaluating the influence of Chain-of-Thought state analysis and teammate intention modeling within the Planner module was conducted on the Cramped Room layout in Overcooked-AI:

    • Level 1 Prompts (Direct Planning): Outputting a high-level skill directly without state analysis or teammate intention prediction achieved an average return of 100100.
    • Level 2 Prompts (Analysis Only): Introducing textual scene analysis prior to skill generation increased the average return to 184184 (an 84.0%84.0\% improvement over Level 1).
    • Level 3 Prompts (Analysis and Intention Prediction): Incorporating both textual scene analysis and explicit teammate intention inference with belief revision achieved an average return of 204204 (a 10.9%10.9\% improvement over Level 2 and a 104.0%104.0\% improvement over Level 1).

    These results demonstrate that structured intermediate Chain-of-Thought analysis provides essential context for multi-agent decision-making, and explicitly predicting partner intentions further enhances coordination efficiency.

  10. Knowl 10 — Ablation of the Verificator Module in Feedback-Driven Reasoning

    empirical result

    In an ablation experiment where the Verificator module was removed—forcing ProAgent to execute high-level skills generated by the Planner open-loop without precondition checks or reflection feedback—the skill execution success rate over 100 interaction steps dropped from normal operation down to 20%20\%. This confirms that real-time precondition verification and failure-driven re-planning loops are critical for correcting LLM reasoning errors and ensuring valid action selection in dynamic environments.

Coverage note — All major contributions—the ProAgent architecture, individual modules (Planner, Verificator, Controller, Memory, Belief Revision), experimental evaluations across AI-AI and Human-AI setups, and ablation studies—have been fully extracted. Standard related work descriptions of RL baselines and specific prompt texts were omitted.

References

  1. 1.Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Fu, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Ho, D.; Hsu, J.; Ibarz, J.; Ichter, B.; Irpan, A.; Jang, E.; Ruano, R. J.; Jeffrey, K.; Jesmonth, S.; Joshi, N. J.; Julian, R.; Kalashnikov, D.; Kuang, Y.; Lee, K.-H.; Levine, S.; Lu, Y.; Luu, L.; Parada, C.; Pastor, P.; Quiambao, J.; Rao, K.; Rettinghouse, J.; Reyes, D.; Sermanet, P.; Sievers, N.; Tan, C.; Toshev, A.; Vanhoucke, V.; Xia, F.; Xiao, T.; Xu, P.; Xu, S.; Yan, M.; and Zeng, A. 2022. Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. arXiv:2204.01691.
  2. 2.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language Models are Few-Shot Learners. In Advances in neural information processing systems, volume 33, 1877–1901.
  3. 3.Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; Nori, H.; Palangi, H.; Ribeiro, M. T.; and Zhang, Y. 2023. Sparks of Artificial General Intelligence: Early Experiments with GPT-4. arXiv:2303.12712.
  4. 4.Carroll, M.; Shah, R.; Ho, M. K.; Griffiths, T.; Seshia, S.; Abbeel, P.; and Dragan, A. 2019. On the Utility of Learning about Humans for Human-AI Coordination. In Advances in neural information processing systems, volume 32.
  5. 5.Ding, Z.; Zhang, W.; Yue, J.; Wang, X.; Huang, T.; and Lu, Z. 2023. Entity Divider with Language Grounding in Multi-Agent Reinforcement Learning. In International Conference on Machine Learning, 8103–8119. PMLR.
  6. 6.Du, Y.; Watkins, O.; Wang, Z.; Colas, C.; Darrell, T.; Abbeel, P.; Gupta, A.; and Andreas, J. 2023. Guiding Pretraining in Reinforcement Learning with Large Language Models. arXiv:2302.06692.
  7. 7.Fan, L.; Wang, G.; Jiang, Y.; Mandlekar, A.; Yang, Y.; Zhu, H.; Tang, A.; Huang, D.-A.; Zhu, Y.; and Anandkumar, A. 2022. MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge. In NIPS Processing Systems Datasets and Benchmarks Track.
  8. 8.Gronauer, S.; and Diepold, K. 2022. Multi-Agent Deep Reinforcement Learning: A survey. Artificial Intelligence Review, 1–49.
  9. 9.Hanjie, A. W.; Zhong, V. Y.; and Narasimhan, K. 2021. Grounding Language to Entities and Dynamics for Generalization in Reinforcement Learning. In International Conference on Machine Learning, 4051–4062. PMLR.
  10. 10.Hao, S.; Gu, Y.; Ma, H.; Hong, J. J.; Wang, Z.; Wang, D. Z.; and Hu, Z. 2023. Reasoning with Language Model is Planning with World Model. arXiv:2305.14992.
  11. 11.Hu, H.; and Foerster, J. N. 2020. Simplified Action Decoder for Deep Multi-Agent Reinforcement Learning. In International Conference on Learning Representations.
  12. 12.Hu, H.; Lerer, A.; Cui, B.; Pineda, L.; Brown, N.; and Foerster, J. 2021a. Off-Belief Learning. In International Conference on Machine Learning, 4369–4379. PMLR.
  13. 13.Hu, H.; and Sadigh, D. 2023. Language Instructed Reinforcement Learning for Human-AI Coordination. In Proceedings of the 40th International Conference on Machine Learning. PMLR.
  14. 14.Hu, S.; Zhu, F.; Chang, X.; and Liang, X. 2021b. UPDeT: Universal Multi-agent Reinforcement Learning via Policy Decoupling with Transformers. arXiv:2101.08001.
  15. 15.Huang, J.; and Chang, K. C.-C. 2023. Towards Reasoning in Large Language Models: A Survey. arXiv:2212.10403.
  16. 16.Jaderberg, M.; Dalibard, V.; Osindero, S.; Czarnecki, W. M.; Donahue, J.; Razavi, A.; Vinyals, O.; Green, T.; Dunning, I.; Simonyan, K.; Fernando, C.; and Kavukcuoglu, K. 2017. Population Based Training of Neural Networks. arXiv:1711.09846.
  17. 17.Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022. Large Language Models are Zero-Shot Reasoners. In Advances in neural information processing systems, volume 35, 22199–22213.
  18. 18.Li, W.; Qiao, D.; Wang, B.; Wang, X.; Jin, B.; and Zha, H. 2023a. Semantically Aligned Task Decomposition in Multi-Agent Reinforcement Learning. arXiv:2305.10865.
  19. 19.Li, Y.; Zhang, S.; Sun, J.; Du, Y.; Wen, Y.; Wang, X.; and Pan, W. 2023b. Cooperative Open-ended Learning Framework for Zero-shot Coordination. In Proceedings of the 40th International Conference on Machine Learning. PMLR.
  20. 20.Li, Y.; Zhang, S.; Sun, J.; Zhang, W.; Du, Y.; Wen, Y.; Wang, X.; and Pan, W. 2024. Tackling Cooperative Incompatibility for Zero-Shot Human-AI Coordination. arXiv:2306.03034.
  21. 21.Liang, J.; Huang, W.; Xia, F.; Xu, P.; Hausman, K.; Ichter, B.; Florence, P.; and Zeng, A. 2023. Code as Policies: Language Model Programs for Embodied Control. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 9493–9500. IEEE.
  22. 22.Lucas, K.; and Allen, R. E. 2022. Any-Play: An Intrinsic Augmentation for Zero-Shot Coordination. In International Foundation for Autonomous Agents and Multiagent Systems, 853–861.
  23. 23.Lupu, A.; Cui, B.; Hu, H.; and Foerster, J. 2021. Trajectory Diversity for Zero-Shot Coordination. In International conference on machine learning, 7204–7213. PMLR.
  24. 24.Meng, L.; Wen, M.; Yang, Y.; Le, C.; Li, X.; Zhang, W.; Wen, Y.; Zhang, H.; Wang, J.; and Xu, B. 2022. Offline Pre-trained Multi-Agent Decision Transformer: One Big Sequence Model Tackles All SMAC Tasks. arXiv:2112.02845.
  25. 25.Mialon, G.; Dessì, R.; Lomeli, M.; Nalmpantis, C.; Pasunuru, R.; Raileanu, R.; Roziere, B.; Schick, T.; Dwivedi-Yu, J.; Celikyilmaz, A.; Grave, E.; LeCun, Y.; and Scialom, T. 2023. Augmented Language Models: a Survey. arXiv:2302.07842.
  26. 26.Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P. F.; Leike, J.; and Lowe, R. 2022. Training Language Models to Follow Instructions with Human Feedback. In Advances in Neural Information Processing Systems, volume 35, 27730–27744.
  27. 27.Paul, D.; Ismayilzada, M.; Peyrard, M.; Borges, B.; Bosselut, A.; West, R.; and Faltings, B. 2023. REFINER: Reasoning Feedback on Intermediate Representations. arXiv:2304.01904.
  28. 28.Rashid, T.; Samvelyan, M.; Schroeder, C.; Farquhar, G.; Foerster, J.; and Whiteson, S. 2018. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. In International Conference on Machine Learning, 4295–4304. PMLR.
  29. 29.Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K. R.; and Yao, S. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. In Thirty-seventh Conference on Neural Information Processing Systems.
  30. 30.Singh, I.; Blukis, V.; Mousavian, A.; Goyal, A.; Xu, D.; Tremblay, J.; Fox, D.; Thomason, J.; and Garg, A. 2023. ProgPrompt: Generating Situated Robot Task Plans using Large Language Models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 11523–11530. IEEE.
  31. 31.Strouse, D.; McKee, K.; Botvinick, M.; Hughes, E.; and Everett, R. 2021. Collaborating with Humans without Human Data. In Advances in Neural Information Processing Systems, volume 34, 14502–14515.
  32. 32.Tesauro, G. 1994. TD-Gammon, a self-teaching backgammon program, achieves master-level play. Neural computation, 6(2): 215–219.
  33. 33.Wang, X.; Wei, J.; Schuurmans, D.; Le, Q. V.; Chi, E. H.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023a. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Representations.
  34. 34.Wang, Y.; Zhong, F.; Xu, J.; and Wang, Y. 2021. ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of Mind. In International Conference on Learning Representations.
  35. 35.Wang, Z.; Cai, S.; Chen, G.; Liu, A.; Ma, X.; and Liang, Y. 2023b. Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task Agents. In Thirty-seventh Conference on Neural Information Processing Systems.
  36. 36.Wang, Z.; Cai, S.; Liu, A.; Jin, Y.; Hou, J.; Zhang, B.; Lin, H.; He, Z.; Zheng, Z.; Yang, Y.; Ma, X.; and Liang, Y. 2023c. JARVIS-1: Open-World Multi-Task Agents with Memory-Augmented Multimodal Language Models. arXiv:2311.05997.
  37. 37.Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; brian ichter; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems, volume 35, 24824–24837.
  38. 38.Welleck, S.; Lu, X.; West, P.; Brahman, F.; Shen, T.; Khashabi, D.; and Choi, Y. 2023. Generating Sequences by Learning to Self-Correct. In The Eleventh International Conference on Learning Representations.
  39. 39.Wen, M.; Kuba, J.; Lin, R.; Zhang, W.; Wen, Y.; Wang, J.; and Yang, Y. 2022. Multi-Agent Reinforcement Learning is a Sequence Modeling Problem. In Advances in Neural Information Processing Systems, volume 35, 16509–16521.
  40. 40.Wu, S. A.; Wang, R. E.; Evans, J. A.; Tenenbaum, J. B.; Parkes, D. C.; and Kleiman-Weiner, M. 2021. Too Many Cooks: Bayesian Inference for Coordinating Multi-Agent Collaboration. Topics in Cognitive Science, 13(2): 414–432.
  41. 41.Yang, Y.; and Wang, J. 2021. An Overview of Multi-Agent Reinforcement Learning from Game Theoretical Perspective. arXiv:2011.00583.
  42. 42.Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K. R.; and Cao, Y. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations.
  43. 43.Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; and Wu, Y. 2022. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Advances in Neural Information Processing Systems, volume 35, 24611–24624.
  44. 44.Zhang, H.; Du, W.; Shan, J.; Zhou, Q.; Du, Y.; Tenenbaum, J. B.; Shu, T.; and Gan, C. 2023. Building Cooperative Embodied Agents Modularly with Large Language Models. arXiv:2307.02485.
  45. 45.Zhang, K.; Yang, Z.; and Başar, T. 2021. Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms. Handbook of reinforcement learning and control, 321–384.
  46. 46.Zhao, R.; Song, J.; Yuan, Y.; Hu, H.; Gao, Y.; Wu, Y.; Sun, Z.; and Yang, W. 2023. Maximum Entropy Population Based Training for Zero-Shot Human-AI Coordination. In Proceedings of the AAAI Conference on Artificial Intelligence, 5, 6145–6153.
  47. 47.Zhong, Y.; Kuba, J. G.; Feng, X.; Hu, S.; Ji, J.; and Yang, Y. 2023. Heterogeneous-Agent Reinforcement Learning. arXiv:2304.09870.
  48. 48.Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; Scales, N.; Wang, X.; Schuurmans, D.; Cui, C.; Bousquet, O.; Le, Q. V.; and Chi, E. H. 2023. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. In The Eleventh International Conference on Learning Representations.

Citation

MLA
Zhang, C., et al. “ProAgent: Building Proactive Cooperative Agents with Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2308.11339v3.
APA
Zhang, C., Yang, K., Hu, S., Wang, Z., Li, G., Sun, Y., Zhang, C., Zhang, Z., Liu, A., Zhu, S.-C., Chang, X., Zhang, J., Yin, F., Liang, Y., & Yang, Y. (2023). ProAgent: Building Proactive Cooperative Agents with Large Language Models. arXiv. http://arxiv.org/abs/2308.11339v3
Chicago
Zhang, C., K. Yang, S. Hu, et al. 2023. “ProAgent: Building Proactive Cooperative Agents with Large Language Models”. arXiv. http://arxiv.org/abs/2308.11339v3.
Harvard
Zhang, C. et al. (2023) “ProAgent: Building Proactive Cooperative Agents with Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.11339v3.
Vancouver
1. Zhang C, Yang K, Hu S, et al (2023) ProAgent: Building Proactive Cooperative Agents with Large Language Models. arXiv

BibTeX

@article{zhang2023proagent,
  title = {ProAgent: Building Proactive Cooperative Agents with Large Language Models},
  author = {Zhang, Ceyao and Yang, Kaijie and Hu, Siyi and Wang, Zihao and Li, Guanghe and Sun, Yihang and Zhang, Cheng and Zhang, Zhaowei and Liu, Anji and Zhu, Song-Chun and Chang, Xiaojun and Zhang, Junge and Yin, Feng and Liang, Yitao and Yang, Yaodong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.11339v3},
  eprint = {2308.11339}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF