ProAgent: Building Proactive Cooperative Agents with Large Language Models
Ceyao ZhangKaijie YangSiyi HuZihao WangGuanghe LiYihang SunCheng ZhangZhaowei ZhangAnji LiuSong-Chun Zhu
Proposes an LLM-based cooperative agent framework that dynamically infers teammate intentions and updates beliefs without prior training, outperforming standard self-play and population-based methods in zero-shot coordination on Overcooked-AI.
Building artificial intelligence systems capable of collaborating smoothly with unfamiliar partners remains a critical challenge across multi-agent environments. Current systems depend heavily on learning-based approaches that require extensive training with diverse teammates. When deployed alongside novel artificial agents or human partners in zero-shot coordination scenarios—where agents must cooperate without prior joint experience—these traditional methods often fail to adapt their strategies, demand excessive computational resources for training, and operate as uninterpretable black boxes.
The article introduces and evaluates ProAgent, a modular framework that leverages large language models to create proactive, adaptive agents for cooperative multi-agent tasks without requiring task-specific training or fine-tuning. The framework aims to demonstrate that large language models can interpret environmental states, infer teammate intentions, and dynamically plan and revise cooperative actions in zero-shot settings.
To evaluate this approach, the researchers conducted extensive simulations in Overcooked-AI, a standard multi-agent benchmark requiring two partners to coordinate task allocation under spatial constraints. The framework was structured into four modular components: a Planner that reasons through intermediate steps to infer partner intentions, a Verificator that validates planned skills and provides failure feedback, a Controller that translates high-level skills into executable low-level movements, and a Memory module supported by a Belief Revision mechanism to update predictions against actual partner actions. ProAgent was evaluated across five classical kitchen layouts against five established zero-shot learning methods—including self-play and population-based training baselines—when paired with diverse AI partners and human proxy models.
The findings show that ProAgent systematically outperforms traditional baselines in multi-agent cooperation. When playing as Player 0 across all five test layouts, ProAgent consistently achieved the highest average scores among all evaluated methods, demonstrating strong zero-shot adaptability without prior joint training. When collaborating with human proxy models, ProAgent outperformed the baselines in four out of five environments, demonstrating an average performance improvement exceeding 10% over current state-of-the-art approaches. Furthermore, ablation experiments confirmed that inferring partner intentions significantly enhances coordination efficiency: direct planning without intention modeling reduced reward scores by over 50% in testing, while removing the failure verification module caused execution success rates to plummet from reliable levels down to 20%.
These results demonstrate that large language models can bridge the strategic coordination gap in multi-agent systems using common-sense reasoning and step-by-step intention inference. By replacing computationally intensive, brittle reinforcement learning pipelines with modular, language-grounded planning, organizations can deploy cooperative agents that adapt to unfamiliar partners in real time. Because the decision-making pipeline operates in natural language, it also provides clear operational interpretability, reducing systemic risks and simplifying compliance monitoring in safety-critical human-machine collaboration.
Organizations developing cooperative autonomous systems should consider modular, language-model-based coordination architectures rather than relying exclusively on population-based training. Before full-scale operational deployment, practitioners should conduct pilot studies paired with live human operators to confirm interaction dynamics beyond synthetic proxy models. Engineering teams should also integrate robust internal verification mechanisms, as the findings show that automated failure analysis is indispensable for preventing execution errors.
While the empirical results are robust across standardized benchmarks, the evaluation is subject to notable boundary conditions. Human collaboration was evaluated using behavior cloning proxy models rather than live, varied human subjects, and the low-level movement controller relied partly on simplified search rules. Confidence in the framework's high-level reasoning and zero-shot adaptability is high, but stakeholders should exercise caution when extrapolating these simulated findings to physical systems with high latency or complex sensory inputs.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). This comprehensive survey outlines the foundational brain, perception, and action architectures for LLM-based autonomous agents that ProAgent adapts for multi-agent coordination.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). It establishes the core architectural modules (memory, planning, and profiling) underlying LLM autonomous agents that ProAgent extends into proactive, belief-updating team environments.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). This paper presents multi-agent LLM conversation programming, establishing key concepts of dynamic multi-agent interaction that ProAgent builds upon for zero-shot cooperative coordination.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). It demonstrates how LLM-based agents maintain memory, synthesize reflections, and execute plans in social settings, providing the grounding for ProAgent's belief-updating and intention inference.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). This work introduces verbal reinforcement learning and self-reflection in LLM agents, an essential mechanism ProAgent leverages for dynamic behavioral adaptation and belief updating.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). It introduces the ReAct paradigm combining reasoning traces with environment actions, which provides the foundational decision-making loop used by ProAgent.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). It establishes closed-loop embodied planning via environmental feedback in language models, directly informing ProAgent's interactive state analysis and adaptation.
- Paper: Cooperative Multi-Agent Learning: The State of the Art, Liviu Panait et al. (2005). This survey provides fundamental background on cooperative multi-agent learning and the non-stationarity challenges of interacting with unfamiliar teammates that ProAgent aims to solve.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This survey provides a comprehensive framework generalizing agentic reasoning, self-evolving mechanisms, and collective multi-agent collaboration pioneered by systems like ProAgent.
- Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). It systematically investigates the scaling properties and trade-offs of multi-agent coordination architectures compared to single-agent baselines in complex agentic tasks.
- Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). It extends hand-designed LLM agent frameworks like ProAgent by using meta-agents to automate the programmatic discovery and optimization of entire agentic systems.
- Paper: Intelligent AI Delegation, Nenad Tomašev et al. (2026). It builds upon multi-agent coordination protocols to formalize dynamic authority transfer, intent clarity, and accountability in complex AI-to-AI and human-AI delegations.
- Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). It advances dynamic agent adaptation by systematically evolving context playbooks and strategies from interaction trajectories without parameter updates.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). It investigates learning automated multi-agent orchestration directly in natural language using reinforcement learning over heterogeneous worker models.
