LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay
Yihuai LanZhiqiang HuLei WangYang WangDeheng YePeilin ZhaoEe-Peng LimHui XiongHao Wang
Presents a multi-agent framework using the social deduction game Avalon to systematically evaluate how large language model agents exhibit complex social behaviors such as deception, leadership, persuasion, and teamwork under incomplete information.
As artificial intelligence systems are increasingly deployed to automate decisions and simulate human environments, understanding how autonomous language agents interact in competitive and complex social settings has become critical. Most prior research focused narrowly on cooperative, positive interactions, leaving a substantial gap in understanding how autonomous agents manage negative or adversarial social dynamics such as deception, conflict, and persuasion. The article addresses this gap by investigating the strategic social behaviors of large language model agents within the multi-agent deduction game Avalon, evaluating how structured decision architectures enable agents to navigate incomplete information, collaborate with allies, and confront adversaries.
To study these dynamics, the authors designed a six-module architecture incorporating memory storage and summarization, environmental analysis, strategic planning, action selection, response generation, and iterative experience learning. The framework was evaluated primarily using OpenAI’s GPT-3.5 backend in six-player matches against established baseline agents across multiple 10-game series, testing both good and evil faction configurations. The investigation tracked quantitative performance metrics—such as win rates, quest participation, and voting patterns—alongside behavioral indicators evaluating leadership, persuasion, camouflage, teamwork, confrontation, and information sharing.
The findings show that the proposed multi-agent framework significantly outperforms the baseline. The designed agents achieved a 90% winning rate when playing as the good faction and a 100% winning rate as the evil faction against baseline opponents. Evil-faction agents were substantially more aggressive and effective, securing a 40.3% quest engagement rate and an 84.0% failure voting rate compared to 33.1% and 36.5% for the baseline. Ablation testing revealed that removing the analysis and strategy learning modules caused winning rates to drop by up to 30% to 50%, highlighting their essential role in understanding competitor intent. Furthermore, the agents demonstrated emergent social behaviors: evil roles spontaneously adopted deceptive camouflage 10% to 15% of the time without explicit instructions, while good roles achieved leader approval rates exceeding 80% and learned to selectively confront suspected adversaries as game rounds progressed.
These results indicate that structured reasoning and feedback mechanisms enable language models to autonomously execute sophisticated social strategies, including deception, covert sabotage, and alignment deduction. From an operational and risk standpoint, this demonstrates that language agents can effectively coordinate and manipulate information in multi-agent environments, posing new challenges for safety, alignment monitoring, and detecting autonomous deceptive behaviors in deployment. The framework successfully demonstrates that breaking down reasoning into discrete cognitive steps significantly enhances autonomous strategic performance compared to baseline prompting approaches.
Organizations seeking to implement autonomous multi-agent systems should integrate explicit analysis and learning loops into agent architectures to improve strategic robustness. However, decision-makers must consider trade-offs regarding computational cost and latency, as the multi-step module design requires frequent model querying. Additionally, experiments testing smaller open-source models like LLaMA-2 revealed a 25.1% drop in valid response compliance compared to GPT-3.5 (59.9% versus 85.0%), indicating that simpler models currently struggle with complex multi-step deductive rules. Key limitations include small game sample sizes, elevated operational costs, and occasional suboptimal behaviors such as excessive initial identity disclosure. Decision-makers should view these findings with moderate confidence as a viable proof-of-concept for strategic agent design while conducting broader pilot testing and safety audits before deploying similar architectures in high-stakes environments.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). This seminal paper introduces the core architecture of memory retrieval, reflection, and strategic planning that underpins modern LLM-based social agent simulations.
- Paper: Theory of Mind for Multi-Agent Collaboration via Large Language Models, Huao Li et al. (2023). It provides foundational insights into how language models perform Theory of Mind reasoning and mental-state attribution in cooperative text games.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). This work establishes the communicative role-playing paradigm that allows autonomous LLM agents to interact and coordinate without continuous human intervention.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). It synthesizes the architectural taxonomy—profiling, memory, planning, and action—essential for constructing modular LLM decision-making agents.
- Paper: Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments, Ryan Lowe et al. (2017). This foundational paper establishes key principles for modeling simultaneous cooperation and competition among autonomous agents in mixed environments.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). This paper critically evaluates the limitations and information-asymmetry challenges of LLM social simulations revealed by game environments like Avalon.
- Paper: Code World Models for General Game Playing, Wolfgang Lehrach et al. (2026). It advances beyond pure conversational prompting in imperfect-information games by having LLMs synthesize executable code world models paired with search algorithms.
- Paper: Towards a Science of Scaling Agent Systems, Yubin Kim et al. (2025). This study rigorously investigates the performance and coordination trade-offs of multi-agent architectures across varying task complexities and agent team structures.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). It provides a broad framework categorizing foundational planning, self-evolving reflection, and collective multi-agent interaction paradigms.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). It examines how post-trained reasoning models internalize multi-perspective social deliberation into their chain-of-thought processes.
