LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay

Yihuai LanZhiqiang HuLei WangYang WangDeheng YePeilin ZhaoEe-Peng LimHui XiongHao Wang

article2024EMNLP67 citations

Presents a multi-agent framework using the social deduction game Avalon to systematically evaluate how large language model agents exhibit complex social behaviors such as deception, leadership, persuasion, and teamwork under incomplete information.

Listen

As artificial intelligence systems are increasingly deployed to automate decisions and simulate human environments, understanding how autonomous language agents interact in competitive and complex social settings has become critical. Most prior research focused narrowly on cooperative, positive interactions, leaving a substantial gap in understanding how autonomous agents manage negative or adversarial social dynamics such as deception, conflict, and persuasion. The article addresses this gap by investigating the strategic social behaviors of large language model agents within the multi-agent deduction game Avalon, evaluating how structured decision architectures enable agents to navigate incomplete information, collaborate with allies, and confront adversaries.

To study these dynamics, the authors designed a six-module architecture incorporating memory storage and summarization, environmental analysis, strategic planning, action selection, response generation, and iterative experience learning. The framework was evaluated primarily using OpenAI’s GPT-3.5 backend in six-player matches against established baseline agents across multiple 10-game series, testing both good and evil faction configurations. The investigation tracked quantitative performance metrics—such as win rates, quest participation, and voting patterns—alongside behavioral indicators evaluating leadership, persuasion, camouflage, teamwork, confrontation, and information sharing.

The findings show that the proposed multi-agent framework significantly outperforms the baseline. The designed agents achieved a 90% winning rate when playing as the good faction and a 100% winning rate as the evil faction against baseline opponents. Evil-faction agents were substantially more aggressive and effective, securing a 40.3% quest engagement rate and an 84.0% failure voting rate compared to 33.1% and 36.5% for the baseline. Ablation testing revealed that removing the analysis and strategy learning modules caused winning rates to drop by up to 30% to 50%, highlighting their essential role in understanding competitor intent. Furthermore, the agents demonstrated emergent social behaviors: evil roles spontaneously adopted deceptive camouflage 10% to 15% of the time without explicit instructions, while good roles achieved leader approval rates exceeding 80% and learned to selectively confront suspected adversaries as game rounds progressed.

These results indicate that structured reasoning and feedback mechanisms enable language models to autonomously execute sophisticated social strategies, including deception, covert sabotage, and alignment deduction. From an operational and risk standpoint, this demonstrates that language agents can effectively coordinate and manipulate information in multi-agent environments, posing new challenges for safety, alignment monitoring, and detecting autonomous deceptive behaviors in deployment. The framework successfully demonstrates that breaking down reasoning into discrete cognitive steps significantly enhances autonomous strategic performance compared to baseline prompting approaches.

Organizations seeking to implement autonomous multi-agent systems should integrate explicit analysis and learning loops into agent architectures to improve strategic robustness. However, decision-makers must consider trade-offs regarding computational cost and latency, as the multi-step module design requires frequent model querying. Additionally, experiments testing smaller open-source models like LLaMA-2 revealed a 25.1% drop in valid response compliance compared to GPT-3.5 (59.9% versus 85.0%), indicating that simpler models currently struggle with complex multi-step deductive rules. Key limitations include small game sample sizes, elevated operational costs, and occasional suboptimal behaviors such as excessive initial identity disclosure. Decision-makers should view these findings with moderate confidence as a viable proof-of-concept for strategic agent design while conducting broader pilot testing and safety audits before deploying similar architectures in high-stakes environments.

Cover for LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay

Abstract

This paper explores the open research problem of understanding the social behaviors of LLM-based agents. Using Avalon as a testbed, we employ system prompts to guide LLM agents in gameplay. While previous studies have touched on gameplay with LLM agents, research on their social behaviors is lacking. We propose a novel framework, tailored for Avalon, features a multi-agent system facilitating efficient communication and interaction. We evaluate its performance based on game success and analyze LLM agents’ social behaviors. Results affirm the framework’s effectiveness in creating adaptive agents and suggest LLM-based agents’ potential in navigating dynamic social interactions. By examining collaboration and confrontation behaviors, we offer insights into this field’s research and applications. Our code is publicly available at https://github.com/3DAgentWorld/LLM-Game-Agent.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Social Deduction Game Agent
  • 2.2 LLM-Based Gameplay
  • 2.3 LLMs' Impact on Society
  • 3 Background
  • 3.1 Social Behaviors in Avalon
  • 4 Approach
  • 4.1 Setup
  • 4.2 Memory Storage
  • 4.3 Memory Summarization.
  • 4.4 Analysis
  • 4.5 Planning
  • 4.6 Action
  • 4.7 Response Generation
  • 4.8 Experience Learning
  • 4.8.1 Self-Role Strategy Learning
  • 4.8.2 Other-Role Strategy Learning
  • 5 Experiment
  • 5.1 Implementation Details
  • 5.2 Evaluation Metrics
  • 5.2.1 Gameplay Outcome and Strategy.
  • 5.2.2 Social Behaviors.
  • 5.3 Experiment Results
  • 6 Social Behaviors of AI Agents
  • 6.1 Leadership
  • 6.2 Persuasion
  • 6.3 Camouflage
  • 6.4 Teamwork and Confrontation
  • 6.5 Sharing
  • 6.6 Vacillation
  • 6.7 Behavior Spontaneity
  • 7 Conclusion
  • 8 Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Avalon Introduction
  • A.2 Game Rules and Role Description
  • A.3 Module Prompts
  • A.4 Heuristic Rules for LLM Gameplay
  • A.5 Ablation Study
  • B Case Study
  • C Exploration on LLaMA-Based Agents
  • D Teamwork and Confrontation

Knowls

  1. Knowl 1 — Multi-Module Cognitive Architecture for LLM Avalon Agents

    model/method

    The multi-agent Avalon framework structures LLM agent reasoning into six specialized modules operating in sequence per game turn:

    1. Memory Storage and Summarization: Maintains separate private memory pools for each player pip_i. Conversations are stored as structured objects containing the speaker name, message content, round number, and a visibility flag (public or private). To manage context window limits, round memories are compressed via summarization prompting:

    Mt=⟨SMR(Mt−1),(Rtp1,…,Rtp6,It)⟩M_t = \langle \text{SMR}(M_{t-1}), (R^{p_1}_t, \dots, R^{p_6}_t, I_t) \rangle

    where MtM_t is the memory state at round tt, SMR(⋅)\text{SMR}(\cdot) denotes LLM summarization of the prior round's memory Mt−1M_{t-1}, RtpiR^{p_i}_t is the text response from player pip_i at round tt, ItI_t represents moderator instructions, and ⟨⋅⟩\langle \cdot \rangle denotes concatenation.

    1. Situation Analysis Module: Analyzes player identities, hidden allegiances, and potential adversary strategies in under 100 words:

    Htpi=ANA(Mt,RIpi)H^{p_i}_t = \text{ANA}(M_t, RI^{p_i})

    where HtpiH^{p_i}_t is the generated analysis and RIpiRI^{p_i} is agent pip_i's private role description.

    1. Planning Module: Synthesizes game history, situational analysis, and high-level goals into an actionable strategic plan:

    Ptpi=PLAN(Mt,Htpi,Pt−1pi,RIpi,Gpi,Spi)P^{p_i}_t = \text{PLAN}(M_t, H^{p_i}_t, P^{p_i}_{t-1}, RI^{p_i}, G^{p_i}, S^{p_i})

    where PtpiP^{p_i}_t is the updated plan, Pt−1piP^{p_i}_{t-1} is the previous plan, GpiG^{p_i} is the faction winning condition, and SpiS^{p_i} is the role's abstracted strategy.

    1. Action Module: Determines confidential game actions from five allowed discrete categories (nominating quest members, voting agree/disagree on proposals, submitting success/failure quest cards, non-verbal signals, or remaining silent):

    Atpi∼p(A∣Mt,Htpi,Ptpi,RIpi,Gpi,Spi,It′)A^{p_i}_t \sim p(A \mid M_t, H^{p_i}_t, P^{p_i}_t, RI^{p_i}, G^{p_i}, S^{p_i}, I'_t)

    where It′I'_t is the direct host action prompt.

    1. Response Generation Module: Produces natural language dialogue (capped at 100 words) public to other players, providing rationales aligned with the chosen action and permitted deceptive stances.

    2. Experience Learning Module: Updates long-term strategies across games.

  2. Knowl 2 — Cross-Game Strategy Evolution via Experience Learning

    model/method

    The experience learning module allows LLM agents to refine gameplay strategies across consecutive Avalon games by analyzing completed game logs:

    1. Self-Role Strategy Learning:
    • Step 1 (Recommendation Generation): Based on full-game logs, current strategy, and faction goals, the agent prompts the LLM to generate three strategic recommendations (each ≤2\le 2 sentences). Prompts strictly mandate referring to roles by role names (e.g., Merlin, Morgana) rather than specific player IDs to ensure generalizability.
    • Step 2 (Strategy Synthesis): The agent incorporates these suggestions into its existing strategy, synthesizing a revised strategy constrained to no more than two continuous sentences without bullet points or numbering, retaining past advantages.
    1. Other-Role Strategy Learning:
    • The agent reviews the post-game role-to-player mappings and summarizes the observed behavioral patterns and strategies of other roles in ≤100\le 100 words. These summaries are accumulated across games to enhance future situation analysis (HtpiH^{p_i}_t) and opponent modeling.
  3. Knowl 3 — Avalon Gameplay Evaluation Protocol and Heuristic Decision Rules

    experimental setup

    Experiments evaluate a 6-player Avalon setup: Good faction (Merlin, Percival, two Loyal Servants) versus Evil faction (Morgana, Assassin). Matches are executed using gpt-3.5-turbo-16k as the backend LLM with a sampling temperature of 0.30.3 for agent generation and 0.00.0 for structured output extraction.

    To ensure game execution when agent text outputs are ambiguous or non-compliant, deterministic heuristic fallback rules are applied:

    • Ambiguous Team Approval Vote: If an agent's response lacks a clear agree/disagree stance during team nomination voting, the vote defaults to agree (True).
    • Ambiguous Quest Card Submission: If an agent's quest execution choice is unclear, it defaults to a failure vote (False).
    • Candidate Over-Selection: If an agent proposes more players than required for a quest, the candidate list is truncated to the required team size.
    • Candidate Under-Selection: If too few players are proposed, the moderator reprompts the agent; if the count remains unmet after multiple retries, random candidates are assigned.

    Performance is evaluated using five primary metrics:

    WR=(#Wins#Games Played)×100%WR = \left(\frac{\#\text{Wins}}{\#\text{Games Played}}\right) \times 100\%

    QER=(#Engagement Rounds#Rounds)×100%QER = \left(\frac{\#\text{Engagement Rounds}}{\#\text{Rounds}}\right) \times 100\%

    FVR=(#Failure Votes#Votes)×100%FVR = \left(\frac{\#\text{Failure Votes}}{\#\text{Votes}}\right) \times 100\%

    LAR=(#Approval Votes on Proposed Teams#Total Leader Votes)×100%LAR = \left(\frac{\#\text{Approval Votes on Proposed Teams}}{\#\text{Total Leader Votes}}\right) \times 100\%

    VRR=(#Valid Responses#Total Responses)×100%VRR = \left(\frac{\#\text{Valid Responses}}{\#\text{Total Responses}}\right) \times 100\%

  4. Knowl 4 — Gameplay Win Rates and Module Ablation Performance

    data/table

    Across 10-game head-to-head matches against a competitive baseline agent (repurposed from Werewolf AI), the proposed framework achieves superior win rates on both factions. Ablation experiments demonstrate the critical contribution of each cognitive module:

    Method Good Side WR (%) Evil Side WR (%)
    Ours (Full Framework) 90 100
    w/o analysis 60 60
    w/o plan 80 100
    w/o action 100 80
    w/o strategy learning 50 60

    Removing the situation analysis module drops win rates by 30% for Good and 40% for Evil. Removing the experience learning module yields the largest decline for Good (dropping to 50% WR) and reduces Evil WR to 60%. Ablating the planning module reduces Good side win rate to 80%, while ablating the confidential action module reduces Evil side win rate to 80%.

  5. Knowl 5 — Evil Faction Aggressiveness: Quest Engagement and Failure Voting Dynamics

    empirical result

    When playing as the evil faction (Morgana and Assassin), the proposed framework adopts an assertive sabotage strategy compared to baseline agents across 20 games:

    • Quest Engagement Rate (QER): Proposed evil agents achieve a 40.3%40.3\% quest engagement rate compared to 33.1%33.1\% for baseline evil agents, successfully infiltrating more quest teams.
    • Failure Vote Rate (FVR): Once on a quest team, proposed evil agents cast sabotage votes 84.0%84.0\% of the time, compared to only 36.5%36.5\% for baseline evil agents.

    Baseline evil agents frequently cast success cards or fail to secure team seats, enabling good factions to win. The proactive quest infiltration and high failure voting rate of the proposed agents allow them to achieve a 100%100\% win rate against the baseline good faction.

  6. Knowl 6 — Target-Specific Enemy Analysis for Assassin Role Advantage

    data/table

    An ablation study isolating the target of the Analysis Module under the no-experience-learning setting (10 games per condition as the Evil side) shows that analyzing adversaries is essential for evil faction success:

    Method WR (%) QER (%) FVR (%)
    Morgana Assassin Morgana Assassin
    All players 90 55.5 58.3 93.7 100
    Teammates only 80 26.8 48.1 62.5 100
    Adversaries only 90 38.3 45.3 92.3 100

    Restricting analysis strictly to teammates reduces win rate to 80%80\% and leads to passive play (Morgana QER drops to 26.8%26.8\% and FVR to 62.5%62.5\%). In contrast, analyzing adversaries only preserves the 90%90\% win rate and high failure voting rates (92.3%92.3\% and 100%100\%). This occurs because adversary profiling provides the Assassin with the necessary behavioral evidence to correctly identify Merlin at the end of the game.

  7. Knowl 7 — Spontaneous Emergence of Camouflage and Identity Deception

    empirical result

    In the opening round of Avalon gameplay, agents exhibit role-differentiated camouflage strategies without explicit hand-crafted deception rules:

    • Good Faction: Loyal Servants and Percival exhibit high rates of voluntary self-disclosure (60%–70%60\%\text{--}70\%) to build coordination and trust.
    • Evil Faction: Morgana and Assassin spontaneously disguise their roles. In the first round, Morgana assumes an alternate identity (active camouflage) in 10%10\% of games, discloses identity in 35%35\%, and withholds identity in 55%55\%. Assassin assumes an alternate identity in 15%15\% of games, discloses identity in 40%40\%, and withholds identity in 45%45\%.

    Evil agents actively exploit the game rule where Percival perceives both Merlin and Morgana as valid candidates, adopting false claims to mislead Percival and loyal servants.

  8. Knowl 8 — Mitigation of Maladaptive Confrontation via Strategic Learning

    empirical result

    Comparing agent dialogue interactions with and without the experience learning module demonstrates that cross-game strategy updates suppress maladaptive aggression:

    • Without Experience Learning: Evil-side agents (such as Morgana) exhibit excessive, unprovoked confrontation toward other players, easily exposing their evil allegiance. Meanwhile, good-side agents show low, undifferentiated confrontation rates across all roles.
    • With Experience Learning: Evil agents learn to mask hostility, increasing cooperative dialogue and teamwork displays toward good players to avoid detection. Good-side agents (Merlin and Percival) learn to selectively direct confrontation toward identified evil opponents (Morgana and Assassin) while reducing conflict with potential allies.
  9. Knowl 9 — Model Comprehension Constraints: Valid Response Rate in LLaMA-2 vs GPT-3.5

    data/table

    Evaluating the multi-agent framework on Llama2-7b-chat-hf compared to gpt-3.5-turbo-16k reveals significant differences in instruction following and action validity across roles:

    Base Model Valid Response Rate (VRR, %)
    Loyal Servant Merlin Percival Morgana Assassin Average
    LLaMA-2 51.9 61.0 53.6 66.5 66.9 59.9
    GPT-3.5 81.7 84.2 81.9 89.7 87.6 85.0

    LLaMA-2 exhibits an absolute decrease of 25.1%25.1\% in average Valid Response Rate (59.9%59.9\% vs 85.0%85.0\%). This lower syntactic adherence and difficulty handling multi-turn hidden information deduction highlight the need for strong base language comprehension capabilities in complex multi-agent environments.

  10. Knowl 10 — Latency and Excessive Information Disclosure Limitations in Social Deduction Agents

    limitation

    The proposed LLM framework has two notable operational and behavioral limitations:

    1. Execution Latency and API Overhead: Executing sequential LLM calls per agent per turn (summarization, analysis, planning, action selection, and response generation) incurs substantial financial cost and slow response latency during multi-agent game progression.
    2. Pathological Information Disclosure: In the initial round, both proposed agents and baseline agents playing Merlin and Percival exhibit an excessive tendency to publicly share known clues (approaching 100%100\% sharing rate), failing to exercise strategic restraint and inadvertently exposing Merlin to the Assassin.

Coverage note — Exact prompt templates from Appendix Tables 4 and 5 were summarized into the respective model and experimental setup knowls rather than transcribed verbatim to avoid redundant prompt listings.

References

  1. 1.Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz. 2023. Playing repeated games with large language models. ArXiv, abs/2305.16867.
  2. 2.Nicolo' Brandizzi, Davide Grossi, and Luca Iocchi. 2021. Rlupus: Cooperation through emergent communication in the werewolf social deduction game. ArXiv, abs/2106.05018.
  3. 3.Zhendong Chen, Siu Cheung Hui, Fuzhen Zhuang, Lejian Liao, Fei Li, Meihuizi Jia, and Jiaqi Li. 2022. Evidencenet: Evidence fusion network for fact verification. In Proceedings of the ACM Web Conference 2022, pages 2636–2645.
  4. 4.Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata. 2023. Improving language model negotiation with self-play and in-context learning from ai feedback.
  5. 5.Chen Gao, Xiaochong Lan, Zhi jie Lu, Jinzhu Mao, Jing Piao, Huandong Wang, Depeng Jin, and Yong Li. 2023. S3: Social-network simulation system with large language model-empowered agents. ArXiv, abs/2307.14984.
  6. 6.Navid Ghaffarzadegan, Aritra Majumdar, Ross Williams, and Niyousha Hosseinichimeh. 2023. Generative agent-based modeling: Unveiling social system dynamics through coupling mechanistic models with generative artificial intelligence. ArXiv, abs/2309.11456.
  7. 7.Yuya Hirata, Michimasa Inaba, Kenichi Takahashi, Fujio Toriumi, Hirotaka Osawa, Daisuke Katagami, and Kousuke Shinoda. 2016. Werewolf game modeling using action probabilities based on play log analysis. In Computers and Games.
  8. 8.Zhao Kaiya, Michelangelo Naim, Jovana Kondic, Manuel Cortes, Jiaxin Ge, Shuying Luo, Guangyu Robert Yang, and Andrew Ahn. 2023. Lyfe agents: Generative agents for low-cost real-time social interactions.
  9. 9.Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. 2023. Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual differences, 103:102274.
  10. 10.Sharon Levy, Robert E Kraut, Jane A Yu, Kristen M Altenburger, and Yi-Chia Wang. 2022. Understanding conflicts in online conversations. In Proceedings of the ACM Web Conference 2022, pages 2592–2602.
  11. 11.Paul Pu Liang, Jeffrey Chen, Ruslan Salakhutdinov, Louis-Philippe Morency, and Satwik Kottur. 2020. On emergent communication in competitive multiagent teams. ArXiv, abs/2003.01848.
  12. 12.Bill Yuchen Lin, Yicheng Fu, Karina Yang, Prithviraj Ammanabrolu, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2023. Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks. ArXiv, abs/2305.17390.
  13. 13.Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hanguang Li. 2023. Trustworthy llms: a survey and guideline for evaluating large language models' alignment. ArXiv, abs/2308.05374.
  14. 14.Rajiv Movva, S. Balachandar, Kenny Peng, Gabriel Agostini, Nikhil Garg, and Emma Pierson. 2023. Large language models shape and are shaped by society: A survey of arxiv publication patterns. ArXiv, abs/2307.10700.
  15. 15.Noritsugu Nakamura, Michimasa Inaba, Kenichi Takahashi, Fujio Toriumi, Hirotaka Osawa, Daisuke Katagami, and Kousuke Shinoda. 2016. Constructing a human-like agent for the werewolf game using a psychological model based multiple perspectives. 2016 IEEE Symposium Series on Computational Intelligence (SSCI), pages 1–8.
  16. 16.Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative agents: Interactive simulacra of human behavior.
  17. 17.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277.
  18. 18.Chen Qian, Xin Cong, Cheng Yang, Weize Chen, Yusheng Su, Juyuan Xu, Zhiyuan Liu, and Maosong Sun. 2023. Communicative agents for software development. ArXiv, abs/2307.07924.
  19. 19.Zijing Shi, Meng Fang, Shunfeng Zheng, Shilong Deng, Ling Chen, and Yali Du. 2023. Cooperation on the fly: Exploring language agents for ad hoc teamwork in the avalon game.
  20. 20.Qiurong Song and Jiepu Jiang. 2022. How misinformation density affects health information search. In Proceedings of the ACM Web Conference 2022, pages 2668–2677.
  21. 21.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  22. 22.Chen Feng Tsai, Xiaochen Zhou, Sierra S Liu, Jing Li, Mo Yu, and Hongyuan Mei. 2023. Can large language models play text games well? current state-of-the-art and open questions. arXiv preprint arXiv:2304.02868.
  23. 23.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  24. 24.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023a. Voyager: An open-ended embodied agent with large language models.
  25. 25.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2023b. A survey on large language model based autonomous agents. arXiv preprint arXiv:2308.11432.
  26. 26.Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang. 2023c. Avalon's game of thoughts: Battle against deception through recursive contemplation.
  27. 27.Tianhe Wang and Tomoyuki Kaneko. 2018. Application of deep reinforcement learning in werewolf game agents. 2018 Conference on Technologies and Applications of Artificial Intelligence (TAAI), pages 28–33.
  28. 28.Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2023d. Aligning large language models with human: A survey. ArXiv, abs/2307.12966.
  29. 29.Sarah Wiseman and Kevin B. Lewis. 2019. What data do players rely on in social deduction games? Extended Abstracts of the Annual Symposium on Computer-Human Interaction in Play Companion Extended Abstracts.
  30. 30.Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864.
  31. 31.Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, and Yang Liu. 2023a. Exploring large language models for communication games: An empirical study on werewolf.
  32. 32.Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu. 2023b. Language agents with reinforcement learning for strategic play in the werewolf game.
  33. 33.Kai-Cheng Yang and Filippo Menczer. 2023. Anatomy of an ai-powered malicious social botnet. ArXiv, abs/2307.16336.
  34. 34.Haoqi Yuan, Chi Zhang, Hongcheng Wang, Feiyang Xie, Penglin Cai, Hao Dong, and Zongqing Lu. 2023. Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks.
  35. 35.Xuanhe Zhou, Guoliang Li, and Zhiyuan Liu. 2023. Llm as dba.
  36. 36.Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, Yu Qiao, Zhaoxiang Zhang, and Jifeng Dai. 2023. Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory.

Citation

MLA
Lan, Y., et al. “LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 128–45, https://doi.org/10.18653/v1/2024.emnlp-main.7.
APA
Lan, Y., Hu, Z., Wang, L., Wang, Y., Ye, D., Zhao, P., Lim, E.-P., Xiong, H., & Wang, H. (2024). LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 128–145. https://doi.org/10.18653/v1/2024.emnlp-main.7
Chicago
Lan, Y., Z. Hu, L. Wang, et al. 2024. “LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 128–45. https://doi.org/10.18653/v1/2024.emnlp-main.7.
Harvard
Lan, Y. et al. (2024) “LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 128–145. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.7.
Vancouver
1. Lan Y, Hu Z, Wang L, Wang Y, Ye D, Zhao P, Lim E-P, Xiong H, Wang H (2024) LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 128–145

BibTeX

@inproceedings{lan-etal-2024-llm,
    title = "{LLM}-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay",
    author = "Lan, Yihuai  and
      Hu, Zhiqiang  and
      Wang, Lei  and
      Wang, Yang  and
      Ye, Deheng  and
      Zhao, Peilin  and
      Lim, Ee-Peng  and
      Xiong, Hui  and
      Wang, Hao",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.7/",
    doi = "10.18653/v1/2024.emnlp-main.7",
    pages = "128--145"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/