Theory of Mind for Multi-Agent Collaboration via Large Language Models

Huao LiYu Quan ChongSimon StepputtisJoseph CampbellDana HughesCharles LewisKatia P. Sycara

article2023EMNLP140 citations

Evaluates large language model agents on cooperative multi-agent tasks requiring Theory of Mind, revealing that tracking explicit belief states substantially corrects their long-horizon planning errors and state hallucinations.

Listen

As large language models increasingly enter collaborative and autonomous multi-agent environments, understanding their ability to work together and model other agents' perspectives is critical. Effective collaboration requires Theory of Mind—the cognitive capacity to infer teammates' hidden knowledge, intentions, and beliefs. The article evaluates the collaborative planning and Theory of Mind capabilities of large language model-based agents in interactive, multi-agent search and rescue missions.

To conduct this evaluation, the researchers developed a simulated text-based cooperative game where three decentralized agents must explore an environment of five interconnected rooms, locate five color-coded bombs, and coordinate specific tool actions to safely defuse them. The study tested OpenAI's ChatGPT and GPT-4 models under different prompting conditions, comparing them against established baselines: a state-of-the-art Multi-Agent Reinforcement Learning algorithm trained over 45 million timesteps, a Conflict-Based Search planning baseline, and random actions. Beyond task completion, the study systematically evaluated agents on three tiers of Theory of Mind inference—introspection, first-order belief estimation, and complex second-order belief estimation—during dynamic interactions.

The findings demonstrate that advanced language models can autonomously exhibit collaborative behaviors, such as delegating roles and sharing critical information, without task-specific training. Standard GPT-4 achieved a perfect score of 90 points, requiring an average of 28.3 rounds to complete the mission, whereas ChatGPT failed to finish, averaging only 43.3 points over 30 rounds. However, baseline GPT-4 exhibited systematic failures, including hallucinations about game states and difficulties managing long-horizon constraints, resulting in invalid actions. To resolve this, the researchers incorporated explicit, text-based belief state representations into prompts. This modification reduced invalid actions by roughly 50.7% and improved task efficiency by about 130%, reducing completion time to 12.3 rounds—approaching the reinforcement learning baseline's 11.0 rounds. In Theory of Mind assessments, GPT-4 with belief representations scored 97.2% in introspection, 80.1% in first-order inferences, and 69.4% in second-order inferences, substantially outperforming ChatGPT across all levels.

These results show that large language models possess emergent social intelligence and can serve as zero-shot planners comparable to specialized reinforcement learning systems. Structured belief tracking is essential to mitigate operational risks such as misinformation cascades, where an agent's hallucination quickly spreads false beliefs across a team. Organizations deploying autonomous multi-agent systems should integrate explicit internal state representations into agent architectures to preserve reasoning accuracy over extended tasks.

Decision-makers should interpret these results within the context of the study's boundaries, as tests were conducted in a simplified, five-room simulation with homogeneous agents and human-annotated evaluations. Further validation in larger environments, heterogeneous teams, and hybrid human-agent settings is necessary before deploying fully autonomous multi-agent systems in high-risk operational environments.

arXiv: 2310.10701
Cover for Theory of Mind for Multi-Agent Collaboration via Large Language Models

Abstract

While Large Language Models (LLMs) have demonstrated impressive accomplishments in both reasoning and planning, their abilities in multi-agent collaborations remains largely unexplored. This study evaluates LLM-based agents in a multi-agent cooperative text game with Theory of Mind (ToM) inference tasks, comparing their performance with Multi-Agent Reinforcement Learning (MARL) and planning-based baselines. We observed evidence of emergent collaborative behaviors and high-order Theory of Mind capabilities among LLM-based agents. Our results reveal limitations in LLM-based agents' planning optimization due to systematic failures in managing long-horizon contexts and hallucination about the task state. We explore the use of explicit belief state representations to mitigate these issues, finding that it enhances task performance and the accuracy of ToM inferences for LLM-based agents.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large language models
  • 2.2 Theory of Mind
  • 2.3 Multi-agent collaboration
  • 3 Multi-agent Collaboration Tasks
  • 3.1 Task environment
  • 3.2 Text game interface
  • 4 LLM-based Embodied Agents
  • 4.1 Multi-agent communication
  • 4.2 Belief state
  • 5 Experiments
  • 5.1 Setups
  • 5.2 Baselines
  • 5.3 Theory of mind inferences
  • 6 Results
  • 6.1 Task performance
  • 6.2 Basic embodied interactions
  • 6.3 Emergent collaborative behaviors
  • 6.4 LLM's systematic failures
  • 6.4.1 Long-horizon contexts
  • 6.4.2 Hallucination
  • 6.5 Theory of Mind Inference
  • 6.5.1 Case study
  • 6.5.2 Inference under false belief
  • 7 Discussions
  • 8 Conclusions
  • Limitations
  • 9 Acknowledgements
  • References
  • Appendix
  • A Prompts
  • A.1 Task context
  • A.2 Initial belief state
  • B Environment feedback for Error correction
  • C Theory of Mind Questions
  • C.1 Introspection
  • C.2 First-order ToM
  • C.3 Second-order ToM

Knowls

  1. Knowl 1 — Search-and-rescue collaboration environment

    experimental setup

    The paper evaluates three cooperative agents—Alpha, Bravo, and Charlie—in a text-based search-and-rescue mission. The agents must explore a connected five-room graph, locate five hidden bombs, inspect each bomb to discover its color-coded phase sequence, and apply wire cutters in the correct order. The rooms are labeled 0, 3, 5, 6, and 8; room 0 connects to rooms 3, 5, 6, and 8, room 5 connects to room 6, and room 8 connects to rooms 3 and 6. Alpha has red and green cutters, Bravo has green and blue cutters, and Charlie has blue and red cutters.

    In each round, an agent can move to an adjacent room, inspect a bomb in its current room, or apply one of its cutters to a bomb in that room. The team receives 10 points for every correctly processed bomb phase. The mission contains two one-phase bombs, two two-phase bombs, and one three-phase bomb, so the maximum score is 90 points. Each agent observes only its current room, local agent status, teammate locations, available tools, and communicated messages. A trial ends when all bombs are defused, 30 rounds have elapsed, or the team enters a deadlock by repeatedly producing the same outputs.

  2. Knowl 2 — Decentralized LLM embodied-agent interface

    model/method

    The LLM agents use OpenAI gpt-3.5-turbo-0301, called ChatGPT, or gpt-4-0314. At every timestamp, the three agents take turns receiving natural-language observations, selecting an environment action, and sending a text message to both teammates. An agent cannot directly observe another agent's actions or their consequences unless that information is communicated. The interface converts observations into templated natural-language descriptions and converts requested actions back into abstract game actions through keyword matching. Invalid or unintelligible outputs receive rule-based feedback so that the agent can attempt an error correction.

    The agents are prompted with the game rules and their current observations. Their interaction histories remain in the prompt until the model input limit is reached; in the reported setup, the agents retain the rules and the preceding two rounds, totaling 4096 tokens. The model temperature is set to zero. This produces a fully decentralized, zero-shot team in which information about bomb states, actions, and intentions must be propagated through communication.

  3. Knowl 3 — Explicit textual belief-state prompting

    model/method

    The paper adds an explicit belief-state representation to mitigate the loss of long-term task information from the LLM context. After receiving an environmental observation, each agent is prompted to update a textual record of its beliefs about mission-relevant facts, such as room contents and bomb locations, phase sequences, and state changes. The updated record is preserved in the agent's interaction history and supplied during subsequent action planning.

    The prompt includes an initial example specifying the format of a belief state, but the update rules are zero-shot: the LLM must revise the record using its observations, the mission context, and common-sense reasoning. For example, an agent changes the belief about a bomb's sequence from unknown to red after inspecting it and retains that information in later rounds. The method is intended to function as an external structured memory and to re-emphasize facts that might otherwise be buried in a long interaction history.

  4. Knowl 4 — Interactive three-level Theory of Mind evaluation

    model/method

    During the mission, the LLM agents answer Theory of Mind questions after actions that may change the world state or another agent's beliefs. The evaluation has three levels: introspection asks whether an agent knows its own relevant state; first-order Theory of Mind asks whether it knows another agent's hidden state; and second-order Theory of Mind asks whether it knows that another agent knows something about the first agent's state.

    Questions concern facts such as room contents, bomb sequences, bomb-state changes, and whether a phase has been defused. Human annotators determine correctness from the globally observable interaction and communication history. Their criteria include whether the target agent was present and could observe the consequence, whether it had previously visited the relevant room, and whether the consequence was communicated. Cases involving ambiguous communication or false beliefs are adjudicated by the annotators, making the benchmark an interactive and dynamic extension of Sally–Anne-style tests rather than a static text-only test.

  5. Knowl 5 — Experimental conditions and comparison baselines

    experimental setup

    The main ablation varies the LLM and whether the explicit belief state is provided, producing four conditions: ChatGPT without belief state, GPT-4 without belief state, and GPT-4 with belief state; the paper describes the model and belief-state factors as a four-condition design, with ChatGPT and GPT-4 each evaluated with and without belief representation. Environments are reset with randomized starting locations, room connections, bomb distributions, and bomb sequences. LLM agents use temperature zero and three repeated trials are performed for stability.

    The comparison includes a random policy, Multi-Agent Proximal Policy Optimization (MAPPO), and a Conflict-Based Search (CBS) planner. MAPPO uses a recurrent stateful actor-critic with shared actor and critic networks, is trained with default StarCraft Multi-Agent Challenge hyperparameters, and receives an additional shaping reward of +1+1 for a correct cutter application and −1-1 for causing an explosion. CBS centrally assigns tasks and collision-free paths while respecting precedence and temporal constraints; when all five bombs are treated as one subtask, it is complete and optimal with respect to the score under the planner's centralized full-information setting.

  6. Knowl 6 — Task-performance comparison

    data/table

    The performance comparison reported on page 5 evaluates team score, rounds required to finish, and the proportion of LLM outputs that encode valid game actions. A score of 90 means that all nine bomb phases were successfully processed. The results show that GPT-4 reaches the maximum score, while adding an explicit belief state preserves the score and greatly reduces completion time. ChatGPT does not finish the mission within the 30-round limit. The learned MAPPO baseline is substantially faster than GPT-4 without belief state but slower than the centralized CBS planner.

    Could not parse LaTeX table

    The numbers after pm\\pm are one standard deviation. The results demonstrate that zero-shot GPT-4 teams can solve the cooperative task at the same maximum score as trained MAPPO and CBS, although GPT-4 without belief state is much less efficient than both; explicit belief prompting brings its completion time close to MAPPO.

  7. Knowl 7 — Systematic planning failures and belief-state mitigation

    empirical result

    The paper identifies two recurring failure modes in LLM-based teams. First, long-horizon context failures cause agents to overlook rules and facts that appeared earlier in the prompt, leading to invalid actions such as moving to a non-adjacent room or using a cutter that the agent does not possess. Second, hallucinations produce actions that are syntactically valid but infeasible under the actual task state, such as searching for an already defused bomb or claiming a bomb sequence that has never been inspected. The authors attribute these hallucinations mainly to partial observations and the absence of an explicit, persistent belief representation.

    For GPT-4, adding the textual belief state increases valid-action output from 71.8% to 86.1%, which the paper characterizes as a 50.7% reduction in invalid actions. Completion improves from 28.3 to 12.3 rounds, characterized by the paper as a 130% increase in team efficiency. The page 6 interaction diagram illustrates the long-context failure as an attempted move to a non-adjacent room and contrasts it with belief-supported reasoning that retains bomb and room information.

  8. Knowl 8 — Theory of Mind accuracy results

    data/table

    The ToM evaluation compares natural-language answers from the LLM agents with human-annotated ground truth derived from the complete mission history. Explicit belief prompting improves accuracy at every evaluated level, with the largest absolute gains for introspection and first-order inference. GPT-4 also outperforms ChatGPT at all three levels, and GPT-4 with belief state answers second-order questions correctly in nearly 70% of cases.

    Could not parse LaTeX table

    The results indicate that LLM agents can track their own relevant knowledge and, to a lesser extent, teammates' knowledge in an evolving mission. They also show that second-order reasoning remains substantially harder than introspection, especially for ChatGPT.

  9. Knowl 9 — Emergent collaborative behaviors

    empirical result

    Qualitative analysis of team trajectories found collaborative behaviors that were not explicitly programmed or task-specifically trained. Agents used messages to delegate subtasks, request assistance, resolve conflicts, and share task information. In the interaction example shown on page 6, Alpha voluntarily assumes a leadership role by instructing Bravo to move to one room and Charlie to move to another while Alpha inspects a bomb. The behavior emerges despite decentralized observations and zero-shot prompting.

    The authors interpret these behaviors as evidence that the LLMs can apply social and teamwork patterns acquired from language data to embodied multi-agent coordination. These behaviors help explain why GPT-4 teams attain the maximum mission score even though they do not use centralized coordination or task-specific reinforcement learning.

  10. Knowl 10 — Scope and limitations of the evaluation

    limitation

    The evaluation is limited to two OpenAI chat models, a relatively small environment with five rooms and five bombs, and homogeneous teams of three LLM agents. It does not establish how the results would change with newer or differently trained models, larger environments, more restrictive task rules, heterogeneous agents, or human–agent teams. The authors specifically identify trust, transparency, and human-agent co-training as untested issues for future human-centered evaluation.

    The ToM ground truth is also an approximation: human annotators with a global view of the mission infer what an agent should know under a rational-human interpretation. This can be ambiguous when agents communicate inaccurate intentions or maintain false beliefs. The current belief state represents only each agent's own world knowledge, not explicit beliefs about other agents. The paper therefore proposes, but does not evaluate, extending the representation to first-order and second-order beliefs about teammates and using those maintained belief states as a more direct ToM evaluation target.

Coverage note — No substantial contributed material was omitted; implementation-level prompt templates and appendix error-message variants were excluded because they do not add independent load-bearing findings beyond the interface and belief-state method.

References

  1. 1.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691.
  2. 2.Chris L Baker, Julian Jara-Ettinger, Rebecca Saxe, and Joshua B Tenenbaum. 2017. Rational quantitative attribution of beliefs, desires and percepts in human mentalizing. Nature Human Behaviour, 1(4):0064.
  3. 3.Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016. Openai gym.
  4. 4.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Eean R Crawford and Jeffery A Lepine. 2013. A configural theory of team processes: Accounting for the structure of taskwork and teamwork. Academy of Management Review, 38(1):32–48.
  7. 7.Xiaocong Fan and John Yen. 2004. Modeling and simulating human teamwork behaviors using intelligent agents. Physics of life reviews, 1(3):173–201.
  8. 8.Thilo Hagendorff. 2023. Machine psychology: Investigating emergent capabilities and behavior in large language models using psychological methods. arXiv preprint arXiv:2303.13988.
  9. 9.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pages 9118–9147. PMLR.
  10. 10.Boaz Keysar, Shuhong Lin, and Dale J Barr. 2003. Limits on theory of mind use in adults. Cognition, 89(1):25–41.
  11. 11.Michal Kosinski. 2023. Theory of mind may have spontaneously emerged in large language models. arXiv preprint arXiv:2302.02083.
  12. 12.Huao Li, Ini Oguntola, Dana Hughes, Michael Lewis, and Katia Sycara. 2022. Theory of mind modeling in search and rescue teams. In 2022 31st IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), pages 483–489. IEEE.
  13. 13.Terence X Lim, Sidney Tio, and Desmond C Ong. 2020. Improving multi-agent cooperation using theory of mind. arXiv preprint arXiv:2007.15703.
  14. 14.Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone. 2023. Llm+ p: Empowering large language models with optimal planning proficiency. arXiv preprint arXiv:2304.11477.
  15. 15.Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. 2023. Dissociating language and thought in large language models: a cognitive perspective. arXiv preprint arXiv:2301.06627.
  16. 16.Scott A Miller. 2009. Children’s understanding of second-order mental states. Psychological bulletin, 135(5):749.
  17. 17.Shima Rahimi Moghaddam and Christopher J Honey. 2023. Boosting theory-of-mind performance in large language models via prompting. arXiv preprint arXiv:2304.11490.
  18. 18.Ini Oguntola, Joseph Campbell, Simon Stepputtis, and Katia Sycara. 2023. Theory of mind as intrinsic motivation for multi-agent reinforcement learning. arXiv preprint arXiv:2307.01158.
  19. 19.OpenAI. 2023. Gpt-4 technical report.
  20. 20.Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. arXiv preprint arXiv:2304.03442.
  21. 21.Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W Crandall, Nicholas A Christakis, Iain D Couzin, Matthew O Jackson, et al. 2019. Machine behaviour. Nature, 568(7753):477–486.
  22. 22.Christoph Riedl, Young Ji Kim, Pranav Gupta, Thomas W Malone, and Anita Williams Woolley. 2021. Quantifying collective intelligence in human groups. Proceedings of the National Academy of Sciences, 118(21):e2005737118.
  23. 23.Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson. 2019. The StarCraft Multi-Agent Challenge. CoRR, abs/1902.04043.
  24. 24.Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi. 2023. Neural theory-of-mind? on the limits of social intelligence in large lms.
  25. 25.Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov. 2023. Minding language models’(lack of) theory of mind: A plug-and-play multi-character belief tracker. arXiv preprint arXiv:2306.00924.
  26. 26.Guni Sharon, Roni Stern, Ariel Felner, and Nathan R Sturtevant. 2015. Conflict-based search for optimal multi-agent pathfinding. Artificial Intelligence, 219:40–66.
  27. 27.Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. 2017. Value-decomposition networks for cooperative multi-agent learning. arXiv preprint arXiv:1706.05296.
  28. 28.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  29. 29.Tomer Ullman. 2023. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399.
  30. 30.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023a. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291.
  31. 31.Zihao Wang, Shaofei Cai, Anji Liu, Xiaojian Ma, and Yitao Liang. 2023b. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv preprint arXiv:2302.01560.
  32. 32.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  33. 33.Jessica Williams, Stephen M Fiore, and Florian Jentsch. 2022. Supporting artificial social intelligence with theory of mind. Frontiers in artificial intelligence, 5.
  34. 34.Yaqi Xie, Chen Yu, Tongyao Zhu, Jinbin Bai, Ze Gong, and Harold Soh. 2023. Translating natural language to planning goals with large-language models. arXiv preprint arXiv:2302.05128.
  35. 35.Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022. The surprising effectiveness of ppo in cooperative multi-agent games. Advances in Neural Information Processing Systems, 35:24611–24624.
  36. 36.Luyao Yuan, Zipeng Fu, Linqi Zhou, Kexin Yang, and Song-Chun Zhu. 2021. Emergence of theory of mind collaboration in multiagent systems. arXiv preprint arXiv:2110.00121.
  37. 37.Jun Zhang, Trey Hedden, and Adrian Chia. 2012. Perspective-taking and depth of theory-of-mind reasoning in sequential-move games. Cognitive science, 36(3):560–573.
  38. 38.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.

Citation

MLA
Li, H., et al. “Theory of Mind for Multi-Agent Collaboration via Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 180–92, https://doi.org/10.18653/v1/2023.emnlp-main.13.
APA
Li, H., Chong, Y., Stepputtis, S., Campbell, J. P., Hughes, D., Lewis, C., & Sycara, K. (2023). Theory of Mind for Multi-Agent Collaboration via Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 180–192. https://doi.org/10.18653/v1/2023.emnlp-main.13
Chicago
Li, H., Y. Chong, S. Stepputtis, et al. 2023. “Theory of Mind for Multi-Agent Collaboration via Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 180–92. https://doi.org/10.18653/v1/2023.emnlp-main.13.
Harvard
Li, H. et al. (2023) “Theory of Mind for Multi-Agent Collaboration via Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 180–192. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.13.
Vancouver
1. Li H, Chong Y, Stepputtis S, Campbell JP, Hughes D, Lewis C, Sycara K (2023) Theory of Mind for Multi-Agent Collaboration via Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 180–192

BibTeX

@inproceedings{li-etal-2023-theory,
    title = "Theory of Mind for Multi-Agent Collaboration via Large Language Models",
    author = "Li, Huao  and
      Chong, Yu  and
      Stepputtis, Simon  and
      Campbell, Joseph  and
      Hughes, Dana  and
      Lewis, Charles  and
      Sycara, Katia",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.13/",
    doi = "10.18653/v1/2023.emnlp-main.13",
    pages = "180--192"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/