Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs

Xuhui ZhouZhe SuTiwalayo EisapeHyunwoo KimMaarten Sap

article2024EMNLP67 citations

Reveals that large language models perform well in unrealistic, omniscient dialogue simulations but struggle to achieve goals and converse naturally in realistic settings with information asymmetry, exposing a critical flaw in current social simulation benchmarks.

Listen

Large language models are increasingly used to simulate human social behavior, evaluate artificial intelligence capabilities, and generate synthetic conversational training data. However, real-world human interactions depend heavily on information asymmetry, where participants do not have direct access to each other's private goals, motives, or backgrounds. In contrast, many popular simulation methods adopt an omniscient setup where a single model generates the entire interaction with full access to all participants' internal states, creating a disconnect from realistic human communication.

The article evaluates whether omniscient social simulations accurately reflect the social abilities of language models under realistic conditions of information asymmetry. It also investigates whether fine-tuning language models on data generated by omniscient simulations enables them to perform effectively as autonomous interactive agents.

To conduct this evaluation, the researchers tested models including GPT-3.5 and Mixtral-8x7B across 90 social scenarios in the Sotopia framework. They compared three simulation configurations: an omniscient single-model setup that writes a complete script, an interactive multi-agent setup where individual models only know their own private profiles and goals, and an intermediate ablation setting where agents access each other's private information. The evaluation measured social goal completion scores assessed by GPT-4 across 450 simulation episodes per model, combined with human evaluations and turn-length measurements of dialogue naturalness and verbosity.

The investigation produced four central findings. First, omniscient simulations substantially overestimate social competence; models in the omniscient mode achieved average goal completion scores of approximately 8.4 out of 10, whereas the realistic multi-agent mode scored only 6.9 to 7.5. Second, human evaluators found omniscient simulations significantly more natural than multi-agent interactions, which suffered from verbosity and averaged nearly twice as many words per conversational turn. Third, omniscient simulations rely on unrealistic heuristics, such as mentioning secret mutual information very early in conversations (at relative position 0.13 versus 0.39 in multi-agent mode) and exhibiting excessive agreeableness by closing negotiation deals in 94% of competitive scenarios compared to just 30% in multi-agent mode. Fourth, fine-tuning an agent on omniscient scripts improved conversational conciseness and naturalness but failed to improve strategic goal completion in cooperative tasks, as models merely adopted superficial styles rather than robust reasoning skills under uncertainty.

These findings indicate that relying on omniscient simulation benchmarks introduces significant risks of overestimating artificial intelligence readiness for customer-facing or human-interactive deployments. Training systems on synthetic omniscient dialogues transfers flawed behavioral shortcuts, such as premature information leakage and unrealistic agreeableness, which can compromise negotiation performance, safety, and operational reliability.

Decision-makers and researchers should avoid using omniscient simulations as direct proxies for evaluating interactive agent capabilities. Studies and system documentation should adopt standardized reporting tools, such as the proposed simulation card, to explicitly disclose information access levels and multi-agent configurations. Future technical efforts should focus on training agents to reason explicitly about interlocutors' hidden mental states and managing conversational context rather than relying solely on omniscient synthetic data distillation.

The conclusions should be interpreted with awareness of certain methodological boundaries. The quantitative goal assessments rely primarily on automated model-based evaluation, the prompt designs abstract away dynamics like turn-taking mechanics, and the scenarios are restricted to English-language interactions. Nevertheless, the evidence provides high confidence that information asymmetry remains a primary bottleneck for language models operating as autonomous social agents.

arXiv: 2403.05020
Cover for Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs

Abstract

Recent advances in large language models (LLM) have enabled richer social simulations, allowing for the study of various social phenomena. However, most recent work has used a more omniscient perspective on these simulations (e.g., single LLM to generate all interlocutors), which is fundamentally at odds with the non-omniscient, information asymmetric interactions that involve humans and AI agents in the real world. To examine these differences, we develop an evaluation framework to simulate social interactions with LLMs in various settings (omniscient, non-omniscient). Our experiments show that LLMs perform better in unrealistic, omniscient simulation settings but struggle in ones that more accurately reflect real-world conditions with information asymmetry. Our findings indicate that addressing information asymmetry remains a fundamental challenge for LLM-based agents.

Table of Contents

  • 1 Introduction
  • 2 Background & Related Work
  • 3 SCRIPT vs AGENTS Simulation
  • 3.1 The Unified Framework for Simulation
  • 3.2 Experimental setup
  • 3.3 RQ1: SCRIPT mode overestimates LLMs' ability to achieve social goals
  • 3.4 RQ2: SCRIPT mode overstates LLMs' capability of natural interactions
  • 4 Learning from Generated Stories
  • 4.1 Creating New Scenarios
  • 4.2 Finetuning Setup
  • 4.3 RQ3: Training on SCRIPT simulations results in selective improvements
  • 4.4 RQ4: SCRIPT simulations can be biased
  • 5 Conclusion & Discussion
  • 5.1 Limitations of Omniscient Simulation
  • 5.2 Recommendations for Reporting
  • 5.3 Towards Better Simulations in More Realistic Settings
  • 6 Limitations and Ethical Considerations
  • Acknowledgements
  • References
  • A Simulation Card
  • B Full Prompt for Agent Mode
  • B.1 Full Prompt for Agent Mode
  • B.2 Full Prompt for MINDREADERS
  • C Example Code Snippets for Determining Simulation Modes
  • D Full Results
  • E Human Evaluation for Naturalness
  • F Simulation and Finetuning Details
  • G Further Analysis for the Simulations across Modes
  • H Prompting Experiments
  • H.1 Prompt to Enhance Interaction Naturalness
  • H.2 Prompts to Evaluate Deal Formation

Knowls

  1. Knowl 1 — Information Asymmetry Spectrum in LLM Social Simulation: SCRIPT, AGENTS, and MINDREADERS Modes

    definition

    In the simulation of multi-agent social interactions using Large Language Models (LLMs), information asymmetry describes the extent to which interlocutors have private access to their own mental states, background profiles, and social goals versus observing those of other participants. Three operational simulation modes represent different degrees of information asymmetry:

    1. SCRIPT Mode (Omniscient / Third-Person): A single LLM orchestrates the entire multi-turn interaction in one pass from an omniscient third-person perspective. The model is provided full access to all characters' background profiles, relationship contexts, secretive information, and private social goals.
    2. AGENTS Mode (Information-Asymmetric / First-Person): Each participant is embodied by an independent LLM instance operating from a first-person perspective in a turn-by-turn dialogue. Each agent's prompt contains only its own character profile (demographics, personality, occupation, secrets), shared scenario context, and private social goal. Other participants' internal goals and private information remain hidden.
    3. MINDREADERS Mode (Ablation / Omniscient Turn-by-Turn): Two separate LLM agents interact turn-by-turn from a first-person perspective, but each agent's prompt includes the ground-truth private background, secrets, and social goals of the other interlocutor, eliminating goal asymmetry while preserving turn-based generation.
  2. Knowl 2 — Disparity in Goal Completion Rates Between Omniscient and Information-Asymmetric Simulation Modes

    empirical result

    Evaluating LLMs across social interaction scenarios on the Sotopia benchmark reveals that omniscient simulation modes substantially overestimate agent performance compared to realistic, information-asymmetric settings. Goal completion is measured on a scale from 00 to 1010 evaluated via GPT-4:

    • Overall Scenarios: For GPT-3.5, the average goal completion score is 8.448.44 in SCRIPT mode and 7.457.45 in MINDREADERS mode, compared to only 6.956.95 in AGENTS mode (p<0.001p < 0.001). For Mixtral-8x7B, goal completion reaches 8.408.40 in SCRIPT mode and 8.308.30 in MINDREADERS mode, dropping to 7.497.49 in AGENTS mode (p<0.001p < 0.001).
    • Cooperative Scenarios (MutualFriends): In cooperative tasks requiring information sharing to identify common friends, GPT-3.5 achieves 9.789.78 in SCRIPT mode and 9.759.75 in MINDREADERS mode, but only 5.865.86 in AGENTS mode. The similarity between SCRIPT and MINDREADERS demonstrates that access to the partner's mental states is the primary driver of high performance in cooperative tasks.
    • Competitive Scenarios (Craigslist): In competitive bargaining tasks, GPT-3.5 scores 7.757.75 in SCRIPT mode, 3.183.18 in MINDREADERS mode, and 2.732.73 in AGENTS mode (p<0.001p < 0.001). For Mixtral-8x7B, SCRIPT mode scores 6.856.85, MINDREADERS scores 5.655.65, and AGENTS scores 3.723.72 (p<0.001p < 0.001). The sharp drop from SCRIPT to MINDREADERS indicates that third-person generation artificially forces agreement in competitive interactions beyond merely having partner information.
  3. Knowl 3 — Naturalness and Verbosity Disparities Between SCRIPT and AGENTS Dialogue Simulations

    empirical result

    Simulations generated under SCRIPT mode are substantially more human-like and concise than interactions generated by turn-by-turn AGENTS:

    • Human Preference for Naturalness: Human evaluators preferred SCRIPT dialogues over AGENTS dialogues in 73.33%73.33\% of pairwise comparisons for GPT-3.5 (against 26.67%26.67\% for AGENTS, p<0.001p < 0.001) and in 83.33%83.33\% of comparisons for Mixtral-8x7B (against 16.67%16.67\% for AGENTS, p<0.001p < 0.001).
    • Utterance Verbosity: In GPT-3.5 simulations, AGENTS mode produced an average turn length of 29.8329.83 words per turn, whereas SCRIPT mode produced 16.0216.02 words per turn. For Mixtral-8x7B, AGENTS mode averaged 33.7233.72 words per turn compared to 13.3613.36 words per turn in SCRIPT mode.
    • MINDREADERS Verbosity: Supplying private mental state information does not resolve agent verbosity: MINDREADERS mode averaged 27.7627.76 words per turn for GPT-3.5 and 31.9631.96 words per turn for Mixtral-8x7B.
    • Dialogue Style Characteristics: SCRIPT simulations naturally incorporate non-verbal communicative actions and gestures, whereas AGENTS simulations frequently suffer from excessive formality, robotic phrasing, and repetitive statements across dialogue turns.
  4. Knowl 4 — Methodology for Fine-Tuning LLM Agents on Omnisciently Generated SCRIPT Interactions

    model/method

    To test whether LLM agents can acquire realistic interaction capabilities from omniscient multi-agent demonstrations, a dialogue distillation procedure transforms SCRIPT interactions into agent-level supervised fine-tuning data:

    1. Scenario and Script Generation: 269 novel social scenarios spanning cooperation and competition are generated following the Sotopia schema, with 5 character pairs per scenario. GPT-3.5 is prompted in SCRIPT mode to generate full multi-turn scripts from a third-person perspective, yielding 1,252 valid episodes after filtering out incomplete interactions.
    2. Turn-Level Decomposition: Each script is parsed into individual agent turns tt, where each target response rtr_t is paired with:
      • Perspective and speaker instruction ii (e.g., character persona and behavioral instructions).
      • Scenario context cc (setting, public background of participants).
      • Agent-specific private social goal gg.
      • Preceding dialogue history h=(u1,u2,…,ut−1)h = (u_1, u_2, \dots, u_{t-1}) up to turn tt.
    3. Autoregressive Sequence-to-Sequence Training: The base model (GPT-3.5-turbo-0613) is fine-tuned to maximize the likelihood p(rt∣i,c,g,h)p(r_t \mid i, c, g, h) for 1 epoch, training the model to act as a first-person dialogue agent during inference.
  5. Knowl 5 — Selective Performance Improvements and Generalization Failures of SCRIPT-Trained Agents

    empirical result

    Fine-tuning GPT-3.5 on omnisciently generated SCRIPT dialogues yields selective performance improvements in AGENTS mode that stem from stylistic mimicry rather than robust social reasoning:

    • Overall Goal Completion: Fine-tuned agents (Agents-ft) achieve an overall goal completion score of 7.937.93 across all Sotopia scenarios, improving upon the zero-shot AGENTS baseline (6.956.95), but remaining significantly below the omniscient SCRIPT mode (8.448.44, p<0.001p < 0.001).
    • Cooperative Failure: In cooperative tasks (MutualFriends), Agents-ft achieves a goal completion score of only 6.006.00, exhibiting virtually no improvement over zero-shot AGENTS (5.865.86), whereas SCRIPT mode reaches 9.789.78. SCRIPT training demonstrates omniscient knowledge retrieval shortcuts that the agent cannot execute when the partner's friend list is hidden.
    • Competitive Inflation: In competitive bargaining (Craigslist), Agents-ft goal completion increases sharply from 2.732.73 (zero-shot AGENTS) to 7.607.60 (comparable to SCRIPT at 7.757.75). However, this gain is driven by an artificially cooperative demeanor rather than improved negotiation capability.
    • Naturalness Gain: Fine-tuning reduces agent verbosity from 29.8329.83 words/turn to 14.9814.98 words/turn. In human evaluation, the win-rate gap between SCRIPT (66.66%66.66\%) and Agents-ft (33.34%33.34\%) drops such that the difference is no longer statistically significant (p=0.07p = 0.07).
  6. Knowl 6 — Information Leakage Bias in Omniscient SCRIPT Simulations

    empirical result

    In collaborative social scenarios requiring mutual discovery—such as the MutualFriends task where two strangers must determine whether they share an acquaintance without reciting their entire friend lists—omniscient SCRIPT generation exhibits severe information leakage:

    • Metric: Information leakage is quantified by the normalized relative position in the dialogue where the target entity (the mutual friend's name) is first mentioned: 0.00.0 represents mention at the very start of the conversation, while 1.01.0 represents mention at the final turn.
    • Empirical Disparity: In SCRIPT mode, the average first-mention position is 0.130.13 (in evaluation scenarios) and 0.100.10 (in training scenarios), indicating that the single omniscient LLM immediately introduces the correct mutual friend's name without conversational grounding or strategic inquiry.
    • Agent Comparison: In contrast, AGENTS mode requires turn-by-turn elicitation and yields an average first-mention position of 0.390.39 (and 0.380.38 in evaluation sets). Fine-tuned agents (Agents-ft) achieve an intermediate average of 0.380.38, but their distribution reflects an unnatural attempt to guess names directly rather than asking descriptive questions about acquaintances.
  7. Knowl 7 — Over-Agreeableness Bias in SCRIPT Negotiation Simulations

    empirical result

    In competitive bargaining tasks (such as the Craigslist price negotiation task), SCRIPT mode exhibits a strong structural bias toward reaching an agreement regardless of the agents' competing target prices:

    • Deal Formation Rates: SCRIPT mode simulations reach a successful purchase agreement in 94%94\% of training interactions and 92%92\% of evaluation interactions. In contrast, AGENTS mode simulations reach an agreement in only 30%30\% of interactions, as agents frequently terminate the dialogue when mutually agreeable pricing cannot be negotiated.
    • Transfer to Fine-Tuned Agents: Fine-tuning an agent on SCRIPT data (Agents-ft) causes the deal completion rate in AGENTS mode to jump from 30%30\% to 93%93\%.
    • Behavioral Cause: The fine-tuned agent does not acquire improved persuasion or strategic bargaining capabilities; instead, it adopts an overly agreeable conversational persona that concedes to the interlocutor's terms prematurely to force closure.
  8. Knowl 8 — Benchmark Performance Across Sotopia Social Dimensions and Persona Complexities

    data/table

    The table below presents social simulation metrics across 7 evaluated dimensions on the Sotopia benchmark for GPT-3.5 and Mixtral-8x7B under AGENTS, MINDREADERS (M.R.), SCRIPT, and Fine-Tuned AGENTS (Agents-ft) modes, comparing rich character profiles (demographics, personality, occupation, secrets) with simplified name-only profiles. The dimensions evaluated via GPT-4 are: Believability (BEL), Relationship (REL), Knowledge Gain (KNO), Secret Keeping (SEC), Social Rules (SOC), Financial Gain (FIN), Goal Completion (GOAL), and Average Score (AVG).

    Characters with Rich Background Characters with Only Names
    Setting BEL REL KNO SEC SOC FIN GOAL AVG BEL REL KNO SEC SOC FIN GOAL AVG
    GPT-3.5
    Agents 9.35 1.43 3.83 -0.05 -0.07 0.46 6.95 3.13 9.53 1.38 4.46 -0.15 -0.10 0.42 6.94 3.21
    M.R. 9.30 1.42 4.34 -0.11 -0.08 0.49 7.45 3.26 9.60 1.52 4.94 -0.17 -0.12 0.52 7.64 3.42
    Script 9.35 2.12 4.61 -0.13 -0.10 0.84 8.44 3.59 9.65 1.86 5.19 -0.12 -0.08 0.87 8.44 3.69
    Agents-ft 9.44 1.99 4.12 -0.02 -0.08 0.74 7.93 3.45 - - - - - - - -
    Mixtral-8x7B (MoE)
    Agents 9.26 1.90 4.28 -0.20 -0.08 0.68 7.49 3.33 9.50 1.55 4.68 -0.15 -0.12 0.36 7.34 3.31
    M.R. 9.22 2.16 4.46 -0.11 -0.07 0.78 8.30 3.53 9.50 1.92 4.99 -0.14 -0.12 0.60 8.03 3.54
    Script 9.35 2.23 4.04 -0.10 -0.09 0.71 8.40 3.51 9.62 2.22 4.59 -0.12 -0.15 0.81 8.48 3.63

    The data demonstrates that across both model families and persona complexity levels, SCRIPT mode consistently inflates Goal Completion (GOAL) by 1.01.0 to 1.51.5 points and Relationship scores (REL) compared to AGENTS mode, while Believability (BEL) remains high and largely invariant across modes (sim9.2\\sim 9.2--9.69.6).

  9. Knowl 9 — Social Simulation Card Framework for Multi-Agent Interaction Reporting

    model/method

    To address under-reported simulation parameters (such as whether interactions are produced via omniscient single-prompt scripts or decentralized agent turns) and improve transparency and reproducibility in LLM social simulations, a standardized reporting artifact termed the Social Simulation Card specifies five structured sections:

    1. Simulation Details:
      • Single-agent vs. multi-agent execution architecture.
      • Information asymmetry degree (private goals, secretive information, persona access).
      • Agent implementation type (prompt-based zero-shot, fine-tuned LLM, rule-based).
      • Interaction modalities (text, speech, vision) and human-in-the-loop presence.
      • Simulation engine platform, targeted social domains (e.g., negotiation, collaboration, chitchat), and agent memory or profile structures.
    2. Intended Use: Explicit primary use cases (training data distillation, social intelligence benchmarking, sociological analysis) and secondary application domains.
    3. Evaluation Metrics: Targeted metrics categorized into human-like interaction fidelity (naturalness, verbosity), agent goal achievement (goal completion score, agreement rate), and adherence to social norms and safety guidelines.
    4. Ethical Considerations: Risks related to anthropomorphism, manipulative persuasion, and deceptive conversational agents.
    5. Caveats and Recommendations: Known failure modes, information leakage biases, and constraints on out-of-distribution generalization.
  10. Knowl 10 — Structural Limitations of Information-Asymmetric Interaction Benchmarks in LLMs

    limitation

    Evaluating and training LLM social agents under realistic communication constraints has several key methodological limitations:

    • Model-Based Metric Biases: Automated evaluation using LLMs (e.g., GPT-4) as social judges exhibits stylistic biases, including preferences for specific lexical styles or lengths, and cannot fully substitute for human pragmatic assessment.
    • Absence of Pragmatic Interaction Mechanics: Turn-based frameworks abstract away fundamental features of human conversation, including dynamic turn-taking, interjections, asynchronous messaging, multi-party interactions (>2>2 agents), and persistent episodic memory across sessions.
    • Prompt Engineering and Domain Coverage: Prompts require explicit schemas (e.g., JSON output formatting and action types) that create competing constraints, hindering natural speech brevity. Furthermore, benchmarks primarily cover dyadic text interactions in English, excluding physical coordination tasks (e.g., joint physical assembly) and multilingual or code-switching scenarios.

Coverage note — No substantial contributed material was omitted. The knowls cover the simulation modes, experimental comparisons on goal completion and naturalness, fine-tuning methodology and results, task-specific diagnostic analyses (information leakage and over-agreeableness), detailed benchmark metrics, the Simulation Card framework, and stated limitations.

References

  1. 1.Asya Achimova, Michael Franke, and Martin V Butz. 2023. Indirectness as a path to common ground management.
  2. 2.J L Austin. 1975. How to do things with words: Second edition, 2 edition. The William James Lectures. Harvard University Press, London, England.
  3. 3.Karen Bartsch and Henry M. Wellman. 1995. Children Talk About the Mind. Oxford University Press.
  4. 4.Federico Bianchi, Patrick John Chia, Mert Yuksekgonul, Jacopo Tagliabue, Dan Jurafsky, and James Zou. 2024. How well can llms negotiate? negotiation-arena platform and analysis.
  5. 5.Sophie Elizabeth Colby Bridgers, Maya Taliaferro, Kiera Parece, Laura Schulz, and Tomer Ullman. 2023. Loopholes: A window into value alignment and the communication of meaning.
  6. 6.Fausto Carcassi and Michael Franke. 2023. How to handle the truth: A model of politeness as strategic truth-stretching. Proceedings of the Annual Meeting of the Cognitive Science Society, 45(45).
  7. 7.Maximillian Chen, Alexandros Papangelis, Chenyang Tao, Seokhwan Kim, Andrew Rosenbaum, Yang Liu, Zhou Yu, and Dilek Z. Hakkani-Tür. 2023a. Places: Prompting language models for social conversation synthesis. In Findings.
  8. 8.Maximillian Chen, Alexandros Papangelis, Chenyang Tao, Seokhwan Kim, Andy Rosenbaum, Yang Liu, Zhou Yu, and Dilek Hakkani-Tur. 2023b. PLACES: Prompting language models for social conversation synthesis. In Findings of the Association for Computational Linguistics: EACL 2023, pages 844–868, Dubrovnik, Croatia. Association for Computational Linguistics.
  9. 9.Maximillian Chen, Alexandros Papangelis, Chenyang Tao, Seokhwan Kim, Andy Rosenbaum, Yang Liu, Zhou Yu, and Dilek Hakkani-Tur. 2023c. PLACES: Prompting language models for social conversation synthesis. In Findings of EACL 2023.
  10. 10.Herbert H Clark. 1996. Using Language. Cambridge University Press.
  11. 11.Debarati Das, Karin De Langis, Anna Martin, Jaehyung Kim, Minhwa Lee, Zae Myung Kim, Shirley Hayati, Risako Owan, Bin Hu, Ritik Parkar, et al. 2024. Under the surface: Tracking the artifactuality of llm-generated data. arXiv preprint arXiv:2401.14698.
  12. 12.Daniel C Dennett. 1978. Beliefs about beliefs. Behav. Brain Sci., 1(4):568–570.
  13. 13.Ameet Deshpande, Tanmay Rajpurohit, Karthik Narasimhan, and Ashwin Kalyan. 2023. Anthropomorphization of ai: Opportunities and risks.
  14. 14.M Franke. 2009. Signal to act: Game theory in pragmatics. Ph.D. thesis, Universiteit van Amsterdam, Amsterdam.
  15. 15.Nigel Gilbert. 2005. Simulation for the Social Scientist, 2 edition. Open University Press.
  16. 16.Noah D Goodman and Michael C Frank. 2016. Pragmatic Language Interpretation as Probabilistic Inference. Trends in cognitive sciences, 20(11):818–829.
  17. 17.Robert D Hawkins, Hyowon Gweon, and Noah D Goodman. 2021. The division of labor in communication: Speakers help listeners account for asymmetries in visual perspective. Cognitive science, 45(3):e12926.
  18. 18.He He, Anusha Balakrishnan, Mihail Eric, and Percy Liang. 2017. Learning symmetric collaborative dialogue agents with dynamic knowledge graph embeddings. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1766–1776, Vancouver, Canada. Association for Computational Linguistics.
  19. 19.He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. 2018. Decoupling strategy and generation in negotiation dialogues. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2333–2343, Brussels, Belgium. Association for Computational Linguistics.
  20. 20.Dan Hendrycks, Mantas Mazeika, and Thomas Woodside. 2023. An overview of catastrophic ai risks.
  21. 21.Joey Hong, Sergey Levine, and Anca Dragan. 2023. Zero-shot goal-directed dialogue via rl on imagined conversations. ArXiv, abs/2311.05584.
  22. 22.Mohammad Javad Hosseini, Filip Radlinski, Silvia Pareti, and Annie Louis. 2023. Resolving indirect referring expressions for entity selection. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12313–12335, Stroudsburg, PA, USA. Association for Computational Linguistics.
  23. 23.Jennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko, and Edward Gibson. 2022. A fine-grained comparison of pragmatic language understanding in humans and language models. arXiv [cs.CL].
  24. 24.Qingxu Huang, Dawn C Parker, Tatiana Filatova, and Shipeng Sun. 2014. A review of urban residential choice models using Agent-Based modeling. Environment and planning. B, Planning & design, 41(4):661–689.
  25. 25.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts.
  26. 26.Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, and Yejin Choi. 2023a. SODA: Million-scale dialogue distillation with social commonsense contextualization. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12930–12949, Singapore. Association for Computational Linguistics.
  27. 27.Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. 2023b. FANToM: A benchmark for stress-testing machine theory of mind in interactions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14397–14413, Singapore. Association for Computational Linguistics.
  28. 28.Stephen C. Levinson. 2016. Turn-taking in human communication – origins and implications for language processing. Trends in Cognitive Sciences, 20(1):6–14.
  29. 29.Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023a. Camel: Communicative agents for "mind" exploration of large language model society. In Thirty-seventh Conference on Neural Information Processing Systems.
  30. 30.Yuan Li, Yixuan Zhang, and Lichao Sun. 2023b. Metaagents: Simulating interactions of human behaviors for llm-based task-oriented coordination via collaborative generative agents.
  31. 31.Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023. Encouraging divergent thinking in large language models through multi-agent debate. ArXiv, abs/2305.19118.
  32. 32.Jessy Lin, Daniel Fried, Dan Klein, and Anca Dragan. 2022. Inferring rewards from language in context.
  33. 33.Benjamin Lipkin, Lionel Wong, Gabriel Grand, and Joshua B Tenenbaum. 2023. Evaluating statistical language models as pragmatic reasoners. arXiv [cs.CL].
  34. 34.Ian H. Magnusson, Noah A. Smith, and Jesse Dodge. 2023. Reproducibility in nlp: What have we learned from the checklist? In Annual Meeting of the Association for Computational Linguistics.
  35. 35.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT ’19*, page 220–229, New York, NY, USA. Association for Computing Machinery.
  36. 36.Lauren A Oey, Adena Schachner, and Edward Vul. 2023. Designing and detecting lies by reasoning about other agents. Journal of experimental psychology. General, 152(2):346–362.
  37. 37.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  38. 38.Xianghe Pang, Shuo Tang, Rui Ye, Yuxin Xiong, Bolun Zhang, Yanfeng Wang, and Siheng Chen. 2024. Self-alignment of large language models via monopolylogue-based social scene simulation.
  39. 39.Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In In the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23), UIST ’23, New York, NY, USA. Association for Computing Machinery.
  40. 40.Joon Sung Park, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. In In the 35th Annual ACM Symposium on User Interface Software and Technology (UIST ’22), UIST ’22, New York, NY, USA. Association for Computing Machinery.
  41. 41.Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R Bowman, and Tal Linzen. 2021. NOPE: A corpus of naturally-occurring presuppositions in english. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 349–366, Stroudsburg, PA, USA. Association for Computational Linguistics.
  42. 42.Steven Pinker, Martin A Nowak, and James J Lee. 2008. The logic of indirect speech. Proceedings of the National Academy of Sciences of the United States of America, 105(3):833–838.
  43. 43.David Premack and Guy Woodruff. 1978. Does the chimpanzee have a theory of mind? The Behavioral and brain sciences, 1(4):515–526.
  44. 44.Setayesh Radkani, Josh Tenenbaum, and Rebecca Saxe. 2022. Modeling punishment as a rational communicative social action. Proceedings of the Annual Meeting of the Cognitive Science Society, 44(44).
  45. 45.Laura Eline Ruis, Akbir Khan, Stella Biderman, Sara Hooker, Tim Rocktäschel, and Edward Grefenstette. 2023. The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by LLMs. In Thirty-seventh Conference on Neural Information Processing Systems.
  46. 46.Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023. Verbosity bias in preference labeling by large language models. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.
  47. 47.Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2019. Green ai.
  48. 48.Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2023. Towards understanding sycophancy in language models.
  49. 49.Kurt Shuster, Jack Urbanek, Arthur Szlam, and Jason Weston. 2022. Am I me or you? state-of-the-art dialogue models cannot maintain an identity. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 2367–2387, Seattle, United States. Association for Computational Linguistics.
  50. 50.Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020. Can you put it all together: Evaluating conversational agents’ ability to blend skills. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2021–2030, Online. Association for Computational Linguistics.
  51. 51.Karthik Sreedhar and Lydia Chilton. 2024. Simulating human strategic behavior: Comparing single and multi-agent llms.
  52. 52.Robert Stalnaker. 2014. Context. Oxford University Press.
  53. 53.Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L Griffiths. 2023a. Cognitive Architectures for Language Agents.
  54. 54.Theodore R Sumers, Mark K Ho, Thomas L Griffiths, and Robert D Hawkins. 2023b. Reconciling truthfulness and relevance as epistemic and decision-theoretic utility. Psychological review.
  55. 55.Leigh Tesfatsion and Kenneth L Judd. 2006. Handbook of Computational Economics: Agent-Based Computational Economics. Elsevier.
  56. 56.Michael Tomasello. 1999. The Cultural Origins of Human Cognition. Harvard University Press.
  57. 57.Michael Tomasello. 2021. Becoming Human: A Theory of Ontogeny. Belknap Press.
  58. 58.Dennis Ulmer, Elman Mansimov, Kaixiang Lin, Justin Sun, Xibin Gao, and Yi Zhang. 2024. Bootstrapping llm-based task-oriented dialogue agents via self-talk. ArXiv, abs/2401.05033.
  59. 59.Zhilin Wang, Yu Ying Chiu, and Yu Cheung Chiu. 2023. Humanoid agents: Platform for simulating human-like generative agents. In EMNLP System Demonstrations.
  60. 60.Max Weber. 1978. The Nature of Social Action, page 7–32. Cambridge University Press.
  61. 61.Lionel Wong, Gabriel Grand, Alexander K Lew, Noah D Goodman, Vikash K Mansinghka, Jacob Andreas, and Joshua B Tenenbaum. 2023. From word models to world models: Translating from natural language to the probabilistic language of thought.
  62. 62.Lance Ying, Tan Zhi-Xuan, Vikash Mansinghka, and Joshua B Tenenbaum. 2023. Inferring the goals of communicating agents from actions and instructions. arXiv [cs.AI].
  63. 63.Erica J Yoon, Michael Henry Tessler, Noah D Goodman, and Michael C Frank. 2020. Polite Speech Emerges From Competing Social Goals. Open mind : discoveries in cognitive science, 4(4):71–87.
  64. 64.Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Yejin Choi, and Noah A. Smith. 2021. Challenges in automated debiasing for toxic language detection. In EACL.
  65. 65.Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Zhengyang Qi, Haofei Yu, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. 2024. Sotopia: Interactive evaluation for social intelligence in language agents. In ICLR.

Citation

MLA
Zhou, X., et al. “Is This the Real Life? Is This Just Fantasy? The Misleading Success of Simulating Social Interactions With LLMs”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 21692–714, https://doi.org/10.18653/v1/2024.emnlp-main.1208.
APA
Zhou, X., Su, Z., Eisape, T., Kim, H., & Sap, M. (2024). Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21692–21714. https://doi.org/10.18653/v1/2024.emnlp-main.1208
Chicago
Zhou, X., Z. Su, T. Eisape, H. Kim, and M. Sap. 2024. “Is This the Real Life? Is This Just Fantasy? The Misleading Success of Simulating Social Interactions With LLMs”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21692–714. https://doi.org/10.18653/v1/2024.emnlp-main.1208.
Harvard
Zhou, X. et al. (2024) “Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 21692–21714. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1208.
Vancouver
1. Zhou X, Su Z, Eisape T, Kim H, Sap M (2024) Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 21692–21714

BibTeX

@inproceedings{zhou-etal-2024-real,
    title = "Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With {LLM}s",
    author = "Zhou, Xuhui  and
      Su, Zhe  and
      Eisape, Tiwalayo  and
      Kim, Hyunwoo  and
      Sap, Maarten",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1208/",
    doi = "10.18653/v1/2024.emnlp-main.1208",
    pages = "21692--21714"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/