Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
Xuhui ZhouZhe SuTiwalayo EisapeHyunwoo KimMaarten Sap
Reveals that large language models perform well in unrealistic, omniscient dialogue simulations but struggle to achieve goals and converse naturally in realistic settings with information asymmetry, exposing a critical flaw in current social simulation benchmarks.
Large language models are increasingly used to simulate human social behavior, evaluate artificial intelligence capabilities, and generate synthetic conversational training data. However, real-world human interactions depend heavily on information asymmetry, where participants do not have direct access to each other's private goals, motives, or backgrounds. In contrast, many popular simulation methods adopt an omniscient setup where a single model generates the entire interaction with full access to all participants' internal states, creating a disconnect from realistic human communication.
The article evaluates whether omniscient social simulations accurately reflect the social abilities of language models under realistic conditions of information asymmetry. It also investigates whether fine-tuning language models on data generated by omniscient simulations enables them to perform effectively as autonomous interactive agents.
To conduct this evaluation, the researchers tested models including GPT-3.5 and Mixtral-8x7B across 90 social scenarios in the Sotopia framework. They compared three simulation configurations: an omniscient single-model setup that writes a complete script, an interactive multi-agent setup where individual models only know their own private profiles and goals, and an intermediate ablation setting where agents access each other's private information. The evaluation measured social goal completion scores assessed by GPT-4 across 450 simulation episodes per model, combined with human evaluations and turn-length measurements of dialogue naturalness and verbosity.
The investigation produced four central findings. First, omniscient simulations substantially overestimate social competence; models in the omniscient mode achieved average goal completion scores of approximately 8.4 out of 10, whereas the realistic multi-agent mode scored only 6.9 to 7.5. Second, human evaluators found omniscient simulations significantly more natural than multi-agent interactions, which suffered from verbosity and averaged nearly twice as many words per conversational turn. Third, omniscient simulations rely on unrealistic heuristics, such as mentioning secret mutual information very early in conversations (at relative position 0.13 versus 0.39 in multi-agent mode) and exhibiting excessive agreeableness by closing negotiation deals in 94% of competitive scenarios compared to just 30% in multi-agent mode. Fourth, fine-tuning an agent on omniscient scripts improved conversational conciseness and naturalness but failed to improve strategic goal completion in cooperative tasks, as models merely adopted superficial styles rather than robust reasoning skills under uncertainty.
These findings indicate that relying on omniscient simulation benchmarks introduces significant risks of overestimating artificial intelligence readiness for customer-facing or human-interactive deployments. Training systems on synthetic omniscient dialogues transfers flawed behavioral shortcuts, such as premature information leakage and unrealistic agreeableness, which can compromise negotiation performance, safety, and operational reliability.
Decision-makers and researchers should avoid using omniscient simulations as direct proxies for evaluating interactive agent capabilities. Studies and system documentation should adopt standardized reporting tools, such as the proposed simulation card, to explicitly disclose information access levels and multi-agent configurations. Future technical efforts should focus on training agents to reason explicitly about interlocutors' hidden mental states and managing conversational context rather than relying solely on omniscient synthetic data distillation.
The conclusions should be interpreted with awareness of certain methodological boundaries. The quantitative goal assessments rely primarily on automated model-based evaluation, the prompt designs abstract away dynamics like turn-taking mechanics, and the scenarios are restricted to English-language interactions. Nevertheless, the evidence provides high confidence that information asymmetry remains a primary bottleneck for language models operating as autonomous social agents.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). It introduces communicative agent role-playing setups for social simulation, establishing the paradigm whose omniscient evaluation vulnerabilities are critiqued in the source paper.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). It details how language models reason about hidden mental states and private beliefs using causal graphs, which directly motivates the source's investigation of information asymmetry.
- Paper: Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker, Melanie Sclar et al. (2023). It establishes belief-tracking mechanisms under divergent character knowledge, highlighting the fundamental challenge of managing non-omniscient perspective states.
- Paper: Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies, Gati V. Aher et al. (2023). It presents foundational methodologies for simulating human social and economic behaviors using large language models, providing the experimental baseline the source critically re-evaluates.
- Paper: Theory of Mind for Multi-Agent Collaboration via Large Language Models, Huao Li et al. (2023). It explores Theory of Mind inference specifically within multi-agent collaborative environments, serving as a direct prerequisite for understanding multi-agent social interactions.
- Paper: Towards Understanding Sycophancy in Language Models, Mrinank Sharma et al. (2023). It documents the systemic agreeableness and sycophancy in aligned language models that explain the unrealistic deal-closing heuristics uncovered in the source's simulations.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). It formalizes the LLM-as-a-judge paradigm and its evaluation biases, which underpins the automated GPT-4 social goal scoring methodology used throughout the source.
- Paper: LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay, Yihuai Lan et al. (2024). It applies multi-agent interaction architectures to deduction games like Avalon, extending the source's insights on strategic social reasoning under severe information asymmetry and adversarial deception.
- Paper: Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation, Xianghe Pang et al. (2024). It implements monopolylogue-based social scene simulation for self-alignment, offering a concrete practical application of multi-character simulation that contends with the dynamics examined in the source.
- Paper: MMToM-QA: Multimodal Theory of Mind Question Answering, Chuanyang Jin et al. (2024). It extends Theory of Mind evaluation into multimodal video and dialogue settings, broadening the assessment of hidden human intent beyond text-only social scenarios.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). It provides a comprehensive architectural survey of agentic reasoning and multi-agent systems, integrating the source's warnings regarding information access and simulation cards into broader agent design.
