Understanding Social Reasoning in Language Models with Language Models
Kanishk GandhiJan-Philipp FränkenTobias GerstenbergNoah D. Goodman
Introduces a causal-template framework and the 5,000-scenario BigToM benchmark to rigorously evaluate Theory-of-Mind reasoning in large language models, revealing that advanced models like GPT-4 exhibit human-like social inference patterns while remaining less reliable.
As large language models become increasingly integrated into daily applications, evaluating their social reasoning abilities is essential for safe, reliable human-AI collaboration. Humans rely on Theory of Mind, the ability to track and infer the hidden mental states, beliefs, and desires of others to predict their actions. Prior assessments of machine social reasoning produced conflicting results because earlier tests were often ambiguous, lacked robust control conditions, or relied on narrow psychology tasks vulnerable to data leakage.
The article aims to evaluate the social reasoning capabilities of various large language models using a scalable, procedurally generated evaluation framework based on causal models. Specifically, it introduces a new benchmark called BigToM to establish whether models exhibit reliable, human-like mental state inference across controlled scenarios.
To construct this benchmark, the researchers used a three-stage method: establishing an abstract causal graph of agent beliefs, desires, and actions; prompting GPT-4 to populate concrete scenario variables; and stitching these variables into 5,000 fluent test items across 25 control conditions. The items probe forward inferences (predicting beliefs or actions from observations) and backward inferences (deducing hidden beliefs from observed actions). Both human experts and crowdworkers evaluated the generated dataset, rating its clarity and coherence higher than existing crowdsourced benchmarks and comparable to expert-written tests. The benchmark was then used to evaluate multiple advanced language models against human baselines.
The evaluation revealed several key findings regarding model capabilities. First, GPT-4 demonstrated social reasoning patterns that closely mirror human reasoning, achieving 90% to 97% combined accuracy on forward belief tasks and up to 100% on forward action predictions when prompted with single examples. Second, earlier and competing models struggled substantially, frequently failing on false-belief tasks where an agent's internal knowledge diverges from actual reality. Third, explicitly stating an agent's initial belief often biased models toward anchoring on outdated information rather than updating beliefs after environmental changes. Fourth, backward belief inference—inferring hidden beliefs purely from observed actions—proved to be the most difficult challenge for all systems; while humans achieved 72% to 82% accuracy, unassisted models scored far lower, with GPT-4 reaching only 40% combined accuracy in zero-shot settings.
These findings indicate that while advanced systems like GPT-4 possess nascent social reasoning capabilities, relying on language models for complex social coordination carries substantial risk. In deployment settings requiring accurate interpretation of human intent and unstated assumptions, most models remain brittle and prone to error. Decision-makers should not assume automated social comprehension without rigorous validation.
Organizations deploying conversational agents should implement strict guardrails and provide explicit contextual demonstrations, which consistently improve inference reliability over unassisted prompting. Future development should prioritize dynamic evaluation benchmarks and real-time interactive simulations rather than relying solely on static vignettes. Although synthetic test generation introduces minor risks of inherited model biases, the benchmark's systematic controls and cross-model validation provide high confidence in these findings.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). This paper establishes the foundational BIG-bench evaluation suite and human-baseline methodology that directly inspired BigToM's procedural benchmarking framework for assessing complex model reasoning.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). This work isolates challenging reasoning subtasks where standard language model evaluations fail, motivating BigToM's focus on rigorously controlled social reasoning benchmarks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This foundational study demonstrates how step-by-step reasoning prompts elicit complex problem-solving in large language models, providing key context for evaluating multi-step Theory-of-Mind inferences.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). This benchmark reveals how large language models mimic human falsehoods and misconceptions, establishing crucial groundwork for evaluating whether models genuinely track mental states versus superficial cues.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). This paper introduces controlled diagnostic evaluation datasets to catch shortcut heuristics in language models, directly informing the procedural template and control design of BigToM.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). This paper extends the exploration of multi-perspective social and cognitive dynamics by analyzing how advanced reasoning models simulate internal multi-agent dialogues to solve complex problems.
- Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). This survey builds upon Theory-of-Mind and mental-state modeling by evaluating large language models' capacity for higher-order metacognition and self-monitoring.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This work formalizes and stress-tests the LLM-as-a-judge paradigm that underpins automated, model-generated evaluation pipelines like BigToM.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This study critiques and refines automated LLM-based evaluation frameworks by exposing where model judges get misled by surface-level polish rather than true task adherence.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). This benchmark advances rigorous reasoning evaluation by using model-assisted generation and expanded distractors to overcome evaluation saturation and heuristic guessing.
- Paper: DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots, Jared Moore et al. (2026). This evaluation applies social reasoning and mental-state interaction benchmarks to measure how chatbots reinforce human psychological distress and delusional loops in real interactions.
