Understanding Social Reasoning in Language Models with Language Models

Kanishk GandhiJan-Philipp FränkenTobias GerstenbergNoah D. Goodman

article2023NeurIPS207 citations

Introduces a causal-template framework and the 5,000-scenario BigToM benchmark to rigorously evaluate Theory-of-Mind reasoning in large language models, revealing that advanced models like GPT-4 exhibit human-like social inference patterns while remaining less reliable.

Listen

As large language models become increasingly integrated into daily applications, evaluating their social reasoning abilities is essential for safe, reliable human-AI collaboration. Humans rely on Theory of Mind, the ability to track and infer the hidden mental states, beliefs, and desires of others to predict their actions. Prior assessments of machine social reasoning produced conflicting results because earlier tests were often ambiguous, lacked robust control conditions, or relied on narrow psychology tasks vulnerable to data leakage.

The article aims to evaluate the social reasoning capabilities of various large language models using a scalable, procedurally generated evaluation framework based on causal models. Specifically, it introduces a new benchmark called BigToM to establish whether models exhibit reliable, human-like mental state inference across controlled scenarios.

To construct this benchmark, the researchers used a three-stage method: establishing an abstract causal graph of agent beliefs, desires, and actions; prompting GPT-4 to populate concrete scenario variables; and stitching these variables into 5,000 fluent test items across 25 control conditions. The items probe forward inferences (predicting beliefs or actions from observations) and backward inferences (deducing hidden beliefs from observed actions). Both human experts and crowdworkers evaluated the generated dataset, rating its clarity and coherence higher than existing crowdsourced benchmarks and comparable to expert-written tests. The benchmark was then used to evaluate multiple advanced language models against human baselines.

The evaluation revealed several key findings regarding model capabilities. First, GPT-4 demonstrated social reasoning patterns that closely mirror human reasoning, achieving 90% to 97% combined accuracy on forward belief tasks and up to 100% on forward action predictions when prompted with single examples. Second, earlier and competing models struggled substantially, frequently failing on false-belief tasks where an agent's internal knowledge diverges from actual reality. Third, explicitly stating an agent's initial belief often biased models toward anchoring on outdated information rather than updating beliefs after environmental changes. Fourth, backward belief inference—inferring hidden beliefs purely from observed actions—proved to be the most difficult challenge for all systems; while humans achieved 72% to 82% accuracy, unassisted models scored far lower, with GPT-4 reaching only 40% combined accuracy in zero-shot settings.

These findings indicate that while advanced systems like GPT-4 possess nascent social reasoning capabilities, relying on language models for complex social coordination carries substantial risk. In deployment settings requiring accurate interpretation of human intent and unstated assumptions, most models remain brittle and prone to error. Decision-makers should not assume automated social comprehension without rigorous validation.

Organizations deploying conversational agents should implement strict guardrails and provide explicit contextual demonstrations, which consistently improve inference reliability over unassisted prompting. Future development should prioritize dynamic evaluation benchmarks and real-time interactive simulations rather than relying solely on static vignettes. Although synthetic test generation introduces minor risks of inherited model biases, the benchmark's systematic controls and cross-model validation provide high confidence in these findings.

arXiv: 2306.15448
Cover for Understanding Social Reasoning in Language Models with Language Models

Abstract

As Large Language Models (LLMs) become increasingly integrated into our everyday lives, understanding their ability to comprehend human mental states becomes critical for ensuring effective interactions. However, despite the recent attempts to assess the Theory-of-Mind (ToM) reasoning capabilities of LLMs, the degree to which these models can align with human ToM remains a nuanced topic of exploration. This is primarily due to two distinct challenges: (1) the presence of inconsistent results from previous evaluations, and (2) concerns surrounding the validity of existing evaluation methodologies. To address these challenges, we present a novel framework for procedurally generating evaluations with LLMs by populating causal templates. Using our framework, we create a new social reasoning benchmark (BigToM) for LLMs which consists of 25 controls and 5,000 model-written evaluations. We find that human participants rate the quality of our benchmark higher than previous crowd-sourced evaluations and comparable to expert-written evaluations. Using BigToM, we evaluate the social reasoning capabilities of a variety of LLMs and compare model performances with human performance. Our results suggest that GPT4 has ToM capabilities that mirror human inference patterns, though less reliable, while other LLMs struggle.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Model-Written Evaluations with Causal Templates
  • 3.1 Stage 1: Building a Causal Template
  • 3.2 Stage 2: Populating Causal Templates With Language Models
  • 3.3 Stage 3: Composing Test Items from Template Variables
  • 3.4 Quality of Generated Data
  • 4 Experiments
  • 4.1 Results and Discussion
  • 5 Discussion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Procedural Generation of Theory-of-Mind Evaluations from Causal Templates

    model/method

    To evaluate Theory-of-Mind (ToM) reasoning without human-authored test sets or noisy unstructured crowdsourced data, social scenarios are generated via a three-stage causal template framework:

    1. Causal Template Formulation: An abstract causal graphical model of mental state reasoning is constructed containing variables for background context, agent desires, agent actions, agent percepts, agent beliefs, environmental causal events, and subsequent percepts and actions.
    2. Template Variable Population: A large language model (gpt-4-0314, with sampling temperature T=0.5T = 0.5, default hyperparameters, and 3 few-shot in-context demonstrations) is prompted to instantiate each variable in the abstract graph with exactly one natural language sentence. The prompt requires generating the prior situation, a causal event altering the world state, alternative agent percepts (observing vs. not observing the event), alternative actions, target questions, and correct/incorrect answer options.
    3. Test Item Composition (Stitching): Concrete multiple-choice test items consisting of a story, a question, and two candidate answers (s,q,a)(s, q, a) are assembled by systematically concatenating specific variable subsets from the populated causal template according to the target experimental condition (such as true-belief vs. false-belief or forward vs. backward inference).
  2. Knowl 2 — Causal Formulations of Theory-of-Mind Reasoning Tasks

    equation

    In a causal model of social agency where an agent's mental state and behavior are defined over random variables for Desire DD, Percept PP, Belief BB, and Action AA, Theory-of-Mind queries are formalized as distinct causal inferences:

    1. Initial Percept to Initial Belief: Probes the baseline transition from an agent's initial sensory perception P0P_0 (and initial action A0A_0) to the initial belief B0B_0: P(B0∣P0)P(B_0 \mid P_0)

    2. Forward Belief Inference: Evaluates the probability distribution over an agent's updated belief state BB given the agent's perception PP of a world-altering causal event: P(B∣P)P(B \mid P)

    3. Forward Action Inference: Evaluates the prediction of an agent's subsequent action AA from observed percept PP and desire DD, requiring marginalization over the latent belief state BB: P(A∣P,D)=∑BP(A∣P,D,B)P(A \mid P, D) = \sum_{B} P(A \mid P, D, B)

    4. Backward Belief Inference: Evaluates the inverse inference of an agent's latent belief BB given an observed action AA and desire DD, requiring joint marginalization over latent percepts PP and beliefs: P(B∣A,D)∝∑PP(A∣D,P,B)P(B∣P)P(P)P(B \mid A, D) \propto \sum_{P} P(A \mid D, P, B) P(B \mid P) P(P) where the joint likelihood of action AA given desire DD expands as: ∑P∑BP(A∣D,P,B)\sum_{P} \sum_{B} P(A \mid D, P, B)

  3. Knowl 3 — BigToM Benchmark Specification

    experimental setup

    The BigToM benchmark is a controlled dataset designed to evaluate social and Theory-of-Mind reasoning in Large Language Models across 5,000 test items generated from 200 base causal templates populated by GPT-4. Each template yields 25 distinct experimental conditions by combining variables and intervening on scenario components:

    • Inference Dimensions: Forward Belief inference (predicting belief from percepts), Forward Action inference (predicting action from percepts and desires via latent beliefs), and Backward Belief inference (inferring beliefs from observed actions).
    • Epistemic States: True Belief (TB, where the agent perceives the world-altering causal event) vs. False Belief (FB, where the agent is unaware of the causal change).
    • Prior Context Variations: 'With Initial Belief' (the agent's starting belief prior to the causal event is explicitly stated in the scenario context) vs. 'Without Initial Belief' (the initial belief must be inferred directly from the agent's prior actions and percepts).
    • Control Conditions: Control scenarios where the world-altering 'Causal Event' is replaced with a 'Random Event' that leaves the environment state unchanged, isolating ToM reasoning from general context interference.

    Every test item is presented as a reading comprehension scenario accompanied by a target question and two mutually exclusive answer options.

  4. Knowl 4 — False-Belief Contingent Evaluation Metric

    definition

    In Theory-of-Mind evaluation, an evaluation metric denotes success on a false-belief condition only if the model also answers the corresponding true-belief condition correctly, represented as: Accuracy(TB∧FB)\text{Accuracy}(\text{TB} \land \text{FB})

    Evaluating false-belief accuracy unconditionally can yield false positives if a model systematically fails to comprehend environmental state updates or defaults to historical state priors. Conditioning false-belief success on true-belief accuracy ensures that correct false-belief predictions reflect genuine attribution of unaligned mental states rather than language comprehension failure regarding the underlying physical change.

  5. Knowl 5 — GPT-4 Performance on the BigToM Benchmark

    data/table

    The table below details the performance of GPT-4 (gpt-4-0314, temperature 0) across Theory-of-Mind inference conditions on the BigToM benchmark under four prompting regimes: zero-shot (0-shot), zero-shot Chain-of-Thought (0-shot-cot), one-shot (1-shot), and one-shot Chain-of-Thought (1-shot-cot). Tasks are evaluated both without an explicitly stated initial belief (†\dagger) and with an explicitly stated initial belief (‡\ddagger).

    Condition Contingency 0-shot 0-shot-cot 1-shot 1-shot-cot
    Fwd. Belief TB .99†/.91‡.99^\dagger / .91^\ddagger .99†/.99‡.99^\dagger / .99^\ddagger .99†/.97‡.99^\dagger / .97^\ddagger 1.00†/.97‡1.00^\dagger / .97^\ddagger
    FB .98†/.99‡.98^\dagger / .99^\ddagger .99†/.99‡.99^\dagger / .99^\ddagger .99†/.99‡.99^\dagger / .99^\ddagger .99†/.99‡.99^\dagger / .99^\ddagger
    TB∧FB\text{TB} \land \text{FB} .97†/.90‡.97^\dagger / .90^\ddagger .98†/.98‡.98^\dagger / .98^\ddagger .97†/.96‡.97^\dagger / .96^\ddagger .99†/.96‡.99^\dagger / .96^\ddagger
    Fwd. Action TB .98†/.98‡.98^\dagger / .98^\ddagger .99†/.99‡.99^\dagger / .99^\ddagger 1.00†/1.00‡1.00^\dagger / 1.00^\ddagger 1.00†/1.00‡1.00^\dagger / 1.00^\ddagger
    FB .81†/.92‡.81^\dagger / .92^\ddagger .88†/.96‡.88^\dagger / .96^\ddagger .98†/1.00‡.98^\dagger / 1.00^\ddagger 1.00†/.99‡1.00^\dagger / .99^\ddagger
    TB∧FB\text{TB} \land \text{FB} .79†/.90‡.79^\dagger / .90^\ddagger .87†/.95‡.87^\dagger / .95^\ddagger .98†/1.00‡.98^\dagger / 1.00^\ddagger 1.00†/.99‡1.00^\dagger / .99^\ddagger
    Bwd. Belief TB .86†/.62‡.86^\dagger / .62^\ddagger .84†/.76‡.84^\dagger / .76^\ddagger .68†/.57‡.68^\dagger / .57^\ddagger .83†/.81‡.83^\dagger / .81^\ddagger
    FB .53†/.77‡.53^\dagger / .77^\ddagger .54†/.63‡.54^\dagger / .63^\ddagger .85†/.92‡.85^\dagger / .92^\ddagger .75†/.85‡.75^\dagger / .85^\ddagger
    TB∧FB\text{TB} \land \text{FB} .40†/.40‡.40^\dagger / .40^\ddagger .38†/.40‡.38^\dagger / .40^\ddagger .53†/.49‡.53^\dagger / .49^\ddagger .58†/.65‡.58^\dagger / .65^\ddagger

    GPT-4 achieves near-ceiling performance on Forward Belief and Forward Action tasks (exceeding 0.90 across most prompt setups in TB∧FB\text{TB} \land \text{FB}). However, on Backward Belief inference, GPT-4 drops to a zero-shot TB∧FB\text{TB} \land \text{FB} accuracy of 0.40 (both with and without initial belief), reflecting severe difficulty when required to invert causal actions to infer latent beliefs and unobserved perceptions.

  6. Knowl 6 — Comparative Theory-of-Mind Capabilities Across Language Models and Humans

    empirical result

    Evaluations across language models (LLaMA-65B, text-davinci-003, gpt-3.5-turbo, Claude-v1.3, Claude-2, and gpt-4-0314) and human baselines on the BigToM benchmark demonstrate clear performance stratification:

    1. Initial Percept-to-Belief Mapping: All tested models reliably understand that an agent's initial physical actions and perceptions give rise to corresponding initial beliefs.
    2. Forward Inferences: Most models other than GPT-4 and Claude struggle with False Belief conditions in Forward Belief and Forward Action tasks. When an initial belief is explicitly stated in the scenario context, smaller/earlier models strongly anchor on this stated prior, failing to update beliefs after world-altering causal events. GPT-4 closely matches or slightly exceeds human performance on forward action prediction.
    3. Backward Belief Inference: Inverting observed actions to infer latent mental states is the most challenging task for all evaluators. Human participants achieve 82%82\% accuracy on True Belief and 72%72\% on False Belief conditions. Non-GPT-4 models perform substantially below chance (<50%<50\%) in contingent accuracy (TB∧FB\text{TB} \land \text{FB}), consistently attributing false beliefs to agents even when their observed actions indicate correct true beliefs. GPT-4 is the only model displaying human-like inference dynamics on backward belief tasks, though its zero-shot contingent accuracy (40%40\%) remains well below human level.
  7. Knowl 7 — Quality and Structural Validity of Model-Written Theory-of-Mind Datasets

    empirical result

    Validation of LLM-generated Theory-of-Mind scenarios using expert evaluation and human subject ratings established the following quality benchmarks:

    • Expert Structural Compliance: Two expert annotators independently evaluating 100 model-written templates across 25 conditions (2,500 test items) found 93.94%93.94\% agreement on structural compliance (Question 1 mean ratings: 0.9190.919 [95%95\% CI: 0.8590.859--0.9700.970] and 0.9600.960 [95%95\% CI: 0.9190.919--0.9900.990]). Ratings for whether scenarios accurately tested desired ToM behaviors (Question 2 on a 1--5 Likert scale) averaged 4.334.33 (95%95\% CI: 4.134.13--4.534.53) and 4.354.35 (95%95\% CI: 4.184.18--4.524.52) with a median of 5.
    • Human Participant Comparison: Human ratings of 200 BigToM items against 50 items from the crowdsourced socialIQa benchmark and 50 human expert-written ToM items revealed that BigToM scored highest across understandability, question-answer coherence, un-ambiguity, and aggregate quality. Bayesian linear mixed-effects regression showed BigToM test items were significantly less ambiguous and more coherent than crowdsourced items and performed comparably to or better than expert human-written items.
  8. Knowl 8 — Effects of In-Context Prompting Paradigms on Social Reasoning Performance

    empirical result

    Comparing prompting strategies (0-shot, 0-shot Chain-of-Thought [CoT], 1-shot, and 1-shot CoT using a single Forward Belief False Belief exemplar) reveals distinct effects on ToM performance:

    • 0-Shot CoT: Zero-shot CoT prompting does not yield consistent performance gains over standard 0-shot prompting across ToM conditions and models.
    • 1-Shot Demonstration: Providing a single input-output example consistently improves performance across all evaluated LLMs and inference conditions.
    • 1-Shot CoT: Adding a single CoT demonstration consistently yields the largest performance improvements across conditions (e.g., raising GPT-4 Backward Belief TB∧FB\text{TB} \land \text{FB} accuracy from 0.400.40 to 0.650.65). However, this gain largely reflects template-mimicry of the step-by-step reasoning structure rather than necessarily indicating an intrinsic enhancement in latent Theory-of-Mind capacity.
  9. Knowl 9 — Evaluation Circularity Robustness and Inherent Biases of Model-Generated Benchmarks

    limitation

    Generating evaluation datasets using the same model family being tested introduces potential limitations that are bounded as follows:

    1. Circularity: The generation process tests whether models understand isolated inferential steps when provided structured situation facts. When tested on a benchmark generated by an independent model (Claude-2), gpt-4-0314 achieved equivalent performance and outperformed Claude-2 on Claude-2's own generated data, verifying that benchmark performance is not an artifact of self-generation circularity.
    2. Stereotyping and Representation Biases: LLM-generated scenarios inherit normative social associations and stereotypes present in pretraining corpora, which can cause imbalances in context and role distributions unless guided by steering instructions.
    3. Common-Sense Semantic Errors: Approximately 3%3\% of LLM-generated templates exhibit minor common-sense inconsistencies.
    4. Syntactic Uniformity: Templated prompt stitching produces repetitive syntactic sentence structures across conditions, reducing grammatical diversity.

Coverage note — None was omitted; all key formalisms, generation stages, experimental setups, human/expert validations, comparative model results, and stated limitations are fully covered.

References

  1. 1.Chris L Baker, Noah D Goodman, and Joshua B Tenenbaum. Theory-based social goal inference. In Proceedings of the thirtieth annual conference of the cognitive science society, pages 1447–1452. Cognitive Science Society Austin, TX, 2008.
  2. 2.Chris L Baker, Julian Jara-Ettinger, Rebecca Saxe, and Joshua B Tenenbaum. Rational quantitative attribution of beliefs, desires and percepts in human mentalizing. Nature Human Behaviour, 1(4):0064, 2017.
  3. 3.Simon Baron-Cohen, Alan M Leslie, and Uta Frith. Does the autistic child have a “theory of mind”? Cognition, 21(1):37–46, 1985.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  5. 5.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  6. 6.Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. Language models show human-like content effects on reasoning. arXiv preprint arXiv:2207.07051, 2022.
  7. 7.David Dodell-Feder, Jorie Koster-Hale, Marina Bedny, and Rebecca Saxe. fmri item analysis in a theory of mind task. neuroimage, 55(2):705–712, 2011.
  8. 8.Avia Efrat and Omer Levy. The turking test: Can language models understand instructions? arXiv preprint arXiv:2010.11982, 2020.
  9. 9.Jan-Philipp Fränken, Simon Valentin, Christopher G Lucas, and Neil Bramley. Naïve information aggregation in human social learning. PsyArXiv, 2023.
  10. 10.Chris Frith and Uta Frith. Theory of mind. Current biology, 15(17):R644–R645, 2005.
  11. 11.Kanishk Gandhi, Gala Stojnic, Brenden M Lake, and Moira R Dillon. Baby intuitions benchmark (bib): Discerning the goals, preferences, and actions of others. Advances in Neural Information Processing Systems, 34:9963–9976, 2021.
  12. 12.Georgi Gerganov. llama.cpp. https://github.com/ggerganov/llama.cpp, 2023.
  13. 13.György Gergely and Gergely Csibra. Teleological reasoning in infancy: The naıve theory of rational action. Trends in cognitive sciences, 7(7):287–292, 2003.
  14. 14.Noah D Goodman, Chris L Baker, and Joshua B Tenenbaum. Cause and intent: Social reasoning in causal learning. In Proceedings of the 31st annual conference of the cognitive science society, pages 2759–2764. Cognitive Science Society Austin, TX, 2009.
  15. 15.Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan. Cooperative inverse reinforcement learning. Advances in neural information processing systems, 29, 2016.
  16. 16.Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. arXiv preprint arXiv:2203.09509, 2022.
  17. 17.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  18. 18.Michal Kosinski. Theory of mind may have spontaneously emerged in large language models. arXiv preprint arXiv:2302.02083, 2023.
  19. 19.Eliza Kosoy, Adrian Liu, Jasmine L Collins, David Chan, Jessica B Hamrick, Nan Rosemary Ke, Sandy Huang, Bryanna Kaufmann, John Canny, and Alison Gopnik. Learning causal overhypotheses through exploration in children and computational models. In Conference on Causal Learning and Reasoning, pages 390–406. PMLR, 2022.
  20. 20.Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building machines that learn and think like people. Behavioral and brain sciences, 40:e253, 2017.
  21. 21.Matt Le, Y-Lan Boureau, and Maximilian Nickel. Revisiting the evaluation of theory of mind through question answering. In Conference on Empirical Methods in Natural Language Processing, 2019.
  22. 22.Alan M Leslie, Ori Friedman, and Tim P German. Core mechanisms in ‘theory of mind’. Trends in cognitive sciences, 8(12):528–533, 2004.
  23. 23.Xiaomeng Ma, Lingyu Gao, and Qihui Xu. Tomchallenges: A principle-guided dataset and diverse evaluation tasks for exploring theory of mind. arXiv preprint arXiv:2305.15068, 2023.
  24. 24.Shima Rahimi Moghaddam and Christopher J Honey. Boosting theory-of-mind performance in large language models via prompting. arXiv preprint arXiv:2304.11490, 2023.
  25. 25.Kristine H Onishi and Renée Baillargeon. Do 15-month-old infants understand false beliefs? science, 308(5719):255–258, 2005.
  26. 26.Stefan Palan and Christian Schitter. Prolific. ac—a subject pool for online experiments. Journal of Behavioral and Experimental Finance, 17:22–27, 2018.
  27. 27.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. arXiv preprint arXiv:2202.03286, 2022.
  28. 28.Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. Discovering language model behaviors with model-written evaluations. arXiv preprint arXiv:2212.09251, 2022.
  29. 29.Josef Perner, Susan R Leekam, and Heinz Wimmer. Three-year-olds’ difficulty with false belief: The case for a conceptual deficit. British journal of developmental psychology, 5(2):125–137, 1987.
  30. 30.Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick. Machine theory of mind. In International conference on machine learning, pages 4218–4227. PMLR, 2018.
  31. 31.Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus. Modeling others using oneself in multi-agent reinforcement learning. In International conference on machine learning, pages 4257–4266. PMLR, 2018.
  32. 32.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728, 2019.
  33. 33.Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi. Neural theory-of-mind? on the limits of social intelligence in large lms. arXiv preprint arXiv:2210.13312, 2022.
  34. 34.Timo Schick and Hinrich Schütze. Generating datasets with pretrained language models. arXiv preprint arXiv:2104.07540, 2021.
  35. 35.Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz. Clever hans or neural theory of mind? stress testing social reasoning in large language models. arXiv preprint arXiv:2305.14763, 2023.
  36. 36.Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin Smith, Shari Liu, Dan Gutfreund, Elizabeth Spelke, Joshua Tenenbaum, and Tomer Ullman. Agent: A benchmark for core psychological reasoning. In International Conference on Machine Learning, pages 9614–9625. PMLR, 2021.
  37. 37.Felix A Sosa, Tomer Ullman, Joshua B Tenenbaum, Samuel J Gershman, and Tobias Gerstenberg. Moral dynamics: Grounding moral judgment in intuitive physics and intuitive psychology. Cognition, 217:104890, 2021.
  38. 38.Elizabeth S. Spelke. 279Core Knowledge and Conceptual Change: A Perspective on Social Cognition. In Core Knowledge and Conceptual Change. Oxford University Press, 09 2016. ISBN 9780190467630. doi: 10.1093/acprof:oso/9780190467630.003.0016. URL https://doi.org/10.1093/acprof:oso/9780190467630.003.0016.
  39. 39.Gala Stojnić, Kanishk Gandhi, Shannon Yasuda, Brenden M Lake, and Moira R Dillon. Commonsense psychology in human infants and machines. Cognition, 235:105406, 2023.
  40. 40.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  41. 41.Sean Trott, Cameron Jones, Tyler Chang, James Michaelov, and Benjamin Bergen. Do large language models know what humans know? arXiv preprint arXiv:2209.01515, 2022.
  42. 42.Tomer Ullman. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399, 2023.
  43. 43.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022.
  44. 44.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  45. 45.Henry M Wellman. The child’s theory of mind. The MIT Press, 1992.
  46. 46.Heinz Wimmer and Josef Perner. Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception. Cognition, 13(1):103–128, 1983.
  47. 47.Sarah A. Wu, Shruti Sridhar, and Tobias Gerstenberg. A computational model of responsibility judgments from counterfactual simulations and intention inferences. In Micah B. Goldwater, Florencia Anggoro, Brett Hayes, and Desmond C Ong, editors, Proceedings of the 45th Annual Conference of the Cognitive Science Society, 2023. URL https://psyarxiv.com/uwdbr/.

Citation

MLA
Gandhi, K., et al. “Understanding Social Reasoning in Language Models with Language Models”. arXiv, 2023, http://arxiv.org/abs/2306.15448v2.
APA
Gandhi, K., Fränken, J.-P., Gerstenberg, T., & Goodman, N. D. (2023). Understanding Social Reasoning in Language Models with Language Models. arXiv. http://arxiv.org/abs/2306.15448v2
Chicago
Gandhi, K., J.-P. Fränken, T. Gerstenberg, and N. D. Goodman. 2023. “Understanding Social Reasoning in Language Models with Language Models”. arXiv. http://arxiv.org/abs/2306.15448v2.
Harvard
Gandhi, K. et al. (2023) “Understanding Social Reasoning in Language Models with Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.15448v2.
Vancouver
1. Gandhi K, Fränken J-P, Gerstenberg T, Goodman ND (2023) Understanding Social Reasoning in Language Models with Language Models. arXiv

BibTeX

@article{gandhi2023understanding,
  title = {Understanding Social Reasoning in Language Models with Language Models},
  author = {Gandhi, Kanishk and Fränken, Jan-Philipp and Gerstenberg, Tobias and Goodman, Noah D.},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.15448v2},
  eprint = {2306.15448}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors