Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker

Melanie SclarSachin KumarPeter WestAlane SuhrYejin ChoiYulia Tsvetkov

article2023ACL111 citations

Introduces SYMBOLICTOM, a training-free decoding algorithm that tracks multi-character belief states using explicit symbolic graphs to dramatically boost the zero-shot theory-of-mind reasoning capabilities of off-the-shelf language models.

Listen

Large language models often struggle to understand human mental states, an essential cognitive capability known as theory of mind that allows individuals to track others' differing beliefs, perspectives, and potential misconceptions. As artificial intelligence systems are increasingly deployed in tutoring, dialogue, and collaborative environments, the inability to accurately model multi-party knowledge and false beliefs poses significant risks to effective communication and reliable natural language understanding.

The article develops and demonstrates SymbolicToM, an inference-time method that equips off-the-shelf neural language models with structured, multi-character belief-tracking capabilities without requiring additional training or supervised fine-tuning. The researchers evaluated this approach to determine whether explicit symbolic representations can resolve complex social reasoning tasks and overcome the severe brittleness typical of supervised methods.

To achieve this, the approach processes chronological stories by decomposing narrative text into simple subtasks using standard tools for natural language inference, information extraction, and language generation. It constructs an omniscient global context graph alongside explicit local belief graphs that recursively model what each character believes about the world and what they believe other characters know. When a perspective-based question is asked, the system identifies the relevant entity, retrieves only the corresponding belief subgraph, and supplies this focused textual context to a base language model for zero-shot question answering. The evaluation measured reading comprehension across diverse models on the established ToMi benchmark and several newly modified, out-of-distribution test sets.

The findings show that SymbolicToM substantially boosts baseline model accuracy across all cognitive reasoning tasks. On false-belief questions, the method increased accuracy by substantial margins, such as a 62 percentage point gain for Flan-T5-XL on first-order questions and a 78 percentage point gain for GPT-3.5 on second-order questions. Across broader benchmark evaluations, GPT-3's overall average accuracy rose by 38 percentage points to reach 92 percent. Furthermore, while heavily fine-tuned supervised baselines deteriorated sharply when faced with story variations and linguistic paraphrasing—dropping by up to 54 percentage points—SymbolicToM maintained superior accuracy and demonstrated successful scaling to complex third-order reasoning tasks.

These results indicate that simply scaling up model size is insufficient to master the implicit symbolic nuances of human social intelligence. Instead, pairing foundational language models with explicit, dynamic symbolic structures offers a practical and modular path toward reliable multi-agent reasoning. This neuro-symbolic framework reduces system error rates in narrative comprehension without incurring the heavy computational costs of fine-tuning dedicated models.

Organizations developing customer-facing conversational agents or advanced text analysis tools should consider hybrid neuro-symbolic architectures rather than relying solely on raw model prompting or brittle supervised fine-tuning. Moving forward, developers should invest in expanding benchmark datasets beyond rigid template narratives to capture richer physical commonsense and complex real-world social dynamics before deploying such mechanisms in critical decision-making environments.

The primary limitations of this study stem from its reliance on public benchmark datasets based on structured Sally-Anne false-belief tests, which lack realistic spatial physical constraints and assume strictly chronological narratives. While reader confidence in the core performance gains can remain high across standard benchmark settings, caution is warranted when applying these findings to unstructured, non-linear stories or unconstrained multi-party human dialogues where cascading information extraction errors may occur.

arXiv: 2306.00924
Cover for Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker

Abstract

Theory of Mind (ToM)—the ability to reason about the mental states of other people—is a key element of our social intelligence. Yet, despite their ever more impressive performance, large-scale neural language models still lack basic theory of mind capabilities out-of-the-box. We posit that simply scaling up models will not imbue them with theory of mind due to the inherently symbolic and implicit nature of the phenomenon, and instead investigate an alternative: can we design a decoding-time algorithm that enhances theory of mind of off-the-shelf neural language models without explicit supervision? We present SYMBOLICTOM, a plug-and-play approach to reason about the belief states of multiple characters in reading comprehension tasks via explicit symbolic representation. More concretely, our approach tracks each entity’s beliefs, their estimation of other entities’ beliefs, and higher-order levels of reasoning, all through graphical representations, allowing for more precise and interpretable reasoning than previous approaches. Empirical results on the well-known ToMi benchmark (Le et al., 2019) demonstrate that SYMBOLICTOM dramatically enhances off-the-shelf neural networks’ theory of mind in a zero-shot setting while showing robust out-of-distribution performance compared to supervised baselines. Our work also reveals spurious patterns in existing theory of mind benchmarks, emphasizing the importance of out-of-distribution evaluation and methods that do not overfit a particular dataset.

Table of Contents

  • 1 Introduction
  • 2 Motivation and Background
  • 3 Methods
  • 3.1 SYMBOLICTOM: Algorithm Overview
  • 3.2 Computing the Belief Graphs B p 1 ...p k
  • 3.2.1 Detecting Witnesses, Updating Graphs, and Propagating Knowledge
  • 3.3 Notes on Memory Efficiency
  • 4 Fundamental Issues in Existing ToM Datasets
  • 5 Experiments
  • 5.1 In-Domain Evaluation
  • 5.2 Story Structure Robustness Test Sets
  • 5.3 Paraphrasing Robustness Evaluation
  • 6 Related Work
  • 7 Conclusions
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Additional Details on SYMBOLICTOM
  • A.1 Detailed Description of Information Contained in Global Context G
  • A.2 Prompts for Resulting State Extraction
  • A.3 Solving PROCESSQUESTION using GPT3
  • B Details on Out-Of-Domain Evaluation
  • B.1 Linguistic Diversity Per ToMi Template
  • B.1.1 All Paraphrases of PersonX entered the RoomY.
  • B.1.2 All Paraphrases of PersonX exited the RoomY.
  • B.1.3 All Paraphrases of The Object1 is in the Container1.
  • B.1.4 All Paraphrases of PersonX moved the Object1 to the Container1.
  • B.1.5 All Paraphrases of PersonX is in the RoomY.
  • B.1.6 All Paraphrases of positive distractor sentences
  • B.1.7 All Paraphrases of positive negative sentences (PersonX hates ObjectY)
  • B.2 Structure of Story Structure Robustness Test Sets
  • B.2.1 Double Room False-Belief Episode
  • B.2.2 Three Active Characters Story
  • B.2.3 True-Belief Interaction, Falsified by Unwitnessed Third-Person Story
  • B.2.4 Four Containers with Multiple Movements
  • C Expanded Results
  • C.1 Ablating FILTERBASEDONQUESTION from SYMBOLICTOM
  • C.2 Third-Order Theory of Mind Evaluation
  • D.2's Third-Order ToM Questions
  • C.3 Alternative RESULTINGSTATE Implementations
  • C.4 Detailed Result Tables

Knowls

  1. Knowl 1 — SYMBOLICTOM represents recursively nested character beliefs as graphs

    model/method

    SYMBOLICTOM is a training-free, decoding-time method for answering reading-comprehension questions that require theory of mind. It maintains a separate symbolic graph for each nested belief state rather than using one graph for the entire story.

    For a sequence of characters p1,…,pkp_1,\ldots,p_k, where 1≤k≤m1 \leq k \leq m and mm is a user-specified maximum reasoning depth, Bp1,…,pkB_{p_1,\ldots,p_k} denotes what p1p_1 believes that p2p_2 believes that ⋯\cdots pkp_k believes about the current world state. Each belief state is a directed labeled graph: nodes are entities such as characters, objects, rooms, and containers, and edges are positive relations extracted from the story, such as (celery,is in,basket)(\text{celery},\text{is in},\text{basket}).

    The method also maintains a global graph GG containing the observable true world state. The local belief graphs differ from GG because a character’s graph is updated only with information available to that character. If a character witnesses a later change, outdated relations in that character’s graph are removed and the new relation is added. This gives SYMBOLICTOM an explicit object-permanence bias: a relation remains in a graph until the method infers a contradictory state.

    The representation treats repeated self-modeling as redundant: what pkp_k thinks that pkp_k thinks is identified with what pkp_k thinks. Thus a lower-order state can be represented at depth mm as

    Bp1,…,pk≡Bp1,…,pk,pk,…,pk⏟m−k repetitionsB_{p_1,\ldots,p_k} \equiv B_{p_1,\ldots,p_k,\underbrace{p_k,\ldots,p_k}_{m-k\text{ repetitions}}}.

    The resulting collection of graphs supports questions about reality, a character’s beliefs, and arbitrarily nested beliefs. All reported experiments use m=2m=2, because existing theory-of-mind datasets evaluate at most second-order reasoning, although the representation itself is not restricted to that depth.

  2. Knowl 2 — Witness-based sequential construction of global and local belief graphs

    algorithm

    The belief-tracking procedure takes a chronologically ordered story as input and returns the global world-state graph GG together with local graphs Bp1,…,pmB_{p_1,\ldots,p_m} for every length-mm sequence of characters. A character is a witness of a sentence when the character belongs to the same connected component of the updated global graph as the newly inserted state relations. The method assumes that this connected-component criterion identifies who can observe the described event.

    For each sentence, the procedure first updates GG from an omniscient perspective. It detects global edges contradicted by the sentence using natural-language inference, asks a language model to express the resulting state after the sentence, extracts nodes and positive subject–relation–object triples from that state with OpenIE, adds the triples to GG, identifies witnesses before deleting contradicted edges, and then removes the invalid edges.

    Each local graph is updated with the same new positive triples only when every character in its belief-state index is a witness. The method then propagates implicit knowledge: for each witnessing character, it copies into the local graph all edges in that character’s connected component of GG, representing facts learned by entering or observing a location. Edges in the local graph that are absent from that current component are removed as outdated beliefs.

    Input: Chronological sentences s1,…,sns_1,\ldots,s_n and maximum depth mm
    Output: Global graph GG and local belief graphs Bp1,…,pmB_{p_1,\ldots,p_m}
    Initialize GG and all local graphs as empty
    for each sentence ss do
        Detect edges in GG contradicted by ss
        Ask a language model for the positive resulting state after ss
        Extract positive entity-relation-entity triples from that state
        Add the resulting nodes and edges to GG
        Let WW be the characters in connected components containing the new edges
        Remove the contradicted edges from GG
        for every character tuple (p1,…,pm)(p_1,\ldots,p_m) in WmW^m do
            Add the resulting triples to Bp1,…,pmB_{p_1,\ldots,p_m}
            Copy into that local graph all edges in the relevant connected component of GG
            Remove local edges absent from that connected component
            Remove edges contradicted by ss
        end for
    end for
    return GG and all local belief graphs

    The global update uses WANLI for contradiction detection and OpenIE for triple extraction in the experiments. Only positive relations are inserted; negated relations are represented operationally by deleting contradictory positive edges. The number of stored graphs grows exponentially with mm, so memory is the principal scalability limitation; the authors use m=2m=2 in all main experiments.

  3. Knowl 3 — Question answering by graph selection and zero-shot language-model querying

    model/method

    SYMBOLICTOM answers a question by converting a nested mental-state question into a direct world-state question and selecting the corresponding belief graph. Given a question mentioning characters p1,…,pkp_1,\ldots,p_k in order, the method heuristically detects those characters and rewrites the question so that it asks directly about the relevant entity or relation in the world state. For example, a question about where Bob thinks Alice will search for an object is converted into a question about where the object is, while querying BBob,AliceB_{\text{Bob},\text{Alice}}.

    Let S′S' be the ordered subset of original story sentences whose edges occur in the selected graph. The method feeds S′S' and the rewritten question to an off-the-shelf language model in a zero-shot prompt, treating the selected graph as the real world for that question. An optional final filter produces S′′S'' by retaining only graph edges for which at least one endpoint is an entity mentioned in the original question. The final answer is generated from S′′S'' and the rewritten question.

    This decomposition separates symbolic access control—deciding which facts are available from the selected character perspective—from neural language generation. It requires no task-specific training or fine-tuning. On the templatic questions in the evaluated datasets, the heuristic question rewriting correctly rephrases all questions.

  4. Knowl 4 — Correction of ToMi ambiguities and mislabeled examples

    model/method

    The authors identify and correct two problems in the ToMi theory-of-mind reading-comprehension benchmark. First, ToMi can leave the locations of containers physically ambiguous. A reader may need to infer several unstated spatial relations, and the answer can change with commonsense associations of the nouns. In one human study, replacing the object and container names with (hat,box)(\text{hat},\text{box}) produced 80% accuracy, whereas (apple,pantry)(\text{apple},\text{pantry}) produced 20% accuracy, even though the intended ToMi label was unchanged.

    To remove this unintended commonsense burden, the authors automatically insert a sentence specifying the location of each container immediately after relevant object-location or object-movement primitives. They also correct mislabeled second-order questions reported in prior analyses. The resulting corrected ToMi version is used for evaluation and is released with the authors’ out-of-distribution resources.

    Second, the authors construct three structural robustness sets without introducing new action types or linguistic paraphrases. These sets alter the ordering, number of active characters, or number of object movements so that models cannot rely on the original Sally–Anne template. The sets are: two concatenated false-belief episodes involving the same characters; a three-character story in which different characters witness different object movements; and a four-container story with repeated movements by two characters.

  5. Knowl 5 — In-domain ToMi performance across base language models

    data/table

    The main in-domain evaluation compares zero-shot off-the-shelf language models with the same models augmented by SYMBOLICTOM on 100 examples of each ToMi question type: first-order true belief, first-order false belief, second-order true belief, second-order false belief, reality, and memory. Values in brackets are the corresponding out-of-the-box accuracies; the final two rows are supervised baselines trained or developed for ToM. SYMBOLICTOM substantially improves false-belief reasoning across model families while preserving near-perfect memory performance. For example, GPT3-Davinci rises from 0.25 to 0.96 on first-order false belief and from 0.26 to 0.90 on second-order false belief; GPT3.5 rises from 0.66 to 0.95 and from 0.09 to 0.87, respectively.

    Model 1st TB 1st FB 2nd TB 2nd FB Reality Memory
    Macaw-3B 0.86 [0.50] 0.79 [0.33] 0.86 [0.34] 0.84 [0.17] 0.10 [0.14] 0.95 [0.91]
    GPT3-Curie 0.77 [0.42] 0.82 [0.35] 0.73 [0.26] 0.89 [0.26] 0.61 [0.69] 0.99 [0.86]
    GPT3-Davinci 0.96 [0.75] 0.96 [0.25] 0.93 [0.14] 0.90 [0.26] 0.77 [0.86] 0.98 [0.98]
    Flan-T5-XL 0.98 [0.97] 0.80 [0.18] 0.98 [0.68] 0.78 [0.56] 0.73 [0.97] 1.00 [1.00]
    Flan-T5-XXL 0.98 [0.84] 0.95 [0.67] 1.00 [0.76] 0.90 [0.39] 0.13 [0.63] 1.00 [1.00]
    LLaMA-7B 0.82 [0.32] 0.95 [0.66] 0.66 [0.31] 0.72 [0.41] 0.87 [0.37] 1.00 [0.83]
    LLaMA-13B 0.82 [0.60] 0.86 [0.67] 0.70 [0.53] 0.62 [0.77] 0.87 [0.48] 1.00 [0.90]
    GPT3.5 0.97 [0.76] 0.95 [0.66] 0.99 [0.02] 0.87 [0.09] 0.98 [1.00] 0.99 [0.80]
    GPT4 0.98 [0.83] 0.94 [0.73] 0.98 [0.36] 0.89 [0.64] 0.94 [1.00] 1.00 [1.00]
    Finetuned GPT3 0.95 0.99 0.97 1.00 1.00 1.00
    TTT-learned 0.84 1.00 0.82 0.88 1.00 1.00

    The experiments use Macaw-3B, GPT3-Curie and GPT3-Davinci, Flan-T5-XL and Flan-T5-XXL, LLaMA-7B and LLaMA-13B, GPT3.5, and GPT4. The symbolic method is inference-time only and uses WANLI for NLI contradiction detection and AllenNLP OpenIE for relation extraction.

  6. Knowl 6 — Structural robustness against modified ToMi stories

    data/table

    The three modified test sets evaluate whether a model has learned theory-of-mind reasoning or merely memorized ToMi’s original story structure. Each column reports precision over all questions from 100 stories. The supervised baselines were trained on original ToMi; all other systems are zero-shot. SYMBOLICTOM improves nearly every off-the-shelf model and generally exceeds both supervised baselines, especially on the more complex multi-character and multi-movement settings.

    D1D_1 concatenates two false-belief episodes involving the same two characters. D2D_2 uses three active characters who witness different movements. D3D_3 contains multiple movements of one object across four containers.

    Model D1D_1 D2D_2 D3D_3
    Off-the-shelf models
    Macaw-3B 8 12 30
    Flan-T5-XL 86 51 68
    Flan-T5-XXL 69 59 52
    GPT3-Curie 37 39 57
    GPT3-Davinci 20 25 39
    GPT3.5 1 0 48
    GPT4 58 62 97
    LLaMA-7B 17 17 17
    LLaMA-13B 26 36 37
    SYMBOLICTOM + off-the-shelf models
    Macaw-3B 89 (+81) 71 (+60) 70 (+41)
    Flan-T5-XL 76 (-10) 96 (+46) 100 (+33)
    Flan-T5-XXL 93 (+24) 100 (+41) 100 (+49)
    GPT3-Curie 84 (+48) 81 (+42) 73 (+16)
    GPT3-Davinci 92 (+73) 91 (+66) 90 (+50)
    GPT3.5 100 (+99) 100 (+99) 99 (+51)
    GPT4 100 (+42) 100 (+38) 100 (+4)
    LLaMA-7B 99 (+82) 92 (+75) 88 (+71)
    LLaMA-13B 78 (+52) 84 (+48) 84 (+47)
    Supervised models
    TTT 49 65 78
    Finetuned GPT3 51 68 32

    The results show strong structural overfitting by supervised models: finetuned GPT3 falls to 32 on D3D_3, while GPT3.5 with SYMBOLICTOM reaches 100, 100, and 99 on D1D_1, D2D_2, and D3D_3. SYMBOLICTOM also makes smaller models robust, such as raising Macaw-3B from 8 to 89 on D1D_1 and from 12 to 71 on D2D_2.

  7. Knowl 7 — Robustness to linguistic paraphrases

    empirical result

    The authors create ParaphrasedToMi by rewording every ToMi template with GPT3-Davinci, varying names, objects, rooms, and descriptions of entering, exiting, locating, and moving objects, followed by manual removal of incorrect paraphrases. The resulting stories express the same underlying events with substantially more varied and less direct language.

    Supervised models transfer poorly: TTT’s average accuracy decreases by 54 percentage points relative to original ToMi, with losses across all question types, while finetuned GPT3 loses approximately 40 percentage points on false-belief questions. Zero-shot models also degrade, but adding SYMBOLICTOM continues to produce large gains on theory-of-mind questions and performs significantly better than supervised TTT across all ToM question types.

    Paraphrasing exposes errors in the symbolic front end. NLI mistakes can prevent obsolete edges from being removed, and resulting-state or OpenIE mistakes can prevent correct edges from being inserted. GPT3 was the most reliable implementation of the resulting-state extraction subtask under paraphrases, so the main experiments use GPT3 for that subtask regardless of the final question-answering model. The authors report that LLaMA, GPT3.5, and GPT4 can perform even better with alternative implementations of resulting-state extraction.

  8. Knowl 8 — Third-order theory-of-mind evaluation using second-order graphs

    data/table

    The three-character structural test set D2D_2 is extended with third-order questions such as where p1p_1 thinks that p2p_2 thinks that p1p_1 will search for an object. Each story receives eight third-order questions. Although the questions theoretically require depth k=3k=3, SYMBOLICTOM uses the same k=2k=2 graph representation as in the main experiments. The authors hypothesize that second-order graphs already prevent the language model from attending to facts inaccessible to the relevant characters.

    Precision over 100 stories is:

    Model Third-order ToM precision
    Off-the-shelf models
    Macaw-3B 13
    Flan-T5-XL 32
    Flan-T5-XXL 62
    GPT3-Curie 28
    GPT3-Davinci 19
    GPT3.5 8
    GPT4 26
    LLaMA-7B 22
    LLaMA-13B 39
    SYMBOLICTOM + off-the-shelf models
    Macaw-3B 85 (+72)
    Flan-T5-XL 97 (+65)
    Flan-T5-XXL 100 (+38)
    GPT3-Curie 89 (+61)
    GPT3-Davinci 90 (+71)
    GPT3.5 100 (+91)
    GPT4 100 (+73)
    LLaMA-7B 90 (+68)
    LLaMA-13B 95 (+57)
    Supervised models
    TTT 52
    Finetuned GPT3 76

    SYMBOLICTOM substantially outperforms both the off-the-shelf models and the ToMi-trained supervised baselines. For example, GPT3.5 rises from 8 to 100 and GPT4 from 26 to 100. This experiment is presented as an initial demonstration of third-order reasoning rather than a comprehensive third-order benchmark.

  9. Knowl 9 — Effect of filtering graph evidence by question entities

    empirical result

    The final question-based filter retains from the selected belief graph only edges whose endpoints include an entity mentioned in the question. Its effect is model-dependent because shorter evidence can remove distracting facts but can also amplify answer biases.

    Across all question types, applying the filter increases average accuracy by 7 points for Macaw-3B, 3.5 points for GPT3-Davinci, 12.8 points for Flan-T5-XXL, and 15 points for LLaMA-7B. It decreases average accuracy by 5.3 points for Flan-T5-XL and 4 points for GPT4. The filter can particularly worsen Flan-T5 models’ reality performance, because a very short context containing one container can induce a bias toward answering with the room rather than the valid container.

    Regardless of whether the filter is used, SYMBOLICTOM improves theory-of-mind question performance for every evaluated model. The filter is necessary for Flan-T5-XXL to exceed its base-model accuracy on ToM questions, but the authors’ ablation indicates that it can often be omitted without losing the central gains.

  10. Knowl 10 — Stated limitations of SYMBOLICTOM

    limitation

    SYMBOLICTOM assumes that stories are narrated chronologically. Human-written narratives can use flashbacks, anaphoric temporal references, or other nonchronological structures, which the sequential update procedure does not explicitly resolve.

    The method depends on off-the-shelf language models for resulting-state extraction, NLI contradiction detection, and OpenIE triple extraction. Errors in these components can propagate through the global and local graphs; this problem becomes more visible on paraphrased stories. The authors suggest that stronger language models could reduce these errors but do not evaluate that option fully because of budget constraints.

    Witness detection also assumes that all people in the same connected spatial component as a changed relation witnessed the event. Existing theory-of-mind datasets contain no realistic notion of distance, so being specified as present in a location is treated as sufficient for observation. In realistic text, being somewhere as large as a country would not imply witnessing every event there. More sophisticated physical-commonsense reasoning or direct language-model judgments over paths between people and events would be needed.

    Finally, available benchmarks largely instantiate Sally–Anne-style object-location stories and omit richer interactions such as dialogue, realistic spatial distance, and varied social situations. Consequently, the demonstrated robustness is limited to the reading-comprehension setting and the interaction types represented by these datasets.

Coverage note — No substantial contributed material was omitted; implementation prompts and the full paraphrase inventory were excluded because they support the main method and robustness result rather than constituting separate load-bearing contributions.

References

  1. 1.Ashutosh Adhikari, Xingdi Yuan, Marc-Alexandre Côté, Mikuláš Zelinka, Marc-Antoine Rondeau, Romain Laroche, Pascal Poupart, Jian Tang, Adam Trischler, and Will Hamilton. 2020. Learning dynamic belief graphs to generalize on text-based games. Advances in Neural Information Processing Systems, 33:3045–3057.
  2. 2.Arjun R. Akula, Keze Wang, Changsong Liu, Sari Saba-Sadiya, Hongjing Lu, Sinisa Todorovic, Joyce Chai, and Song-Chun Zhu. 2022. Cx-tom: Counterfactual explanations with theory-of-mind for enhancing human trust in image recognition models. iScience, 25(1):103581.
  3. 3.Prithviraj Ammanabrolu and Mark Riedl. 2021. Learning knowledge graph-based world models of textual environments. In Advances in Neural Information Processing Systems.
  4. 4.Akshatha Arodi and Jackie Chi Kit Cheung. 2021. Textual time travel: A temporally informed approach to theory of mind. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4162–4172, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  5. 5.Cristian-Paul Bara, CH-Wang Sky, and Joyce Chai. 2021. Mindcraft: Theory of mind modeling for situated dialogue in collaborative tasks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1112–1125.
  6. 6.Simon Baron-Cohen, Alan M Leslie, and Uta Frith. 1985. Does the autistic child have a “theory of mind”? Cognition, 21(1):37–46.
  7. 7.Marta Białecka-Pikul, Anna Kołodziejczyk, and Sandra Bosacki. 2017. Advanced theory of mind in adolescence: Do age, gender and friendship style play a role? Journal of Adolescence, 56:145–156.
  8. 8.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  9. 9.James Carney, Rafael Wlodarski, and Robin Dunbar. 2014. Inference or enaction? the impact of genre on the narrative processing of other minds. PloS one, 9(12):e114172.
  10. 10.Hyung Won Chung, Le Hou, Shayne Longpre, Barrett Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  11. 11.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2023. Selection-inference: Exploiting large language models for interpretable logical reasoning. In The Eleventh International Conference on Learning Representations.
  12. 12.Hossein Rajaby Faghihi and Parisa Kordjamshidi. 2021. Time-stamped language model: Teaching language models to understand the flow of events. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4560–4570.
  13. 13.Meta Fundamental AI Research Diplomacy Team (FAIR), Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra. 2022. Human-level play in the game of <i>diplomacy</i> by combining language models with strategic reasoning. Science, 378(6624):1067–1074.
  14. 14.C.D. Frith, D.M. Wolpert, Uta Frith, and Christopher D. Frith. 2003. Development and neurophysiology of mentalizing. Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences, 358(1431):459–473.
  15. 15.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018. AllenNLP: A deep semantic natural language processing platform. In Proceedings of Workshop for NLP Open Source Software (NLP-OSS), pages 1–6, Melbourne, Australia. Association for Computational Linguistics.
  16. 16.Erin Grant, Aida Nematzadeh, and Thomas L. Griffiths. 2017. How can memory-augmented neural networks pass a false-belief task? Cognitive Science.
  17. 17.Mikael Henaff, Jason Weston, Arthur Szlam, Antoine Bordes, and Yann LeCun. 2017. Tracking the world state with recurrent entity networks. In International Conference on Learning Representations.
  18. 18.Léo Jacqmin, Lina M Rojas Barahona, and Benoit Favre. 2022. “do you follow me?”: A survey of recent approaches in dialogue state tracking. In Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 336–350.
  19. 19.Peter Jansen. 2022. A systematic survey of text worlds as embodied natural language environments. In The Third Wordplay: When Language Meets Games Workshop.
  20. 20.Seyed Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, and Deepak Ramachandran. 2022. Lambada: Backward chaining for automated reasoning in natural language. arXiv preprint arXiv:2212.13894.
  21. 21.Matthew Le, Y-Lan Boureau, and Maximilian Nickel. 2019. Revisiting the evaluation of theory of mind through question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5872–5877.
  22. 22.Alan M Leslie, Ori Friedman, and Tim P German. 2004. Core mechanisms in ‘theory of mind’. Trends in cognitive sciences, 8(12):528–533.
  23. 23.Paula Leverage, Howard Mancing, and Richard Schweickert. 2010. Theory of mind and literature. Purdue University Press.
  24. 24.Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022. WANLI: Worker and ai collaboration for natural language inference dataset creation. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 6826–6847, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  25. 25.Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Tom Griffiths. 2018. Evaluating theory of mind in question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2392–2400, Brussels, Belgium. Association for Computational Linguistics.
  26. 26.Maxwell Nye, Michael Tessler, Josh Tenenbaum, and Brenden M Lake. 2021. Improving coherence and consistency in neural sequence models with dual-system, neuro-symbolic reasoning. Advances in Neural Information Processing Systems, 34:25192–25204.
  27. 27.OpenAI. 2022. ChatGPT: Optimizing language models for dialogue.
  28. 28.OpenAI. 2023. GPT-4 technical report.
  29. 29.Christopher Osterhaus, Susanne Koerber, and Beate Sodian. 2016. Scaling of advanced theory-of-mind tasks. Child development, 87(6):1971–1991.
  30. 30.David Premack and Guy Woodruff. 1978. Does the chimpanzee have a theory of mind? Behavioral and brain sciences, 1(4):515–526.
  31. 31.Liang Qiu, Yizhou Zhao, Yuan Liang, Pan Lu, Weiyan Shi, Zhou Yu, and Song-Chun Zhu. 2022. Towards socially intelligent agents with mental state transition and human value. In Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 146–158, Edinburgh, UK. Association for Computational Linguistics.
  32. 32.Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick. 2018. Machine theory of mind. In International conference on machine learning, pages 4218–4227. PMLR.
  33. 33.Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi. 2022. Neural theory-of-mind? on the limits of social intelligence in large lms. In Proceedings of the Association for Computational Linguistics: EMNLP 2022, page 3762–3780, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  34. 34.Melanie Sclar, Graham Neubig, and Yonatan Bisk. 2022. Symmetric machine theory of mind. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 19450–19466. PMLR.
  35. 35.Simone G Shamay-Tsoory, Hagai Harari, Judith Aharon-Peretz, and Yechiel Levkovitz. 2010. The role of the orbitofrontal cortex in affective theory of mind deficits in criminal offenders with psychopathic tendencies. Cortex, 46(5):668–677.
  36. 36.Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz. 2023. Clever hans or neural theory of mind? stress testing social reasoning in large language models.
  37. 37.Gabriel Stanovsky, Julian Michael, Luke Zettlemoyer, and Ido Dagan. 2018. Supervised open information extraction. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 885–895.
  38. 38.Oyvind Tafjord and Peter Clark. 2021. General-purpose question-answering with macaw. arXiv preprint arXiv:2109.02593.
  39. 39.Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2021. ProofWriter: Generating implications, proofs, and abductive statements over natural language. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3621–3634, Online. Association for Computational Linguistics.
  40. 40.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  41. 41.Tomer Ullman. 2023. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399.
  42. 42.Annalisa Valle, Davide Massaro, Ilaria Castelli, and Antonella Marchetti. 2015. Theory of mind development in adolescence and early adulthood: The growing complexity of recursive thinking ability. Europe’s journal of psychology, 11(1):112.
  43. 43.Max J van Duijn, Ineke Sluiter, and Arie Verhagen. 2015. When narrative takes over: The representation of embedded mindstates in shakespeare’s othello. Language and Literature, 24(2):148–166.
  44. 44.Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel. 2021. Towards mutual theory of mind in human-ai interaction: How language reflects what students perceive about a virtual teaching assistant. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–14.
  45. 45.Yuanfei Wang, fangwei zhong, Jing Xu, and Yizhou Wang. 2022. Tom2c: Target-oriented multi-agent communication and cooperation with theory of mind. In International Conference on Learning Representations.
  46. 46.Henry M Wellman. 2014. Making minds: How theory of mind develops. Oxford University Press.
  47. 47.Heinz Wimmer and Josef Perner. 1983. Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception. Cognition, 13(1):103–128.
  48. 48.Ping Yu, Tianlu Wang, Olga Golovneva, Badr Alkhamissy, Gargi Ghosh, Mona Diab, and Asli Celikyilmaz. 2022. Alert: Adapting language models to reasoning tasks. arXiv preprint arXiv:2212.08286.
  49. 49.Pei Zhou, Andrew Zhu, Jennifer Hu, Jay Pujara, Xiang Ren, Chris Callison-Burch, Yejin Choi, and Prithviraj Ammanabrolu. 2022. An ai dungeon master’s guide: Learning to converse and guide with intents and theory-of-mind in dungeons and dragons. arXiv preprint arXiv:2212.10060.
  50. 50.Hao Zhu, Graham Neubig, and Yonatan Bisk. 2021. Few-shot language coordination by modeling theory of mind. In International Conference on Machine Learning, pages 12901–12911. PMLR.
  51. 51.Lisa Zunshine. 2006. Why we read fiction: Theory of mind and the novel. Ohio State University Press.

Citation

MLA
Sclar, M., et al. “Minding Language Models’ (Lack Of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 13960–80, https://doi.org/10.18653/v1/2023.acl-long.780.
APA
Sclar, M., Kumar, S., West, P., Suhr, A., Choi, Y., & Tsvetkov, Y. (2023). Minding Language Models’ (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13960–13980. https://doi.org/10.18653/v1/2023.acl-long.780
Chicago
Sclar, M., S. Kumar, P. West, A. Suhr, Y. Choi, and Y. Tsvetkov. 2023. “Minding Language Models’ (Lack Of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13960–80. https://doi.org/10.18653/v1/2023.acl-long.780.
Harvard
Sclar, M. et al. (2023) “Minding Language Models’ (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13960–13980. Available at: https://doi.org/10.18653/v1/2023.acl-long.780.
Vancouver
1. Sclar M, Kumar S, West P, Suhr A, Choi Y, Tsvetkov Y (2023) Minding Language Models’ (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13960–13980

BibTeX

@inproceedings{sclar-etal-2023-minding,
    title = "Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker",
    author = "Sclar, Melanie  and
      Kumar, Sachin  and
      West, Peter  and
      Suhr, Alane  and
      Choi, Yejin  and
      Tsvetkov, Yulia",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.780/",
    doi = "10.18653/v1/2023.acl-long.780",
    pages = "13960--13980"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/