Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker
Melanie SclarSachin KumarPeter WestAlane SuhrYejin ChoiYulia Tsvetkov
Introduces SYMBOLICTOM, a training-free decoding algorithm that tracks multi-character belief states using explicit symbolic graphs to dramatically boost the zero-shot theory-of-mind reasoning capabilities of off-the-shelf language models.
Large language models often struggle to understand human mental states, an essential cognitive capability known as theory of mind that allows individuals to track others' differing beliefs, perspectives, and potential misconceptions. As artificial intelligence systems are increasingly deployed in tutoring, dialogue, and collaborative environments, the inability to accurately model multi-party knowledge and false beliefs poses significant risks to effective communication and reliable natural language understanding.
The article develops and demonstrates SymbolicToM, an inference-time method that equips off-the-shelf neural language models with structured, multi-character belief-tracking capabilities without requiring additional training or supervised fine-tuning. The researchers evaluated this approach to determine whether explicit symbolic representations can resolve complex social reasoning tasks and overcome the severe brittleness typical of supervised methods.
To achieve this, the approach processes chronological stories by decomposing narrative text into simple subtasks using standard tools for natural language inference, information extraction, and language generation. It constructs an omniscient global context graph alongside explicit local belief graphs that recursively model what each character believes about the world and what they believe other characters know. When a perspective-based question is asked, the system identifies the relevant entity, retrieves only the corresponding belief subgraph, and supplies this focused textual context to a base language model for zero-shot question answering. The evaluation measured reading comprehension across diverse models on the established ToMi benchmark and several newly modified, out-of-distribution test sets.
The findings show that SymbolicToM substantially boosts baseline model accuracy across all cognitive reasoning tasks. On false-belief questions, the method increased accuracy by substantial margins, such as a 62 percentage point gain for Flan-T5-XL on first-order questions and a 78 percentage point gain for GPT-3.5 on second-order questions. Across broader benchmark evaluations, GPT-3's overall average accuracy rose by 38 percentage points to reach 92 percent. Furthermore, while heavily fine-tuned supervised baselines deteriorated sharply when faced with story variations and linguistic paraphrasing—dropping by up to 54 percentage points—SymbolicToM maintained superior accuracy and demonstrated successful scaling to complex third-order reasoning tasks.
These results indicate that simply scaling up model size is insufficient to master the implicit symbolic nuances of human social intelligence. Instead, pairing foundational language models with explicit, dynamic symbolic structures offers a practical and modular path toward reliable multi-agent reasoning. This neuro-symbolic framework reduces system error rates in narrative comprehension without incurring the heavy computational costs of fine-tuning dedicated models.
Organizations developing customer-facing conversational agents or advanced text analysis tools should consider hybrid neuro-symbolic architectures rather than relying solely on raw model prompting or brittle supervised fine-tuning. Moving forward, developers should invest in expanding benchmark datasets beyond rigid template narratives to capture richer physical commonsense and complex real-world social dynamics before deploying such mechanisms in critical decision-making environments.
The primary limitations of this study stem from its reliance on public benchmark datasets based on structured Sally-Anne false-belief tests, which lack realistic spatial physical constraints and assume strictly chronological narratives. While reader confidence in the core performance gains can remain high across standard benchmark settings, caution is warranted when applying these findings to unstructured, non-linear stories or unconstrained multi-party human dialogues where cascading information extraction errors may occur.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces chain-of-thought prompting to elicit multi-step reasoning in large language models, providing the core prompting foundation that SymbolicToM aims to structure and overcome.
- Paper: Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations, Jaehun Jung et al. (2022). Presents an inference-time neuro-symbolic framework that combines recursive model explanations with formal satisfiability solving to eliminate contradictions, prefiguring the symbolic graph-based reasoning in SymbolicToM.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). Provides a comprehensive architectural roadmap for integrating explicit knowledge graphs with language models at inference time, grounding the graph-tracking design of SymbolicToM.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). Demonstrates the capabilities and failure modes of chain-of-thought prompting across complex symbolic and multi-step reasoning benchmarks like BIG-Bench Hard.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). Establishes how standard neural language models rely on brittle heuristics rather than true semantic inference, highlighting the vulnerability that SymbolicToM resolves through symbolic belief graphs.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). Introduces BigToM, a scalable causal-graph evaluation benchmark that broadens the assessment of forward and backward Theory of Mind inference beyond the template-based false-belief datasets evaluated in SymbolicToM.
- Paper: MMToM-QA: Multimodal Theory of Mind Question Answering, Chuanyang Jin et al. (2024). Extends Theory of Mind modeling and belief tracking from purely textual narratives into multimodal contexts integrating video and text question answering.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). Critiques the limitations of single-LLM omniscient social simulations and investigates information-asymmetric multi-agent interactions where explicit belief tracking is essential.
- Paper: Theory of Mind for Multi-Agent Collaboration via Large Language Models, Huao Li et al. (2023). Applies Theory of Mind inferences to dynamic multi-agent collaborative environments and interactive text games.
- Paper: LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay, Yihuai Lan et al. (2024). Explores complex multi-agent social deduction involving deception and hidden roles, testing the limits of mental state and belief tracking in adversarial gameplay.
- Paper: Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation, Xianghe Pang et al. (2024). Leverages multi-character social scene simulation and role-play to enforce social norms and self-align language model behavior.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Evaluates the long-term temporal and conversational memory tracking of language models over extended multi-session dialogues.
