MMToM-QA: Multimodal Theory of Mind Question Answering
Chuanyang JinYutong WuJing CaoJiannan XiangYen-Ling KuoZhiting HuTomer D. UllmanAntonio TorralbaJoshua B. TenenbaumTianmin Shu
Introduces MMToM-QA, the first benchmark for evaluating machine Theory of Mind across combined video and text inputs, alongside a model that pairs Bayesian inverse planning with language models to infer human goals and beliefs far more effectively than current multimodal systems.
Developing artificial intelligence systems capable of safe and effective collaboration with humans—such as assistive robotics, autonomous vehicles, and automated tutors—requires Theory of Mind, which is the ability to infer hidden human mental states such as beliefs, goals, and intentions from observed behavior. Existing evaluations have largely assessed language or vision models on narrow, single-modality tasks, which fail to capture how humans naturally synthesize visual observations with contextual background to track changing mental states over time.
The article introduces the Multimodal Theory of Mind Question Answering (MMToM-QA) benchmark to rigorously assess how well AI models infer human goals and beliefs from combined visual and textual information. To address current performance gaps, the article also proposes and evaluates a hybrid computational framework called Bayesian Inverse Planning Accelerated by Language Models (BIP-ALM).
To construct the benchmark, researchers synthesized 134 household activity videos containing 600 two-choice inference questions across seven categories of belief and goal reasoning, complemented by 1,000 synthetic behavioral scenarios for model training. The proposed BIP-ALM method integrates visual and text data into unified symbolic representations and uses language models to efficiently estimate the likelihood of actions under competing mental-state hypotheses. The evaluation compares BIP-ALM against human baselines and prominent baseline models, including GPT-4 and GPT-4V, across multimodal, text-only, and video-only conditions.
The investigation produced three critical findings regarding machine social intelligence. First, human participants achieve approximately 93% accuracy on multimodal tasks, demonstrating that combined textual and visual information enables robust mental inference. Second, current state-of-the-art multimodal foundation models struggle substantially, with GPT-4V achieving an overall multimodal accuracy of only 44% and performing near random chance on complex tasks involving false beliefs and dynamic goal updates. Third, BIP-ALM significantly outperforms these baselines, achieving 75.3% to 76.7% multimodal accuracy and demonstrating superior generalization to real human actions in unseen environments (achieving up to 77% accuracy on a specialized human-subject test set).
These results demonstrate that standard foundation models lack genuine causal models of human behavior, frequently confusing objective physical reality with an individual’s subjective beliefs. Incorporating explicit model-based mental reasoning into AI architectures resolves critical safety and performance risks for interactive systems, which otherwise cannot reliably anticipate human intent or track belief updates.
Decision-makers building interactive or autonomous AI systems should avoid relying purely on end-to-end foundation models for social reasoning and instead adopt hybrid architectures that combine structured symbolic planning with language model flexibility. Future development should focus on extending multimodal benchmarks and inverse-planning frameworks to richer social phenomena, including human emotions, desires, and multi-agent interactions.
Confidence in these findings is supported by controlled human baseline validations and generalization tests on real user trajectories. However, users should note key boundary conditions: the current benchmark operates exclusively within simulated household object-search tasks, uses ground-truth visual perception inputs, and simplifies physical reasoning through discrete symbolic representations.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). Introduces BigToM, a causal-model benchmark evaluating Theory of Mind in language models that directly informs MMToM-QA's expansion into multimodal settings.
- Paper: Building Machines that Learn and Think Like People, Josh Tenenbaum (2018). Presents foundational computational cognitive frameworks for human-like mental state inference and Bayesian inverse planning that underpin the BIP-ALM architecture.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Establishes agentic simulation architectures for modeling human behavior and mental reflections in interactive simulated environments.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). Introduces dynamic video understanding benchmarks for temporal reasoning and action recognition that motivate video-based multimodal evaluation in MMToM-QA.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). Provides a comprehensive evaluation benchmark for multimodal foundation models across integrated perception and reasoning capabilities.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). Establishes standardized evaluation protocols for complex multi-discipline perception and deliberate reasoning in multimodal models.
- Paper: Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies, Gati V. Aher et al. (2023). Develops experimental methodologies for simulating human decision-making and psychological dynamics using large language models.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). Critically examines the limits and informational asymmetries of using LLMs to simulate social interactions and human behavior.
- Paper: CogBench: a large language model walks into a psychology lab, Julian Coda-Forno et al. (2024). Applies cognitive psychology paradigms to benchmark underlying artificial decision-making, planning, and metacognitive traits in foundation models.
- Paper: Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation, Xianghe Pang et al. (2024). Leverages multi-agent social scene simulation and role-playing to achieve autonomous model self-alignment and social reasoning.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). Explores System 2 deliberate search and meta chain-of-thought techniques to enhance complex multi-step reasoning in models.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Provides an extensive evaluation of multimodal LLMs on comprehensive video analysis across diverse temporal scenarios.
- Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). Synthesizes frameworks and benchmarks for assessing metacognition, self-monitoring, and cognitive boundaries in LLMs.
