MMToM-QA: Multimodal Theory of Mind Question Answering

Chuanyang JinYutong WuJing CaoJiannan XiangYen-Ling KuoZhiting HuTomer D. UllmanAntonio TorralbaJoshua B. TenenbaumTianmin Shu

article2024ACL80 citations

Introduces MMToM-QA, the first benchmark for evaluating machine Theory of Mind across combined video and text inputs, alongside a model that pairs Bayesian inverse planning with language models to infer human goals and beliefs far more effectively than current multimodal systems.

Listen

Developing artificial intelligence systems capable of safe and effective collaboration with humans—such as assistive robotics, autonomous vehicles, and automated tutors—requires Theory of Mind, which is the ability to infer hidden human mental states such as beliefs, goals, and intentions from observed behavior. Existing evaluations have largely assessed language or vision models on narrow, single-modality tasks, which fail to capture how humans naturally synthesize visual observations with contextual background to track changing mental states over time.

The article introduces the Multimodal Theory of Mind Question Answering (MMToM-QA) benchmark to rigorously assess how well AI models infer human goals and beliefs from combined visual and textual information. To address current performance gaps, the article also proposes and evaluates a hybrid computational framework called Bayesian Inverse Planning Accelerated by Language Models (BIP-ALM).

To construct the benchmark, researchers synthesized 134 household activity videos containing 600 two-choice inference questions across seven categories of belief and goal reasoning, complemented by 1,000 synthetic behavioral scenarios for model training. The proposed BIP-ALM method integrates visual and text data into unified symbolic representations and uses language models to efficiently estimate the likelihood of actions under competing mental-state hypotheses. The evaluation compares BIP-ALM against human baselines and prominent baseline models, including GPT-4 and GPT-4V, across multimodal, text-only, and video-only conditions.

The investigation produced three critical findings regarding machine social intelligence. First, human participants achieve approximately 93% accuracy on multimodal tasks, demonstrating that combined textual and visual information enables robust mental inference. Second, current state-of-the-art multimodal foundation models struggle substantially, with GPT-4V achieving an overall multimodal accuracy of only 44% and performing near random chance on complex tasks involving false beliefs and dynamic goal updates. Third, BIP-ALM significantly outperforms these baselines, achieving 75.3% to 76.7% multimodal accuracy and demonstrating superior generalization to real human actions in unseen environments (achieving up to 77% accuracy on a specialized human-subject test set).

These results demonstrate that standard foundation models lack genuine causal models of human behavior, frequently confusing objective physical reality with an individual’s subjective beliefs. Incorporating explicit model-based mental reasoning into AI architectures resolves critical safety and performance risks for interactive systems, which otherwise cannot reliably anticipate human intent or track belief updates.

Decision-makers building interactive or autonomous AI systems should avoid relying purely on end-to-end foundation models for social reasoning and instead adopt hybrid architectures that combine structured symbolic planning with language model flexibility. Future development should focus on extending multimodal benchmarks and inverse-planning frameworks to richer social phenomena, including human emotions, desires, and multi-agent interactions.

Confidence in these findings is supported by controlled human baseline validations and generalization tests on real user trajectories. However, users should note key boundary conditions: the current benchmark operates exclusively within simulated household object-search tasks, uses ground-truth visual perception inputs, and simplifies physical reasoning through discrete symbolic representations.

arXiv: 2401.08743
Cover for MMToM-QA: Multimodal Theory of Mind Question Answering

Abstract

Theory of Mind (ToM), the ability to understand people’s mental states, is an essential ingredient for developing machines with human-level social intelligence. Recent machine learning models, particularly large language models, seem to show some aspects of ToM understanding. However, existing ToM benchmarks use unimodal datasets – either video or text. Human ToM, on the other hand, is more than video or text understanding. People can flexibly reason about another person’s mind based on conceptual representations (e.g., goals, beliefs, plans) extracted from any available data. To address this, we introduce a multimodal Theory of Mind question answering (MMToM-QA) benchmark. MMToM-QA comprehensively evaluates machine ToM both on multimodal data and on different kinds of unimodal data about a person’s activity in a household environment. To engineer multimodal ToM capacity, we propose a novel method, BIP-ALM (Bayesian Inverse Planning Accelerated by Language Models). BIP-ALM extracts unified representations from multimodal data and utilizes language models for scalable Bayesian inverse planning. We conducted a systematic comparison of human performance, BIP-ALM, and state-of-the-art models, including GPT-4. The experiments demonstrate that large language models and large multimodal models still lack robust ToM capacity. BIP-ALM, on the other hand, shows promising results, by leveraging the power of both model-based mental inference and language models.¹

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 MMToM-QA Benchmark
  • 3.1 Overview
  • 3.2 Question Types
  • 3.3 Procedural Generation
  • 3.4 Evaluation Protocol
  • 4 The BIP-ALM Model
  • 4.1 Unified Symbolic Representations
  • 4.2 Inverse Symbolic Planner
  • 5 Experiments
  • 5.1 Baselines
  • 5.2 Results
  • 6 Discussion & Conclusion
  • Ethics Statement
  • Acknowledgements
  • References
  • A Comparison of Theory of Mind Benchmarks
  • B Benchmark Details
  • B.1 More Quantitative Results
  • B.2 Qualitative Results
  • B.3 Discussion on SimToM and SymbolicToM
  • B.4 More Details About the Human Experiment
  • Consent
  • B.5 Available Data
  • B.6 Benchmark Statistics
  • B.7 Details of the Procedural Generation
  • B.8 Utilizing GPT-4 for Enhanced Text Generation
  • C BIP-ALM Implementation Details
  • C.1 Visual Perception
  • C.2 Text Parsing
  • C.3 Representation Fusion
  • C.4 Belief Representation and Update
  • C.5 Prompt for Language Models in Inverse Symbolic Planner
  • C.6 Training Details
  • D Implementation Details of Baselines
  • E Full Version of the Example Questions in Figure 2
  • E.1 Type 1.1 Example
  • E.2 Type 1.2 Example
  • E.3 Type 1.3 Example
  • E.4 Type 2.1 Example
  • E.5 Type 2.2 Example
  • E.6 Type 2.3 Example
  • E.7 Type 2.4 Example

Knowls

  1. Knowl 1 — MMToM-QA Multimodal Theory of Mind Benchmark

    definition

    The Multimodal Theory of Mind Question Answering (MMToM-QA) benchmark is designed to evaluate machine Theory of Mind (ToM) by jointly testing goal inference and belief inference from multimodal and unimodal inputs describing everyday human activities in household environments.

    The benchmark consists of:

    • Test Set: 134 video streams of people searching for objects across apartments (averaging 1,462 frames and 36 actions per video, generated in the VirtualHome-Social simulator). Each video is paired with full-scene and action text descriptions (averaging 1,595 tokens). From these, 600 two-alternative forced-choice questions are constructed (300 belief questions and 300 goal questions).
    • Training Set: 1,000 procedurally synthesized human activity videos in household environments with ground-truth symbolic annotations of scene graphs, objects, agent beliefs, and goals, with no question-answering examples provided.
    • Generalization (Human) Test Set: 40 videos and 120 questions generated from 2 unseen apartment environments where human participants controlled avatars to achieve designated goals.

    Evaluation is zero-shot across three input modalities:

    1. Multimodal QA: Both video and text descriptions are provided.
    2. Text QA: Only textual descriptions of the scene and actions are provided.
    3. Video QA: Only video frames (RGB-D) are provided.
  2. Knowl 2 — Question Taxonomy of MMToM-QA

    definition

    MMToM-QA classifies questions into two broad categories containing seven specific question types, each requiring conditional reasoning between goals and beliefs:

    1. Belief Inference (300 questions total, 100 per type):

    • Type 1.1 (True belief, short-term): An agent is on the verge of opening a container holding a hypothetical goal object that has not been seen yet. The question tests whether the agent believes the object is in that container (consistent with the current action and matching the true world state).
    • Type 1.2 (False belief, short-term): An agent is about to open a container, but the hypothetical goal object is not inside in reality. The question tests whether the model recognizes the agent's false belief that the object is inside based on the observed action.
    • Type 1.3 (Belief tracking, long-term): An agent passes by a container without checking it and continues searching elsewhere without returning. The question tests whether the model infers that the agent believes the object is not in that bypassed container based on long-term trajectory history.

    2. Goal Inference (300 questions total, 75 per type):

    • Type 2.1 (Goal given true belief): An agent walks toward a container containing an unobserved candidate goal object, having already bypassed another candidate object. The question tests inferring the unknown goal given that the agent has true beliefs about the environment.
    • Type 2.2 (Goal given false belief): An agent walks toward a container holding object XX, but the prompt specifies that the agent believes XX is not inside. The question tests inferring that the agent's actual goal is an alternative object suspected to be in that container.
    • Type 2.3 (Goal given updated belief): An agent opens a container and immediately closes it without taking anything. The question tests inferring that the agent updated their belief upon checking and that their goal is an unseen object rather than any item present inside the container.
    • Type 2.4 (Goal given future actions): An agent moves toward a region on a path that could lead to a distant unobserved object. The question tests inferring the goal by anticipating future action trajectories based on continuous spatial layouts and agent orientation.
  3. Knowl 3 — BIP-ALM Architecture for Multimodal Theory of Mind

    model/method

    Bayesian Inverse Planning Accelerated by Language Models (BIP-ALM) is a multimodal Theory of Mind framework that combines symbolic inverse planning with language model policy amortization. The system operates in four stages:

    1. Visual Perception: Transforms RGB-D video frames into symbolic representations by constructing a voxel map from point clouds, estimating 3D bounding boxes and 3D human poses, and constructing dynamic scene graphs with predicate relations such as In(object, container), Close(agent, object), and container states (open or closed).
    2. Text Parsing: Uses a large language model (GPT-4) to parse the textual prompt into symbolic initial state predicates, a sequence of action commands (e.g., walk towards kitchen, open fridge), and two opposing hypotheses H1=⟨g1,b1t⟩H_1 = \langle g_1, b_1^t \rangle and H2=⟨g2,b2t⟩H_2 = \langle g_2, b_2^t \rangle regarding the agent's goal gg and belief state btb^t.
    3. Multimodal Representation Fusion: Merges symbolic predicates extracted from visual scene graphs and text. To resolve perception noise, contradictory visual predicates are overridden by text predicates. Video frames are segmented into discrete time intervals tt aligned with action steps, generating unified state sequences s1:ts^{1:t} and action sequences a1:t−1a^{1:t-1}.
    4. Inverse Symbolic Planner: Uses a fine-tuned language model to amortize agent policy evaluation π(at∣g,bt)\pi(a^t \mid g, b^t) and computes the posterior likelihood ratio between H1H_1 and H2H_2 under Bayesian inverse planning.
  4. Knowl 4 — Bayesian Inverse Planning Posterior Odds with Amortized Action Likelihoods

    equation

    Assuming an agent's behavior follows a Partially Observable Markov Decision Process (POMDP) ⟨S,A,T,G,R,Ω,O,γ⟩\langle S, A, T, G, R, \Omega, O, \gamma \rangle with deterministic state transitions, the posterior probability of a goal g∈Gg \in G and belief state btb^t at time step tt given observed state sequence s1:ts^{1:t} and action sequence a1:ta^{1:t} is given by:

    P(g,bt∣s1:t,a1:t−1)∝[∏τ=1tπ(aτ∣g,bτ)P(bτ∣bτ−1,sτ)]P(b0)P(g)P(g, b^t \mid s^{1:t}, a^{1:t-1}) \propto \left[ \prod_{\tau=1}^t \pi(a^\tau \mid g, b^\tau) P(b^\tau \mid b^{\tau-1}, s^\tau) \right] P(b^0) P(g)

    where π(aτ∣g,bτ)\pi(a^\tau \mid g, b^\tau) is the agent's policy, P(bτ∣bτ−1,sτ)P(b^\tau \mid b^{\tau-1}, s^\tau) is the belief transition distribution, and P(b0)P(b^0) and P(g)P(g) are uniform priors over initial beliefs and goals.

    To compare two candidate hypotheses H1=⟨g1,b1t⟩H_1 = \langle g_1, b_1^t \rangle and H2=⟨g2,b2t⟩H_2 = \langle g_2, b_2^t \rangle at query time tt, BIP-ALM calculates the relative likelihood ratio:

    P(g1,b1t∣s1:t,a1:t)P(g2,b2t∣s1:t,a1:t)=π(at∣g1,b1t)P(b1t∣b^t−1,st)π(at∣g2,b2t)P(b2t∣b^t−1,st)⋅∏τ=1t−1π(aτ∣g1,b^τ)∏τ=1t−1π(aτ∣g2,b^τ)\frac{P(g_1, b_1^t \mid s^{1:t}, a^{1:t})}{P(g_2, b_2^t \mid s^{1:t}, a^{1:t})} = \frac{\pi(a^t \mid g_1, b_1^t) P(b_1^t \mid \hat{b}^{t-1}, s^t)}{\pi(a^t \mid g_2, b_2^t) P(b_2^t \mid \hat{b}^{t-1}, s^t)} \cdot \frac{\prod_{\tau=1}^{t-1} \pi(a^\tau \mid g_1, \hat{b}^\tau)}{\prod_{\tau=1}^{t-1} \pi(a^\tau \mid g_2, \hat{b}^\tau)}

    where b^τ=b^τ(sτ)\hat{b}^\tau = \hat{b}^\tau(s^\tau) denotes the estimated belief at prior step τ<t\tau < t, updated deterministically by accumulating the agent's partial observations oτ=O(sτ)o^\tau = O(s^\tau).

  5. Knowl 5 — Language Model Policy Amortization and Belief Querying in BIP-ALM

    model/method

    BIP-ALM evaluates the action likelihood π(at∣g,bt)\pi(a^t \mid g, b^t) without costly online planning by prompting an autoregressive language model fine-tuned on symbolic state-belief-goal-action trajectories.

    Prompt Format:

    goal: {hypothetical goal}
    state: {state}
    belief (possible locations the person suspects the {hypothetical goal} could be):
    {hypothetical or predicted belief}
    action:
    

    The probability of the observed action ata^t is obtained directly from the output token log probabilities of the language model conditioned on the prompt.

    Belief Representation and Negation:

    • An agent's belief about an object's location is represented symbolically as a list of candidate container/surface names where the object has non-zero probability.
    • To represent a negative belief (e.g., that goal object gg is not at location LL, denoted ¬In(g,L)\neg \text{In}(g, L)), location LL is removed from the list of suspected locations in the prompt.
    • To assess whether the agent believes gg is at location LL versus not at LL, BIP-ALM evaluates the action likelihood under the exclusion of LL versus the exclusion of alternative locations L′L'. If removing LL results in the lowest action likelihood compared to removing alternative locations, the model infers that the agent believes the object is at LL.
  6. Knowl 6 — Procedural Question Generation via Ideal Observer Hypothesis Elimination

    model/method

    The MMToM-QA benchmark synthesizes ground-truth question-answer pairs using an automated three-stage procedural generation pipeline:

    1. Environment and Trajectory Synthesis: Sample an apartment layout, initial world state, and agent goal. A POMDP planner generates an action sequence for an agent searching for the goal object. RGB-D video frames, instance segmentations, 3D poses, and ground-truth scene graphs are rendered in VirtualHome-Social.
    2. Ideal Observer Hypothesis Filtering: At each step tt, an ideal observer tracking model maintains the set of all plausible ⟨goal,belief⟩\langle \text{goal}, \text{belief} \rangle pairs. It updates this set by: (a) eliminating beliefs incompatible with the agent's perceptual observations oto^t, and (b) simulating future actions via the agent's planner to eliminate hypotheses whose predicted actions contradict the agent's observed behavior. For each sampled question type, the pipeline selects one valid hypothesis from the ideal observer set and one invalid hypothesis eliminated by the model.
    3. Natural Language Translation: Symbolic state and action traces are rendered into templates and paraphrased using GPT-4 with constrained prompts to ensure grammatical variety without modifying semantic facts or object locations.
  7. Knowl 7 — Performance Comparison on the MMToM-QA Benchmark

    data/table

    The table below presents zero-shot accuracy (%) across human evaluators, large multimodal models (LMMs), large language models (LLMs), prompting baselines, and BIP-ALM across the seven question types and three input modality conditions on the 600-question test set:

    Method Belief Inference Goal Inference All
    1.1 1.2 1.3 All 2.1 2.2 2.3 2.4 All
    Multimodal
    Human 95.8 96.7 100 97.5 90.0 91.7 83.3 88.9 88.5 93.0
    InstructBLIP 62.0 52.0 32.0 48.7 46.7 29.3 42.7 60.0 44.7 46.7
    Video-LLaMA 2 36.0 38.0 52.0 42.0 36.0 41.3 30.7 45.3 38.3 40.2
    LLaVA 46.0 14.0 69.0 43.0 65.3 22.7 40.0 48.0 44.0 43.5
    GPT-4V 94.0 13.0 59.0 55.3 56.0 26.7 4.0 52.0 34.7 44.0
    BIP-ALM w/ GPT-J 90.0 69.0 86.0 81.7 68.0 78.7 56.0 73.3 69.0 75.3
    BIP-ALM w/ LLaMA 2 88.0 68.0 85.0 80.3 62.7 77.3 72.0 80.0 73.3 76.7
    Text only
    Human 96.0 95.8 81.3 91.0 85.8 76.7 65.0 68.3 74.0 82.5
    GPT-4 97.0 12.0 77.0 62.0 48.0 42.7 2.7 42.7 34.0 48.0
    GPT-3.5 81.0 11.0 39.0 43.7 46.7 16.0 21.3 48.0 33.0 38.3
    GPT-J 56.0 53.0 38.0 49.0 52.0 50.7 50.7 56.0 52.3 59.7
    LLaMA 2 64.0 55.0 50.0 56.3 49.3 48.0 41.3 38.7 44.3 50.3
    SimToM w/ GPT-4 96.0 15.0 82.0 64.3 61.3 44.0 2.7 54.7 40.7 52.5
    SymbolicToM w/ GPT-4 100 61.0 74.0 78.3 73.3 66.7 0.0 50.7 47.7 63.0
    BIP-ALM w/ GPT-J 88.0 69.0 88.0 81.7 77.3 68.0 30.7 70.7 61.7 71.7
    BIP-ALM w/ LLaMA 2 89.0 68.0 90.0 82.3 54.7 66.7 50.7 62.7 58.7 70.5
    Video only
    Human 69.1 64.3 86.4 73.3 58.5 60.0 76.7 63.3 64.6 68.9
    InstructBLIP 56.0 50.0 42.0 49.3 56.0 45.3 54.7 53.3 52.3 50.8
    Video-LLaMA 2 24.0 32.0 67.0 41.0 50.7 45.3 56.0 52.0 51.0 46.0
    LLaVA 33.0 15.0 69.0 39.0 44.0 24.0 56.0 57.3 45.3 42.2
    GPT-4V 64.0 34.0 39.0 45.7 54.7 26.7 48.0 56.0 46.3 46.0
    BIP-ALM w/ GPT-J 63.0 57.0 72.0 64.0 45.3 62.7 50.7 62.7 55.3 59.7
    BIP-ALM w/ LLaMA 2 69.0 63.0 60.0 64.0 62.7 54.7 53.3 62.7 58.3 61.2

    Humans achieve 93.0% overall multimodal accuracy. Standard LLMs and LMMs score near the 50.0% chance level (GPT-4V: 44.0%, GPT-4: 48.0%). BIP-ALM models significantly outperform all baselines, reaching 76.7% (multimodal), 70.5% (text), and 61.2% (video).

  8. Knowl 8 — Failure Modes of Direct LLM and LMM Reasoning on Theory of Mind

    empirical result

    Evaluation on MMToM-QA reveals specific, systematic failure modes in large language models and multimodal foundation models:

    1. True State vs. Mental Belief Confusion: GPT-4 and GPT-4V perform at or near human levels on Type 1.1 (True belief, short-term: 97.0% text, 94.0% MM), where the agent's belief coincides with physical reality. However, performance collapses on Type 1.2 (False belief, short-term: 12.0% text, 13.0% MM), demonstrating that these models rely on retrieving the true world state rather than attributing a distinct mental state to the agent.
    2. Inability to Track Dynamic Belief Updates: On Type 2.3 (Goal given updated belief), GPT-4 scores 2.7% and GPT-4V scores 4.0%. When an agent inspects a container and immediately closes it without taking an object, standard models assume the goal must be an object present inside the container, failing to infer that the agent updated their belief and is pursuing an unobserved item.
    3. Ineffectiveness of Prompting and Scale: Few-shot prompting, chain-of-thought (CoT) prompting, and model scaling (e.g., comparing 7B vs 13B models) fail to consistently elevate baseline LMM/LLM performance above the chance level across question types.
  9. Knowl 9 — Generalization of BIP-ALM to Human Behavioral Data in Unseen Environments

    empirical result

    To evaluate generalization beyond procedurally synthesized agents, models were evaluated on a specialized generalization test set containing 40 videos and 120 questions recorded from 3 human participants navigating 2 completely unseen apartment environments.

    • Multimodal Condition: BIP-ALM w/ LLaMA 2 achieves 77.0% overall accuracy and BIP-ALM w/ GPT-J achieves 76.6% overall accuracy. In contrast, GPT-4V achieves 48.9%, InstructBLIP achieves 49.1%, LLaVA achieves 50.9%, and Video-LLaMA 2 achieves 36.9%.
    • Text-Only Condition: BIP-ALM w/ LLaMA 2 achieves 67.9% and BIP-ALM w/ GPT-J achieves 62.6%, whereas GPT-4 achieves 50.9%, SymbolicToM w/ GPT-4 achieves 53.3%, and SimToM w/ GPT-4 achieves 56.7%.
    • Video-Only Condition: BIP-ALM w/ LLaMA 2 achieves 63.1% and BIP-ALM w/ GPT-J achieves 59.2%, whereas GPT-4V achieves 44.8% and InstructBLIP achieves 49.1%.

    These results demonstrate that the model-based inverse planning structure of BIP-ALM generalizes to real human action sequences and unseen physical layouts without task-specific retraining on human data.

  10. Knowl 10 — Ablation of Language Model Fine-Tuning in BIP-ALM

    data/table

    An ablation study evaluated the contribution of fine-tuning the underlying language models (GPT-J 6B and LLaMA 2 7B via LoRA on 20,000 state-belief-goal-action samples) compared to using frozen pretrained language models in the BIP-ALM pipeline:

    Method Belief Inference Goal Inference All
    1.1 1.2 1.3 All 2.1 2.2 2.3 2.4 All
    Multimodal (MM)
    GPT-J (w/o FT) 84.0 63.0 92.0 79.7 58.7 64.0 16.0 65.3 51.0 65.3
    GPT-J (FT) 90.0 69.0 86.0 81.7 68.0 78.7 56.0 73.3 69.0 75.3
    LLaMA 2 (w/o FT) 56.0 46.0 96.0 66.0 66.7 48.0 29.3 69.3 53.3 59.7
    LLaMA 2 (FT) 88.0 68.0 85.0 80.3 62.7 77.3 72.0 80.0 73.3 76.7
    Text only
    GPT-J (w/o FT) 76.0 61.0 90.0 75.7 44.0 58.7 26.7 56.0 46.3 61.0
    GPT-J (FT) 88.0 69.0 88.0 81.7 77.3 68.0 30.7 70.7 61.7 71.7
    LLaMA 2 (w/o FT) 66.0 53.0 98.0 72.3 57.3 41.3 30.7 65.3 48.7 60.5
    LLaMA 2 (FT) 89.0 68.0 90.0 82.3 54.7 66.7 50.7 62.7 58.7 70.5
    Video only
    GPT-J (w/o FT) 57.0 36.0 77.0 56.7 56.0 60.0 36.0 54.7 51.7 54.2
    GPT-J (FT) 63.0 57.0 72.0 64.0 45.3 62.7 50.7 62.7 55.3 59.7
    LLaMA 2 (w/o FT) 51.0 33.0 75.0 53.0 45.3 72.0 41.3 50.7 52.3 52.7
    LLaMA 2 (FT) 69.0 63.0 60.0 64.0 62.7 54.7 53.3 62.7 58.3 61.2

    Even without fine-tuning (w/o FT), BIP-ALM with a 6B/7B pretrained language model achieves 65.3% (GPT-J) and 59.7% (LLaMA 2) overall multimodal accuracy, outperforming direct prompting of GPT-4V (44.0%). Fine-tuning substantially improves goal inference, raising Type 2.3 accuracy from 16.0% to 56.0% for GPT-J and from 29.3% to 72.0% for LLaMA 2.

  11. Knowl 11 — Limitations of BIP-ALM and MMToM-QA

    limitation

    The MMToM-QA benchmark and BIP-ALM framework have four primary limitations:

    1. Domain Scope: The benchmark is currently restricted to household object search and retrieval activities, omitting other cognitive ToM dimensions such as desires, affective/emotional states, moral reasoning, social norms, and physical constraints.
    2. Missing State Hallucination in Vision: BIP-ALM cannot infer or imagine hidden environmental state information that is absent from visual observations due to occlusions or camera limits unless explicitly supplied in the text modality.
    3. Continuous Spatial Reasoning: BIP-ALM operates over discrete symbolic predicates and symbolic actions. It cannot accurately project continuous spatial trajectories, heading directions, and geometry required for questions involving distant future navigation paths (e.g., Type 2.4 questions).
    4. Language Model Policy Noise: Action likelihood estimation relies on language model output probabilities, which occasionally suffer from planning inaccuracies and hallucinations inherent to generative language models.

Coverage note — None was omitted. All key contributed aspects—including the MMToM-QA benchmark, question taxonomy, procedural generation pipeline, BIP-ALM framework, mathematical POMDP/BIP formulation, policy amortization prompting, experimental results, failure analyses, generalization evaluations, ablations, and stated limitations—are fully covered.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433.
  2. 2.Chris L Baker, Julian Jara-Ettinger, Rebecca Saxe, and Joshua B Tenenbaum. 2017. Rational quantitative attribution of beliefs, desires and percepts in human mentalizing. Nature Human Behaviour, 1(4):1–10.
  3. 3.Chris L Baker, Rebecca Saxe, and Joshua B Tenenbaum. 2009. Action understanding as inverse planning. Cognition, 113(3):329–349.
  4. 4.Valts Blukis, Chris Paxton, Dieter Fox, Animesh Garg, and Yoav Artzi. 2022. A persistent spatial semantic representation for high-level natural language instruction execution. In Conference on Robot Learning, pages 706–717. PMLR.
  5. 5.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712.
  6. 6.Rohan Chandra, Aniket Bera, and Dinesh Manocha. 2020. Stylepredict: Machine theory of mind for human driver behavior from trajectories. arXiv preprint arXiv:2011.04816.
  7. 7.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning.
  8. 8.Kerstin Dautenhahn. 2007. Socially intelligent robots: dimensions of human–robot interaction. Philosophical transactions of the royal society B: Biological sciences, 362(1480):679–704.
  9. 9.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. 2023. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394.
  10. 10.Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah D Goodman. 2023. Understanding social reasoning in language models with language models. arXiv preprint arXiv:2306.15448.
  11. 11.Kanishk Gandhi, Gala Stojnic, Brenden M Lake, and Moira R Dillon. 2021. Baby intuitions benchmark (bib): Discerning the goals, preferences, and actions of others. Advances in Neural Information Processing Systems, 34:9963–9976.
  12. 12.Andrew S. Gordon. 2016. Commonsense interpretation of triangle behavior. In AAAI Conference on Artificial Intelligence.
  13. 13.Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan. 2016. Cooperative inverse reinforcement learning. In Advances in neural information processing systems.
  14. 14.Yinghui He, Yufan Wu, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. 2023. Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models. arXiv preprint arXiv:2310.16755.
  15. 15.John Hewitt and Michael Cohen. 2021. Exploring roberta’s theory of mind through textual entailment.
  16. 16.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  17. 17.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pages 9118–9147. PMLR.
  18. 18.Julian Jara-Ettinger. 2019. Theory of mind as inverse reinforcement learning. Current Opinion in Behavioral Sciences, 29:105–110.
  19. 19.Julian Jara-Ettinger, Hyowon Gweon, Laura E Schulz, and Joshua B Tenenbaum. 2016. The naïve utility calculus: Computational principles underlying commonsense psychology. Trends in cognitive sciences, 20(8):589–604.
  20. 20.Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra. 1998. Planning and acting in partially observable stochastic domains. Artificial intelligence, 101(1-2):99–134.
  21. 21.Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. 2023. Fantom: A benchmark for stress-testing machine theory of mind in interactions. arXiv preprint arXiv:2310.15421.
  22. 22.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  23. 23.Michal Kosinski. 2023. Theory of mind may have spontaneously emerged in large language models. arXiv preprint arXiv:2302.02083.
  24. 24.Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. 2017. Building machines that learn and think like people. Behavioral and brain sciences, 40:e253.
  25. 25.Matt Le, Y-Lan Boureau, and Maximilian Nickel. 2019. Revisiting the evaluation of theory of mind through question answering. In Conference on Empirical Methods in Natural Language Processing.
  26. 26.Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, et al. 2023. M3it: A large-scale dataset towards multi-modal multilingual instruction tuning. arXiv preprint arXiv:2306.04387.
  27. 27.Shuang Li, Xavier Puig, Chris Paxton, Yilun Du, Clinton Wang, Linxi Fan, Tao Chen, De-An Huang, Ekin Akyürek, Anima Anandkumar, et al. 2022. Pretrained language models for interactive decision-making. Advances in Neural Information Processing Systems, 35:31199–31212.
  28. 28.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  29. 29.Shima Rahimi Moghaddam and Christopher J Honey. 2023. Boosting theory-of-mind performance in large language models via prompting. arXiv preprint arXiv:2304.11490.
  30. 30.Aida Nematzadeh, Kaylee Burns, Erin Grant, Alison Gopnik, and Thomas L Griffiths. 2018. Evaluating theory of mind in question answering. arXiv preprint arXiv:1808.09352.
  31. 31.Aviv Netanyahu, Tianmin Shu, Boris Katz, Andrei Barbu, and Joshua B Tenenbaum. 2021. Phase: Physically-grounded abstract social events for machine social perception. In Proceedings of the aaai conference on artificial intelligence, volume 35, pages 845–853.
  32. 32.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  33. 33.Maithili Patel and Sonia Chernova. 2022. Proactive robot assistance via spatio-temporal object modeling. arXiv preprint arXiv:2211.15501.
  34. 34.Xavier Puig, Tianmin Shu, Shuang Li, Zilin Wang, Yuan-Hong Liao, Joshua B Tenenbaum, Sanja Fidler, and Antonio Torralba. 2020. Watch-and-help: A challenge for social perception and human-ai collaboration. arXiv preprint arXiv:2010.09890.
  35. 35.Xavier Puig, Tianmin Shu, Joshua B Tenenbaum, and Antonio Torralba. 2023. Nopa: Neurally-guided online probabilistic assistance for building socially intelligent home assistants. arXiv preprint arXiv:2301.05223.
  36. 36.Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick. 2018. Machine theory of mind. In International conference on machine learning, pages 4218–4227. PMLR.
  37. 37.Kate Sanders, David Etter, Reno Kriz, and Benjamin Van Durme. 2023. Multivent: Multilingual videos of events with aligned natural text. arXiv preprint arXiv:2307.03153.
  38. 38.Maarten Sap, Ronan LeBras, Daniel Fried, and Yejin Choi. 2022. Neural theory-of-mind? on the limits of social intelligence in large lms. arXiv preprint arXiv:2210.13312.
  39. 39.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728.
  40. 40.Rebecca Saxe. 2012. The happiness of the fish: Evidence for a common theory of one’s own and others’ actions. In Handbook of Imagination and Mental Simulation, pages 257–309. Psychology Press.
  41. 41.Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov. 2023. Minding language models’ (lack of) theory of mind: A plug-and-play multi-character belief tracker.
  42. 42.Melanie Sclar, Graham Neubig, and Yonatan Bisk. 2022. Symmetric machine theory of mind. In International Conference on Machine Learning, pages 19450–19466. PMLR.
  43. 43.Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz. 2023. Clever hans or neural theory of mind? stress testing social reasoning in large language models. arXiv preprint arXiv:2305.14763.
  44. 44.Tianmin Shu, Abhishek Bhandwaldar, Chuang Gan, Kevin Smith, Shari Liu, Dan Gutfreund, Elizabeth Spelke, Joshua Tenenbaum, and Tomer Ullman. 2021. Agent: A benchmark for core psychological reasoning. In International Conference on Machine Learning, pages 9614–9625. PMLR.
  45. 45.Hrituraj Singh, Anshul Nasery, Denil Mehta, Aishwarya Agarwal, Jatin Lamba, and Balaji Vasan Srinivasan. 2021. Mimoqa: Multimodal input multimodal output question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5317–5332.
  46. 46.Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. 2021. Multimodalqa: Complex question answering over text, tables and images. arXiv preprint arXiv:2104.06039.
  47. 47.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  48. 48.Tomer Ullman. 2023. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399.
  49. 49.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  50. 50.Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel. 2021. Towards mutual theory of mind in human-ai interaction: How language reflects what students perceive about a virtual teaching assistant. In Proceedings of the 2021 CHI conference on human factors in computing systems, pages 1–14.
  51. 51.Alex Wilf, Sihyun Shawn Lee, Paul Pu Liang, and Louis-Philippe Morency. 2023. Think twice: Perspective-taking improves large language models’ theory-of-mind capabilities. arXiv preprint arXiv:2311.10227, arXiv:2311.10227v1.
  52. 52.Heinz Wimmer and Josef Perner. 1983. Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception. Cognition, 13(1):103–128.
  53. 53.Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency. 2019. Social-iq: A question answering benchmark for artificial social intelligence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8807–8817.
  54. 54.Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6720–6731.
  55. 55.Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.

Citation

MLA
Jin, C., et al. “MMToM-QA: Multimodal Theory of Mind Question Answering”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 16077–102, https://doi.org/10.18653/v1/2024.acl-long.851.
APA
Jin, C., Wu, Y., Cao, J., Xiang, J., Kuo, Y.-L., Hu, Z., Ullman, T., Torralba, A., Tenenbaum, J., & Shu, T. (2024). MMToM-QA: Multimodal Theory of Mind Question Answering. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16077–16102. https://doi.org/10.18653/v1/2024.acl-long.851
Chicago
Jin, C., Y. Wu, J. Cao, et al. 2024. “MMToM-QA: Multimodal Theory of Mind Question Answering”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16077–102. https://doi.org/10.18653/v1/2024.acl-long.851.
Harvard
Jin, C. et al. (2024) “MMToM-QA: Multimodal Theory of Mind Question Answering”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 16077–16102. Available at: https://doi.org/10.18653/v1/2024.acl-long.851.
Vancouver
1. Jin C, Wu Y, Cao J, Xiang J, Kuo Y-L, Hu Z, Ullman T, Torralba A, Tenenbaum J, Shu T (2024) MMToM-QA: Multimodal Theory of Mind Question Answering. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 16077–16102

BibTeX

@inproceedings{jin-etal-2024-mmtom,
    title = "{MMT}o{M}-{QA}: Multimodal Theory of Mind Question Answering",
    author = "Jin, Chuanyang  and
      Wu, Yutong  and
      Cao, Jing  and
      Xiang, Jiannan  and
      Kuo, Yen-Ling  and
      Hu, Zhiting  and
      Ullman, Tomer  and
      Torralba, Antonio  and
      Tenenbaum, Joshua  and
      Shu, Tianmin",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.851/",
    doi = "10.18653/v1/2024.acl-long.851",
    pages = "16077--16102"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/