Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

Eunkyu ParkWesley DengGunhee KimMotahhare EslamiMaarten Sap

article2025arXiv3 citations

Proposes Cognitive Chain-of-Thought (CoCoT), a three-stage framework decomposing multimodal reasoning into perception, situation inference, and social norm application, which consistently improves vision-language models across social reasoning benchmarks and enables models to internalize structured reasoning through fine-tuning.

Listen

Modern vision-language models frequently fail when interpreting social situations, speaker intentions, and cultural norms from visual scenes. Standard step-by-step reasoning prompts conflate physical observation with abstract social interpretation, causing models to miss visual cues, misread context, or generate plausible-sounding but ungrounded conclusions. The article addresses this operational gap by demonstrating that multimodal reasoning improves when structured into explicit, cognitively grounded stages.

The article introduces and evaluates Cognitive Chain-of-Thought, a structured reasoning framework designed to improve how models interpret social interactions and safety constraints in images and video. The framework breaks reasoning into three sequential stages: Perception (cataloging only directly observable physical evidence), Situation (building a relational model of interactions and goals), and Norm (applying pragmatic and social rules to make a final judgment).

To establish credibility across diverse conditions, the analysis tested seven proprietary, open-source, and specialized reasoning models across four standard benchmarks covering intent disambiguation, theory of mind, social commonsense, and safety instruction compliance. The evaluation compared standard direct answering against unstructured step-by-step reasoning, scene-graph-augmented reasoning, and the three-stage framework. Furthermore, researchers fine-tuned open-source models on structured reasoning traces and conducted blind human evaluations on 100 sample outputs using 87 crowd-workers.

The findings show consistent performance advantages across all evaluated domains. First, the structured framework produced average accuracy gains of 4.6% to 5.9% over direct answers across benchmarks, with intent disambiguation jumping up to 14.4% on select standard models. Second, the method achieved its largest improvements on subtle, indirect social tasks, surging accuracy by 12.8% on non-literal communication tasks where standard unstructured reasoning degraded performance. Third, in safety evaluations, the framework reduced the failure rate of rejecting harmful visual instructions from 28.3% down to 14.9%. Fourth, fine-tuning smaller open-source models on structured reasoning traces improved baseline accuracy by 2.3% to 7.6% even when no structural prompts were used during testing, proving that models internalized the reasoning process. Finally, human evaluators preferred the structured reasoning traces over unstructured alternatives by a margin of 61.8% to 38.2%, citing superior logical flow.

These results demonstrate that enforcing separate observation and interpretation steps prevents models from anchoring on premature assumptions or bypassing visual facts. In practical applications, this structure reduces the risk of inappropriate model outputs in sensitive user interactions, improves system safety compliance without requiring heavy computing overhead, and provides clearer audit trails by isolating whether an error occurred during visual perception, situational framing, or normative evaluation.

Organizations developing or deploying multimodal systems should adopt structured, stage-separated reasoning prompts for socially grounded and safety-critical tasks. Teams seeking lower operational latency can fine-tune lighter open-source models on structured traces to capture accuracy gains without adding runtime prompt overhead. Before widespread implementation in non-Western environments, technical teams must conduct pilot audits to ensure that the embedded normative rules align with local cultural standards, as the evaluated models primarily reflect Western social norms.

Decision-makers should view these findings with high confidence regarding the demonstrated performance gains across standard benchmarks. However, caution is advised because structured prompting does not eliminate hallucinations, guarantee faithful internal logic, or translate directly to purely non-social tasks such as mathematical proofs or software coding.

Cover for Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

Abstract

Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must perceive, understand, and judge all at once; bridging perception with norm-grounded reasoning. Recent work has introduced structured reasoning for multi-turn agent planning and visual QA, decomposing tasks into sequential sub-goals. To extend this to single-shot multimodal social reasoning, we introduce Cognitive Chain-of-Thought (CoCoT), a reasoning framework that structures vision-language-model (VLM) reasoning through three cognitively inspired stages: Perception (extract grounded facts), Situation (infer situations), and Norm (applying social norms). Evaluation across multiple distinct tasks such as multimodal intent disambiguation, multimodal theory of mind, social commonsense reasoning, and safety instruction following, shows consistent improvements (5.9% to 4.6% on average). We further explore the utility of CoCoT for improving models' reasoning through training and show that supervised fine-tuning on CoCoT-structured traces yields 5-6% improvements without explicit CoCoT prompting at inference, demonstrating that models internalize the structured reasoning pattern rather than merely following instructions. We show that structuring model reasoning through cognitively grounded stages enhances interpretability and social alignment, laying the groundwork for more reliable multimodal systems.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Cognitive Chain-of-Thought (CoCoT)
  • 3.1 Supervised Fine-Tuning
  • 4 Experiments
  • 4.1 Evaluation of Multimodal Intent Disambiguation on VAGUE
  • 4.2 Evaluation of Multimodal Theory of Mind on MoMentS
  • 4.3 Evaluation of Multimodal Commonsense Reasoning on M3CoT
  • 4.4 Evaluation of Safety Robustness on VLGuard
  • 4.5 Analyses
  • 4.5.1 Task-Dependent Stage Ablations
  • 4.5.2 Where Does CoCoT Help?
  • 4.5.3 Faithfulness of the Perception Stage
  • 4.6 Results of Supervised Fine-Tuning for Transferability
  • 4.7 Human Evaluation
  • 5 Conclusion
  • 6 Ethics Statement
  • 7 Acknowledgments
  • References
  • A Appendix
  • A.1 Full Accuracy Scores for VAGUE, MoMentS and M3CoT
  • A.2 Prompt Templates for CoCoT Inference
  • A.3 Ablation on Cognitive Stages for VLGuard
  • A.4 SFT Details
  • A.4.1 Trace Generation
  • A.4.2 Training Details
  • A.5 Qualitative Examples: VAGUE
  • A.6 Qualitative Examples: MoMentS
  • A.7 Qualitative Examples: M3CoT
  • A.8 Qualitative Examples: VLGuard
  • A.9 Human Evaluation Interface and Annotation Protocol
  • A.9.1 Annotation Questions

Knowls

  1. Knowl 1 — Cognitive Chain-of-Thought (CoCoT) Reasoning Framework

    model/method

    Cognitive Chain-of-Thought (CoCoT) is a structured reasoning framework for vision-language models (VLMs) performing single-shot multimodal social reasoning. Grounded in cognitive science principles of grounded cognition and incremental sense-making, CoCoT structures inference into three sequential stages:

    1. Perception: Grounding reasoning strictly in directly observable, verifiable visual evidence present in the image or video frames (e.g., physical actions, object states, body postures, spatial arrangements, and facial expressions). The model is explicitly constrained not to infer mental states, emotions, intentions, or social meanings at this stage.
    2. Situation: Constructing a coherent situation model by integrating the perceptual evidence from Stage 1 with situational and contextual knowledge (e.g., social scripts, interaction patterns, relational dynamics, goal inference, and information asymmetry regarding what each agent can plausibly know). The model connects observable cues to the situational context without making final normative judgments.
    3. Norm: Performing normative and pragmatic evaluation to select the most plausible interpretation or answer constrained by the Perception →\to Situation chain. This stage applies social conventions, conversational implicature, pragmatic principles (such as sarcasm or indirectness), and safety or moral norms.

    CoCoT can be deployed either via zero-shot prompt scaffolding at inference time or through supervised fine-tuning on validated CoCoT traces.

  2. Knowl 2 — CoCoT Trace Generation and Validation Pipeline

    algorithm

    To create training traces for fine-tuning models on structured social reasoning, CoCoT uses a teacher generation and automated filtering pipeline that prevents observation-inference conflation.

    Input: Training set D = {(x_image, q, C, y*)} where x_image is the visual input, q is the question, C is the candidate choices, y* is the ground-truth answer; Teacher VLM T; Lexicon of blocked mental-state terms L.
    Output: Validated dataset D_CoCoT of structured reasoning traces.
    D_CoCoT = empty_set
    for each sample (x_image, q, C, y*) in D:
        valid = False
        attempts = 0
        while valid is False and attempts < 3:
            attempts = attempts + 1
            trace = T.generate(x_image, q, C, y*, prompt="CoCoT generation prompt enforcing [Perception], [Situation], [Norm] arriving at y*")
            
            // Criterion 1: Mental-state leakage filter
            P_text = extract_stage(trace, "Perception")
            has_leakage = False
            for term in L:
                if term in lowercase(P_text):
                    has_leakage = True
                    break
            if has_leakage:
                continue
            
            // Criterion 2: Stage ordering check
            if not (contains_in_order(trace, ["[Perception]", "[Situation]", "[Norm]"])):
                continue
            
            // Criterion 3: Grounding check
            S_text = extract_stage(trace, "Situation")
            entities_in_P = extract_entities(P_text)
            if not (any entity in S_text for entity in entities_in_P):
                continue
            
            // Criterion 4: Target answer correctness
            if not arrives_at_answer(trace, y*):
                continue
            
            valid = True
            D_CoCoT.add((x_image, q, trace, y*))
    return D_CoCoT

    The mental-state lexicon LL specifically filters intentional states (intends, wants, hopes), epistemic states (believes, thinks, assumes), affective states (angry, worried, happy), pragmatic inferences (sarcastic, ironic, pretending), teleological constructions (in order to, deliberately, on purpose), and affective dispositions (feels, feeling).

  3. Knowl 3 — Multimodal Social Reasoning Benchmark Performance of CoCoT

    empirical result

    Across multiple social reasoning benchmarks, CoCoT prompting consistently outperforms Direct prompting, standard Chain-of-Thought (CoT), and Compositional Chain-of-Thought (CCoT):

    Model VAGUE (Accuracy %) MoMentS (Accuracy %)
    Direct Δ\DeltaCoT Δ\DeltaCCoT Δ\DeltaCoCoT Direct Δ\DeltaCoT Δ\DeltaCCoT Δ\DeltaCoCoT
    GPT-4o 61.60 -1.60 -11.48 +5.83 70.68 -2.56 +0.11 +1.58
    Claude-3.5-Sonnet 62.92 -3.40 -10.90 +4.38 63.75 -2.40 -1.75 -0.80
    Gemini-2.5-Pro 53.25 +5.07 -6.55 +14.37 55.00 -2.00 +0.50 +5.00
    OpenAI o1-full 59.82 +2.75 -10.98 +1.77 67.60 -3.41 -1.10 +1.55
    OpenAI o3-mini 38.15 +2.06 -3.72 +5.49 56.00 -3.06 -0.50 +2.00
    GPT-5.2 (High) 59.21 +0.78 -11.90 +2.81 62.95 -0.52 -0.63 -0.40
    Gemini-3.0-Flash 73.05 -0.78 -13.05 +4.35 64.80 +2.00 +3.70 +6.51
    LLaVA-OneVision-7B 49.73 +3.58 -5.43 +4.92 45.82 +2.39 -0.32 +0.40
    Qwen2.5-VL-7B 38.52 -0.03 -4.50 +5.90 49.00 -11.91 -8.13 -7.80

    On the M3CoTM^3\text{CoT} social-commonsense subset, CoCoT yields +11.99%+11.99\% on GPT-4o (57.23%→69.22%57.23\% \to 69.22\%) and +5.90%+5.90\% on Claude-3.5-Sonnet (65.10%→71.00%65.10\% \to 71.00\%). Standard CoT often reduces performance compared to Direct answering on these benchmarks, while CCoT (which injects generic visual scene graphs) causes severe accuracy drops (e.g., −11.48%-11.48\% on GPT-4o on VAGUE), showing that generic compositional parsing without grounded cognitive stages introduces noise in subtle social reasoning tasks.

  4. Knowl 4 — Task-Dependent Stage Ablations: Monotonic vs. V-Shaped Recovery

    empirical result

    Ablations examining the cumulative addition of stages (PP: Perception only; P+SP+S: Perception + Situation; Full CoCoT: P+S+NormP+S+Norm) reveal two distinct dynamics depending on task epistemic structure:

    1. Monotonic Progression (VAGUE - Multimodal Intent Disambiguation): Accuracy increases steadily as cognitive abstraction increases: Direct (61.6%)→P only (63.5%)→P+S only (65.5%)→CoCoT (67.4%)\text{Direct } (61.6\%) \to P\text{ only } (63.5\%) \to P+S\text{ only } (65.5\%) \to \text{CoCoT } (67.4\%) This indicates balanced, additive contributions where direct visual-to-pragmatic mappings benefit continuously from intermediate grounding.

    2. V-Shaped Synergistic Recovery (MoMentS - Theory of Mind): Accuracy degrades when intermediate stages are introduced without the final normative stage: Direct (70.7%)→P only (68.5%)→P+S only (66.4%)→CoCoT (72.3%)\text{Direct } (70.7\%) \to P\text{ only } (68.5\%) \to P+S\text{ only } (66.4\%) \to \text{CoCoT } (72.3\%) In Theory of Mind tasks, the Situation stage alone expands representational capacity (hypothesizing unobservable mental states or informational access), which introduces anchoring biases toward partial, misleading interpretations. The final Norm stage acts as a selection filter that resolves ambiguity and aligns the situation model with plausible social conventions.

  5. Knowl 5 — Theory-of-Mind Performance Disaggregated Across Cognitive Dimensions

    empirical result

    Evaluation across the seven Theory of Mind categories in the MoMentS benchmark reveals that CoCoT provides the largest improvements on tasks requiring models to resolve discrepancies between literal visual observations and underlying social meaning:

    • Non-literal Communication: Accuracy improves by +3.2%+3.2\% over Direct and +12.8%+12.8\% over standard CoT (rising from 56.3%56.3\% Direct and 46.7%46.7\% CoT to 59.3%59.3\% CoCoT).
    • Intentions: Accuracy improves by +1.6%+1.6\% over Direct and +8.6%+8.6\% over CoT.
    • Beliefs: Improves by +0.7%+0.7\% over Direct and +3.5%+3.5\% over CoT.
    • Emotions: Improves by +2.3%+2.3\% over Direct and +3.2%+3.2\% over CoT.
    • Knowledge, Percepts, Desires: Marginal or mixed changes (Knowledge: +1.6%+1.6\% over Direct, +0.3%+0.3\% over CoT; Percepts: +1.6%+1.6\% over Direct, −0.7%-0.7\% over CoT; Desires: −0.6%-0.6\% over Direct, +1.3%+1.3\% over CoT).

    Because categories like Percepts and Knowledge can be solved via direct visual extraction, explicit situation modeling offers limited marginal utility, whereas tasks with non-literal communication or latent intent require the Situation stage to bridge the gap between perception and hidden mental states.

  6. Knowl 6 — Internalization of CoCoT Reasoning via Supervised Fine-Tuning

    empirical result

    To test whether CoCoT's reasoning structure can be internalized rather than prompted, open-source models (LLaVA-OneVision-7B and Qwen2.5-VL-Instruct-7B) were fine-tuned on CoCoT-generated traces using QLoRA (4-bit NF4, rank 16, α=32\alpha=32, batch size 16, 3 epochs, cosine schedule) and evaluated using direct prompting (with zero CoCoT stage cues at inference time):

    Model Setting VAGUE M3^3CoT
    Acc. (%) Δ\Delta Acc. (%) Δ\Delta
    LLaVA-OneVision-7B Direct 49.73 — 48.00 —
    Direct + CoCoT-SFT 52.53 +3.80 49.72 +1.72
    Qwen2.5-VL-Instruct-7B Direct 38.52 — 56.54 —
    Direct + CoCoT-SFT 43.12 +5.60 58.80 +2.26

    The retention of +1.7%+1.7\% to +5.6%+5.6\% accuracy improvements in direct inference confirms that models internalize the underlying three-stage cognitive structure as an implicit reasoning pattern rather than merely mimicking output token templates.

  7. Knowl 7 — Perception Stage Faithfulness Audit via LLM-Judge and Object Detection

    empirical result

    The faithfulness of the visual grounding in CoCoT's Perception stage was audited across 200 traces per model/benchmark using two complementary mechanisms:

    1. LLM-Judge Faith@known: Decomposing Perception paragraphs into atomic visual claims via GPT-5.2 and computing: Faith@known=YESYES+NO\text{Faith@known} = \frac{\text{YES}}{\text{YES} + \text{NO}} where confident atomic claims are evaluated against the image/frames.
    2. Grounding DINO Precision: Extracting concrete noun phrases with spaCy and querying open-vocabulary detector Grounding DINO, marking a phrase grounded if any bounding box achieves confidence >0.30> 0.30.
    Model Benchmark Faith@known (%) G-DINO Precision (%) # Claims/Item # Phrases/Item
    GPT-4o VAGUE 91.9 81.2 6.7 10.5
    GPT-4o MoMentS 88.1 72.5 6.6 9.9
    Qwen2.5-VL-7B VAGUE 90.0 86.2 5.4 7.3
    Qwen2.5-VL-7B MoMentS 88.1 71.8 6.1 7.3

    Perception grounding is consistently more reliable on static images (VAGUE) than on video frames (MoMentS), explaining why incomplete chains (PP-only or P+SP+S) suffer on video ToM tasks until the Norm stage provides holistic resolution.

  8. Knowl 8 — Safety Robustness and Harm Mitigation on VLGuard

    empirical result

    On the multimodal safety instruction benchmark VLGuard (tested on GPT-4o across 1K image-text pairs), CoCoT reduces the Attack Success Rate (ASR ↓\downarrow, measuring the proportion of harmful instructions executed):

    Data Subset CoT Moral CoT CCoT CoCoT
    Safe_Unsafe (Safe image, unsafe query) 28.3% 19.0% 46.4% 14.9%
    Unsafe (Unsafe image) 29.4% 25.8% 37.6% 13.4%

    Ablating CoCoT stages demonstrates a trade-off between safety (ASR) and over-conservativeness (False Rejection Rate, FRR ↑\uparrow, on safe queries paired with unsafe-looking images):

    • Full CoCoT: ASR=14.9%\text{ASR} = 14.9\%, FRR=22.4%\text{FRR} = 22.4\%
    • No Perception: ASR=15.6%\text{ASR} = 15.6\%, FRR=24.4%\text{FRR} = 24.4\%
    • No Situation: ASR=13.6%\text{ASR} = 13.6\%, FRR=28.1%\text{FRR} = 28.1\%
    • Norm Only: ASR=19.2%\text{ASR} = 19.2\%, FRR=14.9%\text{FRR} = 14.9\%

    Full CoCoT balances safety enforcement with avoiding spurious refusals by conditioning normative harm judgments on grounded situational analysis.

  9. Knowl 9 — Human Evaluation of CoCoT Reasoning Traces

    empirical result

    In a human evaluation study comprising 300 pairwise comparisons across 100 stratified question-trace pairs evaluated by 87 crowd-workers (Fleiss' κ=0.48\kappa = 0.48), CoCoT was strongly preferred over standard Chain-of-Thought (CoT):

    • Overall Preference Rate: 61.8%61.8\% for CoCoT vs. 38.2%38.2\% for standard CoT (+23.6%+23.6\% margin).
    • Faithfulness (1–5 Likert): 3.943.94 (CoCoT) vs. 3.713.71 (CoT), Δ=+0.23\Delta = +0.23.
    • Logical Coherence (1–5 Likert): 3.773.77 (CoCoT) vs. 3.293.29 (CoT), Δ=+0.48\Delta = +0.48.
    • Social Knowledge (1–5 Likert): 3.823.82 (CoCoT) vs. 3.643.64 (CoT), Δ=+0.18\Delta = +0.18.

    The largest rating gain appears in logical coherence, indicating that the explicit transition from visual evidence to situation modeling to norm application produces clearer, more inspectable inference paths for human oversight.

  10. Knowl 10 — Limitations of CoCoT in Multimodal Social Reasoning

    limitation

    CoCoT has several structural and domain limitations:

    1. Unfaithful Internal Reasoning & Post-Hoc Rationalization: While CoCoT forces external text traces into three stages, it does not guarantee that the VLM's internal computational mechanism faithfully reflects this sequence, retaining the risk of post-hoc rationalization.
    2. Lack of Domain Generalizability: CoCoT's specific cognitive progression is tailored to social, pragmatic, and socionormative reasoning, and does not generalize directly to formal symbolic domains such as mathematical theorem proving or code synthesis.
    3. Hallucination Propagation: CoCoT does not eliminate perceptual or factual hallucinations; errors introduced in the Perception stage can cascade through Situation and Norm stages, making confabulated inferences appear structured and plausible.
    4. Cultural and Normative Bias: CoCoT's normative evaluation assumes shared social conventions, but underlying training corpora and models are predominantly aligned with Western cultural norms, risking misjudgments in diverse non-Western social contexts.

Coverage note — None was omitted; all key framework details, algorithms, main benchmark comparisons, ablation findings, SFT results, safety evaluations, faithfulness checks, human evaluations, and limitations are fully covered.

References

  1. 1.Mohamed R. Amer, Timothy J. Shields, Behjat Siddiquie, Amir Tamrakar, Ajay Divakaran, and Sek M. Chai. Deep multimodal fusion: A hybrid approach. International Journal of Computer Vision, 126:440–456, 2017.
  2. 2.Anthropic. Claude 3.5 sonnet, jun 2024. URL https://www.anthropic.com/news/claude-3-5-sonnet.
  3. 3.Ian A. Apperly and Stephen Butterfill. Do humans have two systems to track beliefs and belief-like states? Psychological Review, 116(4):953–970, 2009. doi: 10.1037/a0016923.
  4. 4.Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems, pp. 1–16, 2021.
  5. 5.Lawrence W Barsalou. Grounded cognition. Annual Review of Psychology, 59(1):617–645, 2008. doi: 10.1146/annurev.psych.59.103006.093639.
  6. 6.Lawrence W Barsalou. Challenges and opportunities for grounding cognition. Journal of Cognition, 3(1):31, 2020. doi: 10.5334/joc.116.
  7. 7.Lawrence W Barsalou et al. Perceptions of perceptual symbols. Behavioral and brain sciences, 22(4):637–660, 1999.
  8. 8.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of the AAAI conference on artificial intelligence, pp. 17682–17690, 2024.
  9. 9.Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che. M3CoT: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8199–8221. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.acl-long.446.
  10. 10.Andy Clark. Whatever next? predictive brains, situated agents, and the future of cognitive science. Behavioral and brain sciences, 36(3):181–204, 2013.
  11. 11.Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025.
  12. 12.Google DeepMind. Gemini api: Thinking mode, 2024. URL https://ai.google.dev/gemini-api/docs/thinking.
  13. 13.Karl Friston. The free-energy principle: a unified brain theory? Nature reviews neuroscience, 11(2):127–138, 2010.
  14. 14.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720, 2022.
  15. 15.Alon Hafri, Anna Papafragou, and John Trueswell. Getting the gist of events: Recognition of two-participant actions from brief displays. Journal of Experimental Psychology: General, 142:880–905, 09 2012. doi: 10.1037/a0030045.
  16. 16.Dongzhi Jiang, Renrui Zhang, Ziyu Guo, Yanwei Li, Yu Qi, Xinyan Chen, Liuhui Wang, Jianhan Jin, Claire Guo, Shen Yan, Bo Zhang, Chaoyou Fu, Peng Gao, and Hongsheng Li. MME-cot: Benchmarking chain-of-thought in large multimodal models for reasoning quality, robustness, and efficiency. In Forty-second International Conference on Machine Learning, 2025.
  17. 17.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations, 2023.
  18. 18.Gary Klein, Brian Moon, and Robert R Hoffman. Making sense of sensemaking 1: Alternative perspectives. IEEE intelligent systems, 21(4):70–73, 2006.
  19. 19.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  20. 20.Erin L Krupka and Roberto A Weber. Identifying social norms using coordination games: Why does dictator game sharing vary? Journal of the European Economic Association, 11(3):495–524, 2013.
  21. 21.Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson E. Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, John Kernion, et al. Measuring faithfulness in chain-of-thought reasoning. arXiv.org, 2023. doi: 10.48550/arXiv.2307.13702.
  22. 22.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. LLaVA-onevision: Easy visual task transfer. Transactions on Machine Learning Research, 2025.
  23. 23.Yuan-Hong Liao, Sven Elflein, Liu He, Laura Leal-Taixe, Yejin Choi, Sanja Fidler, and David ´Acuna. Longperceptualthoughts: Distilling system-2 reasoning for system-1 perception. In Second Conference on Language Modeling, 2025.
  24. 24.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023.
  25. 25.Leena Mathur, Paul Pu Liang, and Louis-Philippe Morency. Advancing social intelligence in AI agents: Technical challenges and open questions. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 20541–20560. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.emnlp-main.1143.
  26. 26.Leena Mathur, Marian Qian, Paul Pu Liang, and Louis-Philippe Morency. Social genome: Grounded social reasoning abilities of multimodal models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 24879–24902. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.emnlp-main.1264.
  27. 27.Nick McKenna, Tianyi Li, Liang Cheng, Mohammad Hosseini, Mark Johnson, and Mark Steedman. Sources of hallucination by large language models on inference tasks. In Findings of the association for computational linguistics: EMNLP 2023, pp. 2758–2774, 2023.
  28. 28.Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. Compositional chain of thought prompting for large multimodal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024.
  29. 29.Heejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung, and Youngjae Yu. Vague: visual contexts clarify ambiguous expressions. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1537–1547. IEEE, 2025.
  30. 30.Albert Newen, Leon De Bruin, and Shaun Gallagher. The Oxford handbook of 4E cognition. Oxford University Press, 2018.
  31. 31.OpenAI. Hello gpt-4o, 2024a. URL https://openai.com/index/hello-gpt-4o/.
  32. 32.OpenAI. Introducing openai o1-preview, 2024b. URL https://openai.com/index/introducing-openai-o1-preview/.
  33. 33.Liuba Papeo, Timo Stein, and Salvador Soto-Faraco. The two-body inversion effect. Psychological Science, 28:369 – 379, 2017. URL https://api.semanticscholar.org/CorpusID:43986600.
  34. 34.Eunkyu Park, Wesley Hanwen Deng, Vasudha Varadarajan, Mingxi Yan, Gunhee Kim, Maarten Sap, and Motahhare Eslami. Critical or compliant? the double-edged sword of reasoning in chain-of-thought explanations. arXiv preprint arXiv:2511.12001, 2025.
  35. 35.Philip Robbins and Murat Aydede. The Cambridge handbook of situated cognition. Cambridge University Press, 2008.
  36. 36.Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang. Visual chain of thought: bridging logical gaps with multimodal infillings. arXiv preprint arXiv:2305.02317, 2023.
  37. 37.Wolff-Michael Roth and Alfredo Jornet. Situated cognition. Wiley Interdisciplinary Reviews: Cognitive Science, 4(5):463–478, 2013. doi: 10.1007/978-3-031-39744-8.
  38. 38.Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 4454–4470, 2023.
  39. 39.Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. Synthetic prompting: Generating chain-of-thought demonstrations for large language models. In ICML, pp. 30706–30775, 2023.
  40. 40.Dan Sperber and Deirdre Wilson. Relevance: Communication and cognition, volume 142. Harvard University Press Cambridge, MA, 1986.
  41. 41.Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems, 36:74952–74965, 2023.
  42. 42.Emilio Villa-Cueva, S M Masrur Ahmed, Rendi Chevi, Jan Christian Blaise Cruz, Kareem Elzeky, Fermin Cristobal, Alham Fikri Aji, Skyler Wang, Rada Mihalcea, and Thamar Solorio. MoMentS: A comprehensive multimodal benchmark for theory of mind. In Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 22591–22611, 2025.
  43. 43.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), pp. 2609–2634, 2023a.
  44. 44.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, 2023b.
  45. 45.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (eds.), Advances in Neural Information Processing Systems, 2022a.
  46. 46.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022b.
  47. 47.Karl E Weick and Karl E Weick. Sensemaking in organizations, volume 3. Sage publications Thousand Oaks, CA, 1995.
  48. 48.Margaret Wilson. Six views of embodied cognition. Psychonomic bulletin & review, 9(4):625–636, 2002.
  49. 49.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671, 2023.
  50. 50.Xuyang Wu, Jinming Nian, Ting-Ruen Wei, Zhiqiang Tao, Hsin-Tai Wu, and Yi Fang. Does reasoning introduce bias? a study of social bias evaluation and mitigation in LLM reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 18534–18555, 2025.
  51. 51.Bei Yan, Jie Zhang, Zhiyuan Chen, Shiguang Shan, and Xilin Chen. Mm-moralbench: A multimodal moral evaluation benchmark for large vision-language models. Pattern Recognition, 179:113624, 2024. doi: 10.1016/j.patcog.2026.113624.
  52. 52.Qwen An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxin Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yi-Chao Zhang, Yunyang Wan, Yuqi Liu, Zeyu Cui, Zhenru Zhang, Zihan Qiu, Shanghaoran Quan, and Zekun Wang. Qwen2.5 technical report. ArXiv, abs/2412.15115, 2024.
  53. 53.Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv preprint arXiv:2303.11381, 2023.
  54. 54.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  55. 55.Xi Ye and Greg Durrett. The unreliability of explanations in few-shot prompting for textual reasoning. Advances in Neural Information Processing Systems, 35:30378–30392, 2022.
  56. 56.Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. Understanding the effect of accuracy on trust in machine learning models. In Proceedings of the 2019 CHI conference on human factors in computing systems, pp. 1–12, 2019.
  57. 57.Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making. In Proceedings of the 2020 conference on fairness, accountability, and transparency, pp. 295–305, 2020.
  58. 58.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations, 2023.
  59. 59.Zhuosheng Zhang, Aston Zhang, Mu Li, hai zhao, George Karypis, and Alex Smola. Multimodal chain-of-thought reasoning in language models. Transactions on Machine Learning Research, 2024.
  60. 60.Denny Zhou, Nathanael Scharli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi. Least-to-most prompting enables complex reasoning in large language models. In The Eleventh International Conference on Learning Representations, 2023.
  61. 61.Xuhui Zhou, Jiarui Liu, Akhila Yerukola, Hyunwoo Kim, and Maarten Sap. Social world models. arXiv preprint arXiv:2509.00559, 2025.
  62. 62.Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales. Safety fine-tuning at (almost) no cost: a baseline for vision large language models. In Proceedings of the 41st International Conference on Machine Learning, ICML’24, 2024.

Citation

MLA
Park, E., et al. “Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning About Social Situations”. arXiv, 2025, http://arxiv.org/abs/2507.20409v3.
APA
Park, E., Deng, W. H., Kim, G., Eslami, M., & Sap, M. (2025). Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations. arXiv. http://arxiv.org/abs/2507.20409v3
Chicago
Park, E., W. H. Deng, G. Kim, M. Eslami, and M. Sap. 2025. “Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning About Social Situations”. arXiv. http://arxiv.org/abs/2507.20409v3.
Harvard
Park, E. et al. (2025) “Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2507.20409v3.
Vancouver
1. Park E, Deng WH, Kim G, Eslami M, Sap M (2025) Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations. arXiv

BibTeX

@article{park2025cognitive,
  title = {Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations},
  author = {Park, Eunkyu and Deng, Wesley Hanwen and Kim, Gunhee and Eslami, Motahhare and Sap, Maarten},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2507.20409v3},
  eprint = {2507.20409}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/