Social-R1: Towards Human-like Social Reasoning in LLMs cover

Social-R1: Towards Human-like Social Reasoning in LLMs

Jincenzi Wu 1^{1}1
1^{1}1 The Chinese University of Hong Kong
Yuxuan Lei 2^{2}2
2^{2}2 Microsoft Research Asia
Jianxun Lian 2^{2}2
2^{2}2 Microsoft Research Asia
Yitian Huang 2^{2}2
2^{2}2 Microsoft Research Asia
Lexin Zhou 3^{3}3
3^{3}3 Princeton University
Haotian Li 2^{2}2
2^{2}2 Microsoft Research Asia
Xing Xie 2^{2}2
2^{2}2 Microsoft Research Asia
Helen Meng 1^{1}1
1^{1}1 The Chinese University of Hong Kong
Correspondence to: Jianxun Lian [email protected]

Abstract

While large language models demonstrate remarkable capabilities across numerous domains, social intelligence—the capacity to perceive social cues, infer mental states, and generate appropriate responses—remains a critical challenge, particularly for enabling effective human-AI collaboration and developing AI that truly serves human needs. Current models often rely on superficial patterns rather than genuine social reasoning. We argue that cultivating human-like social intelligence requires training with challenging cases that resist shortcut solutions. To this end, we introduce ToMBench-Hard, an adversarial benchmark designed to provide hard training examples for social reasoning. Building on this, we propose Social-R1, a reinforcement learning framework that aligns model reasoning with human cognition through multi-dimensional rewards. Unlike outcome-based RL, Social-R1 supervises the entire reasoning process, enforcing structural alignment, logical integrity, and information density. Results show that our approach enables a 4B parameter model to surpass much larger counterparts and generalize robustly across eight diverse benchmarks. These findings demonstrate that challenging training cases with trajectory-level alignment offer a path toward efficient and reliable social intelligence.

1. Introduction

Recent advances in Reinforcement Learning from Verifiable Feedback (RLVF) [1] have significantly improved large language models (LLMs)' performance on formal reasoning tasks like mathematics and programming. However, genuine social intelligence—the capacity to perceive subtle cues, infer latent mental states, and navigate complex interpersonal dynamics—remains a substantial challenge. Unlike formal reasoning with deterministic execution paths, social reasoning is inherently polysemic and context-dependent, lacking the traceable verification that enables effective reward modeling in objective domains.
Despite strong performance on standard benchmarks, current LLMs often rely on shortcut learning rather than authentic social reasoning. Models frequently exhibit a failure mode we term Reasoning Parasitism, characterized by Answer-driven Backfilling—retroactively constructing justifications for predetermined answers rather than deriving inferences through narrative analysis, as illustrated by the example in Figure 1. The inherent fragility of this parasitic performance becomes particularly evident in adversarial or out-of-distribution scenarios. As highlighted by [2, 3], models that excel on standard benchmarks often suffer catastrophic failures when confronted with trivial narrative perturbations. It reveals that current approaches create only a facade of social intelligence while failing to develop robust reasoning capabilities. Our analysis further identifies a critical Interpretation Bottleneck: while models can perceive surface-level social cues, they struggle to map these cues to latent mental states, leading to a "logic reversal" where final answer correctness exceeds the logical integrity of the reasoning process.
**Figure 1:** Social-R1 for Human-like and Efficient Social Reasoning. By integrating SIP-guided rewards into reinforcement learning, Social-R1 mitigates reasoning shortcuts and enforces structured human-like social inference, improving both accuracy and efficiency across model scales. Detailed cases are in Appendix A.

Figure 1: Social-R1 for Human-like and Efficient Social Reasoning. By integrating SIP-guided rewards into reinforcement learning, Social-R1 mitigates reasoning shortcuts and enforces structured human-like social inference, improving both accuracy and efficiency across model scales. Detailed cases are in Appendix A.

We argue that advancing social intelligence requires aligning models' reasoning trajectories with the structured stages of human social inference. Thus, we propose process-based trajectory alignment that cultivates social intelligence as an internalized capability rather than a parasitic performance. Human social reasoning is characterized by high-density information distillation and recursive belief modeling—properties we aim to instill through structured supervision of the reasoning process, as illustrated in the SocialR1-8B example in Figure 1.
In this paper, we introduce Social-R1, a reinforcement learning framework that encourages genuine social reasoning by aligning model trajectories with human cognitive principles: structured, evidence-grounded, and efficient inference. Our approach has two key components. First, we construct ToMBench-Hard, an expert-curated adversarial benchmark designed to expose reasoning shortcuts through socially nuanced perceptual traps. Second, guided by Social Information Processing (SIP) theory ([4]), we develop a multi-dimensional reward system that enforces stage-consistent reasoning progression (RstructR_{\text{struct}}Rstruct​), content integrity (RcontentR_{\text{content}}Rcontent​), and inference efficiency (RlenR_{\text{len}}Rlen​). We demonstrate that this trajectory-level alignment enables more effective and parameter-efficient social intelligence.
We conduct comprehensive experiments across both in-domain and out-of-domain benchmarks related to social intelligence. Results demonstrate that our approach substantially improves model performance on genuine social reasoning tasks. Social-R1 enables smaller models to match or surpass the capabilities of models with significantly larger parameter counts, while showing strong generalization across eight diverse social reasoning evaluations (as illustrated in the right part of Figure 1). These findings indicate that trajectory-level alignment presents a viable and efficient alternative to reliance on pure model scaling for achieving robust social intelligence. The source code and dataset will be released upon paper acceptance. In summary, our contributions are threefold:
  • ToMBench-Hard: A rigorous benchmark with difficulty for social reasoning that exposes shortcut learning in LLMs and mandates genuine cognitive engagement.
  • Social-R1 Framework: A reinforcement learning approach with multi-dimensional rewards that align LLM reasoning trajectories with human social cognition.
  • Performance Superiority: Demonstration that our method enables small models to achieve large-model performance, proving trajectory quality surpasses parameter scaling for social intelligence.

2. Related Work

2.1 Theory-of-Mind Analysis in LLMs

Recent evaluations demonstrate LLMs' growing proficiency on Theory-of-Mind tasks, with models like GPT-4 achieving near-human performance on false-belief tasks and various social cognition assessments [5, 6, 7]. However, broader audits reveal significant limitations in genuine social reasoning capabilities. Studies highlight inconsistent performance across different evaluation frameworks [8], systematic failure modes in comprehensive capability assessments [9], and particular weaknesses in psychologically complex scenarios compared to physical-world contexts [10, 11, 12]. Dynamic evaluation settings further expose models' difficulties with evolving mental states over time [13], while higher-order reasoning tasks reveal sharp performance drops beyond first-level inference [14]. Current improvement attempts remain limited in scope. Prompt-based interventions [11] offer task-specific benefits but lack generalizability, while reinforcement learning approaches [15] show promise but risk overfitting to narrow task patterns without broader training diversity.
From this literature, we identify three key limitations motivating our work: (i) over-reliance on outcome-based evaluation rather than reasoning trajectory analysis; (ii) susceptibility to shallow statistical cues instead of recursive mental-state inference; and (iii) limited generalization across social reasoning benchmarks and related tasks [16, 17]. These gaps necessitate harder adversarial evaluations, trajectory-level supervision, and training regimes designed for transferable social intelligence.

2.2 RL for LLM General Reasoning

Reinforcement learning has emerged as a pivotal technique for enhancing LLM reasoning capabilities, with RLHF establishing the foundation for alignment [18]. Recent large-scale implementations demonstrate that outcome-based rewards can significantly strengthen structured reasoning in mathematical and coding domains [19, 20, 21, 22]. To address sparse reward challenges, process supervision methods provide denser feedback through step-wise trajectory evaluation [23, 24, 25].
Despite these advances, RL remains underexplored for social inference tasks, which require empathy and multi-factor reasoning in unstructured scenarios. Preliminary work [15] shows potential but remains in early development. This highlights two critical gaps: (i) the need for refreshed training materials tailored to social reasoning, and (ii) the opportunity for trajectory-level rewards that guide models toward human-like social inference patterns.

3. The ToMBench-Hard Benchmark

Social intelligence remains a persistent bottleneck in LLMs, with current systems often relying on shortcut learning rather than authentic social reasoning. To address this limitation, we introduce ToMBench-Hard, a diagnostic benchmark specifically designed to expose the shortcut learning behaviors prevalent in current LLMs. The primary goal of ToMBench-Hard is to disentangle genuine social inference from surface-level pattern matching by creating a challenging adversarial environment that mandates structured, human-like reasoning processes.

3.1 Adversarial Data Construction

ToMBench-Hard is grounded in the Abilities in the Theory-of-Mind Space (ATOMS) framework ([26]), providing a comprehensive evaluation covering six core dimensions of social intelligence: Belief, Desire, Emotion, Intention, Knowledge, and Non-literal Communication (NLC). We developed 800 expert-annotated multiple-choice questions through a rigorous annotation protocol involving three independent annotators to ensure cognitive validity and high quality (see Appendix B for details).
To mitigate heuristic-based reasoning strategies such as lexical overlap between questions and options, we introduced ToM-consistent adversarial perturbations ([27, 28]). These perturbations include nuanced manipulations of perceptual access (e.g., unobserved state changes) and asymmetric information (e.g., second-order beliefs), ensuring models cannot succeed through statistical guessing alone. Each sample was carefully designed to require a structured, human-like reasoning process that progresses from cue encoding to mental state interpretation.

3.2 Diagnostic Analysis: Exposing the Shortcut Illusion

To validate the diagnostic utility of ToMBench-Hard, we conducted comprehensive benchmarking across a spectrum of state-of-the-art LLMs, comparing their performance against human expert baselines. The results reveal a significant performance cliff between simple synthetic benchmarks and our adversarial evaluation.
As shown in Table 1, while human experts achieve robust 87% accuracy, frontier models exhibit dramatic performance drops. Notably, models like O3 and Deepseek-R1 achieve near-human performance on simpler benchmarks like ToM-RL ([15]) (reaching 87-88% accuracy), but their accuracy plummets to below 61% on ToMBench-Hard. This performance gap unmasks what we term the shortcut illusion—the phenomenon where high scores on conventional benchmarks reflect template matching artifacts rather than genuine social reasoning capabilities.

Table 1: Results on ToMBench-Hard (ours) and ToM-RL (public).

The substantial performance gap between ToMBench-Hard and ToM-RL suggests that high scores on the latter may not reliably indicate strong theory-of-mind capabilities, as these benchmarks might be insufficiently challenging for LLMs. By systematically exposing these limitations, ToMBench-Hard serves dual purposes: it provides a rigorous evaluation framework for assessing genuine social reasoning capabilities, while also offering a high-quality dataset for aligning LLMs with human-like cognitive processes. To facilitate both training and evaluation, we partition the complete ToMBench-Hard dataset into non-overlapping training and test subsets, enabling the development of our alignment framework in the subsequent sections.

4. The Social-R1 Framework

The Social-R1 framework introduces a novel paradigm for cultivating genuine social intelligence in LLMs by aligning their reasoning processes with human social cognition. Departing from traditional outcome-based reinforcement learning approaches, our method specifically tackles the fundamental challenge of Reasoning Parasitism—a phenomenon where models engage in Answer-driven Backfilling by constructing post-hoc justifications for predetermined answers. Social-R1 counteracts these shortcut behaviors through a multi-dimensional reward system that supervises the entire reasoning trajectory rather than merely rewarding final outcomes. This comprehensive approach transforms social intelligence from a parasitic performance into an internalized capability, ensuring model reasoning embodies the core characteristics of human social cognition: precise, stage-consistent social inference ([4]) and high information density ([29]).

4.1 Multi-Dimensional Reward Design

The core innovation of Social-R1 lies in its comprehensive reward system that aligns model reasoning trajectories with three key characteristics of human social cognition: structured progression, logical integrity, and information density. This multi-dimensional approach provides fine-grained supervision over the reasoning process, addressing different aspects of high-quality social inference.
SIP Structural Alignment (RstructR_{{struct}}Rstruct​). Human social reasoning follows a disciplined progression from perception to response generation ([30]). To instill this cognitive scaffold, we introduce a structural reward RstructR_{{struct}}Rstruct​ that enforces sequential reasoning across the four stages of Social Information Processing (SIP): (1) Encoding Social Cues: Identifying relevant social signals from the narrative, (2) Interpreting Cues: Inferring latent mental states from perceived cues, (3) Clarifying Goals: Determining social objectives and interpersonal intentions, and (4) Response Generation: Selecting appropriate behavioral responses. The RstructR_{{struct}}Rstruct​ reward validates whether intermediate reasoning steps adhere to this stage-wise progression, penalizing premature conclusions and stage skipping. This structural constraint encourages coherent, story-grounded causal inference while mitigating shortcut behaviors like option parasitism, where models anchor reasoning directly on multiple-choice options rather than narrative evidence. We leverage GPT-4o as judge for this reward, with the prompt illustrated in Appendix C.1.
SIP Content Integrity (RcontentR_{{content}}Rcontent​). While structural alignment ensures proper reasoning staging, content integrity guarantees logical rigor within each stage. The content reward RcontentR_{{content}}Rcontent​ audits whether intermediate inferences remain grounded in story-internal evidence, correctly reflecting social cues, intentions, and goals. This reward penalizes three critical failure modes: (1) Erroneous cue encoding: Misidentification of relevant social signals, (2) Flawed interpretation: Incorrect mental state attribution, and (3) Misidentified goals: Inaccurate inference of social objectives. By ensuring each reasoning step maintains evidential support, RcontentR_{{content}}Rcontent​ discourages superficial rationalizations and promotes authentic social sensing derived from narrative context. The detailed implementation of this reward model can be found at Section 5.1.
Inference Efficiency Optimization (RlenR_{{len}}Rlen​). Human social reasoning achieves high information density through selective attention and avoidance of redundant processing ([31, 32]). To emulate this cognitive efficiency, we design RlenR_{{len}}Rlen​ as the product of two complementary components:
Rlen=Rrep(ρ)⋅Rwin(L)R_{{len}} = R_{{rep}}(\rho) \cdot R_{{win}}(L)
The repetition penalty component Rrep(ρ)R_{{rep}}(\rho)Rrep​(ρ) specifically targets circular over-thinking by penalizing excessive n-gram repetition beyond a threshold τ=0.1\tau = 0.1τ=0.1:
Rrep(ρ)={1,ρ≤τexp⁡ ⁣(−β(ρ−τ)),ρ>τR_{{rep}}(\rho) = \begin{cases} 1, & \rho \le \tau \\ \exp\!\left(-\beta (\rho - \tau)\right), & \rho > \tau \end{cases}
where β=8\beta = 8β=8 controls the severity of penalty decay for high repetition ratios. The length window constraint Rwin(L)R_{{win}}(L)Rwin​(L) maintains reasoning trajectories within an empirically optimal range [Lmin⁡,Lmax⁡][L_{\min}, L_{\max}][Lmin​,Lmax​] through smooth gating:
Rwin(L)=σ(L−Lmin⁡k)⋅σ(Lmax⁡−Lk)R_{{win}}(L) = \sigma\left(\frac{L - L_{\min}}{k}\right) \cdot \sigma\left(\frac{L_{\max} - L}{k}\right)
with k=50k = 50k=50 controlling transition smoothness. The bounds Lmin⁡=400L_{\min} = 400Lmin​=400 and Lmax⁡=2500L_{\max} = 2500Lmax​=2500 are derived from strong chain-of-thought baselines, ensuring reasoning remains concise yet comprehensive. This dual-mechanism approach encourages the model to emulate human-like efficiency by avoiding both redundant repetition and excessive verbosity while maintaining substantive social inference.
Verifiable Format Alignment (RfmtR_{{fmt}}Rfmt​) We adopt the format reward from [33] that enforces structured thinking processes. The model is rewarded for producing outputs with predefined XML-style tags (<thinking> and <answer>), which enables deterministic extraction of both reasoning trajectories and final answers while preserving semantic freedom in the reasoning content.

4.2 Reward Synthesis and Learning

The composite reward function integrates all components through a carefully designed synthesis strategy that balances outcome supervision with process-level reasoning signals:
Rtotal=Rfmt⋅(woRout+τ(wstructRstruct+wcontentRcontent))⋅Rlen\begin{split}R_{\text{total}} = R_{\text{fmt}} &\cdot \Bigl( w_o R_{\text{out}} + \tau \bigl( w_{\text{struct}} R_{\text{struct}}\\&+ w_{\text{content}} R_{\text{content}} \bigr) \Bigr) \cdot R_{\text{len}}\end{split}
where RoutR_{{out}}Rout​ denotes the verifiable outcome reward, We implement a curriculum learning strategy where outcome supervision dominates early training phases (wo(t)=2w_o(t) = 2wo​(t)=2), while process-level rewards are progressively emphasized through time-dependent weighting: wstruct(t)=wcontent(t)=1+γtTw_{{struct}}(t) = w_{{content}}(t) = 1 + \gamma\frac{t}{T}wstruct​(t)=wcontent​(t)=1+γTt​. This curriculum ensures stable initial convergence while gradually reinforcing human-like reasoning patterns as training progresses. The optimization employs Group Relative Policy Optimization ([33]), which performs group-relative updates over sampled reasoning trajectories.

5. Experiment

5.1 Experiment Setting

We evaluate our model on eight multiple-choice social benchmarks, including two in-domain benchmarks: the public ToMBench ([9]) and our ToMBench-Hard test set, and six out-of-domain benchmarks: SocialIQA ([34]) for social commonsense reasoning, EmoBench ([16]) for emotion intelligence, MotiveBench ([17]) for social motivation reasoning, SimpleToM ([11]) for examining whether models can consciously infer others’ mental states (MS) and proactively applying such reasoning to behavior inference, Hi-ToM ([35]) for evaluating higher-order Theory-of-Mind reasoning, and TactfulToM ([36]) for testing whether models can interpret white lies and infer the underlying prosocial intent to preserve interpersonal harmony.
SIP Content Integrity Reward Model (RMcontentRM_{content}RMcontent​). To instantiate RcontentR_{content}Rcontent​, we train a dedicated Content Reward Model that assigns a scalar quality score to intermediate SIP-stage reasoning segments. We construct the SocialPairs-20K dataset by sampling NNN trajectories per training instance from various Social-R1 checkpoints. These checkpoints span different training stages, ensuring a diverse pool of reasoning qualities ranging from nascent logic to sophisticated social inference. For each segment, a strong teacher-judge (o3) generates silver-standard stage-wise rationales by conditioning on the gold final answer and the original social context. We then employ a multi-dimensional rubric—focusing on fact-grounding, mental-state attribution accuracy, and stage-specific relevance—to score the sampled segments against the teacher-generated references. This process forms preference pairs (chosen vs. rejected), while rejected segments exhibit missing cues, incorrect mental-state attributions, or wrong goal identification. We validate the model on two held-out sets: (1) an automatic test split of 2k pairs, where RcontentR_{content}Rcontent​ achieves an accuracy of 89.2%; and (2) a human-calibrated subset of 200 pairs annotated by experts, achieving an 87.5% agreement with human labels. (More RcontentR_{content}Rcontent​ can be found in Appendix C.2).
Implementation Details SIP Content Integrity Reward Model (RMcontentRM_{content}RMcontent​) is trained as a pairwise preference reward model. We initialize it from Qwen3-4B and fine-tune with LoRA on the SocialPairs-20K dataset mentioned previously. For policy optimisation, we train two versions of reasoning models from Qwen3-4B and Qwen3-8B on ToMBench-Hard, consisting of 700 training instances and 100 test instances. Reinforcement learning is performed for 600 optimisation steps using VERL{\mathchoice{\text{V{\scriptsize ERL}}}{\text{V{\scriptsize ERL}}}{\text{V{\scriptscriptstyle ERL}}}{\text{VERL}}}VERL ([37]) on 8 NVIDIA A100 (80GB) GPUs. We set the group size to 5, the KL coefficient to 0.04, and the learning rate to 5×10−75 \times 10^{-7}5×10−7. More details are provided in Appendix D.

5.2 Main Results

We apply the Social-R1 framework to two open-source backbones of different scales, Qwen3-4B and Qwen3-8B. The overall results across eight social reasoning benchmarks are reported in Table 2. Social-R1 achieves strong and consistent improvements over the corresponding Qwen baselines, demonstrating that reinforcement learning with challenging social supervision and trajectory-level reward signals can substantially enhance Theory-of-Mind reasoning. Notably, Social-R1-4B surpasses LLaMa3.1-70B across all benchmarks, despite being more than an order of magnitude smaller, highlighting the effectiveness of process-based social alignment beyond parameter scaling. Even more strikingly, Social-R1-8B outperforms DeepSeek-R1 on several benchmarks and achieves stronger overall performance, consistently matching or exceeding much larger baselines such as Qwen3-32B in out-of-domain generalization. More detailed results are provided in Appendix E.

Table 2: In-domain and out-of-domain performance across eight social reasoning benchmarks. Green denotes the best result among our Social-R1 reward variants (ablations), and bold indicates the overall best score.

5.3 Ablation Studies

To systematically evaluate the contribution of each reward component in Social-R1, we conduct comprehensive ablation experiments on both SocialR1-4B and SocialR1-8B variants. Specifically, we remove three key components: the length-control reward (w/o Rlenw/o\ R_{\text{len}}w/o Rlen​), the structural trajectory reward (w/o Rstructw/o\ R_{\text{struct}}w/o Rstruct​), and the content integrity reward (w/o Rcontw/o\ R_{\text{cont}}w/o Rcont​). Additionally, we examine a baseline variant that removes all the progress rewards (only Routonly\ R_{\text{out}}only Rout​). The complete results are presented in Table 2.
Benchmarks The ablation analysis reveals distinct performance patterns associated with each reward component. Removing w/o Rlenw/o\ R_{\text{len}}w/o Rlen​ results in significant performance degradation on Hi-ToM (e.g., 0.7083→0.62670.7083 \rightarrow 0.62670.7083→0.6267 for SocialR1-8B), indicating that uncontrolled reasoning length negatively impacts high-order social inference. The w/o RlenR_{\text{len}}Rlen​ variant exhibits a substantial increase in reasoning verbosity, with the average thinking length rising by approximately +250% compared to SocialR1-8B as shown in Figure 10.
The structural reward RstructR_{\text{struct}}Rstruct​ shows consistent importance across benchmarks, with its removal causing accuracy drops (e.g., from 0.50790.50790.5079 to 0.45580.45580.4558 on TactfulToM). Similarly, RcontR_{\text{cont}}Rcont​ removal leads to performance declines, suggesting its role in maintaining reasoning quality. Notably, the only Routonly\ R_{\text{out}}only Rout​ variant exhibits more severe overall performance deterioration compared to other ablations, demonstrating the importance of incorporating process-level rewards. While these numerical changes demonstrate the overall effectiveness of each component, subsequent in-depth analyses (Section 6.1–Section 6.3) will provide mechanistic evidence through detailed trajectory examinations and perturbation studies to understand how each reward shapes reasoning behavior.

6. In-Depth Analysis

To examine whether Social-R1 yields genuinely internalised social reasoning—rather than Reasoning Parasitism driven by answer-first backfilling—we conduct a mechanistic analysis of its reasoning dynamics. We compare Social-R1-8B against three strong baselines (DeepSeek-R1, DeepSeek-R1-Distill-Llama-70B, and Qwen3-8B), together with controlled reward variants that isolate outcome-only training and ablate key trajectory-level constraints. In line with our reward design, we operationalise human-like social reasoning through three diagnostic signatures: (i) option-independent inference grounded in narrative cues; (ii) stage-consistent SIP trajectories with content integrity, where intermediate beliefs remain logically valid across Encoding, Interpretation, Goal Clarification, and Response Generation; and (iii) selective robustness, where models avoid redundant trajectory bloat under perturbation. Accordingly, our analysis proceeds in three steps: (1) quantifying Reasoning Parasitism as a diagnostic of shortcut reliance; (2) auditing stage-wise cognitive fidelity to identify interpretation bottlenecks; and (3) introducing controlled distractors to assess robustness and conciseness.

6.1 From Parasitism to Independent Inference

**Figure 2:** Option-Mention Density across SIP reasoning stages.

Figure 2: Option-Mention Density across SIP reasoning stages.

gt;80\%$), performance drops sharply—by nearly 25 points—once entering Cue Interpretation, which requires latent mental-state attribution and recursive belief reasoning. More critically, outcome-only training induces a *reasoning reversal* phenomenon. For example, Qwen3-8B trained with $R_{\text{out}}$ often achieves higher final answer accuracy than the correctness of its preceding SIP stages, indicating that surface-level option correlations can yield correct predictions despite incoherent intermediate social inference. This highlights the limitation of outcome supervision: it rewards answer selection without enforcing psychologically valid reasoning trajectories. Crucially, the proposed Content Reward $R_{\text{content}}$ directly targets this failure mode by penalising erroneous cue encoding, flawed mental-state interpretation, and misidentified social goals. Ablation results provide clear causal evidence: removing $R_{\text{content}}$ reduces Interpretation accuracy by 6.2 points (77.5%$\rightarrow$71.3%) and Goal Classification by 6.2 points (75.0%$\rightarrow$68.8%). Ultimately, Social-R1-8B preserves the most stage-consistent and smoothly degrading SIP trajectory, demonstrating that its gains stem from cognitively grounded reasoning rather than shortcut-driven recovery. " data-original-markdown=" A primary failure mode in multiple-choice social reasoning is *Reasoning Parasitism*—a shortcut behaviour where models anchor their deductions on statistical regularities in the answer choices rather than deriving conclusions from story-internal social cues. Figure 2 reveals substantial divergence in reasoning dynamics. DeepSeek-R1, Qwen3-8B, and the outcome-only variant of SocialR1-8B exhibit high option-mention density as early as the Cue Encoding stage (Q1 group in Figure 2), often exceeding five explicit option references per sample. Such premature reliance suggests *answer-conditioned backfilling*: models use the provided choices to retroactively assemble plausible justifications, rather than performing independent social inference. In contrast, Social-R1-8B maintains a largely option-agnostic trajectory. Its option mentions remain minimal and nearly flat throughout cue perception and interpretation, increasing only slightly (to $\sim$1.3) during final Response Generation (Q4). This provides mechanistic evidence that the proposed Social Think Combined Reward effectively suppresses shortcut dependence on answer options, compelling the model to engage in narrative-grounded social deduction. Additional qualitative examples are provided in Appendix F. ### 6.2 Stage-wise Diagnosis To pinpoint where social reasoning most systematically breaks down, we conduct a human audit of 80 instances sampled from eight benchmarks, ensuring high annotation quality and genuine Theory-of-Mind demands. We track success rates across the four Social Information Processing (SIP) stages. Figure 3 exposes a pronounced *Interpretation Bottleneck*. While strong baselines sustain high accuracy in factual cue encoding (
gt;80\%$), performance drops sharply—by nearly 25 points—once entering Cue Interpretation, which requires latent mental-state attribution and recursive belief reasoning. More critically, outcome-only training induces a *reasoning reversal* phenomenon. For example, Qwen3-8B trained with $R_{\text{out}}$ often achieves higher final answer accuracy than the correctness of its preceding SIP stages, indicating that surface-level option correlations can yield correct predictions despite incoherent intermediate social inference. This highlights the limitation of outcome supervision: it rewards answer selection without enforcing psychologically valid reasoning trajectories. Crucially, the proposed Content Reward $R_{\text{content}}$ directly targets this failure mode by penalising erroneous cue encoding, flawed mental-state interpretation, and misidentified social goals. Ablation results provide clear causal evidence: removing $R_{\text{content}}$ reduces Interpretation accuracy by 6.2 points (77.5%$\rightarrow$71.3%) and Goal Classification by 6.2 points (75.0%$\rightarrow$68.8%). Ultimately, Social-R1-8B preserves the most stage-consistent and smoothly degrading SIP trajectory, demonstrating that its gains stem from cognitively grounded reasoning rather than shortcut-driven recovery. " data-source-offset="26737" class="markdown-segment">
A primary failure mode in multiple-choice social reasoning is Reasoning Parasitism—a shortcut behaviour where models anchor their deductions on statistical regularities in the answer choices rather than deriving conclusions from story-internal social cues. Figure 2 reveals substantial divergence in reasoning dynamics. DeepSeek-R1, Qwen3-8B, and the outcome-only variant of SocialR1-8B exhibit high option-mention density as early as the Cue Encoding stage (Q1 group in Figure 2), often exceeding five explicit option references per sample. Such premature reliance suggests answer-conditioned backfilling: models use the provided choices to retroactively assemble plausible justifications, rather than performing independent social inference. In contrast, Social-R1-8B maintains a largely option-agnostic trajectory. Its option mentions remain minimal and nearly flat throughout cue perception and interpretation, increasing only slightly (to ∼\sim∼1.3) during final Response Generation (Q4). This provides mechanistic evidence that the proposed Social Think Combined Reward effectively suppresses shortcut dependence on answer options, compelling the model to engage in narrative-grounded social deduction. Additional qualitative examples are provided in Appendix F.

6.2 Stage-wise Diagnosis

To pinpoint where social reasoning most systematically breaks down, we conduct a human audit of 80 instances sampled from eight benchmarks, ensuring high annotation quality and genuine Theory-of-Mind demands. We track success rates across the four Social Information Processing (SIP) stages. Figure 3 exposes a pronounced Interpretation Bottleneck. While strong baselines sustain high accuracy in factual cue encoding (>80%>80\%>80%), performance drops sharply—by nearly 25 points—once entering Cue Interpretation, which requires latent mental-state attribution and recursive belief reasoning. More critically, outcome-only training induces a reasoning reversal phenomenon. For example, Qwen3-8B trained with RoutR_{\text{out}}Rout​ often achieves higher final answer accuracy than the correctness of its preceding SIP stages, indicating that surface-level option correlations can yield correct predictions despite incoherent intermediate social inference. This highlights the limitation of outcome supervision: it rewards answer selection without enforcing psychologically valid reasoning trajectories. Crucially, the proposed Content Reward RcontentR_{\text{content}}Rcontent​ directly targets this failure mode by penalising erroneous cue encoding, flawed mental-state interpretation, and misidentified social goals. Ablation results provide clear causal evidence: removing RcontentR_{\text{content}}Rcontent​ reduces Interpretation accuracy by 6.2 points (77.5%→\rightarrow→71.3%) and Goal Classification by 6.2 points (75.0%→\rightarrow→68.8%). Ultimately, Social-R1-8B preserves the most stage-consistent and smoothly degrading SIP trajectory, demonstrating that its gains stem from cognitively grounded reasoning rather than shortcut-driven recovery.
**Figure 3:** Stage-wise SIP accuracy across models.

Figure 3: Stage-wise SIP accuracy across models.

Case Study To further illustrate this mechanistic distinction, Figure 4 presents a representative example from ToMBench, where the correct answer critically depends on grounding interpretation in the protagonist’s epistemic access. In this story, Grimmo inhabits an underground world without sky, celestial bodies, or human presence, and therefore cannot plausibly imitate planets, clouds, or dancers. Social-R1-8B correctly encodes this constraint and performs a psychologically valid interpretation, inferring that the imitation must stem from locally observable phenomena (e.g., bioluminescent fungi spinning in the dark), yielding a coherent SIP trajectory. In contrast, reward-ablated variants exhibit the identified failure modes. Social-R1-8B w/o RcontR_{\text{cont}}Rcont​ drifts into ungrounded interpretation, despite explicit narrative exclusion—a breakdown at the Interpretation and Goal stages. Likewise, the model trained with only RoutR_{\text{out}}Rout​ collapses into option-level lexical shortcutting (e.g., "spinning dancer") rather than maintaining story-consistent mental-state reasoning. Qwen3-8B shows a similar reversal behaviour, selecting an answer that reflects generic associations instead of the agent’s accessible knowledge. Overall, this case provides concrete mechanistic evidence that Social-R1’s gains do not arise from answer-first recovery, but from enforcing evidence-grounded interpretation and stage-consistent social inference.
**Figure 4:** Case study highlighting the Interpretation Bottleneck. Detailed cases are provided in Appendix F.

Figure 4: Case study highlighting the Interpretation Bottleneck. Detailed cases are provided in Appendix F.

6.3 Robustness under Perturbation

To assess whether Social-R1 promotes cognitively disciplined robustness rather than brittle performance sustained by overextended reasoning, we conduct a controlled perturbation study. Specifically, we inject story-consistent but decision-irrelevant distractor cues into 40 instances where both SocialR1-8B and DeepSeek-R1 initially succeed. These distractors introduce no new evidence that alters the correct option, enabling a direct comparison of reasoning trajectory stability (Appendix G). Figure 5 illustrates the resulting accuracy–efficiency trade-off. Although SocialR1-8B and DeepSeek-R1 retain comparable post-perturbation accuracy, DeepSeek-R1 does so only by producing substantially longer reasoning trajectories, suggesting reliance on diffuse and overextended inference. In contrast, SocialR1-8B exhibits only mild token drift, indicating more selective attention and cognitively efficient deduction. Reward ablations further provide mechanistic support. Removing RstructR_{\text{struct}}Rstruct​ disrupts stage-wise progression, removing RcontentR_{\text{content}}Rcontent​ undermines evidence-grounded interpretation, and training with only RoutR_{\text{out}}Rout​ yields the most severe robustness collapse. Together, these findings demonstrate that Social-R1 achieves robustness not through increased verbosity, but through enforcing structured, grounded, and concise social reasoning, aligning model inference with human-like selective cognition rather than scale-driven overthinking.
**Figure 5:** Robustness study under story-consistent distractors.

Figure 5: Robustness study under story-consistent distractors.

7. Conclusion

In this work, we introduce ToMBench-Hard, a challenging benchmark that rigorously evaluates the Theory of Mind capabilities in LLMs. Building on this, we propose Social-R1, a reinforcement learning framework that integrates both outcome-level and thinking-level rewards to cultivate human-like social intelligence in LLMs. Our results demonstrate that outcome-based reinforcement learning over ToMBench-Hard already enhances social reasoning, while thinking-level supervision yields further improvements. These findings highlight the importance of supervising not only what a model concludes but also how it reasons, paving the way toward socially intelligent LLMs. Future work may extend this framework to broader domains of social tasks, such as human-AI collaboration and LLM-based simulations for social sciencec.

8. Impact Statement

This research introduces Social-R1, a framework for enhancing social reasoning in LLMs, which could enable more natural human-AI collaboration in applications like education, healthcare, and assistive technologies. However, improved social intelligence also raises ethical concerns, such as the potential for misuse in manipulative systems or the amplification of social biases if not properly aligned with human values. We encourage rigorous oversight and fairness audits to mitigate these risks while leveraging the benefits of robust AI social cognition.

Appendix

A. The Detailed Case in Figure 1

**Figure 6:** The Detailed Case in Figure 1

Figure 6: The Detailed Case in Figure 1

B. ToMBench_Hard

ToMBench_Hard is deliberately curated to increase task difficulty by introducing nuanced distractors and context-dependent reasoning. Inspired by the Abilities in the Theory-of-Mind Space (ATOMS) framework ([26]), each question is designed to probe a distinct aspect of ToM reasoning, detailed definition of each subabilities and dimension can be found in ([26]) . To further increase difficulty ([27, 28]), adversarial variations such as asymmetric access to information, discrepant intentions, and subtle social cues are included.

B.1 Human Annotation

ToMBench_Hard is developed jointly by the author and one psychology graduate student, who construct the scenarios, questions, options and answers. Annotation is carried out by five computer science graduate students (after receiving training) and five social psychology graduate students. Each sample is independently answered by two annotators, and disagreements are discussed and resolved through group review and iterative modification. This procedure ensures both linguistic clarity and psychological validity. The annotation process emphasized consistency across dimensions and aimed to capture nuanced aspects of social reasoning.

B.2 Data Statistic

Table 3: Distribution of fine-grained ToMBench_Hard sub-abilities.

Ability (Total) Sub-ability Count
Intention
(243)
Prediction of actions 111
Intentions explanations 102
Completion of failed actions 18
Discrepant intentions 12
Belief
(186)
Second-order beliefs 67
Location false beliefs 49
Beliefs based action/emotions 39
Identity false beliefs 12
Content false beliefs 11
Sequence false beliefs 8
Emotion
(143)
Typical emotional reactions 59
Atypical emotional reactions 26
Mixed emotions 26
Emotion regulation 12
Hidden emotions 12
Moral emotions 8
Knowledge
(96)
Information-knowledge links 53
Knowledge-pretend play links 21
Knowledge-attention links 12
Percepts-knowledge 10
Desire
(82)
Desires influence on emotions and actions 32
Discrepant desires 18
Desire-action contradiction 14
Multiple desires 9
Desires influence on actions 6
Desires influence on emotions (beliefs) 3
Non-Literal
Communication
(50)
Involuntary lies 10
Faux pas 10
Egocentric lies 8
Humor 8
Irony/Sarcasm 8
White lies 8

B.3 Cases in ToMBench_Hard

To provide an intuitive understanding of the challenges posed by ToMBench-Hard, Figure 7 presents representative examples spanning diverse Theory-of-Mind abilities, including belief tracking, discrepant desires, intention explanation, atypical emotional reactions, knowledge inference, and non-literal humour comprehension. Each instance follows a unified Ability–Story–Question–Answer structure, consisting of a socially grounded narrative context, a multiple-choice question, and the annotated gold option. These cases highlight the benchmark’s structural diversity and nuanced social cues, requiring models to move beyond superficial option matching and motivating our trajectory-level alignment rewards.
**Figure 7:** **Example cases from ToMBench-Hard.**

Figure 7: Example cases from ToMBench-Hard.

B.4 Performance on ToMBench_Hard

Table 4: Performance on ToMBench_Hard

C. Reward Model

C.1 Structure Reward

We provide the detailed GPT-4o judging prompt in Figure 8, which is used to evaluate whether a model’s reasoning trajectory explicitly adheres to the four-stage Social Information Processing (SIP) framework, including Cue Encoding, Cue Interpretation, Goal Clarification, and Response Generation.
**Figure 8:** Prompt template for structural reward evaluation.

Figure 8: Prompt template for structural reward evaluation.

C.2 Content Reward

Table 5: content sample rubric

gt;$ Tier C (strongest supervision against hard negatives). **P1**: Tier A
gt;$ Tier C. **P2**: Tier A
gt;$ Tier B. **P3**: Tier B
gt;$ Tier D. **P4**: late-stage short trajectories
gt;$ early-stage long trajectories, encouraging concise and cognitively disciplined inference. Applying this policy produces **SocialPairs-20K**, consisting of 20K preference pairs for training $RM_{\text{content}}$. **Held-out evaluation.** We construct a held-out test set of 2K preference pairs from ToMBench-Hard test instances using the same pipeline, achieving 89.2% pairwise accuracy. Additionally, we randomly sample 200 pairs for expert human verification, obtaining 87.5% agreement with human preferences. The detailed multi-dimensional scoring rubric used by GPT-5 for stage-wise segment evaluation is provided in Table 5. ### D. Implementation The content reward model (TRM) is initialized from Qwen3-4B-Instruct-2507and trained using supervised fine-tuning for two epochs on four NVIDIA A100 80GB GPUs. The model is trained with pairwise comparisons and, during inference, predicts a scalar reward conditioned on (*scenario, question, options, reasoning*). We evaluate multiple LLM families, including DeepSeek-R1([20]), DeepSeek-R1-Distill-Llama-70B([20])<sup>[^1]</sup>, Llama-3.1-70B-Instruct([38])<sup>[^2]</sup>, Qwen3-4B/8B/32B([22])<sup>[^3]</sup>, GPT-5-2025-08-07([18]), GPT-4o-2024-08-06 ([39]), and o3-2025-04-16. We prompt Qwen3-4B/8B/32B in both Thinking and No-Think setting. " data-original-markdown=" **Teacher references (gold SIP rationales).** We first provide the teacher model (o3) with the gold final answer and the original narrative context, and obtain stage-consistent SIP rationales as silver references for each intermediate SIP stage. **Trajectory sampling from diverse checkpoints.** To ensure a wide spectrum of reasoning quality, we collect candidate trajectories from 10 checkpoints of the *w/o $R_{\text{struct}}$* run at steps $\{30, 90, 120, 180, 270, 360, 420, 510, 570, 600\}$. We sample $K{=}6$ trajectories per instance per checkpoint. Given $N{=}700$ training instances, this yields: 42000 candidate SIP-stage reasoning segments, spanning from early unstable logic to mature social inference. **LLM-based scoring with teacher calibration.** Each sampled segment is scored by GPT-5 using the o3 silver rationale as the reference. The evaluation rubric focuses on: (i) fact-grounding to story-internal evidence, (ii) correctness of mental-state attribution, and (iii) stage-specific relevance within the SIP trajectory. **Tiering and hard-negative identification.** We categorise all segments into five quality tiers: **Tier S (Expert)**: o3 gold-quality samples. **Tier A (Quality Positive)**: $\mathrm{ACC}{=}1$ and $\mathrm{LLM\_Score} \ge 0.8$. **Tier B (Weak Positive)**: $\mathrm{ACC}{=}1$ and $\ge 0.6 \mathrm{LLM\_Score} \le 0.8$. **Tier C (Hard Negative)**: $\mathrm{ACC}{=}1$ while $\mathrm{LLM\_Score} \le 0.6$ . **Tier D (Common Negative)**: $\mathrm{ACC}{=}0$ and $\mathrm{LLM\_Score}$ is very low. We construct preference pairs (*chosen* vs. *rejected*) using the following priority scheme: **P0**: Tier S
gt;$ Tier C (strongest supervision against hard negatives). **P1**: Tier A
gt;$ Tier C. **P2**: Tier A
gt;$ Tier B. **P3**: Tier B
gt;$ Tier D. **P4**: late-stage short trajectories
gt;$ early-stage long trajectories, encouraging concise and cognitively disciplined inference. Applying this policy produces **SocialPairs-20K**, consisting of 20K preference pairs for training $RM_{\text{content}}$. **Held-out evaluation.** We construct a held-out test set of 2K preference pairs from ToMBench-Hard test instances using the same pipeline, achieving 89.2% pairwise accuracy. Additionally, we randomly sample 200 pairs for expert human verification, obtaining 87.5% agreement with human preferences. The detailed multi-dimensional scoring rubric used by GPT-5 for stage-wise segment evaluation is provided in Table 5. ### D. Implementation The content reward model (TRM) is initialized from Qwen3-4B-Instruct-2507and trained using supervised fine-tuning for two epochs on four NVIDIA A100 80GB GPUs. The model is trained with pairwise comparisons and, during inference, predicts a scalar reward conditioned on (*scenario, question, options, reasoning*). We evaluate multiple LLM families, including DeepSeek-R1([20]), DeepSeek-R1-Distill-Llama-70B([20])[^1], Llama-3.1-70B-Instruct([38])[^2], Qwen3-4B/8B/32B([22])[^3], GPT-5-2025-08-07([18]), GPT-4o-2024-08-06 ([39]), and o3-2025-04-16. We prompt Qwen3-4B/8B/32B in both Thinking and No-Think setting. " data-source-offset="39847" class="markdown-segment">
Teacher references (gold SIP rationales). We first provide the teacher model (o3) with the gold final answer and the original narrative context, and obtain stage-consistent SIP rationales as silver references for each intermediate SIP stage.
Trajectory sampling from diverse checkpoints. To ensure a wide spectrum of reasoning quality, we collect candidate trajectories from 10 checkpoints of the w/o RstructR_{\text{struct}}Rstruct​ run at steps {30,90,120,180,270,360,420,510,570,600}\{30, 90, 120, 180, 270, 360, 420, 510, 570, 600\}{30,90,120,180,270,360,420,510,570,600}. We sample K=6K{=}6K=6 trajectories per instance per checkpoint. Given N=700N{=}700N=700 training instances, this yields: 42000 candidate SIP-stage reasoning segments, spanning from early unstable logic to mature social inference.
LLM-based scoring with teacher calibration. Each sampled segment is scored by GPT-5 using the o3 silver rationale as the reference. The evaluation rubric focuses on: (i) fact-grounding to story-internal evidence, (ii) correctness of mental-state attribution, and (iii) stage-specific relevance within the SIP trajectory.
Tiering and hard-negative identification. We categorise all segments into five quality tiers: Tier S (Expert): o3 gold-quality samples. Tier A (Quality Positive): ACC=1\mathrm{ACC}{=}1ACC=1 and LLM_Score≥0.8\mathrm{LLM\_Score} \ge 0.8LLM_Score≥0.8. Tier B (Weak Positive): ACC=1\mathrm{ACC}{=}1ACC=1 and ≥0.6LLM_Score≤0.8\ge 0.6 \mathrm{LLM\_Score} \le 0.8≥0.6LLM_Score≤0.8. Tier C (Hard Negative): ACC=1\mathrm{ACC}{=}1ACC=1 while LLM_Score≤0.6\mathrm{LLM\_Score} \le 0.6LLM_Score≤0.6 . Tier D (Common Negative): ACC=0\mathrm{ACC}{=}0ACC=0 and LLM_Score\mathrm{LLM\_Score}LLM_Score is very low.
We construct preference pairs (chosen vs. rejected) using the following priority scheme: P0: Tier S >>> Tier C (strongest supervision against hard negatives). P1: Tier A >>> Tier C. P2: Tier A >>> Tier B. P3: Tier B >>> Tier D. P4: late-stage short trajectories >>> early-stage long trajectories, encouraging concise and cognitively disciplined inference.
Applying this policy produces SocialPairs-20K, consisting of 20K preference pairs for training RMcontentRM_{\text{content}}RMcontent​.
Held-out evaluation. We construct a held-out test set of 2K preference pairs from ToMBench-Hard test instances using the same pipeline, achieving 89.2% pairwise accuracy. Additionally, we randomly sample 200 pairs for expert human verification, obtaining 87.5% agreement with human preferences.
The detailed multi-dimensional scoring rubric used by GPT-5 for stage-wise segment evaluation is provided in Table 5.

D. Implementation

The content reward model (TRM) is initialized from Qwen3-4B-Instruct-2507and trained using supervised fine-tuning for two epochs on four NVIDIA A100 80GB GPUs. The model is trained with pairwise comparisons and, during inference, predicts a scalar reward conditioned on (scenario, question, options, reasoning).
We evaluate multiple LLM families, including DeepSeek-R1([20]), DeepSeek-R1-Distill-Llama-70B([20])1, Llama-3.1-70B-Instruct([38])2, Qwen3-4B/8B/32B([22])3, GPT-5-2025-08-07([18]), GPT-4o-2024-08-06 ([39]), and o3-2025-04-16. We prompt Qwen3-4B/8B/32B in both Thinking and No-Think setting.
For completeness, we provide additional training dynamics in Appendix, including the evolution of training accuracy (Figure 9), trajectory length (Figure 10), content integrity score (Figure 12), and structural alignment score during optimisation.
**Figure 9:** Training accuracy during training.

Figure 9: Training accuracy during training.

**Figure 10:** Trajectory length during training

Figure 10: Trajectory length during training

**Figure 11:** Content Score During Training

Figure 11: Content Score During Training

**Figure 12:** Structure Score During Training

Figure 12: Structure Score During Training

E. Detailed Results

For completeness, we report the full detailed results across all evaluated benchmarks in Table 6 in Appendix–Table 10, including MotivationBench (Table 6), ToMBench-Hard (Table 11), SimpleToM (Table 7), TactfulToM (Table 9), and EmoBench (Table 10).

Table 6: Performance on MotiveBench

Model Amazon Blog Persona Overall
Closed-sourced LLMs
DeepSeek-R1 0.9000 0.8333 0.8633 0.8655
o3 0.9800 0.9067 0.9333 0.9400
o3_COT 0.9667 0.8933 0.9000 0.9200
GPT-5 0.9200 0.9133 0.8933 0.9089
GPT-5_COT 0.9600 0.9267 0.8933 0.9267
GPT-4o 0.9733 0.9033 0.9133 0.9300
GPT-4o_COT 0.9400 0.8867 0.8667 0.8978
Open-sourced LLMs
Qwen3-4B (Disable) 0.8067 0.7867 0.7267 0.7734
Qwen3-4B 0.9133 0.7933 0.8267 0.8444
Qwen3-8B (Disable) 0.8200 0.8000 0.7600 0.7933
Qwen3-8B 0.7333 0.6600 0.6700 0.6878
Qwen3-32B (Disable) 0.9333 0.8900 0.8533 0.8922
Qwen3-32B 0.9067 0.8867 0.8467 0.8800
LLaMa3.1-70B 0.9000 0.7867 0.8467 0.8445
LLaMa3.1-70B_COT 0.8933 0.8800 0.8867 0.8867
Distill-LLaMa-70B 0.9333 0.8333 0.8867 0.8844
Ours
SocialR1-4B only Rout_{\text{out}} 0.8933 0.7933 0.8733 0.8533
SocialR1-4B w/o Rlen_{\text{len}} 0.8800 0.8133 0.8000 0.8311
SocialR1-4B w/o Rstruct_{\text{struct}} 0.8667 0.7867 0.8400 0.8311
SocialR1-4B w/o Rcont_{\text{cont}} 0.8667 0.7933 0.8300 0.8300
SocialR1-4B Full 0.8733 0.8433 0.8333 0.8500

SocialR1-8B only Rout_{\text{out}}
0.9333 0.8667 0.8667 0.8889
SocialR1-8B w/o Rlen_{\text{len}} 0.9267 0.8333 0.8700 0.8767
SocialR1-8B w/o Rstruct_{\text{struct}} 0.9000 0.8600 0.8767 0.8789
SocialR1-8B w/o Rcont_{\text{cont}} 0.9000 0.8267 0.8433 0.8567
SocialR1-8B Full 0.9000 0.8600 0.8667 0.8756

Table 7: Performance on SimpleToM.

Table 8: Performances on ToMBench

Table 9: Performance on TactfulToM.

Table 10: Performances on EmoBench

Table 11: Performance on ToMBench_Hard Validation Set

F. Case Study

For completeness, we provide the full qualitative case study examples corresponding to Figure 4. Table 12–Table 17 present the complete narrative contexts, questions, answer options, and model-generated SIP trajectories across Social-R1-8B and its reward-ablated variants. These detailed instances allow a closer inspection of the mechanistic failure modes identified in the main text, including ungrounded cue interpretation, goal misidentification, and option-level lexical shortcutting under outcome-only supervision. Together, these supplementary cases offer concrete evidence that Social-R1’s improvements arise from enforcing evidence-grounded interpretation and stage-consistent social reasoning, rather than answer-driven backfilling or superficial option matching.

Table 12: Case Study SocialR1-8B

Table 13: Case Study Social-8B w/o RcR_c ont

Table 14: Case Study SocialR1-8B only RoR_o ut

Table 15: Case Study Qwen3-8B

Table 16: Case Study Deepseek-R1 Part one

Table 17: Case Study Deepseek-R1 Part two

G. Perturbation Analysis

To provide an intuitive illustration of our perturbation protocol, Table 18 presents a representative case in which story-consistent but decision-irrelevant distractor cues are injected into the original narrative while the correct answer remains unchanged.

Table 18: Case about Perturbation

References

[1] Wen et al. (2025). Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs. arXiv preprint arXiv:2506.14245.
[2] Shapira et al. (2024). Clever hans or neural theory of mind? stress testing social reasoning in large language models. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers).
[3] Pang et al. (2025). Do large language models have a theory of mind?. Proceedings of the National Academy of Sciences.
[4] Salancik, Gerald R and Pfeffer, Jeffrey (1978). A social information processing approach to job attitudes and task design. Administrative science quarterly. pp. 224–253.
[5] Kosinski, Michal (2024). Evaluating large language models in theory of mind tasks. Proceedings of the National Academy of Sciences. 121(45). pp. e2405460121.
[6] Street et al. (2024). Llms achieve adult human performance on higher-order theory of mind tasks. arXiv preprint arXiv:2405.18870.
[7] Strachan et al. (2024). Testing theory of mind in large language models and humans. Nature Human Behaviour. 8(7). pp. 1285–1295.
[8] Gandhi et al. (2023). Understanding social reasoning in language models with language models. Advances in Neural Information Processing Systems. 36. pp. 13518–13529.
[9] Chen et al. (2024). Tombench: Benchmarking theory of mind in large language models. arXiv preprint arXiv:2402.15052.
[10] Xu et al. (2024). OpenToM: A comprehensive benchmark for evaluating theory-of-mind reasoning capabilities of large language models. arXiv preprint arXiv:2402.06044.
[11] Gu et al. (2024). Simpletom: Exposing the gap between explicit tom inference and implicit tom application in llms. arXiv preprint arXiv:2410.13648.
[12] Zhou et al. (2023). How far are large language models from agents with theory-of-mind?. arXiv preprint arXiv:2310.03051.
[13] Xiao et al. (2025). Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States. arXiv preprint arXiv:2505.17663.
[14] He et al. (2023). Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models. arXiv preprint arXiv:2310.16755.
[15] Lu et al. (2025). Tom-rl: Reinforcement learning unlocks theory of mind in small llms. arXiv e-prints. pp. arXiv–2504.
[16] Sabour et al. (2024). Emobench: Evaluating the emotional intelligence of large language models. arXiv preprint arXiv:2402.12071.
[17] Yong et al. (2025). MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?. arXiv preprint arXiv:2506.13065.
[18] Ouyang et al. (2022). Training language models to follow instructions with human feedback. Advances in neural information processing systems. 35. pp. 27730–27744.
[19] Jaech et al. (2024). Openai o1 system card. arXiv preprint arXiv:2412.16720.
[20] Guo et al. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
[21] Team et al. (2025). Kimi k1. 5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599.
[22] Yang et al. (2025). Qwen3 technical report. arXiv preprint arXiv:2505.09388.
[23] Lightman et al. (2023). Let's verify step by step. In The Twelfth International Conference on Learning Representations.
[24] Uesato et al. (2022). Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275.
[25] Zhang et al. (2025). ReasonFlux-PRM: Trajectory-Aware Process Rewards for Chain-of-Thought Supervision. arXiv preprint arXiv:2506.54321.
[26] Osterhaus, Christopher and Bosacki, Sandra L (2022). Looking for the lighthouse: A systematic review of advanced theory-of-mind tests beyond preschool. Developmental Review.
[27] Ullman, Tomer (2023). Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399.
[28] Hu et al. (2025). Re-evaluating Theory of Mind evaluation in large language models. Philosophical Transactions B. pp. 20230499.
[29] Sperber, Dan and Wilson, Deirdre (1986). Relevance: Communication and cognition. Harvard University Press Cambridge, MA.
[30] Crick, Nicki R and Dodge, Kenneth A (1994). A review and reformulation of social information-processing mechanisms in children's social adjustment.. Psychological bulletin.
[31] Simon, Herbert A (1955). A behavioral model of rational choice. The quarterly journal of economics.
[32] Gigerenzer, Gerd and Goldstein, Daniel G (1996). Reasoning the fast and frugal way: models of bounded rationality.. Psychological review.
[33] Shao et al. (2024). Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300.
[34] Sap et al. (2019). Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728.
[35] Wu et al. (2023). Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023.
[36] Liu et al. (2025). TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing.
[37] Guangming Sheng et al. (2024). HybridFlow: A Flexible and Efficient RLHF Framework. arXiv preprint arXiv: 2409.19256.
[38] Dubey et al. (2024). The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
[39] Achiam et al. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774.