Social-R1: Towards Human-like Social Reasoning in LLMs
Jincenzi Wu 1^{1}1
1^{1}1 The Chinese University of Hong Kong
Yuxuan Lei 2^{2}2
2^{2}2 Microsoft Research Asia
Jianxun Lian 2^{2}2
2^{2}2 Microsoft Research Asia
Yitian Huang 2^{2}2
2^{2}2 Microsoft Research Asia
Lexin Zhou 3^{3}3
3^{3}3 Princeton University
Haotian Li 2^{2}2
2^{2}2 Microsoft Research Asia
Xing Xie 2^{2}2
2^{2}2 Microsoft Research Asia
Helen Meng 1^{1}1
1^{1}1 The Chinese University of Hong Kong
Abstract
While large language models demonstrate remarkable capabilities across numerous domains, social intelligence—the capacity to perceive social cues, infer mental states, and generate appropriate responses—remains a critical challenge, particularly for enabling effective human-AI collaboration and developing AI that truly serves human needs. Current models often rely on superficial patterns rather than genuine social reasoning. We argue that cultivating human-like social intelligence requires training with challenging cases that resist shortcut solutions. To this end, we introduce ToMBench-Hard, an adversarial benchmark designed to provide hard training examples for social reasoning. Building on this, we propose Social-R1, a reinforcement learning framework that aligns model reasoning with human cognition through multi-dimensional rewards. Unlike outcome-based RL, Social-R1 supervises the entire reasoning process, enforcing structural alignment, logical integrity, and information density. Results show that our approach enables a 4B parameter model to surpass much larger counterparts and generalize robustly across eight diverse benchmarks. These findings demonstrate that challenging training cases with trajectory-level alignment offer a path toward efficient and reliable social intelligence.
1. Introduction
Recent advances in Reinforcement Learning from Verifiable Feedback (RLVF) [1] have significantly improved large language models (LLMs)' performance on formal reasoning tasks like mathematics and programming. However, genuine social intelligence—the capacity to perceive subtle cues, infer latent mental states, and navigate complex interpersonal dynamics—remains a substantial challenge. Unlike formal reasoning with deterministic execution paths, social reasoning is inherently polysemic and context-dependent, lacking the traceable verification that enables effective reward modeling in objective domains.
Despite strong performance on standard benchmarks, current LLMs often rely on shortcut learning rather than authentic social reasoning. Models frequently exhibit a failure mode we term Reasoning Parasitism, characterized by Answer-driven Backfilling—retroactively constructing justifications for predetermined answers rather than deriving inferences through narrative analysis, as illustrated by the example in Figure 1. The inherent fragility of this parasitic performance becomes particularly evident in adversarial or out-of-distribution scenarios. As highlighted by [2, 3], models that excel on standard benchmarks often suffer catastrophic failures when confronted with trivial narrative perturbations. It reveals that current approaches create only a facade of social intelligence while failing to develop robust reasoning capabilities. Our analysis further identifies a critical Interpretation Bottleneck: while models can perceive surface-level social cues, they struggle to map these cues to latent mental states, leading to a "logic reversal" where final answer correctness exceeds the logical integrity of the reasoning process.
We argue that advancing social intelligence requires aligning models' reasoning trajectories with the structured stages of human social inference. Thus, we propose process-based trajectory alignment that cultivates social intelligence as an internalized capability rather than a parasitic performance. Human social reasoning is characterized by high-density information distillation and recursive belief modeling—properties we aim to instill through structured supervision of the reasoning process, as illustrated in the SocialR1-8B example in Figure 1.
In this paper, we introduce Social-R1, a reinforcement learning framework that encourages genuine social reasoning by aligning model trajectories with human cognitive principles: structured, evidence-grounded, and efficient inference. Our approach has two key components. First, we construct ToMBench-Hard, an expert-curated adversarial benchmark designed to expose reasoning shortcuts through socially nuanced perceptual traps. Second, guided by Social Information Processing (SIP) theory ([4]), we develop a multi-dimensional reward system that enforces stage-consistent reasoning progression (RstructR_{\text{struct}}Rstruct), content integrity (RcontentR_{\text{content}}Rcontent), and inference efficiency (RlenR_{\text{len}}Rlen). We demonstrate that this trajectory-level alignment enables more effective and parameter-efficient social intelligence.
We conduct comprehensive experiments across both in-domain and out-of-domain benchmarks related to social intelligence. Results demonstrate that our approach substantially improves model performance on genuine social reasoning tasks. Social-R1 enables smaller models to match or surpass the capabilities of models with significantly larger parameter counts, while showing strong generalization across eight diverse social reasoning evaluations (as illustrated in the right part of Figure 1). These findings indicate that trajectory-level alignment presents a viable and efficient alternative to reliance on pure model scaling for achieving robust social intelligence. The source code and dataset will be released upon paper acceptance. In summary, our contributions are threefold:
- ToMBench-Hard: A rigorous benchmark with difficulty for social reasoning that exposes shortcut learning in LLMs and mandates genuine cognitive engagement.
- Social-R1 Framework: A reinforcement learning approach with multi-dimensional rewards that align LLM reasoning trajectories with human social cognition.
- Performance Superiority: Demonstration that our method enables small models to achieve large-model performance, proving trajectory quality surpasses parameter scaling for social intelligence.
2. Related Work
2.1 Theory-of-Mind Analysis in LLMs
Recent evaluations demonstrate LLMs' growing proficiency on Theory-of-Mind tasks, with models like GPT-4 achieving near-human performance on false-belief tasks and various social cognition assessments [5, 6, 7]. However, broader audits reveal significant limitations in genuine social reasoning capabilities. Studies highlight inconsistent performance across different evaluation frameworks [8], systematic failure modes in comprehensive capability assessments [9], and particular weaknesses in psychologically complex scenarios compared to physical-world contexts [10, 11, 12]. Dynamic evaluation settings further expose models' difficulties with evolving mental states over time [13], while higher-order reasoning tasks reveal sharp performance drops beyond first-level inference [14]. Current improvement attempts remain limited in scope. Prompt-based interventions [11] offer task-specific benefits but lack generalizability, while reinforcement learning approaches [15] show promise but risk overfitting to narrow task patterns without broader training diversity.
From this literature, we identify three key limitations motivating our work: (i) over-reliance on outcome-based evaluation rather than reasoning trajectory analysis; (ii) susceptibility to shallow statistical cues instead of recursive mental-state inference; and (iii) limited generalization across social reasoning benchmarks and related tasks [16, 17]. These gaps necessitate harder adversarial evaluations, trajectory-level supervision, and training regimes designed for transferable social intelligence.
2.2 RL for LLM General Reasoning
Reinforcement learning has emerged as a pivotal technique for enhancing LLM reasoning capabilities, with RLHF establishing the foundation for alignment [18]. Recent large-scale implementations demonstrate that outcome-based rewards can significantly strengthen structured reasoning in mathematical and coding domains [19, 20, 21, 22]. To address sparse reward challenges, process supervision methods provide denser feedback through step-wise trajectory evaluation [23, 24, 25].
Despite these advances, RL remains underexplored for social inference tasks, which require empathy and multi-factor reasoning in unstructured scenarios. Preliminary work [15] shows potential but remains in early development. This highlights two critical gaps: (i) the need for refreshed training materials tailored to social reasoning, and (ii) the opportunity for trajectory-level rewards that guide models toward human-like social inference patterns.
3. The ToMBench-Hard Benchmark
Social intelligence remains a persistent bottleneck in LLMs, with current systems often relying on shortcut learning rather than authentic social reasoning. To address this limitation, we introduce ToMBench-Hard, a diagnostic benchmark specifically designed to expose the shortcut learning behaviors prevalent in current LLMs. The primary goal of ToMBench-Hard is to disentangle genuine social inference from surface-level pattern matching by creating a challenging adversarial environment that mandates structured, human-like reasoning processes.
3.1 Adversarial Data Construction
ToMBench-Hard is grounded in the Abilities in the Theory-of-Mind Space (ATOMS) framework ([26]), providing a comprehensive evaluation covering six core dimensions of social intelligence: Belief, Desire, Emotion, Intention, Knowledge, and Non-literal Communication (NLC). We developed 800 expert-annotated multiple-choice questions through a rigorous annotation protocol involving three independent annotators to ensure cognitive validity and high quality (see Appendix B for details).
To mitigate heuristic-based reasoning strategies such as lexical overlap between questions and options, we introduced ToM-consistent adversarial perturbations ([27, 28]). These perturbations include nuanced manipulations of perceptual access (e.g., unobserved state changes) and asymmetric information (e.g., second-order beliefs), ensuring models cannot succeed through statistical guessing alone. Each sample was carefully designed to require a structured, human-like reasoning process that progresses from cue encoding to mental state interpretation.
3.2 Diagnostic Analysis: Exposing the Shortcut Illusion
To validate the diagnostic utility of ToMBench-Hard, we conducted comprehensive benchmarking across a spectrum of state-of-the-art LLMs, comparing their performance against human expert baselines. The results reveal a significant performance cliff between simple synthetic benchmarks and our adversarial evaluation.
As shown in Table 1, while human experts achieve robust 87% accuracy, frontier models exhibit dramatic performance drops. Notably, models like O3 and Deepseek-R1 achieve near-human performance on simpler benchmarks like ToM-RL ([15]) (reaching 87-88% accuracy), but their accuracy plummets to below 61% on ToMBench-Hard. This performance gap unmasks what we term the shortcut illusion—the phenomenon where high scores on conventional benchmarks reflect template matching artifacts rather than genuine social reasoning capabilities.
The substantial performance gap between ToMBench-Hard and ToM-RL suggests that high scores on the latter may not reliably indicate strong theory-of-mind capabilities, as these benchmarks might be insufficiently challenging for LLMs. By systematically exposing these limitations, ToMBench-Hard serves dual purposes: it provides a rigorous evaluation framework for assessing genuine social reasoning capabilities, while also offering a high-quality dataset for aligning LLMs with human-like cognitive processes. To facilitate both training and evaluation, we partition the complete ToMBench-Hard dataset into non-overlapping training and test subsets, enabling the development of our alignment framework in the subsequent sections.
4. The Social-R1 Framework
The Social-R1 framework introduces a novel paradigm for cultivating genuine social intelligence in LLMs by aligning their reasoning processes with human social cognition. Departing from traditional outcome-based reinforcement learning approaches, our method specifically tackles the fundamental challenge of Reasoning Parasitism—a phenomenon where models engage in Answer-driven Backfilling by constructing post-hoc justifications for predetermined answers. Social-R1 counteracts these shortcut behaviors through a multi-dimensional reward system that supervises the entire reasoning trajectory rather than merely rewarding final outcomes. This comprehensive approach transforms social intelligence from a parasitic performance into an internalized capability, ensuring model reasoning embodies the core characteristics of human social cognition: precise, stage-consistent social inference ([4]) and high information density ([29]).
4.1 Multi-Dimensional Reward Design
The core innovation of Social-R1 lies in its comprehensive reward system that aligns model reasoning trajectories with three key characteristics of human social cognition: structured progression, logical integrity, and information density. This multi-dimensional approach provides fine-grained supervision over the reasoning process, addressing different aspects of high-quality social inference.
SIP Structural Alignment (RstructR_{{struct}}Rstruct). Human social reasoning follows a disciplined progression from perception to response generation ([30]). To instill this cognitive scaffold, we introduce a structural reward RstructR_{{struct}}Rstruct that enforces sequential reasoning across the four stages of Social Information Processing (SIP): (1) Encoding Social Cues: Identifying relevant social signals from the narrative, (2) Interpreting Cues: Inferring latent mental states from perceived cues, (3) Clarifying Goals: Determining social objectives and interpersonal intentions, and (4) Response Generation: Selecting appropriate behavioral responses. The RstructR_{{struct}}Rstruct reward validates whether intermediate reasoning steps adhere to this stage-wise progression, penalizing premature conclusions and stage skipping. This structural constraint encourages coherent, story-grounded causal inference while mitigating shortcut behaviors like option parasitism, where models anchor reasoning directly on multiple-choice options rather than narrative evidence. We leverage GPT-4o as judge for this reward, with the prompt illustrated in Appendix C.1.
SIP Content Integrity (RcontentR_{{content}}Rcontent). While structural alignment ensures proper reasoning staging, content integrity guarantees logical rigor within each stage. The content reward RcontentR_{{content}}Rcontent audits whether intermediate inferences remain grounded in story-internal evidence, correctly reflecting social cues, intentions, and goals. This reward penalizes three critical failure modes: (1) Erroneous cue encoding: Misidentification of relevant social signals, (2) Flawed interpretation: Incorrect mental state attribution, and (3) Misidentified goals: Inaccurate inference of social objectives. By ensuring each reasoning step maintains evidential support, RcontentR_{{content}}Rcontent discourages superficial rationalizations and promotes authentic social sensing derived from narrative context. The detailed implementation of this reward model can be found at Section 5.1.
Inference Efficiency Optimization (RlenR_{{len}}Rlen). Human social reasoning achieves high information density through selective attention and avoidance of redundant processing ([31, 32]). To emulate this cognitive efficiency, we design RlenR_{{len}}Rlen as the product of two complementary components:
The repetition penalty component Rrep(ρ)R_{{rep}}(\rho)Rrep(ρ) specifically targets circular over-thinking by penalizing excessive n-gram repetition beyond a threshold τ=0.1\tau = 0.1τ=0.1:
where β=8\beta = 8β=8 controls the severity of penalty decay for high repetition ratios. The length window constraint Rwin(L)R_{{win}}(L)Rwin(L) maintains reasoning trajectories within an empirically optimal range [Lmin,Lmax][L_{\min}, L_{\max}][Lmin,Lmax] through smooth gating:
with k=50k = 50k=50 controlling transition smoothness. The bounds Lmin=400L_{\min} = 400Lmin=400 and Lmax=2500L_{\max} = 2500Lmax=2500 are derived from strong chain-of-thought baselines, ensuring reasoning remains concise yet comprehensive. This dual-mechanism approach encourages the model to emulate human-like efficiency by avoiding both redundant repetition and excessive verbosity while maintaining substantive social inference.
Verifiable Format Alignment (RfmtR_{{fmt}}Rfmt) We adopt the format reward from [33] that enforces structured thinking processes. The model is rewarded for producing outputs with predefined XML-style tags (<thinking> and <answer>), which enables deterministic extraction of both reasoning trajectories and final answers while preserving semantic freedom in the reasoning content.
4.2 Reward Synthesis and Learning
The composite reward function integrates all components through a carefully designed synthesis strategy that balances outcome supervision with process-level reasoning signals:
where RoutR_{{out}}Rout denotes the verifiable outcome reward, We implement a curriculum learning strategy where outcome supervision dominates early training phases (wo(t)=2w_o(t) = 2wo(t)=2), while process-level rewards are progressively emphasized through time-dependent weighting: wstruct(t)=wcontent(t)=1+γtTw_{{struct}}(t) = w_{{content}}(t) = 1 + \gamma\frac{t}{T}wstruct(t)=wcontent(t)=1+γTt. This curriculum ensures stable initial convergence while gradually reinforcing human-like reasoning patterns as training progresses. The optimization employs Group Relative Policy Optimization ([33]), which performs group-relative updates over sampled reasoning trajectories.
5. Experiment
5.1 Experiment Setting
We evaluate our model on eight multiple-choice social benchmarks, including two in-domain benchmarks: the public ToMBench ([9]) and our ToMBench-Hard test set, and six out-of-domain benchmarks: SocialIQA ([34]) for social commonsense reasoning, EmoBench ([16]) for emotion intelligence, MotiveBench ([17]) for social motivation reasoning, SimpleToM ([11]) for examining whether models can consciously infer others’ mental states (MS) and proactively applying such reasoning to behavior inference, Hi-ToM ([35]) for evaluating higher-order Theory-of-Mind reasoning, and TactfulToM ([36]) for testing whether models can interpret white lies and infer the underlying prosocial intent to preserve interpersonal harmony.
SIP Content Integrity Reward Model (RMcontentRM_{content}RMcontent). To instantiate RcontentR_{content}Rcontent, we train a dedicated Content Reward Model that assigns a scalar quality score to intermediate SIP-stage reasoning segments. We construct the SocialPairs-20K dataset by sampling NNN trajectories per training instance from various Social-R1 checkpoints. These checkpoints span different training stages, ensuring a diverse pool of reasoning qualities ranging from nascent logic to sophisticated social inference. For each segment, a strong teacher-judge (o3) generates silver-standard stage-wise rationales by conditioning on the gold final answer and the original social context. We then employ a multi-dimensional rubric—focusing on fact-grounding, mental-state attribution accuracy, and stage-specific relevance—to score the sampled segments against the teacher-generated references. This process forms preference pairs (chosen vs. rejected), while rejected segments exhibit missing cues, incorrect mental-state attributions, or wrong goal identification. We validate the model on two held-out sets: (1) an automatic test split of 2k pairs, where RcontentR_{content}Rcontent achieves an accuracy of 89.2%; and (2) a human-calibrated subset of 200 pairs annotated by experts, achieving an 87.5% agreement with human labels. (More RcontentR_{content}Rcontent can be found in Appendix C.2).
Implementation Details SIP Content Integrity Reward Model (RMcontentRM_{content}RMcontent) is trained as a pairwise preference reward model. We initialize it from Qwen3-4B and fine-tune with LoRA on the SocialPairs-20K dataset mentioned previously. For policy optimisation, we train two versions of reasoning models from Qwen3-4B and Qwen3-8B on ToMBench-Hard, consisting of 700 training instances and 100 test instances. Reinforcement learning is performed for 600 optimisation steps using VERL{\mathchoice{\text{V{\scriptsize ERL}}}{\text{V{\scriptsize ERL}}}{\text{V{\scriptscriptstyle ERL}}}{\text{VERL}}}VERL ([37]) on 8 NVIDIA A100 (80GB) GPUs. We set the group size to 5, the KL coefficient to 0.04, and the learning rate to 5×10−75 \times 10^{-7}5×10−7. More details are provided in Appendix D.
5.2 Main Results
We apply the Social-R1 framework to two open-source backbones of different scales, Qwen3-4B and Qwen3-8B. The overall results across eight social reasoning benchmarks are reported in Table 2. Social-R1 achieves strong and consistent improvements over the corresponding Qwen baselines, demonstrating that reinforcement learning with challenging social supervision and trajectory-level reward signals can substantially enhance Theory-of-Mind reasoning. Notably, Social-R1-4B surpasses LLaMa3.1-70B across all benchmarks, despite being more than an order of magnitude smaller, highlighting the effectiveness of process-based social alignment beyond parameter scaling. Even more strikingly, Social-R1-8B outperforms DeepSeek-R1 on several benchmarks and achieves stronger overall performance, consistently matching or exceeding much larger baselines such as Qwen3-32B in out-of-domain generalization. More detailed results are provided in Appendix E.
5.3 Ablation Studies
To systematically evaluate the contribution of each reward component in Social-R1, we conduct comprehensive ablation experiments on both SocialR1-4B and SocialR1-8B variants. Specifically, we remove three key components: the length-control reward (w/o Rlenw/o\ R_{\text{len}}w/o Rlen), the structural trajectory reward (w/o Rstructw/o\ R_{\text{struct}}w/o Rstruct), and the content integrity reward (w/o Rcontw/o\ R_{\text{cont}}w/o Rcont). Additionally, we examine a baseline variant that removes all the progress rewards (only Routonly\ R_{\text{out}}only Rout). The complete results are presented in Table 2.
Benchmarks The ablation analysis reveals distinct performance patterns associated with each reward component. Removing w/o Rlenw/o\ R_{\text{len}}w/o Rlen results in significant performance degradation on Hi-ToM (e.g., 0.7083→0.62670.7083 \rightarrow 0.62670.7083→0.6267 for SocialR1-8B), indicating that uncontrolled reasoning length negatively impacts high-order social inference. The w/o RlenR_{\text{len}}Rlen variant exhibits a substantial increase in reasoning verbosity, with the average thinking length rising by approximately +250% compared to SocialR1-8B as shown in Figure 10.
The structural reward RstructR_{\text{struct}}Rstruct shows consistent importance across benchmarks, with its removal causing accuracy drops (e.g., from 0.50790.50790.5079 to 0.45580.45580.4558 on TactfulToM). Similarly, RcontR_{\text{cont}}Rcont removal leads to performance declines, suggesting its role in maintaining reasoning quality. Notably, the only Routonly\ R_{\text{out}}only Rout variant exhibits more severe overall performance deterioration compared to other ablations, demonstrating the importance of incorporating process-level rewards. While these numerical changes demonstrate the overall effectiveness of each component, subsequent in-depth analyses (Section 6.1–Section 6.3) will provide mechanistic evidence through detailed trajectory examinations and perturbation studies to understand how each reward shapes reasoning behavior.
6. In-Depth Analysis
To examine whether Social-R1 yields genuinely internalised social reasoning—rather than Reasoning Parasitism driven by answer-first backfilling—we conduct a mechanistic analysis of its reasoning dynamics. We compare Social-R1-8B against three strong baselines (DeepSeek-R1, DeepSeek-R1-Distill-Llama-70B, and Qwen3-8B), together with controlled reward variants that isolate outcome-only training and ablate key trajectory-level constraints. In line with our reward design, we operationalise human-like social reasoning through three diagnostic signatures: (i) option-independent inference grounded in narrative cues; (ii) stage-consistent SIP trajectories with content integrity, where intermediate beliefs remain logically valid across Encoding, Interpretation, Goal Clarification, and Response Generation; and (iii) selective robustness, where models avoid redundant trajectory bloat under perturbation. Accordingly, our analysis proceeds in three steps: (1) quantifying Reasoning Parasitism as a diagnostic of shortcut reliance; (2) auditing stage-wise cognitive fidelity to identify interpretation bottlenecks; and (3) introducing controlled distractors to assess robustness and conciseness.
6.1 From Parasitism to Independent Inference
gt;80\%$), performance drops sharply—by nearly 25 points—once entering Cue Interpretation, which requires latent mental-state attribution and recursive belief reasoning. More critically, outcome-only training induces a *reasoning reversal* phenomenon. For example, Qwen3-8B trained with $R_{\text{out}}$ often achieves higher final answer accuracy than the correctness of its preceding SIP stages, indicating that surface-level option correlations can yield correct predictions despite incoherent intermediate social inference. This highlights the limitation of outcome supervision: it rewards answer selection without enforcing psychologically valid reasoning trajectories. Crucially, the proposed Content Reward $R_{\text{content}}$ directly targets this failure mode by penalising erroneous cue encoding, flawed mental-state interpretation, and misidentified social goals. Ablation results provide clear causal evidence: removing $R_{\text{content}}$ reduces Interpretation accuracy by 6.2 points (77.5%$\rightarrow$71.3%) and Goal Classification by 6.2 points (75.0%$\rightarrow$68.8%). Ultimately, Social-R1-8B preserves the most stage-consistent and smoothly degrading SIP trajectory, demonstrating that its gains stem from cognitively grounded reasoning rather than shortcut-driven recovery.
" data-original-markdown="
A primary failure mode in multiple-choice social reasoning is *Reasoning Parasitism*—a shortcut behaviour where models anchor their deductions on statistical regularities in the answer choices rather than deriving conclusions from story-internal social cues. Figure 2 reveals substantial divergence in reasoning dynamics. DeepSeek-R1, Qwen3-8B, and the outcome-only variant of SocialR1-8B exhibit high option-mention density as early as the Cue Encoding stage (Q1 group in Figure 2), often exceeding five explicit option references per sample. Such premature reliance suggests *answer-conditioned backfilling*: models use the provided choices to retroactively assemble plausible justifications, rather than performing independent social inference. In contrast, Social-R1-8B maintains a largely option-agnostic trajectory. Its option mentions remain minimal and nearly flat throughout cue perception and interpretation, increasing only slightly (to $\sim$1.3) during final Response Generation (Q4). This provides mechanistic evidence that the proposed Social Think Combined Reward effectively suppresses shortcut dependence on answer options, compelling the model to engage in narrative-grounded social deduction. Additional qualitative examples are provided in Appendix F.
### 6.2 Stage-wise Diagnosis
To pinpoint where social reasoning most systematically breaks down, we conduct a human audit of 80 instances sampled from eight benchmarks, ensuring high annotation quality and genuine Theory-of-Mind demands. We track success rates across the four Social Information Processing (SIP) stages. Figure 3 exposes a pronounced *Interpretation Bottleneck*. While strong baselines sustain high accuracy in factual cue encoding (
gt;80\%$), performance drops sharply—by nearly 25 points—once entering Cue Interpretation, which requires latent mental-state attribution and recursive belief reasoning. More critically, outcome-only training induces a *reasoning reversal* phenomenon. For example, Qwen3-8B trained with $R_{\text{out}}$ often achieves higher final answer accuracy than the correctness of its preceding SIP stages, indicating that surface-level option correlations can yield correct predictions despite incoherent intermediate social inference. This highlights the limitation of outcome supervision: it rewards answer selection without enforcing psychologically valid reasoning trajectories. Crucially, the proposed Content Reward $R_{\text{content}}$ directly targets this failure mode by penalising erroneous cue encoding, flawed mental-state interpretation, and misidentified social goals. Ablation results provide clear causal evidence: removing $R_{\text{content}}$ reduces Interpretation accuracy by 6.2 points (77.5%$\rightarrow$71.3%) and Goal Classification by 6.2 points (75.0%$\rightarrow$68.8%). Ultimately, Social-R1-8B preserves the most stage-consistent and smoothly degrading SIP trajectory, demonstrating that its gains stem from cognitively grounded reasoning rather than shortcut-driven recovery.
" data-source-offset="26737" class="markdown-segment">
A primary failure mode in multiple-choice social reasoning is Reasoning Parasitism—a shortcut behaviour where models anchor their deductions on statistical regularities in the answer choices rather than deriving conclusions from story-internal social cues. Figure 2 reveals substantial divergence in reasoning dynamics. DeepSeek-R1, Qwen3-8B, and the outcome-only variant of SocialR1-8B exhibit high option-mention density as early as the Cue Encoding stage (Q1 group in Figure 2), often exceeding five explicit option references per sample. Such premature reliance suggests answer-conditioned backfilling: models use the provided choices to retroactively assemble plausible justifications, rather than performing independent social inference. In contrast, Social-R1-8B maintains a largely option-agnostic trajectory. Its option mentions remain minimal and nearly flat throughout cue perception and interpretation, increasing only slightly (to ∼\sim∼1.3) during final Response Generation (Q4). This provides mechanistic evidence that the proposed Social Think Combined Reward effectively suppresses shortcut dependence on answer options, compelling the model to engage in narrative-grounded social deduction. Additional qualitative examples are provided in Appendix F.
6.2 Stage-wise Diagnosis
To pinpoint where social reasoning most systematically breaks down, we conduct a human audit of 80 instances sampled from eight benchmarks, ensuring high annotation quality and genuine Theory-of-Mind demands. We track success rates across the four Social Information Processing (SIP) stages. Figure 3 exposes a pronounced Interpretation Bottleneck. While strong baselines sustain high accuracy in factual cue encoding (>80%>80\%>80%), performance drops sharply—by nearly 25 points—once entering Cue Interpretation, which requires latent mental-state attribution and recursive belief reasoning. More critically, outcome-only training induces a reasoning reversal phenomenon. For example, Qwen3-8B trained with RoutR_{\text{out}}Rout often achieves higher final answer accuracy than the correctness of its preceding SIP stages, indicating that surface-level option correlations can yield correct predictions despite incoherent intermediate social inference. This highlights the limitation of outcome supervision: it rewards answer selection without enforcing psychologically valid reasoning trajectories. Crucially, the proposed Content Reward RcontentR_{\text{content}}Rcontent directly targets this failure mode by penalising erroneous cue encoding, flawed mental-state interpretation, and misidentified social goals. Ablation results provide clear causal evidence: removing RcontentR_{\text{content}}Rcontent reduces Interpretation accuracy by 6.2 points (77.5%→\rightarrow→71.3%) and Goal Classification by 6.2 points (75.0%→\rightarrow→68.8%). Ultimately, Social-R1-8B preserves the most stage-consistent and smoothly degrading SIP trajectory, demonstrating that its gains stem from cognitively grounded reasoning rather than shortcut-driven recovery.
gt;$ Tier C (strongest supervision against hard negatives). **P1**: Tier A
gt;$ Tier C. **P2**: Tier A
gt;$ Tier B. **P3**: Tier B
gt;$ Tier D. **P4**: late-stage short trajectories
gt;$ early-stage long trajectories, encouraging concise and cognitively disciplined inference.
Applying this policy produces **SocialPairs-20K**, consisting of 20K preference pairs for training $RM_{\text{content}}$.
**Held-out evaluation.** We construct a held-out test set of 2K preference pairs from ToMBench-Hard test instances using the same pipeline, achieving 89.2% pairwise accuracy. Additionally, we randomly sample 200 pairs for expert human verification, obtaining 87.5% agreement with human preferences.
The detailed multi-dimensional scoring rubric used by GPT-5 for stage-wise segment evaluation is provided in Table 5.
### D. Implementation
The content reward model (TRM) is initialized from Qwen3-4B-Instruct-2507and trained using supervised fine-tuning for two epochs on four NVIDIA A100 80GB GPUs. The model is trained with pairwise comparisons and, during inference, predicts a scalar reward conditioned on (*scenario, question, options, reasoning*).
We evaluate multiple LLM families, including DeepSeek-R1([20]), DeepSeek-R1-Distill-Llama-70B([20])<sup>[^1]</sup>, Llama-3.1-70B-Instruct([38])<sup>[^2]</sup>, Qwen3-4B/8B/32B([22])<sup>[^3]</sup>, GPT-5-2025-08-07([18]), GPT-4o-2024-08-06 ([39]), and o3-2025-04-16. We prompt Qwen3-4B/8B/32B in both Thinking and No-Think setting.
" data-original-markdown="
**Teacher references (gold SIP rationales).** We first provide the teacher model (o3) with the gold final answer and the original narrative context, and obtain stage-consistent SIP rationales as silver references for each intermediate SIP stage.
**Trajectory sampling from diverse checkpoints.** To ensure a wide spectrum of reasoning quality, we collect candidate trajectories from 10 checkpoints of the *w/o $R_{\text{struct}}$* run at steps $\{30, 90, 120, 180, 270, 360, 420, 510, 570, 600\}$. We sample $K{=}6$ trajectories per instance per checkpoint. Given $N{=}700$ training instances, this yields: 42000 candidate SIP-stage reasoning segments, spanning from early unstable logic to mature social inference.
**LLM-based scoring with teacher calibration.** Each sampled segment is scored by GPT-5 using the o3 silver rationale as the reference. The evaluation rubric focuses on: (i) fact-grounding to story-internal evidence, (ii) correctness of mental-state attribution, and (iii) stage-specific relevance within the SIP trajectory.
**Tiering and hard-negative identification.** We categorise all segments into five quality tiers: **Tier S (Expert)**: o3 gold-quality samples. **Tier A (Quality Positive)**: $\mathrm{ACC}{=}1$ and $\mathrm{LLM\_Score} \ge 0.8$. **Tier B (Weak Positive)**: $\mathrm{ACC}{=}1$ and $\ge 0.6 \mathrm{LLM\_Score} \le 0.8$. **Tier C (Hard Negative)**: $\mathrm{ACC}{=}1$ while $\mathrm{LLM\_Score} \le 0.6$ . **Tier D (Common Negative)**: $\mathrm{ACC}{=}0$ and $\mathrm{LLM\_Score}$ is very low.
We construct preference pairs (*chosen* vs. *rejected*) using the following priority scheme: **P0**: Tier S
gt;$ Tier C (strongest supervision against hard negatives). **P1**: Tier A
gt;$ Tier C. **P2**: Tier A
gt;$ Tier B. **P3**: Tier B
gt;$ Tier D. **P4**: late-stage short trajectories
gt;$ early-stage long trajectories, encouraging concise and cognitively disciplined inference.
Applying this policy produces **SocialPairs-20K**, consisting of 20K preference pairs for training $RM_{\text{content}}$.
**Held-out evaluation.** We construct a held-out test set of 2K preference pairs from ToMBench-Hard test instances using the same pipeline, achieving 89.2% pairwise accuracy. Additionally, we randomly sample 200 pairs for expert human verification, obtaining 87.5% agreement with human preferences.
The detailed multi-dimensional scoring rubric used by GPT-5 for stage-wise segment evaluation is provided in Table 5.
### D. Implementation
The content reward model (TRM) is initialized from Qwen3-4B-Instruct-2507and trained using supervised fine-tuning for two epochs on four NVIDIA A100 80GB GPUs. The model is trained with pairwise comparisons and, during inference, predicts a scalar reward conditioned on (*scenario, question, options, reasoning*).
We evaluate multiple LLM families, including DeepSeek-R1([20]), DeepSeek-R1-Distill-Llama-70B([20])[^1], Llama-3.1-70B-Instruct([38])[^2], Qwen3-4B/8B/32B([22])[^3], GPT-5-2025-08-07([18]), GPT-4o-2024-08-06 ([39]), and o3-2025-04-16. We prompt Qwen3-4B/8B/32B in both Thinking and No-Think setting.
" data-source-offset="39847" class="markdown-segment">
We evaluate multiple LLM families, including DeepSeek-R1([20]), DeepSeek-R1-Distill-Llama-70B([20])1, Llama-3.1-70B-Instruct([38])2, Qwen3-4B/8B/32B([22])3, GPT-5-2025-08-07([18]), GPT-4o-2024-08-06 ([39]), and o3-2025-04-16. We prompt Qwen3-4B/8B/32B in both Thinking and No-Think setting.