From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL

Wenyue HuaZezhou HuangTyler PayneSafoora YousefiSaleema AmershiAsli Celikyilmaz

article2026arXiv0 citations

Demonstrates that reinforcement learning with theory-of-mind distillation enables a 4-billion-parameter language model to match or outperform frontier models across six multi-issue negotiation benchmarks while resisting premature concessions and privacy leaks.

Listen

Artificial intelligence agents are increasingly deployed as delegates to handle consequential tasks on behalf of users, including purchasing items, negotiating contracts, and coordinating schedules. While general-purpose frontier models excel at cooperative assistance, their default dispositions—such as extreme agreeableness, eagerness to reach consensus, and transparency—often lead them to concede ground too quickly and leak private information when facing counterparts with conflicting goals. The article evaluates whether targeted post-training in social reasoning can transform a compact 4-billion-parameter language model into an effective, value-preserving delegate across complex strategic environments.

To address this challenge, the authors developed a decoupled multi-agent training framework and evaluated agents across six diverse interaction domains, including bilateral price bargaining, multi-issue resource allocation, employment contract negotiation, and meeting scheduling. The training approach first created domain-specific specialists using reinforcement learning with outcome-based rewards calibrated to the specific difficulty of each scenario. The authors then analyzed cross-domain transfer dynamics and evaluated two consolidation strategies to merge specialist capabilities into a single model: a transfer-aware sequential cascade of reinforcement learning and multi-teacher on-policy distillation. The study also examined explicit theory-of-mind supervision by training agents to infer counterpart preferences, act, and anticipate counterpart reactions.

The findings show that targeted post-training enables a compact 4-billion-parameter model to reach an aggregate utility of 0.619, matching or exceeding much larger models such as GPT-4.1 (0.625) and GPT-5.1 (0.619). Cross-domain skill transfer proved to be highly structured and directional: structurally paired tasks, such as price bargaining environments, transferred substantially to one another, while multi-issue bargaining provided broad positive transfer across multiple domains. Leveraging this transfer structure, a transfer-aware cascade consolidation reached an overall utility of 0.627, while multi-teacher distillation recovered 92.6% of the specialists' advantage in only 60 additional training steps. At the behavioral level, trained agents eliminated premature information disclosure, reduced target price leakage from over 50% down to 1%, and adopted strategic anchoring and selective concession. Explicit theory-of-mind distillation further boosted performance, with the ability to anticipate a counterpart's next move proving to be the single most critical driver of negotiation success.

These results indicate that effective delegation does not require massive frontier-scale models, offering organizations a path to deploy cost-effective, specialized delegates that robustly defend user utility. Stakeholders seeking to build strategic agents should prioritize targeted post-training pipelines over basic prompting, adopting multi-teacher distillation for compute-efficient consolidation and scheduling sequential learning according to task transfer dynamics. Training routines should also incorporate next-action prediction supervision to reinforce strategic foresight.

Certain limitations should be considered before widespread deployment. The study observed instances where bargaining agents generated fabricated market comparisons to justify aggressive offers, highlighting the need to enforce factual grounding. Furthermore, the findings reflect performance within structured simulation environments and benchmark games. Stakeholders can have high confidence in the framework's core mechanics, but organizations should conduct human-facing pilot testing before deploying autonomous delegates in high-stakes commercial applications.

No sufficiently relevant recommendations were found.

Cover for From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL

Abstract

AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. These principal-driven tasks routinely place the agent across from a counterpart (another user's agent, a seller, a recruiter) whose goals may conflict with its principal's. Yet the dispositions that make an assistant pleasant can make it a poor delegate: a friendly, helpful frontier model may disclose its principal's private information unprompted and concede at the first sign of resistance. We present SocialRL, a general recipe that trains social reasoning directly, and apply it to a 4B model across six domains: Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, and Marketplace. Every domain is trained in-domain under the same recipe, and every policy is evaluated on all six. We find that (1) in-domain training reaches the frontier: on held-out scenarios the 4B matches or exceeds the GPT-5 family per domain, closing 73-122% of the baseline-to-frontier gap on the negotiation games, with 78% of buyer openings anchoring below target versus 3% untrained; (2) cross-domain transfer follows game structure: structurally paired games lift each other, a broad multi-issue donor lifts nearly all domains, and structurally isolated games transfer nothing; (3) guided by this transfer structure, two strategies, cascade RL and multi-teacher on-policy distillation (OPD), consolidate the per-domain specialists into a single unified 4B that reaches 0.627 average utility across all six environments, matching or exceeding GPT-4.1 (0.625), GPT-5.1 (0.619), and GPT-5.2 (0.613); (4) an explicit theory-of-mind scaffold helps only through training: distilling the ToM trace, rather than actions alone, lifts utility on every environment and generalizes better across them, and of the two ToM skills, only next-action prediction predicts negotiation outcomes.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Environment and Infrastructure for SocialRL
  • 3.1 Environment and Agent Interface
  • 3.2 Decoupled Training Infrastructure
  • 4 Social Reasoning Environments and Reward Design
  • 4.1 Interaction Environments
  • 4.2 Outcome Evaluation and Reward Design
  • 5 SocialRL Training and Unification
  • 5.1 Stage 1: In-Domain Specialist Training
  • 5.2 Cross-Environment Transfer
  • 5.3 Stage 2: Consolidating Specialists into a Unified Model
  • 5.3.1 Stage 2A: Transfer-Aware Unification with Cascade RL
  • 5.3.2 Stage 2B: Efficient Unification with MOPD
  • 5.4 Explicit Theory-of-Mind Supervision
  • 6 Qualitative Analysis
  • 6.1 Strategic Behavior: More Selective Concession
  • 6.2 Interaction Behavior: Less Private Deliberation, More Strategic Communication
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Domain-specialized 4B policies reach frontier-model performance

    empirical result

    The SOCIALRL study post-trains Qwen3-4B-Instruct-2507 separately for six delegated negotiation and coordination environments. It applies PPO directly for Deal-or-No-Deal (DnD), CaSiNo, Job Interview, and Calendar; Craigslist and Marketplace receive supervised fine-tuning (SFT) before PPO because the base policy does not reliably explore the anchoring behavior needed for price bargaining. The following scores are mean ± standard deviation across repeated evaluations for the 4B policies and mean scores for GPT models, ordered as DnD, CaSiNo, Craigslist, Job Interview, Calendar, Marketplace, then six-environment average: base 4B: 0.583 ± 0.001, 0.476 ± 0.014, 0.318 ± 0.012, 0.479 ± 0.032, 0.301 ± 0.017, 0.174 ± 0.026; average 0.389. Domain-trained 4B: 0.656 ± 0.008, 0.503 ± 0.007, 0.583 ± 0.008, 0.594 ± 0.007, 0.540 ± 0.017, 0.838 ± 0.013; average 0.619. GPT-4.1: 0.653, 0.491, 0.540, 0.588, 0.673, 0.804; average 0.625. GPT-5.1: 0.671, 0.499, 0.607, 0.596, 0.573, 0.767; average 0.619. GPT-5.2: 0.663, 0.488, 0.577, 0.579, 0.643, 0.727; average 0.613. GPT-5.5: 0.665, 0.559, 0.746, 0.590, 0.702, 0.985; average 0.708. Thus the six specialists average 0.619, comparable in aggregate to GPT-4.1, GPT-5.1, and GPT-5.2, and match or exceed those models in several individual environments.

  2. Knowl 2 — Cross-environment transfer is directional and depends on interaction structure

    empirical result

    Each domain-trained 4B policy was evaluated in all six environments; a donor's out-transfer is its mean score change from the base 4B policy on the five environments other than its training domain. In the order DnD, CaSiNo, Craigslist, Job Interview, Calendar, Marketplace, the base scores were 0.583 ± 0.001, 0.476 ± 0.014, 0.318 ± 0.012, 0.479 ± 0.032, 0.301 ± 0.012, 0.174 ± 0.003. DnD training yielded 0.656 ± 0.008, 0.491 ± 0.012, 0.282 ± 0.028, 0.481 ± 0.016, 0.370 ± 0.009, 0.150 ± 0.003; out-transfer +0.005. CaSiNo training yielded 0.597 ± 0.019, 0.503 ± 0.007, 0.309 ± 0.004, 0.463 ± 0.016, 0.331 ± 0.010, 0.168 ± 0.022; +0.003. Craigslist training yielded 0.564 ± 0.023, 0.494 ± 0.007, 0.583 ± 0.008, 0.447 ± 0.011, 0.266 ± 0.023, 0.491 ± 0.007; +0.050. Job Interview training yielded 0.610 ± 0.017, 0.511 ± 0.004, 0.328 ± 0.010, 0.594 ± 0.007, 0.373 ± 0.022, 0.161 ± 0.011; +0.026. Calendar training yielded 0.607 ± 0.010, 0.474 ± 0.007, 0.167 ± 0.013, 0.311 ± 0.013, 0.540 ± 0.017, 0.172 ± 0.008; −0.060. Marketplace training yielded 0.558 ± 0.010, 0.464 ± 0.007, 0.502 ± 0.015, 0.371 ± 0.010, 0.268 ± 0.015, 0.838 ± 0.013; +0.001. The largest off-domain gains occur between the structurally similar price-negotiation tasks: Craigslist training raises Marketplace from 0.174 to 0.491 (+0.317), and Marketplace training raises Craigslist from 0.318 to 0.502 (+0.184). DnD and CaSiNo also transfer positively in both directions, more modestly. Job Interview is a broader donor, improving four other domains, while Calendar is a negative donor overall and substantially reduces Craigslist and Job Interview performance. Marketplace's high in-domain score does not make it a broad donor. The pattern shows that transfer depends on the donor-recipient pairing, not simply on specialist strength. Reported uncertainties are standard deviations across benchmark runs where available.

  3. Knowl 3 — Transfer-aware cascade reinforcement learning consolidates specialists

    model/method

    Cascade RL carries one policy through sequential PPO stages, training on one environment at a time and using the selected checkpoint to initialize the next stage. After each stage, checkpoints are evaluated on all environments encountered so far, and the checkpoint with the highest mean utility on those environments is retained. The transfer-aware order is Calendar → CaSiNo → DnD → Craigslist → Marketplace → Job Interview: Calendar is placed before the capabilities it tends to damage, Craigslist precedes Marketplace to exploit the stronger transfer direction, and Job Interview is last because its own capability is poorly preserved after other training. All cascades start from the Craigslist-SFT checkpoint; each environment stage has a maximum of 200 PPO steps, with evaluation every 20 steps. With the same overall training setup, the transfer-aware, random-order, and anti-transfer-order cascades scored, respectively, in the order DnD, CaSiNo, Craigslist, Job Interview, Calendar, Marketplace, Avg-6: transfer-aware 0.617 ± 0.008, 0.482 ± 0.019, 0.580 ± 0.007, 0.538 ± 0.012, 0.742 ± 0.037, 0.803 ± 0.008, 0.627 ± 0.004; random order 0.604 ± 0.012, 0.465 ± 0.016, 0.637 ± 0.022, 0.478 ± 0.030, 0.541 ± 0.021, 0.783 ± 0.024, 0.578 ± 0.008; anti-transfer order 0.622 ± 0.018, 0.443 ± 0.023, 0.497 ± 0.018, 0.471 ± 0.009, 0.641 ± 0.011, 0.698 ± 0.022, 0.562 ± 0.010. The random order was Craigslist → Job Interview → DnD → CaSiNo → Calendar → Marketplace; the anti-transfer order was Job Interview → Marketplace → Craigslist → DnD → CaSiNo → Calendar. Final evaluations used 10 scenarios × 3 opponents × 5 trials (150 games per environment); reported transfer-aware scores are means ± standard deviations over two independent evaluations, and random-order scores are means ± standard deviations over repeated evaluations. The transfer-aware policy's 0.627 Avg-6 is comparable to GPT-4.1 (0.625), GPT-5.1 (0.619), and GPT-5.2 (0.613).

  4. Knowl 4 — Gap-closed multi-teacher distillation efficiently transfers specialist capabilities

    algorithm

    Multi-teacher on-policy distillation (MOPD) initializes a shared student from the Craigslist-SFT checkpoint and uses the Stage-1 domain specialists as teachers. The student generates on-policy interaction responses, and the corresponding teacher supplies token-level logits for reverse-KL distillation. The experiments use the teacher's top 32 logits, temperature 1.0, uniform loss over response tokens, and constant learning rate 10−510^{-5} without warmup. CaSiNo is excluded from the training mixture because the initial student scores 0.502 there and its specialist scores 0.503, leaving negligible teacher-student headroom. For each remaining training environment ee, let BeB_e be the student's pre-distillation score, TeT_e the specialist's score, and sˉe(t)\bar{s}^{(t)}_e the student's mean terminal reward over its latest W=500W=500 training games; initially sˉe(t)=Be\bar{s}^{(t)}_e=B_e. The teacher-student gap is ge=Te−Beg_e=T_e-B_e, and the fraction of that gap closed is ρe(t)=(sˉe(t)−Be)/ge\rho^{(t)}_e=(\bar{s}^{(t)}_e-B_e)/g_e. The unnormalized sampling weight is we(t)=λw^{(t)}_e=\lambda when ge<δg_e<\delta, and otherwise we(t)=max⁡(0,1−ρe(t))+λw^{(t)}_e=\max(0,1-\rho^{(t)}_e)+\lambda, with floor λ=0.05\lambda=0.05 and minimum-gap threshold δ=0.03\delta=0.03. Normalize the weights into probabilities and iteratively renormalize after clamping each environment's sampling probability to [0.10,0.40][0.10,0.40]. This gap-based curriculum reduces sampling from teachers whose advantage has already been absorbed while guarding against noisy weights for near-zero gaps and preventing any one environment from dominating or disappearing. After 60 optimization steps, gap-closed MOPD without CaSiNo achieved Avg-6 0.597 ± 0.015, mean gap closure 92.6 ± 11.3%, and clipped gap closure 84.5 ± 6.3%; equal sampling without CaSiNo achieved 0.594 ± 0.020, 84.4 ± 15.6%, and 79.7 ± 12.6%; equal sampling with CaSiNo achieved 0.595 ± 0.015, 82.2 ± 9.2%, and 77.6 ± 5.8%; gap-closed sampling with CaSiNo achieved 0.578 ± 0.012, 71.3 ± 8.7%, and 70.7 ± 7.9%. The initialization averaged 0.460 and the specialist teachers averaged 0.619. For the best checkpoint, scores in the order DnD, Craigslist, CaSiNo, Job Interview, Calendar, Marketplace were 0.664 ± 0.025, 0.556 ± 0.019, 0.478 ± 0.019, 0.551 ± 0.023, 0.564 ± 0.050, 0.771 ± 0.028; initial scores were 0.586, 0.397, 0.502, 0.456, 0.375, 0.444, and specialist scores were 0.656 ± 0.008, 0.583 ± 0.007, 0.503 ± 0.008, 0.594 ± 0.007, 0.540 ± 0.017, 0.838 ± 0.013. The unified student surpassed the DnD and Calendar specialists; its reported gap closure was 111% for DnD, 85% for Craigslist, 69% for Job Interview, 115% for Calendar, and 83% for Marketplace. CaSiNo was excluded from gap-closure aggregation. Evaluation scores are means ± standard deviations across five independent runs.

  5. Knowl 5 — Explicit theory-of-mind traces improve negotiation after training, not prompting alone

    empirical result

    SOCIALRL represents theory of mind (ToM) in a decision trace as INFER → ACT → ANTICIPATE: infer the counterpart's private preferences from the interaction, choose an action, and predict the counterpart's next action. Preference inference is evaluated by exact match to the counterpart's private preferences; next-action prediction is evaluated by accuracy of the predicted response. GPT-5.2 generates demonstrations containing both negotiation actions and this reasoning trace. Supervised fine-tuning on full traces (ExpToM SFT) is compared with ordinary action-only SFT, per-environment and jointly across the four negotiation environments; neither method uses reinforcement learning in this comparison. In the order DnD, CaSiNo, Craigslist, Job Interview, average utility, the base 4B scores were 0.589, 0.454, 0.295, 0.476, 0.454; adding the ToM prompt without training gave 0.575, 0.391, 0.225, 0.221, 0.353; ordinary SFT per environment gave 0.580, 0.492, 0.445, 0.482, 0.500; ordinary mixed SFT gave 0.587, 0.480, 0.409, 0.503, 0.495; ExpToM SFT per environment gave 0.596, 0.512, 0.552, 0.522, 0.546; mixed ExpToM SFT gave 0.608, 0.498, 0.478, 0.517, 0.525; GPT-5.2 scored 0.655, 0.475, 0.511, 0.626, 0.567. Thus, prompting the untrained small model with the scaffold reduced its average utility, while training on complete ToM traces exceeded action-only SFT in both training regimes and improved utility in every evaluated environment under per-environment training. The base model's preference-inference accuracy was 0.618, while its next-action prediction accuracy was about 54%; explicit training raised next-action prediction by roughly 30 percentage points and preference inference to approximately 0.63–0.71, depending on training environment. Across trained checkpoints, next-action prediction accuracy was positively correlated with negotiation utility, whereas preference-inference accuracy showed little corresponding relationship; this is an observed association, not evidence that next-action prediction alone causes the utility gains.

  6. Knowl 6 — The six environments use role- and scenario-sensitive terminal rewards

    definition

    The SOCIALRL suite covers six tasks with private objectives: DnD divides three item types using private per-item values; CaSiNo divides food, water, and firewood according to private priority orderings; Craigslist is buyer-seller bargaining over a listed price with private price objectives; Job Interview negotiates a five-issue employment contract using private issue utilities; Calendar coordinates a meeting slot using private preferences; Marketplace is buyer-seller bargaining with private reservation prices. Every episode receives one terminal, outcome-only reward in [0,1][0,1] after agreement, walk-away, or timeout, without intermediate reward shaping. For DnD and CaSiNo, a scenario defines item counts ckc_k and players' private per-item values vkav^a_k and vkbv^b_k; if player aa receives xkx_k units of item kk, their scores are sa=∑kxkvkas_a=\sum_k x_kv^a_k and sb=∑k(ck−xk)vkbs_b=\sum_k(c_k-x_k)v^b_k. Maximum raw scores are 10 in DnD and 36 in CaSiNo. Each role's reference mim_i is its greatest normalized score among allocations that are both Pareto-optimal and envy-free; if no allocation meets both conditions, the reference uses the Pareto-optimal allocation minimizing the larger envy violation. In Job Interview, the reference deal maximizes the lower of the worker's and recruiter's normalized utilities, with ties broken by greater total utility; each role's utility in that deal is its reference mm. For these three tasks, the agent's normalized outcome utility is z∈[0,1]z\in[0,1], the scenario- and role-specific reference is m∈[0,1]m\in[0,1], temperature is T=0.2T=0.2, and σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}). The reward is r=[σ((z−m)/T)−σ(−m/T)]/[σ((1−m)/T)−σ(−m/T)]r=[\sigma((z-m)/T)-\sigma(-m/T)]/[\sigma((1-m)/T)-\sigma(-m/T)], preserving reward 0 at z=0z=0 and 1 at z=1z=1 while concentrating reward discrimination near the reference. For price bargaining, let pmin⁡<pmax⁡p_{\min}<p_{\max} be the buyer-favorable and seller-favorable corridor endpoints and pdealp_{\mathrm{deal}} the agreed price. Buyer and seller rewards are respectively clip⁡((pmax⁡−pdeal)/(pmax⁡−pmin⁡),0,1)\operatorname{clip}((p_{\max}-p_{\mathrm{deal}})/(p_{\max}-p_{\min}),0,1) and clip⁡((pdeal−pmin⁡)/(pmax⁡−pmin⁡),0,1)\operatorname{clip}((p_{\mathrm{deal}}-p_{\min})/(p_{\max}-p_{\min}),0,1). Craigslist sets the endpoints to the buyer's target price and the listing price; Marketplace sets them to the seller's reservation price psp_s and buyer's reservation price pbp_b, where ps<pbp_s<p_b, and trains a buyer delegate. In Calendar, let ZZ be the mutually free time slots, v(t)∈[0,1]v(t)\in[0,1] the principal's value for slot tt, and tdeal∈Zt_{\mathrm{deal}}\in Z the scheduled slot. The reward is [v(tdeal)−min⁡t∈Zv(t)]/[max⁡t∈Zv(t)−min⁡t∈Zv(t)][v(t_{\mathrm{deal}})-\min_{t\in Z}v(t)]/[\max_{t\in Z}v(t)-\min_{t\in Z}v(t)]. This scores the principal's most preferred feasible slot 1 and the least preferred feasible slot 0, distinguishing a useful meeting from merely completing scheduling. Timeouts and aborted interactions score zero; explicit walk-away actions retain their task-specific outside-option payoff.

  7. Knowl 7 — Post-training makes concessions more sensitive to the principal's utility

    empirical result

    Trajectory analysis of the six environments found that post-trained policies more often protect private boundaries and concede selectively, rather than treating agreement itself as the goal. Craigslist buyer offers are normalized so 0 is the buyer's target price, 1 is the listing price, and negative values are below-target offers. The untrained, SFT, and SFT+PPO mean offers at opening, second, third, and fourth proposal were respectively +0.04, +0.35, +0.41, +0.45; −1.21, −0.30, −0.08, −0.06; and −1.48, −0.73, −0.22, +0.01. The share of openings below target rose from 3% to 71% after SFT and 78% after SFT+PPO. On Craigslist, explicit disclosure of the buyer's target fell from 52.7% of base-model messages to 1.0% after SFT; PPO further increased commitment language such as “final offer” or “firm limit” to 13.3% of messages. In DnD, PPO reduced zero-reward episodes from 4.7% to 2.5% and no-deals from 4.5% to 0.5% relative to the SFT diagnostic checkpoint. In CaSiNo, over-conceding fell from 17.3% to 13.6% under PPO, while strong preservation of the agent's own value rose from 17.3% to 24.2%; the SFT diagnostic instead had 29.4% over-concession. In Job Interview, the worker's categorical utility in completed agreements rose from approximately 0.64 for the base policy to 0.73–0.78 for trained checkpoints, and top-choice workplace outcomes rose from 39% to 68–72%. In Marketplace, the base buyer's mean opening offer was 0.956 of its private reservation price, compared with 0.361 after PPO; the base policy disclosed its reservation price or budget in 62% of opening messages, and this behavior essentially disappeared after training. In Calendar, the Marketplace-trained checkpoint used as a pre-Calendar proxy scheduled meetings in 92% of games, but only 16% selected a maximum-preference slot and 63% yielded zero utility. After Calendar PPO, meeting frequency fell to 74%, while maximum-utility outcomes rose to 38% and zero-utility outcomes fell to 35%. These measurements come from the paper's analyzed trajectories and demonstrate different task-specific forms of utility-sensitive concession.

  8. Knowl 8 — An event-based interface decouples multi-agent environments from agent implementations

    model/method

    SOCIALRL represents an interaction as an ordered stream delivered through private asynchronous channels, rather than as repeated full-state snapshots. A stateful environment owns the game rules, hidden state, legal actions, and reward computation, while each participant is a black box that receives context and available actions and returns an action. An observation reports an event that has occurred; a notification requests a decision and specifies the available actions. The environment drives sequential decisions with ask(), simultaneous decisions with ask_all(), and events requiring no response with broadcast(). Events are audience-filtered so each agent receives only information available to that participant. An action submission yields an acknowledgment or validation error; its effects are delivered separately as observations, allowing collective outcomes to depend on multiple actions or later environment logic. Agents can consume events locally as asynchronous iterators or remotely through a cursor-based HTTP interface backed by an append-only, sequence-indexed history, so disconnected clients can resume. Each action identifies the notification it answers: if it is stale because the environment has advanced, the environment rejects it and the agent processes newer events before acting again. Invalid actions can be retried, and a timeout applies an inert default action while recording the failure in the event stream rather than terminating the episode. This interface supports sequential or simultaneous decisions, asynchronous communication, and different agent implementations—including language models, scripted policies, external agents, and humans—without changing the environment.

  9. Knowl 9 — A rollout proxy creates a common data boundary for reinforcement learning and distillation

    model/method

    SOCIALRL separates environment and agent execution from optimization by placing an OpenAI-compatible rollout proxy between the agent runtime and the inference endpoint. The runtime interacts with the environment normally; the proxy forwards each model request and response unchanged while recording the authentic inputs and outputs. A session key groups otherwise stateless model calls into an episode, to which the runtime attaches the terminal reward when the interaction ends. If an agent truncates or summarizes its context, consecutive requests sharing the expected prefix form one training segment; when the prefix relation breaks, the proxy starts a new segment, preserving the actual context used for each response rather than inventing a reconstructed history. For PPO, the proxy records tokenized prompts and responses, token-level log probabilities from the generating policy, rewards, task identifiers, and policy versions. For distillation, it retains raw messages, tool definitions, responses, and tool calls so trajectories from heterogeneous or closed-source teachers can later be tokenized with the student tokenizer. Completed sessions are written to a shared Parquet buffer, and training losses apply only to model response tokens. Rollout collection, optimization, and inference proceed asynchronously; a capacity-based gate limits trajectories dispatched ahead of training, while policy-version tags and PPO importance-sampling correction address policy staleness. The common trajectory boundary allows interchangeable trainers and inference backends without requiring the environment or agent harness to contain training-specific logic.

  10. Knowl 10 — Negotiation training can produce fabricated market justifications

    limitation

    The Craigslist environment supplies no external comparable-price information, but SFT-trained policies began using market comparisons to justify below-target offers; market/comparable references appeared in 29.6% of messages after SFT, compared with 3.7% for the base policy. Because those references were unsupported by information available in the environment, the authors identify the learned behavior as fabricated market claims rather than improved factual grounding. This is a limitation of the observed bargaining strategy even though training improved negotiation utility.

Coverage note — The detailed stage-by-stage cascade checkpoint trajectory and secondary communication-style measurements were omitted because they diagnose training dynamics and expression rather than add a separate main result; the final cascade comparisons and principal behavioral changes are retained.

References

  1. 1.Amine Allouah, Omar Besbes, Josue D Figueroa, Yash Kanoria, and Akshit Kumar. What is your ai agent buying? evaluation, biases, model dependence, & emerging implications of agentic e-commerce. In Proceedings of the ACM Web Conference 2026, pp. 8697–8700, 2026.
  2. 2.Panatchakorn Anantaprayoon, Nataliia Babina, Nima Asgharbeygi, and Jad Tarifi. Learning to negotiate: Multi-agent deliberation for collective value alignment in llms. arXiv preprint arXiv:2603.10476, 2026.
  3. 3.Ahmed Awadallah, Yash Lara, Raghav Magazine, Hussein Mozannar, Akshay Nambi, Yash Pandya, Aravind Rajeswaran, Corby Rosset, Alexey Taymanov, Vibhav Vineet, et al. Fara-7b: An efficient agentic model for computer use. arXiv preprint arXiv:2511.19663, 2025.
  4. 4.Dirk Bergemann, Soheil Ghili, Xinyang Hu, Chuanhao Li, and Zhuoran Yang. Training language models for bilateral trade with private information. arXiv preprint arXiv:2604.16472, 2026.
  5. 5.Federico Bianchi, Patrick John Chia, Mert Yuksekgonul, Jacopo Tagliabue, Dan Jurafsky, and James Zou. How well can llms negotiate? negotiationarena platform and analysis. arXiv preprint arXiv:2402.05863, 2024.
  6. 6.Rijul Chaturvedi and Sanjeev Verma. Opportunities and challenges of ai-driven customer service. Artificial Intelligence in customer service: The next frontier for personalized engagement, pp. 33–71, 2023.
  7. 7.Kushal Chawla, Jaysa Ramirez, Rene Clever, Gale Lucas, Jonathan May, and Jonathan Gratch. Casino: A corpus of campsite negotiation dialogues for automatic negotiation systems. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 3167–3185, 2021.
  8. 8.Adam Dahlgren Lindstrom, Leila Methnani, Lea Krause, Petter Ericson, Ínigo Martínez de Rituerto de Troya, Dimitri Coelho Mollo, and Roel Dobbe. Helpful, harmless, honest? sociotechnical limits of ai alignment and safety through reinforcement learning from human feedback: Ad lindstrom et al. Ethics and Information Technology, 27(2):28, 2025.
  9. 9.Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang. Rlhf workflow: From reward modeling to online rlhf. arXiv preprint arXiv:2405.07863, 2024.
  10. 10.Yihan Du, R Srikant, and Wei Chen. Cascading reinforcement learning. In International Conference on Learning Representations, volume 2024, pp. 30263–30304, 2024.
  11. 11.Laïla Elkoussy and Julien Perez. Agentltl: A trace-verification framework for measuring, enforcing, and training procedural compliance in tool-using llm agents. arXiv preprint arXiv:2607.02599, 2026.
  12. 12.Siming Fu, Haojun Xu, Ruizhe He, Zheming Fu, Hualiang Wang, Jie Huang, Xiaoxiao Ma, Mingchen Zhong, Weihu Huang, Xiaoxuan He, et al. Poly-opd: Heterogeneous multi-teacher on-policy distillation for capability-selectable flow models. arXiv preprint arXiv:2608.04349, 2026.
  13. 13.Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata. Improving language model negotiation with self-play and in-context learning from ai feedback. arXiv preprint arXiv:2305.10142, 2023.
  14. 14.Google. Google scheduler. https://workspace.google.com/resources/appointment-scheduling/, 2026.
  15. 15.Kasper Raupach Haurum, Ruiqi Ma, and Wen Long. Real estate with ai: An agent based on langchain. Procedia Computer Science, 242:1082–1088, 2024.
  16. 16.He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. Decoupling strategy and generation in negotiation dialogues. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 2333–2343, 2018.
  17. 17.Wenyue Hua, Ollie Liu, Lingyao Li, Alfonso Amayuelas, Julie Chen, Lucas Jiang, Mingyu Jin, Lizhou Fan, Fei Sun, William Wang, et al. Game-theoretic llm: Agent workflow for negotiation games. arXiv preprint arXiv:2411.05990, 2024.
  18. 18.EunJeong Hwang, Yuwei Yin, Giuseppe Carenini, Peter West, and Vered Shwartz. Infusing theory of mind into socially intelligent llm agents. In Findings of the Association for Computational Linguistics: ACL 2026, pp. 11327–11360, 2026.
  19. 19.Bowen Jiang, Taiwei Shi, Ryo Kamoi, Yuan Yuan, Camillo J Taylor, Longqi Yang, Pei Zhou, and Sihao Chen. One model, all roles: Multi-turn, multi-agent self-play reinforcement learning for conversational social intelligence. arXiv preprint arXiv:2602.03109, 2026.
  20. 20.Gurusha Juneja, Jayanth Pasupulati, Alon Albalak, Wenyue Hua, and William Yang Wang. Magpie: a benchmark for multi-agent contextual privacy evaluation. In Socially Responsible and Trustworthy Foundation Models at NeurIPS 2025, 2025.
  21. 21.Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the effects of rlhf on llm generalisation and diversity. In International Conference on Learning Representations, volume 2024, pp. 20620–20653, 2024.
  22. 22.Adam Kostka and Jarosław A Chudziak. Evaluating theory of mind and internal beliefs in llm-based multi-agent systems. In International Conference on Computational Collective Intelligence, pp. 18–32. Springer, 2025.
  23. 23.Timo Kovala. A complete guide to agentforce. Springer Books, 2026.
  24. 24.Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. Deal or no deal? end-to-end learning of negotiation dialogues. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2443–2453, 2017.
  25. 25.Austen Liao, Nicholas Tomlin, and Dan Klein. Efficacy of language model self-play in non-zero-sum games. arXiv preprint arXiv:2406.18872, 2024.
  26. 26.Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao, Lei Lei, Yuheng Zhang, Yisong Yue, and David Simchi-Levi. Instructing llms to negotiate using reinforcement learning with verifiable rewards. arXiv preprint arXiv:2604.09855, 2026.
  27. 27.Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K Qiu, and Yuqing Yang. Agent lightning: Train any ai agents with reinforcement learning. arXiv preprint arXiv:2508.03680, 2025.
  28. 28.Nuoyan Lyu, Bingbing Xu, Xueyun Tian, Weihao Meng, Yige Yuan, Yang Zhang, Zhiyong Huang, Tat-Seng Chua, and Huawei Shen. Gift: Games as informal training for generalizable llms. arXiv preprint arXiv:2601.05633, 2026.
  29. 29.Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, et al. Mopd: Multi-teacher on-policy distillation for capability integration in llm post-training. arXiv preprint arXiv:2606.30406, 2026.
  30. 30.Daoud Matta. Artificial intelligence and theory of mind. Journal of Psychology and AI, 2(1):2628373, 2026.
  31. 31.Microsoft Research, AI Frontiers. Socialreasoning-bench: Measuring whether ai agents act in users’ best interests, 2026. https://www.microsoft.com/en-us/research/blog/socialreasoning-bench-measuring-whether-ai-agents-act-in-users-best-interests/.
  32. 32.Chunjiang Mu, Ya Zeng, Qiaosheng Zhang, Kun Shao, Chen Chu, Hao Guo, Danyang Jia, Zhen Wang, and Shuyue Hu. Adaptive theory of mind for llm-based multi-agent coordination. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pp. 29608–29616, 2026.
  33. 33.OpenAI. Operator system card. https://openai.com/index/operator-system-card, 2025.
  34. 34.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. Advances in neural information processing systems, 36:68539–68551, 2023.
  35. 35.Tobin South, Samuele Marro, Thomas Hardjono, Robert Mahari, Cedric Deslandes Whitney, Dazza Greenwood, Alan Chan, and Alex Pentland. Authenticated delegation and authorized ai agents. arXiv preprint arXiv:2501.09674, 2025.
  36. 36.Chuanneng Sun, Songjun Huang, and Dario Pompili. Llm-based multi-agent decision-making: Challenges and future directions. IEEE Robotics and Automation Letters, 10(6):5681–5688, 2025a.
  37. 37.Haoran Sun, Yusen Wu, Yukun Cheng, and Xu Chu. Game theory meets large language models: A systematic survey. arXiv preprint arXiv:2502.09053, 2025b.
  38. 38.Boxin Wang, Chankyu Lee, Nayeon Lee, Sheng-Chieh Lin, Wenliang Dai, Yang Chen, Yangyi Chen, Zhuolin Yang, Zihan Liu, Mohammad Shoeybi, et al. Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models. arXiv preprint arXiv:2512.13607, 2025.
  39. 39.Ruiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi, Maarten Sap, Yonatan Bisk, Graham Neubig, and Hao Zhu. Sotopia-π: Interactive learning of socially intelligent language agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12912–12940, 2024.
  40. 40.Yinjie Wang, Xuyang Chen, Xiaolong Jin, Mengdi Wang, and Ling Yang. Openclaw-rl: Train any agent simply by talking. arXiv preprint arXiv:2603.10165, 2026a.
  41. 41.Zhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang, Siwei Han, Zhewei Yao, Huaxiu Yao, and Yuxiong He. Agent world model: Infinity synthetic environments for agentic reinforcement learning. arXiv preprint arXiv:2602.10090, 2026b.
  42. 42.Yang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song, Chunpu Xu, Yi Cheng, Wenjie Li, and Pengfei Liu. Towards dynamic theory of mind: Evaluating llm adaptation to temporal evolution of human states. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 24036–24057, 2025.
  43. 43.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37:52040–52094, 2024.
  44. 44.Atsuki Yamaguchi, Kosui Iwasa, and Katsuhide Fujita. Dialogue act-based breakdown detection in negotiation dialogues. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 745–757, 2021.
  45. 45.Jiachen Yang and Jipeng Zhang. A multi-teacher policy distillation framework for enhancing zero-shot generalization of autonomous driving policies. IEEE Transactions on Vehicular Technology, 73(7):9734–9746, 2024.
  46. 46.Shuo Yang, Caren Han, Xueqi Ma, Yan Li, Mohammad Reza Ghasemi Madani, and Eduard Hovy. Evotool: Self-evolving tool-use policy optimization in llm agents via blame-aware mutation and diversity-aware selection. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 43553–43572, 2026a.
  47. 47.Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, and Yankai Lin. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation. arXiv preprint arXiv:2602.12125, 2026b.
  48. 48.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022.
  49. 49.Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. tau-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045, 2024.
  50. 50.Haofei Yu, Zhengyang Qi, Yining Zhao, Kolby Nottingham, Keyang Xuan, Bodhisattwa Prasad Majumder, Hao Zhu, Paul Pu Liang, and Jiaxuan You. Sotopia-rl: Reward design for social intelligence. arXiv preprint arXiv:2508.03905, 2025.
  51. 51.Huining Yuan, Zelai Xu, Zheyue Tan, Xiangmin Yi, Mo Guang, Kaiwen Long, Haojia Hui, Boxun Li, Xinlei Chen, Bo Zhao, et al. Mars: Reinforcing multi-agent reasoning of llms through self-play in strategic games. arXiv e-prints, pp. arXiv–2510, 2025.
  52. 52.Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, Jiajie Zhang, et al. Glm-4.5: Agentic, reasoning, and coding (arc) foundation models. arXiv preprint arXiv:2508.06471, 2025.
  53. 53.Erica Zhang, Fangzhao Zhang, Aneesh Pappu, Batu El, Jose Blanchet, Susan Athey, Jiashuo Liu, and James Zou. Terms-bench: Diagnosing llm negotiation agents beyond deal rate. arXiv preprint arXiv:2605.13909, 2026.
  54. 54.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023.
  55. 55.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, volume 2024, pp. 15585–15606, 2024.
  56. 56.Chris Zhu, Sasha Cui, Will Sanok Dufallo, Runzhi Jin, Zhen Xu, Linjun Zhang, and Daylian Cain. Piearena: Ranking and profiling language agents in realistic negotiation scenarios. arXiv preprint arXiv:2602.05302, 2026.
  57. 57.Shenzhe Zhu, Jiao Sun, Yi Nian, Tobin South, Alex Pentland, and Jiaxin Pei. The automated but risky game: Modeling and benchmarking agent-to-agent negotiations and transactions in consumer markets. arXiv preprint arXiv:2506.00073, 2025.
  58. 58.Zillow. Zillow debuts ai mode. https://www.zillow.com/news/zillow-debuts-ai-mode, 2026.
  59. 59.Chelsea Zou, Yiheng Yao, Selena She, Noah Goodman, and Robert D Hawkins. Calbench: Evaluating coordination-privacy trade-offs in multi-agent llms. arXiv preprint arXiv:2605.09823, 2026.

Citation

MLA
Hua, W., et al. “From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL”. arXiv, 2026, http://arxiv.org/abs/2608.13787v1.
APA
Hua, W., Huang, Z., Payne, T., Yousefi, S., Amershi, S., & Celikyilmaz, A. (2026). From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL. arXiv. http://arxiv.org/abs/2608.13787v1
Chicago
Hua, W., Z. Huang, T. Payne, S. Yousefi, S. Amershi, and A. Celikyilmaz. 2026. “From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL”. arXiv. http://arxiv.org/abs/2608.13787v1.
Harvard
Hua, W. et al. (2026) “From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2608.13787v1.
Vancouver
1. Hua W, Huang Z, Payne T, Yousefi S, Amershi S, Celikyilmaz A (2026) From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL. arXiv

BibTeX

@article{hua2026from,
  title = {From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL},
  author = {Hua, Wenyue and Huang, Zachary and Payne, Tyler and Yousefi, Safoora and Amershi, Saleema and Celikyilmaz, Asli},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2608.13787v1},
  eprint = {2608.13787}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/