From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL
Wenyue HuaZezhou HuangTyler PayneSafoora YousefiSaleema AmershiAsli Celikyilmaz
Demonstrates that reinforcement learning with theory-of-mind distillation enables a 4-billion-parameter language model to match or outperform frontier models across six multi-issue negotiation benchmarks while resisting premature concessions and privacy leaks.
Artificial intelligence agents are increasingly deployed as delegates to handle consequential tasks on behalf of users, including purchasing items, negotiating contracts, and coordinating schedules. While general-purpose frontier models excel at cooperative assistance, their default dispositions—such as extreme agreeableness, eagerness to reach consensus, and transparency—often lead them to concede ground too quickly and leak private information when facing counterparts with conflicting goals. The article evaluates whether targeted post-training in social reasoning can transform a compact 4-billion-parameter language model into an effective, value-preserving delegate across complex strategic environments.
To address this challenge, the authors developed a decoupled multi-agent training framework and evaluated agents across six diverse interaction domains, including bilateral price bargaining, multi-issue resource allocation, employment contract negotiation, and meeting scheduling. The training approach first created domain-specific specialists using reinforcement learning with outcome-based rewards calibrated to the specific difficulty of each scenario. The authors then analyzed cross-domain transfer dynamics and evaluated two consolidation strategies to merge specialist capabilities into a single model: a transfer-aware sequential cascade of reinforcement learning and multi-teacher on-policy distillation. The study also examined explicit theory-of-mind supervision by training agents to infer counterpart preferences, act, and anticipate counterpart reactions.
The findings show that targeted post-training enables a compact 4-billion-parameter model to reach an aggregate utility of 0.619, matching or exceeding much larger models such as GPT-4.1 (0.625) and GPT-5.1 (0.619). Cross-domain skill transfer proved to be highly structured and directional: structurally paired tasks, such as price bargaining environments, transferred substantially to one another, while multi-issue bargaining provided broad positive transfer across multiple domains. Leveraging this transfer structure, a transfer-aware cascade consolidation reached an overall utility of 0.627, while multi-teacher distillation recovered 92.6% of the specialists' advantage in only 60 additional training steps. At the behavioral level, trained agents eliminated premature information disclosure, reduced target price leakage from over 50% down to 1%, and adopted strategic anchoring and selective concession. Explicit theory-of-mind distillation further boosted performance, with the ability to anticipate a counterpart's next move proving to be the single most critical driver of negotiation success.
These results indicate that effective delegation does not require massive frontier-scale models, offering organizations a path to deploy cost-effective, specialized delegates that robustly defend user utility. Stakeholders seeking to build strategic agents should prioritize targeted post-training pipelines over basic prompting, adopting multi-teacher distillation for compute-efficient consolidation and scheduling sequential learning according to task transfer dynamics. Training routines should also incorporate next-action prediction supervision to reinforce strategic foresight.
Certain limitations should be considered before widespread deployment. The study observed instances where bargaining agents generated fabricated market comparisons to justify aggressive offers, highlighting the need to enforce factual grounding. Furthermore, the findings reflect performance within structured simulation environments and benchmark games. Stakeholders can have high confidence in the framework's core mechanics, but organizations should conduct human-facing pilot testing before deploying autonomous delegates in high-stakes commercial applications.
- Paper: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning, A list of authors and their affiliations appears at the end of the paper (2025). Its reinforcement-learning recipe for eliciting new behaviors from outcome rewards provides a useful foundation for understanding SocialRL’s post-training approach.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). Its controlled Theory-of-Mind framework clarifies the mental-state inference that SocialRL trains and tests in strategic interactions.
- Paper: Theory of Mind for Multi-Agent Collaboration via Large Language Models, Huao Li et al. (2023). Its evaluation of agents’ belief and intention inference in multi-agent settings prepares readers for SocialRL’s use of social reasoning in interactive tasks.
- Paper: Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs, Xuhui Zhou et al. (2024). Its demonstration that omniscient simulations overstate social competence highlights why SocialRL’s agents must reason under counterpart-specific information and goals.
No sufficiently relevant recommendations were found.
