ProsocialDialog: A Prosocial Backbone for Conversational Agents
Hyunwoo KimYoungjae YuLiwei JiangXiming LuDaniel KhashabiGunhee KimYejin ChoiMaarten Sap
Presents a large-scale multi-turn dataset and models grounded in commonsense social rules to train conversational agents that actively guide users toward prosocial behavior instead of passively agreeing with toxic or unethical inputs.
Conversational artificial intelligence systems frequently fail when interacting with users who introduce toxic, unethical, rude, or dangerous content. Because current chatbots are trained on predominantly agreeable and positive data, they often condone or validate unsafe user remarks. Existing safety mitigations typically rely on mechanical avoidance, such as offering canned deflections or shutting down sensitive topics altogether, which disrupts dialogue flow and can unintentionally marginalize benign discussions. The article addresses this operational and ethical problem by developing methods to teach conversational agents to constructively push back against problematic inputs using social norms.
The main objective of the article is to introduce PROSOCIALDIALOG, a multi-turn dialogue dataset grounded in commonsense social rules, and to demonstrate how these data can be used to detect unsafe conversational contexts and train dialogue systems that respond prosocially. The researchers evaluate two newly developed components: Canary, a safety detection module that infers relevant social rules, and Prost, a conversational model designed to produce constructive, socially responsible dialogue.
To build this resource, the authors employed a human-AI collaborative data collection approach across 58,137 dialogues comprising 331,362 utterances. A large language model drafted problematic conversational scenarios sourced from established morality and bias benchmarks, while crowdworkers proofread exchanges, selected or authored relevant social rules (rules-of-thumb), and wrote constructive, empathetic responses. The dataset incorporates a three-tier safety labeling framework classifying dialogue into casual remarks, contexts needing caution, and severe situations requiring human intervention. Using this dataset along with standard dialogue corpora, the authors trained Canary to predict safety labels and generate social rules, and trained Prost to produce norm-grounded responses.
The key findings show that the proposed models substantially improve safety handling over existing systems. First, in human evaluations, Prost generating responses grounded in rules-of-thumb significantly outperformed standard models like GPT-3, with annotators preferring Prost’s prosociality over GPT-3 by 63.4% to 9.3%. Second, Canary achieved 77.1% accuracy on safety classification and improved social rule generation over standard baselines. Third, in zero-shot evaluations on real-world Reddit toxicity from ToxiChat, Prost produced significantly higher rates of explicit disagreement with toxic content (up to 38.7%) compared to systems like BlenderBot 1 (14.0%) and GPT-3 (11.2%), which frequently agreed with harmful prompts. Finally, prompting off-the-shelf language models with Canary-generated social rules doubled or tripled human preference scores for prosociality and overall quality, closing the performance gap between older base models and instruction-tuned systems.
These findings indicate that integrating explicit social commonsense into dialogue pipelines reduces the risk of models condoning harm without relying on evasive, scripted avoidance. By separating the safety module (Canary) from the conversational backbone (Prost), organizations can update social rules and safety criteria dynamically without retraining entire dialogue systems. Furthermore, the three-tier classification schema provides a clear operational mechanism to escalate critical real-world dangers—such as self-harm or medical emergencies—directly to human intervention rather than relying solely on automated text generation.
Organizations developing or deploying conversational agents should adopt multi-turn safety training that teaches models to constructively address problematic content rather than simply deflecting. Systems should incorporate modular safety layers that ground responses in explicit social rules and establish human escalation pathways for high-risk situations. Future technical work should focus on diversifying social rules beyond predominantly English-speaking, North American cultural norms and reducing instances where the safety module produces irrelevant rules or misclassifies casual exchanges.
The study’s primary limitations stem from the demographic profile of the crowdworkers, who were mostly white, liberal-leaning residents of the United States. Consequently, the captured social norms may reflect majority viewpoints and fail to encompass cross-cultural differences. While confidence in the empirical improvements on the evaluated benchmarks is high, practitioners should exercise caution and conduct domain-specific testing before deploying these models in diverse global settings.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Its automated red-teaming study establishes a safety-testing approach that helps frame ProsocialDialog’s move from detecting harmful responses toward teaching safer ones.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Its benchmark of toxic generation and steering methods provides useful groundwork for ProsocialDialog’s focus on responses to unsafe conversational content.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). Its human-feedback approach to training helpful and harmless assistants clarifies the alignment strategy that ProsocialDialog complements with prosocial dialogue data.
- Paper: ELEPHANT: Measuring and understanding social sycophancy in LLMs, Myra Cheng et al. (2025). It extends the concern about uncritical agreement with users into a benchmark of social sycophancy, including failures to challenge harmful framing and behavior.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). It carries conversational safety into an operational safeguard that classifies both user prompts and model responses against explicit risk categories.
