ProsocialDialog: A Prosocial Backbone for Conversational Agents

Hyunwoo KimYoungjae YuLiwei JiangXiming LuDaniel KhashabiGunhee KimYejin ChoiMaarten Sap

article2022EMNLP171 citations

Presents a large-scale multi-turn dataset and models grounded in commonsense social rules to train conversational agents that actively guide users toward prosocial behavior instead of passively agreeing with toxic or unethical inputs.

Listen

Conversational artificial intelligence systems frequently fail when interacting with users who introduce toxic, unethical, rude, or dangerous content. Because current chatbots are trained on predominantly agreeable and positive data, they often condone or validate unsafe user remarks. Existing safety mitigations typically rely on mechanical avoidance, such as offering canned deflections or shutting down sensitive topics altogether, which disrupts dialogue flow and can unintentionally marginalize benign discussions. The article addresses this operational and ethical problem by developing methods to teach conversational agents to constructively push back against problematic inputs using social norms.

The main objective of the article is to introduce PROSOCIALDIALOG, a multi-turn dialogue dataset grounded in commonsense social rules, and to demonstrate how these data can be used to detect unsafe conversational contexts and train dialogue systems that respond prosocially. The researchers evaluate two newly developed components: Canary, a safety detection module that infers relevant social rules, and Prost, a conversational model designed to produce constructive, socially responsible dialogue.

To build this resource, the authors employed a human-AI collaborative data collection approach across 58,137 dialogues comprising 331,362 utterances. A large language model drafted problematic conversational scenarios sourced from established morality and bias benchmarks, while crowdworkers proofread exchanges, selected or authored relevant social rules (rules-of-thumb), and wrote constructive, empathetic responses. The dataset incorporates a three-tier safety labeling framework classifying dialogue into casual remarks, contexts needing caution, and severe situations requiring human intervention. Using this dataset along with standard dialogue corpora, the authors trained Canary to predict safety labels and generate social rules, and trained Prost to produce norm-grounded responses.

The key findings show that the proposed models substantially improve safety handling over existing systems. First, in human evaluations, Prost generating responses grounded in rules-of-thumb significantly outperformed standard models like GPT-3, with annotators preferring Prost’s prosociality over GPT-3 by 63.4% to 9.3%. Second, Canary achieved 77.1% accuracy on safety classification and improved social rule generation over standard baselines. Third, in zero-shot evaluations on real-world Reddit toxicity from ToxiChat, Prost produced significantly higher rates of explicit disagreement with toxic content (up to 38.7%) compared to systems like BlenderBot 1 (14.0%) and GPT-3 (11.2%), which frequently agreed with harmful prompts. Finally, prompting off-the-shelf language models with Canary-generated social rules doubled or tripled human preference scores for prosociality and overall quality, closing the performance gap between older base models and instruction-tuned systems.

These findings indicate that integrating explicit social commonsense into dialogue pipelines reduces the risk of models condoning harm without relying on evasive, scripted avoidance. By separating the safety module (Canary) from the conversational backbone (Prost), organizations can update social rules and safety criteria dynamically without retraining entire dialogue systems. Furthermore, the three-tier classification schema provides a clear operational mechanism to escalate critical real-world dangers—such as self-harm or medical emergencies—directly to human intervention rather than relying solely on automated text generation.

Organizations developing or deploying conversational agents should adopt multi-turn safety training that teaches models to constructively address problematic content rather than simply deflecting. Systems should incorporate modular safety layers that ground responses in explicit social rules and establish human escalation pathways for high-risk situations. Future technical work should focus on diversifying social rules beyond predominantly English-speaking, North American cultural norms and reducing instances where the safety module produces irrelevant rules or misclassifies casual exchanges.

The study’s primary limitations stem from the demographic profile of the crowdworkers, who were mostly white, liberal-leaning residents of the United States. Consequently, the captured social norms may reflect majority viewpoints and fail to encompass cross-cultural differences. While confidence in the empirical improvements on the evaluated benchmarks is high, practitioners should exercise caution and conduct domain-specific testing before deploying these models in diverse global settings.

Cover for ProsocialDialog: A Prosocial Backbone for Conversational Agents

Abstract

Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them. To address this issue, we introduce ProsocialDialog, the first large-scale multi-turn dialogue dataset to teach conversational agents to respond to problematic content following social norms. Covering diverse unethical, problematic, biased, and toxic situations, ProsocialDialog contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-of-thumb, RoTs). Created via a human-AI collaborative framework, ProsocialDialog consists of 58K dialogues, with 331K utterances, 160K unique RoTs, and 497K dialogue safety labels accompanied by free-form rationales.

With this dataset, we introduce a dialogue safety detection module, Canary, capable of generating RoTs given conversational context, and a socially-informed dialogue agent, Prost. Empirical results show that Prost generates more socially acceptable dialogues compared to other state-of-the-art language and dialogue models in both in-domain and out-of-domain settings. Additionally, Canary effectively guides off-the-shelf language models to generate significantly more prosocial responses. Our work highlights the promise and importance of creating and steering conversational AI to be socially responsible.

Table of Contents

  • 1 Introduction
  • 2 Prosociality and Receptiveness in Conversational Agents
  • 2.1 Prosocial Responses with Rules-of-thumb
  • 2.2 Improving Receptiveness in Dialogues
  • 2.3 Fine-grained and Inclusive Safety Labeling
  • 2.4 Whose Prosociality Is It Anyway?
  • 3 PROSOCIALDIALOG
  • 3.1 Collecting Problematic Situations
  • 3.2 Collecting Dialogues
  • 3.3 Collecting Dialogue Safety Labels
  • 3.4 Analysis of PROSOCIALDIALOG
  • 4 Building Socially Responsible Dialogue Agents with PROSOCIALDIALOG
  • 4.1 Canary: A Dialogue Safety Detection Model Generating RoTs
  • 4.2 Prost: A Prosocial Dialogue Agent Grounded in RoTs
  • 5 Experiments on PROSOCIALDIALOG
  • 5.1 Dialogue Safety Classification & Rule-of-thumb Generation
  • 5.2 Response Generation via Prost
  • 6 Generalizability of Prost and Canary
  • 6.1 Generalizing to Real-world Toxic Phrases
  • 6.2 Improving Prosociality of Pre-trained Language Models with Canary
  • 7 Related Work
  • 8 Conclusion
  • 9 Societal and Ethical Considerations
  • 10 Limitations
  • 11 Acknowledgement
  • References
  • A Details of Constructing PROSOCIALDIALOG
  • A.1 Collecting Problematic Situations
  • A.2 Drafting Dialogue Openers
  • A.3 Collecting Dialogues
  • A.4 Collecting Dialogue Safety Labels
  • A.5 Additional Dataset Statistics
  • A.6 Worker Statistics
  • B Details of Model Training
  • B.1 Canary
  • B.2 Prost
  • B.3 Details of Training Computation
  • C Details of Experiments
  • C.1 Dialogue Safety Classification
  • C.2 Rule-of-thumb Generation
  • C.3 Response Generation Details of human evaluation.
  • D Details of zero-shot experiments
  • D.1 Generalizing to Real-world Toxic Phrases via Prost
  • D.2 Improving Prosociality of Pre-trained Language Models with Canary
  • E Dialogue Dataset Descriptions

Knowls

  1. Knowl 1 — ProsocialDialog Dataset Specification and Statistics

    data/table

    ProsocialDialog is a large-scale multi-turn English dialogue dataset designed to train conversational agents to respond to problematic, unethical, biased, and toxic utterances in a prosocial manner grounded in commonsense social rules, called rules-of-thumb (RoTs).

    The dataset comprises:

    • 58,13758{,}137 dialogues (42,30442{,}304 train, 7,1327{,}132 validation, 8,7018{,}701 test)
    • 331,362331{,}362 utterances (average of 5.75.7 turns per dialogue, average utterance length of 20.020.0 words)
    • 160,295160{,}295 unique rules-of-thumb (74%74\% uniqueness rate across 217,321217{,}321 total RoTs; average RoT length of 9.59.5 words; average of 3.33.3 RoTs per dialogue)
    • 497,043497{,}043 safety annotations with accompanying free-form text rationales (3 human annotations per dialogue context, with inter-annotator agreement of Krippendorff's α=0.49\alpha = 0.49).

    The conversation openings are sourced from three morality-related corpora:

    1. Social Chemistry 101 (62%62\% of initial situations), filtering for first-person actor situations with pressure against or strong pressure for an action (36k36\text{k} filtered situations).
    2. Social Bias Inference Corpus (SBIC) (21%21\% of initial situations), selecting toxic/stereotypical posts with explicit implied biases and length between 1010 and 4040 words (12k12\text{k} filtered posts).
    3. ETHICS Commonsense Morality subset (17%17\% of initial situations), extracting first-person daily scenarios unambiguously labeled as wrong (9.7k9.7\text{k} filtered scenarios).
  2. Knowl 2 — Action-Oriented Three-Tier Safety Classification Schema

    definition

    Rather than labeling dialogue context toxicity statically, the safety schema defines categories based on the required operational response of the conversational agent:

    1. Casual: Non-problematic everyday conversational situations, chitchat, empathetic interactions, civil disagreements, or contexts where the interlocutor has acknowledged past misbehavior and committed to prosocial change. No corrective feedback is required.
    2. Needs Caution: Utterances that are unethical, rude, biased, toxic, lewd, or offensive (violating social norms or commonsense), requiring the agent to push back with constructive, prosocial feedback grounded in rules-of-thumb (RoTs).
    3. Needs Intervention: Contexts involving imminent danger, emergency, illegal actions causing harm, self-harm or suicide, medical crises, threats of violence, or terrorism, requiring immediate human intervention (e.g., calling 911) rather than open-ended chatbot generation.

    To preserve annotator diversity and capture subjective ambiguity across 3 crowd annotations per context, a 5-class context label is constructed:

    • Casual: 3 out of 3 votes for Casual.
    • Possibly Needs Caution: Exactly 1 vote for Needs Caution and 0 votes for Needs Intervention.
    • Probably Needs Caution: Exactly 2 votes for Needs Caution and 0 votes for Needs Intervention.
    • Needs Caution: 3 out of 3 votes for Needs Caution.
    • Needs Intervention: At least 1 vote for Needs Intervention (prioritizing safety recall for emergency situations).
  3. Knowl 3 — Human-AI Collaborative Data Collection Framework for Prosocial Dialogues

    model/method

    To avoid subjecting human crowdworkers to generating toxic content while securing high-quality prosocial feedback, dialogues are generated via a human-AI collaborative pipeline:

    1. Problematic Situation Seeding: Filtered situations from Social Chemistry, ETHICS, and SBIC serve as the starting point.
    2. GPT-3 Dialogue Opening: GPT-3 is prompted in few-shot self-chat to draft the first three turns:
      • Utterance 1: The problematic situation converted into a first-person statement.
      • Utterance 2: An inquisitive, reflective question that rephrases Utterance 1 without premature hostility.
      • Utterance 3: GPT-3's response elaborating on the problematic behavior or bias.
    3. Human Constructive Feedback & RoT Grounding: Mechanical Turk annotators select or write 1–2 rules-of-thumb (RoTs) for problematic contexts and author constructive responses conforming to five conversational receptiveness guidelines: (a) grounding response in selected RoTs; (b) gently advising prosocial behavior; (c) illustrating positive outcomes of socially accepted actions; (d) persuading rather than commanding; and (e) showing empathy.
    4. Multi-Turn Extension: The conversation is fed back to GPT-3 to obtain the next turn, followed by a second round of human prosocial response annotation (up to 6 dialogue turns total).
    5. Proofreading and Validation: Crowdworkers proofread and revise prior turns to ensure conversational consistency (modifying an average of 1.11.1 and 1.71.7 utterances per dialogue in rounds 1 and 2). Two subsequent validation rounds flag accusatory, rude, or incoherent turns for re-annotation.
  4. Knowl 4 — Canary: Dialogue Safety Detection and Rule-of-Thumb Generator

    model/method

    Canary is a sequence-to-sequence safety module based on the T5-large architecture that jointly predicts the safety label ss and generates relevant rules-of-thumb rr given dialogue context cc:

    p(s,r∣c)p(s, r \mid c)

    For non-casual contexts, the target generation text concatenates a special safety classification token with comma-separated RoTs (e.g., __needs_caution__ It is wrong to call 911 just for fun.). For Casual contexts, the target text consists solely of the safety token __casual__.

    Canary is pre-trained on commonsense moral reasoning datasets—with Delphi (trained on 1.7M instances from the Commonsense Norm Bank) yielding the strongest performance, alongside variants pre-trained on Social Chemistry or the Moral Integrity Corpus (MIC). Canary is multi-task fine-tuned using the Adam optimizer (initial learning rate 1×10−51\times 10^{-5}, batch size 24, stopping when validation perplexity does not improve for 5 epochs) across ProsocialDialog and casual conversational corpora with dataset sampling ratio: ProsocialDialog:DailyDialog:EmpatheticDialogues:BlendedSkillTalk=4:1:1:1\text{ProsocialDialog} : \text{DailyDialog} : \text{EmpatheticDialogues} : \text{BlendedSkillTalk} = 4 : 1 : 1 : 1

  5. Knowl 5 — Prost: Prosocial Dialogue Agent Grounded in Rules-of-Thumb

    model/method

    Prost (Prosocial Transformer) is a conversational agent based on the PushShift Transformer (2.7B parameters, 2 encoder layers, 24 decoder layers, 2560 embedding dimensions, 32 attention heads, pre-trained on 1.5B Reddit examples).

    Prost is trained with Maximum Likelihood Estimation (MLE) under two target configurations:

    1. Direct response generation: p(u∣c)p(u \mid c), where uu is the response utterance.
    2. Chain-of-thought RoT-grounded response generation: p(u,r∣c)p(u, r \mid c), where the model sequentially outputs the relevant rule-of-thumb rr followed by the prosocial guiding response uu.

    To ensure the agent does not become overwhelmingly negative while maintaining the ability to counter toxic inputs, Prost is multi-task trained using Adam (initial learning rate 1×10−51\times 10^{-5}, 100 warm-up steps, batch size 32, ≈150K\approx 150\text{K} steps) with the dataset weighting: ProsocialDialog:DailyDialog:TopicalChat:PersonaChat:Wizard of Wikipedia:EmpatheticDialogues:BlendedSkillTalk=9:3:3:3:3:3:1\text{ProsocialDialog} : \text{DailyDialog} : \text{TopicalChat} : \text{PersonaChat} : \text{Wizard of Wikipedia} : \text{EmpatheticDialogues} : \text{BlendedSkillTalk} = 9 : 3 : 3 : 3 : 3 : 3 : 1

  6. Knowl 6 — Performance of Canary on Safety Classification and RoT Generation

    empirical result

    Canary pre-trained on Delphi achieves superior performance in both dialogue safety classification and RoT text generation on the ProsocialDialog benchmark compared to fine-tuned baseline classifiers and vanilla generation models.

    Model Safety Acc. (%) RoT Generation (Test)
    Valid Test BLEU-4 F1 Perplexity
    BAD classifier 72.2 72.1 – – –
    BERT 73.1 72.8 – – –
    NormTransformer – – 10.2 36.1 8.6
    DialoGPT – – 10.0 32.1 8.7
    GPT-2 69.3 68.4 9.6 32.3 8.8
    T5-large 72.4 73.4 16.1 38.9 5.9
    Canary (Social Chemistry) 73.5 73.1 16.3 39.2 5.4
    Canary (MIC) 74.1 74.0 16.2 41.2 5.3
    Canary (Delphi) 77.9 77.1 16.5 43.3 5.3

    These results show that transferring general commonsense moral reasoning representations from Delphi enhances both the detection of unsafe multi-turn dialogue contexts and the relevance and fluency of generated rules-of-thumb.

  7. Knowl 7 — In-Domain Response Generation Performance of Prost

    empirical result

    On the ProsocialDialog test split, grounding dialogue generation in rules-of-thumb improves automatic metric scores and human evaluation preferences over response-only models and general pre-trained language models.

    Automated evaluation metrics:

    • Prost (Response only): BLEU-4 = 3.983.98, F1 = 30.3030.30, Perplexity = 6.316.31
    • Prost (RoT & Response): BLEU-4 = 4.134.13, F1 = 31.1331.13, Perplexity = 6.226.22
    • Prost (Response w/ gold RoT): BLEU-4 = 4.514.51, F1 = 32.7832.78, Perplexity = 6.166.16

    Head-to-head crowdworker evaluation (N=400N=400 sampled test dialogues):

    Comparison Prosocial Engaged Respectful Coherent Overall
    Prost (Response only) 12.9% 12.7% 10.9% 12.7% 21.9%
    Tie 69.8% 70.7% 79.3% 71.6% 48.3%
    Prost (RoT Response) 17.1% 16.4% 9.7% 15.6% 29.6%
    GPT-3 9.3% 12.7% 11.0% 3.1% 10.7%
    Tie 27.3% 37.2% 65.4% 54.4% 14.1%
    Prost (RoT Response) 63.4% 50.1% 23.7% 42.5% 75.2%
    Instruct GPT-3 11.9% 21.3% 12.2% 6.9% 20.2%
    Tie 36.2% 36.5% 69.1% 65.2% 20.7%
    Prost (RoT Response) 51.9% 42.3% 18.8% 27.9% 59.1%
  8. Knowl 8 — Zero-Shot Response Stance and Disagreement on Real-World Toxic Text

    empirical result

    When evaluated zero-shot on 2,000 real-world offensive Reddit conversations from ToxiChat, Prost models demonstrate a substantially higher rate of active disagreement with toxic comments and lower agreement rates compared to open-domain chatbots.

    Model Disagree (%) ↑\uparrow Agree (%) ↓\downarrow Offense (%) ↓\downarrow Bad N-grams (%) ↓\downarrow
    DialoGPT 6.6 13.8 29.6 5.6
    BlenderBot 1 (3B) 14.0 24.2 19.6 7.8
    BlenderBot 2 (3B) 2.0 2.7 12.7 5.3
    GPT-3 11.2 18.6 41.0 26.6
    Instruct GPT-3 3.3 6.7 2.7 6.7
    Prost (Response only) 14.8 7.3 6.0 4.7
    Prost (RoT Response) 38.7 4.6 19.3 13.3

    While newer models such as BlenderBot 2 and Instruct GPT-3 evade toxicity by outputting predominantly neutral responses (95.3%95.3\% and 90.0%90.0\% respectively), Prost (RoT & Response) actively counters toxic claims (38.7%38.7\% disagreement). The elevated offense score (19.3%19.3\%) in Prost (RoT & Response) is largely an artifact of negations in counterspeech (e.g., negations occur in 88%88\% of its outputs), which automated classifiers frequently misclassify as offensive due to lexical overlap.

  9. Knowl 9 — Steering Pretrained Language Models with Canary-Generated Rules-of-Thumb

    model/method

    Canary-generated rules-of-thumb can steer general pre-trained language models (PLMs) such as GPT-3 and Instruct GPT-3 toward prosocial behaviors in a zero-shot setting.

    Given dialogue context cc, Canary first samples candidate RoTs rr when the context is classified as non-casual. The prompt fed to the PLM is structured as: Pr="The following is a conversation between Speaker 1 and Speaker 2. Speaker 2 is trying to gently explain {r}.\n\n Speaker 1: {c}\n Speaker 2:"\mathcal{P}_r = \text{"The following is a conversation between Speaker 1 and Speaker 2. Speaker 2 is trying to gently explain } \{r\}\text{.}\backslash\text{n}\backslash\text{n } \text{Speaker 1: } \{c\} \backslash\text{n } \text{Speaker 2:"}

    In human evaluations on 600 non-casual test contexts:

    • Incorporating Canary RoTs into GPT-3 increases prosociality and overall preference by 2×2\times to 3×3\times compared to vanilla GPT-3 prompts.
    • GPT-3 steered by Canary performs on par overall (41.7%41.7\% vs 39.5%39.5\%) and higher in prosociality (37.2%37.2\% vs 27.7%27.7\%) compared to unguided Instruct GPT-3, effectively bridging the alignment gap without instruction fine-tuning.
    • Control experiments confirm that providing relevant RoTs from Canary is critical: GPT-3 with Canary RoTs is preferred 55.7%55.7\% of the time over GPT-3 prompted with irrelevant or random RoTs (28.4%28.4\%).
  10. Knowl 10 — Demographic, Cultural, and Behavioral Limitations of Prosocial Modeling

    limitation

    Several intrinsic limitations characterize the ProsocialDialog framework and models:

    1. Demographic and Cultural Specificity: The annotator pool consists of 212 Mechanical Turk workers based in the US (97%>1097\% > 10 years) and Canada, skewing heavily White (73%73\%), liberal-leaning (62%62\% vs. 20%20\% conservative), and non-religious (62%62\%). Consequently, annotated rules-of-thumb reflect Western/US dominant cultural norms rather than global consensus or objective moral truth.
    2. Potential Negativity Bias: Because ProsocialDialog deliberately enriches negative responses to counter toxic inputs, training an agent solely on this dataset can yield a negativity-prone chatbot; conversational models must be co-trained with positive dialogue corpora.
    3. Base Model and Generation Errors: Canary occasionally generates irrelevant RoTs or misclassifies casual interactions as requiring intervention. Furthermore, because Prost is built upon PushShift Transformer (pre-trained on unfiltered Reddit data), it retains residual risks of generating toxic or biased responses.

Coverage note — None was omitted; all key contributions—the ProsocialDialog dataset construction, the action-oriented safety schema, the Canary and Prost model architectures, the in-domain and out-of-domain empirical results, and stated limitations—are fully covered.

References

  1. 1.
    1. Emergency. Wex. Accessed April 14, 2022 [Online].
  2. 2.Maurianne Adams, Warren J Blumenfeld, Rosie Castañeda, Heather W Hackman, Madeline L Peters, and Ximena Zúñiga. 2000. Readings for diversity and social justice. Psychology Press.
  3. 3.Ashutosh Baheti, Maarten Sap, Alan Ritter, and Mark Riedl. 2021. Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts. In EMNLP.
  4. 4.C. Daniel Batson and Adam A. Powell. 2003. Altruism and Prosocial Behavior. In Handbook of Psychology, 5th edition. John Wiley & Sons, Inc.
  5. 5.Roy F. Baumeister and Brad J. Bushman. 2017. Social Psychology and Human Nature, 4th edition. Cengage Learning.
  6. 6.Paul Bloom. 2010. How do Morals Change? Nature, 464(7288):490–490.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In NeurIPS.
  8. 8.Frances S Chen, Julia A Minson, and Zakary L Tormala. 2010. Tell Me More: The Effects of Expressed Interest on Receptiveness during Dialog. Journal of Experimental Social Psychology, 46(5):850–853.
  9. 9.Herbert H Clark and Susan E Brennan. 1991. Grounding in communication. In Perspectives on socially shared cognition., pages 127–149. American Psychological Association.
  10. 10.William Collins. 2022. Prosocial. Collins English Dictionary. Accessed March 23, 2022 [Online].
  11. 11.Kate Crawford. 2021. Atlas of AI. Yale University Press.
  12. 12.Leslie A DeChurch and Michelle A Marks. 2001. Maximizing the benefits of task conflict: The role of conflict management. International Journal of Conflict Management.
  13. 13.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A Dataset of Fine-Grained Emotions. In ACL.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL.
  15. 15.Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2022. Safetykit: First aid for measuring safety in open-domain conversational systems. In NAACL.
  16. 16.Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019. Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack. In EMNLP.
  17. 17.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018. Wizard of Wikipedia: Knowledge-Powered Conversational Agents. In ICLR.
  18. 18.Sarah E Finch and Jinho D Choi. 2020. Towards Unified Dialogue System Evaluation: A Comprehensive Analysis of Current Evaluation Protocols. In SIGDial.
  19. 19.Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. Social Chemistry 101: Learning to Reason about Social and Moral Norms. In EMNLP.
  20. 20.Sam Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of EMNLP.
  21. 21.Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019. Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations. In Interspeech.
  22. 22.Jonathan Haidt, Silvia Helena Koller, and Maria G Dias. 1993. Affect, culture, and morality, or is it wrong to eat your dog? Journal of personality and social psychology, 65(4):613.
  23. 23.Dominik Hangartner, Gloria Gennaro, Sary Alasiri, Nicholas Bahrich, Alexandra Bornhoft, Joseph Boucher, Buket Buse Demirci, Laurenz Derksen, Aldo Hall, Matthias Jochum, et al. 2021. Empathy-based Counterspeech can Reduce Racist Hate Speech in a Social Media Field Experiment. Proceedings of the National Academy of Sciences, 118(50).
  24. 24.John Hattie and Helen Timperley. 2007. The power of feedback. Review of educational research, 77(1):81–112.
  25. 25.Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI With Shared Human Values. In ICLR.
  26. 26.Joseph Hoover, Mohammad Atari, Aida Mostafazadeh Davani, Brendan Kennedy, Gwenyth Portillo-Wightman, Leigh Yeh, Drew Kogon, and Morteza Dehghani. 2019. Bound in hatred: The role of group-based morality in acts of hate.
  27. 27.Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, and Aaron Courville. 2021. Understanding by Understanding Not: Modeling Negation in Language Models. In NAACL.
  28. 28.Karen Huang, Michael Yeomans, Alison Wood Brooks, Julia Minson, and Francesca Gino. 2017. It doesn’t Hurt to Ask: Question-asking Increases Liking. Journal of personality and social psychology, 113(3):430.
  29. 29.Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Le Bras Ronan, Maxwell Forbes, Jon Borchardt, Jenny Liang, Oren Etzioni, Maarten Sap, and Yejin Choi. 2021. Delphi: Towards Machine Ethics and Norms. arXiv preprint arXiv:2110.07574.
  30. 30.Diederik P Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. arXiv preprint arXiv:1412.6980.
  31. 31.Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2021. Internet-augmented Dialogue Generation. arXiv preprint arXiv:2107.07566.
  32. 32.Klaus Krippendorff. 2011. Computing Krippendorff’s Alpha-reliability.
  33. 33.Mucahid Kutlu, Tyler McDonnell, Tamer Elsayed, and Matthew Lease. 2020. Annotator rationales for labeling tasks in crowdsourcing. The journal of artificial intelligence research, 69:143–189.
  34. 34.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In IJCNLP.
  35. 35.Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang. 2021. Towards Emotional Support Dialog Systems. In ACL.
  36. 36.Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021. Few-Shot Bot: Prompt-Based Learning for Dialogue Systems. arXiv preprint arXiv:2110.08118.
  37. 37.Shikib Mehri, Jinho Choi, Luis Fernando D’Haro, Jan Deriu, Maxine Eskenazi, Milica Gasic, Kallirroi Georgila, Dilek Hakkani-Tur, Zekang Li, Verena Rieser, Samira Shaikh, David Traum, Yi-Ting Yeh, Zhou Yu, Yizhe Zhang, and Chen Zhang. 2022. Report from the NSF future directions workshop on automatic evaluation of dialog: Research directions and challenges.
  38. 38.A. H. Miller, W. Feng, A. Fisch, J. Lu, D. Batra, A. Bordes, D. Parikh, and J. Weston. 2017. ParlAI: A Dialog Research Software Platform. arXiv:1705.06476.
  39. 39.Nikita Moghe, Siddhartha Arora, Suman Banerjee, and Mitesh M Khapra. 2018. Towards Exploiting Background Knowledge for Building Conversation Systems. In EMNLP.
  40. 40.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories. In NAACL.
  41. 41.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training Language Models to Follow Instructions with Human Feedback. arXiv preprint arXiv:2203.02155.
  42. 42.James W Pennebaker, Ryan L Boyd, Kayla Jordan, and Kate Blackburn. 2015. The Development and Psychometric Properties of LIWC2015. Technical report.
  43. 43.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red Teaming Language Models with Language Models. arXiv preprint arXiv:2202.03286.
  44. 44.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language Models are Unsupervised Multitask Learners. OpenAI blog, 1(8):9.
  45. 45.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21:1–67.
  46. 46.M Afzalur Rahim. 2002. Toward a theory of managing organizational conflict. International journal of conflict management.
  47. 47.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset. In ACL.
  48. 48.Rob Reich, Mehran Sahami, and Jeremy M Weinstein. 2021. System error: Where big tech went wrong and how we can reboot. Hodder & Stoughton.
  49. 49.Sarah T Roberts. 2017. Social media’s silent filter. The Atlantic.
  50. 50.Carl R. Rogers. 1946. Significant Aspects of Client-centered Therapy. American Psychologist, 1(10):415.
  51. 51.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M Smith, et al. 2021. Recipes for Building an Open-Domain Chatbot. In EACL.
  52. 52.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A Smith, and Yejin Choi. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language. In ACL.
  53. 53.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection. In NAACL.
  54. 54.Hilary Silver. 1994. Social exclusion and social solidarity: Three paradigms. Int’l Lab. Rev., 133:531.
  55. 55.Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020. Can You Put it All Together: Evaluating Conversational Agents’ Ability to Blend Skills. In ACL.
  56. 56.Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological Well-Being of content moderators: The emotional labor of commercial moderation and avenues for improving support. In CHI.
  57. 57.Chloe Rose Stuart-Ulin. 2018. Microsoft’s politically correct chatbot is even worse than its racist one. https://qz.com/1340990/microsofts-politically-correct-chat-bot-is-even-worse-than-its-racist-one/. Accessed: 2022-4-28.
  58. 58.Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang. 2022. On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark. In Findings of ACL.
  59. 59.Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2021. A Word on Machine Ethics: A Response to Jiang et al.(2021). arXiv preprint arXiv:2111.04158.
  60. 60.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. LaMDA: Language Models for Dialog Applications. arXiv preprint arXiv:2201.08239.
  61. 61.Jean M. Twenge, Roy F. Baumeister, C. Nathan DeWall, Natalie J. Ciarocco, and J. Michael Bartels. 2007. Social Exclusion Decreases Prosocial Behavior. Journal of Personality and Social Psychology, 92(1):56.
  62. 62.Megan Ung, Jing Xu, and Y-Lan Boureau. 2021. Saferdialogues: Taking feedback gracefully after conversational safety failures. arXiv preprint arXiv:2110.07518.
  63. 63.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. arXiv preprint arXiv:2201.11903.
  64. 64.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and Social Risks of Harm from Language Models. arXiv preprint arXiv:2112.04359.
  65. 65.Anuradha Welivita and Pearl Pu. 2020. A Taxonomy of Empathetic Response Intents in Human Social Conversations. In COLING.
  66. 66.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020. Recipes for Safety in Open-domain Chatbots. arXiv preprint arXiv:2010.07079.
  67. 67.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021. Bot-Adversarial Dialogue for Safe Conversational Agents. In NAACL.
  68. 68.Michael Yeomans, Julia Minson, Hanne Collins, Frances Chen, and Francesca Gino. 2020. Conversational Receptiveness: Improving Engagement with Opposing Views. Organizational Behavior and Human Decision Processes, 160:131–148.
  69. 69.Iris Marion Young. 2014. Five faces of oppression. Rethinking power, pages 174–195.
  70. 70.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing Dialogue Agents: I Have a Dog, Do You Have Pets Too? In ACL.
  71. 71.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. OPT: Open Pre-trained Transformer Language Models. arXiv preprint arXiv:2205.01068.
  72. 72.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020. DialoGPT : Large-Scale Generative Pre-training for Conversational Response Generation. In ACL: System Demonstrations.
  73. 73.Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020. The design and implementation of xiaoice, an empathetic social chatbot. Computational Linguistics, 46(1):53–93.
  74. 74.Pei Zhou, Pegah Jandaghi, Hyundong Cho, Bill Yuchen Lin, Jay Pujara, and Xiang Ren. 2021a. Probing Commonsense Explanation in Dialogue Response Generation. In Findings of EMNLP.
  75. 75.Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Yejin Choi, and Noah A Smith. 2021b. Challenges in Automated Debiasing for Toxic Language Detection. In EACL.
  76. 76.Caleb Ziems, Jane A Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022. The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems. In ACL.

Citation

MLA
Kim, H., et al. “ProsocialDialog: A Prosocial Backbone for Conversational Agents”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 4005–29, https://doi.org/10.18653/v1/2022.emnlp-main.267.
APA
Kim, H., Yu, Y., Jiang, L., Lu, X., Khashabi, D., Kim, G., Choi, Y., & Sap, M. (2022). ProsocialDialog: A Prosocial Backbone for Conversational Agents. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4005–4029. https://doi.org/10.18653/v1/2022.emnlp-main.267
Chicago
Kim, H., Y. Yu, L. Jiang, et al. 2022. “ProsocialDialog: A Prosocial Backbone for Conversational Agents”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4005–29. https://doi.org/10.18653/v1/2022.emnlp-main.267.
Harvard
Kim, H. et al. (2022) “ProsocialDialog: A Prosocial Backbone for Conversational Agents”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4005–4029. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.267.
Vancouver
1. Kim H, Yu Y, Jiang L, Lu X, Khashabi D, Kim G, Choi Y, Sap M (2022) ProsocialDialog: A Prosocial Backbone for Conversational Agents. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4005–4029

BibTeX

@inproceedings{kim-etal-2022-prosocialdialog,
    title = "{P}rosocial{D}ialog: A Prosocial Backbone for Conversational Agents",
    author = "Kim, Hyunwoo  and
      Yu, Youngjae  and
      Jiang, Liwei  and
      Lu, Ximing  and
      Khashabi, Daniel  and
      Kim, Gunhee  and
      Choi, Yejin  and
      Sap, Maarten",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.267/",
    doi = "10.18653/v1/2022.emnlp-main.267",
    pages = "4005--4029"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/