The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems

Caleb ZiemsJane A. YuYi-Chia WangAlon Y. HalevyDiyi Yang

article2022ACL141 citations

Presents a large-scale benchmark of prompt-reply pairs annotated with 99k conversational Rules of Thumb to systematically evaluate, explain, and improve how open-domain dialogue agents handle competing moral assumptions.

Listen

Conversational artificial intelligence systems are increasingly deployed across critical public domains, such as healthcare, education, and customer operations. However, these systems often generate insensitive, harmful, or inconsistent statements learned from broad internet data, directly undermining user trust and raising safety and reputational risks. Standard safeguards like simple word filtering or isolated safety scoring fail because conversational norms are highly contextual, subject to competing social values, and rarely universally agreed upon. To address this challenge, the article introduces a benchmark called the Moral Integrity Corpus (MIC), which provides a structured methodology to transparently evaluate, explain, and moderate the moral and social assumptions embedded in chatbot dialogues.

The research team compiled a diverse benchmark consisting of 38,000 unique human question and chatbot response pairs evaluated across three leading conversational architectures. A trained pool of 186 annotators developed 99,000 distinct natural-language "Rules of Thumb"—concise principles explaining why a given reply is acceptable or problematic—and attached 114,000 structured attribute profiles detailing moral dimensions, perceived severity, and global consensus. Using this dataset, the researchers fine-tuned generative transformer models to automatically draft Rules of Thumb for previously unseen dialogues and trained classification models to categorize the underlying social and ethical attributes of those rules.

The evaluation yielded several key findings regarding model capabilities and benchmark dynamics. First, generative transformer models successfully learned to articulate relevant moral rules, with the best-performing model achieving a high ROUGE-L similarity score of 53 and matching or exceeding human evaluators in fluency and structural well-formedness. Second, models evaluated via beam search achieved human-level relevance scores of 4.03 out of 5, outperforming simple retrieval techniques. Third, despite these high averages, modern generative models remain brittle, producing irrelevant or mismatched moral explanations nearly 28% of the time. Fourth, transfer experiments demonstrated that models trained on narrative text benchmarks fail when applied to open dialogue, proving that conversational settings present unique challenges such as leading questions and conversational pragmatics. Finally, attribute classifiers reliably predicted severity, consensus, and core moral foundations, outperforming human rater baselines on several categorical metrics.

These findings indicate that conversational AI safety cannot rely on static narrative rules or single universal verdicts. Instead, organizations must implement explainable, multi-perspective frameworks capable of interpreting subtle conversational nuances. By providing transparent rationales and alternative, revised answers, the benchmark provides a foundation to steer AI models via reinforcement learning, build nuanced safety classifiers, and design adaptable content moderation systems that respect diverse cultural perspectives.

Decision-makers and developers should leverage the corpus to train automated penalty and steering mechanisms rather than treating AI-generated ethical judgments as authoritative moral advice. Deployment should focus on hybrid moderation workflows where algorithmic rule generators provide transparent diagnostic signals to human operators. Because the underlying training data is limited to English-speaking contributors in the United States, leaders should exercise caution when deploying systems in cross-cultural settings and plan additional pilots to adapt the framework to broader global contexts.

arXiv: 2204.03021GT-SALT/mic

No sufficiently relevant recommendations were found.

Cover for The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems

Abstract

Content Warning: some examples in this paper may be offensive or upsetting.

Conversational agents have come increasingly closer to human competence in open-domain dialogue settings; however, such models can reflect insensitive, hurtful, or entirely incoherent viewpoints that erode a user’s trust in the moral integrity of the system. Moral deviations are difficult to mitigate because moral judgments are not universal, and there may be multiple competing judgments that apply to a situation simultaneously. In this work, we introduce a new resource, not to authoritatively resolve moral ambiguities, but instead to facilitate systematic understanding of the intuitions, values and moral judgments reflected in the utterances of dialogue systems. The Moral Integrity Corpus, MIC, is such a resource, which captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs). Each RoT reflects a particular moral conviction that can explain why a chatbot’s reply may appear acceptable or problematic. We further organize RoTs with a set of 9 moral and social attributes and benchmark performance for attribute classification. Most importantly, we show that current neural language models can automatically generate new RoTs that reasonably describe previously unseen interactions, but they still struggle with certain scenarios. Our findings suggest that MIC will be a useful resource for understanding language models’ implicit moral assumptions and flexibly benchmarking the integrity of conversational agents. To download the data, see https://github.com/GT-SALT/mic

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Moral Annotation Framework
  • 3.1 Rules of Thumb (RoTs)
  • 4 The MORAL INTEGRITY CORPUS
  • 4.1 Collecting Prompt-Reply Pairs
  • 4.2 Annotating RoTs
  • 5 Models
  • 5.1 RoT Generation
  • 5.2 RoT Attribute Classification
  • 6 Results
  • 6.1 RoT Generation Results
  • 6.2 Unique Challenges in MIC
  • 6.3 Attribute Classification Results
  • 7 Discussion and Conclusion
  • Acknowledgements
  • Ethics
  • References
  • A Model Details
  • A.1 Co-opting GPT-Neo as a Chatbot
  • A.2 RoT Attribute Classification
  • B Chatbot Response Filtering
  • C Moral Foundations
  • D Annotation Instructions
  • D.1 RoT Instructions
  • D.2 Moral Foundations Instructions
  • E Ensuring Annotation Quality
  • E.1 Qualification Test
  • E.2 Automatic Quality Checks (Scripting)
  • E.3 Manual Quality Control

Knowls

  1. Knowl 1 — MIC makes chatbot moral assumptions an interpretable dialogue benchmark

    model/method

    The MORAL INTEGRITY CORPUS (MIC) is a benchmark for studying moral and social judgments that people might apply to an open-ended human prompt and a chatbot’s reply. It contains approximately 38,000 unique prompt–reply pairs, 99,000 distinct Rules of Thumb (RoTs), and 114,000 annotation sets. Each RoT makes a possible judgment about a reply interpretable, rather than assigning one definitive moral verdict. The corpus also includes revised answers that are neutral or consistent with the annotator’s RoT. MIC is intended to describe the varied moral assumptions that may be reflected in chatbot behavior, not to establish a universally binding ethical system.

  2. Knowl 2 — RoTs and their annotation attributes define the corpus’s judgment scheme

    definition

    A Rule of Thumb (RoT) is a general, understandable-out-of-context rule about good or bad behavior. It must express a judgment and an action, and be specific enough to connect to the situation without merely restating its details. Several RoTs, including conflicting ones, may apply to the same prompt–reply pair.

    Each RoT can have four kinds of annotation: Reply Alignment records whether the chatbot’s reply agrees with, disagrees with, or is neither aligned nor misaligned with the rule; Global Consensus estimates how widely people globally agree, using categories nobody (<1%), rare (5%–25%), controversial (about 50%), most (75%–90%), and all (>99%); Violation Severity rates the seriousness of not following the rule from 1 (fine) to 5 (worst); and Moral Foundations assigns any number, including none, of six labels: care, fairness, liberty, loyalty, authority, and sanctity. Annotators also provide a revised answer that is neutral or aligns with their RoT.

  3. Knowl 3 — Corpus construction filters opinion prompts and chatbot replies for normative content

    model/method

    The authors began with nearly five million public r/AskReddit posts. They retained prompts whose question and top Reddit answer each contained at least one term from the Expanded Moral Foundations Dictionary and at least one strongly subjective term, yielding 217,700 prompts. Each prompt was sent with greedy decoding to BlenderBot (2.7B parameters), DialoGPT Medium, and GPT-Neo.

    Replies were first filtered for a term from the Expanded Moral Foundations Dictionary. An ALBERT-base-v2 sentence-pair classifier then filtered prompt–reply pairs for both sufficiency—the reply was understandable, specific, and relevant—and moral content—it expressed an idea or behavior someone could judge as right or wrong. The classifier retained pairs when both predicted probabilities exceeded 0.5. The resulting candidate pool contained 30,880 BlenderBot pairs, 11,521 DialoGPT pairs, and 51,141 GPT-Neo pairs; a random subset was annotated for MIC.

  4. Knowl 4 — Human annotation pairs each rule with judgments and a proposed alternative reply

    experimental setup

    For each annotated prompt–reply pair, three annotators independently wrote one RoT and answered the RoT’s alignment, global-consensus, violation-severity, and moral-foundation questions. They also wrote a revised answer that was neutral or aligned with their rule. The 186 annotators were recruited through Amazon Mechanical Turk, had to be located in the United States, and had to answer at least six of seven qualification questions correctly. Training included definitions, examples, and an interactive search tool for RoTs from SOCIAL-CHEM-101.

  5. Knowl 5 — RoT generation benchmarks conditional generation on held-out dialogue pairs

    experimental setup

    MIC’s generation task is to produce a RoT rr given a prompt qq and chatbot reply aa, modeling p(r∣q,a)p(r\mid q,a). The authors fine-tuned GPT-2, BART, and T5, using the same pair-level 80/10/10 train–development–test split and ensuring that no prompt–reply pair crossed splits. They tested 1, 2, 3, or 5 training epochs with batch size 16 and learning rate 3×10−53\times10^{-5}, selecting the epoch by development-set BLEU. At inference, they compared greedy decoding, beam search with 3 beams, and nucleus sampling with p=0.9p=0.9; for beam search and nucleus sampling, they generated three hypotheses and selected the highest-scoring one. Random-RoT and SBERT-based retrieval were comparison baselines.

    Automatic evaluation used ROUGE-1, ROUGE-2, ROUGE-L, BLEU, BERTScore, and average generated length. Because each pair had three reference RoTs, each automatic metric used the maximum score against those references. Human evaluation assessed well-formedness, fluency, and relevance; three workers rated each generation, with 200 generations evaluated per model type, including human-written reference RoTs.

  6. Knowl 6 — Beam-search T5 and GPT-2 produce the strongest RoT generation results

    empirical result

    On MIC test data, beam-search T5 led the reported automatic generation scores: ROUGE-1 53.89, ROUGE-2 33.68, ROUGE-L 52.62, BLEU 24.85, and BERTScore 93.52. Its average length was 8.86 tokens; human ratings were 0.86 well-formedness, 4.51 fluency, and 4.02 relevance. Beam-search GPT-2 scored ROUGE-1 52.86, ROUGE-2 32.35, ROUGE-L 51.57, BLEU 23.44, and BERTScore 93.45; its human ratings were 0.89 well-formedness, 4.57 fluency, and 4.03 relevance.

    For comparison, SBERT retrieval scored ROUGE-1 34.72, ROUGE-2 14.83, ROUGE-L 33.07, BLEU 11.79, and BERTScore 90.98, with relevance 3.65. Human reference RoTs received well-formedness 0.83, fluency 4.55, and relevance 4.03. BART with nucleus sampling had the highest reported model fluency, 4.67, but relevance of 2.30. Thus, the strongest automatic scores did not mean generation was reliable in every case: even the best T5 model received a relevance rating below 2 nearly 28% of the time.

  7. Knowl 7 — ALBERT classifies RoT attributes with moderate performance

    empirical result

    The authors trained separate BERT and ALBERT classifiers for RoT attributes. Reply Alignment was a three-way sentence-pair classification task using the RoT and prompt–reply text. For the other attributes, the input was the RoT alone: Severity and Consensus were modeled as ordinal regression with mean squared error, and the rarest Consensus labels (nobody, rare, controversial) were merged into a controversial class; Moral Foundations were multilabel predictions trained with binary cross-entropy.

    ALBERT achieved Severity correlation r=0.59r=0.59 and MSE 1.011.01, Consensus correlation r=0.44r=0.44 and MSE 45.245.2, and Alignment F1 76.0. Its Moral Foundations F1 scores were: Care 75.3, Fairness 59.6, Liberty 58.0, Loyalty 62.7, Authority 54.3, and Sanctity 40.8. BERT’s corresponding scores were Severity r=0.53r=0.53, MSE 1.13; Consensus r=0.41r=0.41, MSE 47.7; Alignment F1 76.0; and foundation F1 of Care 73.4, Fairness 56.2, Liberty 54.1, Loyalty 59.9, Authority 52.1, and Sanctity 37.0. Sanctity was the weakest foundation classification for both models.

  8. Knowl 8 — Dialogue moral judgments present challenges beyond narrative norm classification

    empirical result

    A GPT-2 RoT generator trained on SOCIAL-CHEM-101 was evaluated on MIC under domain shift. It scored ROUGE-1 28.65, ROUGE-2 9.42, ROUGE-L 26.48, BLEU 6.77, and BERTScore 89.36; human ratings were 0.64 well-formedness, 4.30 fluency, and 3.68 relevance. It did not outperform the MIC-trained generation baselines, indicating that a model trained on narrative situations did not transfer as well to chatbot dialogue.

    The authors identify four distinctive difficulties in the dialogue setting: a reply may combine competing viewpoints; a chatbot may produce an unexpected or non-cooperative moral violation; a reply that seems acceptable in isolation may be inappropriate given the prompt’s pragmatics; and a strategic or adversarial prompt may force a response into conflict with a norm. These cases make it difficult to judge a reply without considering its conversational context.

  9. Knowl 9 — Worker viewpoints vary and agreement on RoT attributes is limited

    empirical result

    The 186-person annotator pool was primarily liberal. The authors report that 25% of workers were conservative-leaning and that 22% of annotations came from conservative-leaning workers. In the abbreviated Moral Foundations Questionnaire, liberal-leaning workers emphasized Care and Fairness more than the other foundations, whereas conservative-leaning workers valued the foundations more evenly.

    For attribute annotation, reported Krippendorff’s α\alpha was 0.27 for Alignment, 0.10 for Consensus, and 0.12 for Severity; for Moral Foundations it ranged from 0.20 for Sanctity to 0.46 for Loyalty. The reported intraclass correlation coefficients were 0.58 for Alignment, 0.49 for Consensus, and 0.62 for Severity; foundation ICCs ranged from 0.42 for Sanctity to 0.72 for Loyalty. The authors characterize agreement as fair to moderate and note that differences in annotators’ calibration may contribute to lower agreement on the Consensus and Severity scales.

  10. Knowl 10 — MIC judgments are culturally and temporally bounded, not universal moral advice

    limitation

    MIC reflects a particular population and source of prompts: English-speaking U.S. annotators living in the 21st century, recruited through MTurk, and questions from Reddit, whose users skew toward younger and middle-aged men. The paper notes that MTurk workers may also differ from the general population in education, religiosity, and employment. Consequently, the corpus does not establish moral judgments that apply universally across cultures or periods; the authors identify extending the framework to other cultures and regions as future work and note that moral judgments can shift over time.

    The authors frame RoTs as descriptions of possible human judgments and of assumptions latent in language-model outputs—not as universally binding rules or moral advice. They describe the resource and findings as intended for research, with any later moderation decisions left to domain experts.

Coverage note — The detailed chatbot-filter classifier training results and the appendix’s individual annotation quality-control checks are omitted because they are supporting implementation details rather than central benchmark findings.

References

  1. 1.Gavin Abercrombie, Amanda Cercas Curry, Mugdha Pandya, and Verena Rieser. 2021. Alexa, Google, Siri: What are your pronouns? gender and anthropomorphism in the design and perception of conversational assistants. In Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing, pages 24–33, Online. Association for Computational Linguistics.
  2. 2.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020. Towards a human-like open-domain chatbot. ArXiv preprint, abs/2001.09977.
  3. 3.Fahad Alaieri and André Vellino. 2016. Ethical decision making in robots: Autonomy, trust and responsibility. In International conference on social robotics, pages 159–168. Springer.
  4. 4.Ashley Amaya, Ruben Bach, Florian Keusch, and Frauke Kreuter. 2021. New data sources in social science research: things to know before working with reddit data. Social science computer review, 39(5):943–960.
  5. 5.Simone Balloccu, Ehud Reiter, Matteo G Collu, Federico Sanna, Manuela Sanguinetti, and Maurizio Atzori. 2021. Unaddressed challenges in persuasive dieting chatbots. In Adjunct Proceedings of the 29th ACM Conference on User Modeling, Adaptation and Personalization, pages 392–395.
  6. 6.Rodrigo Bavaresco, Diórgenes Silveira, Eduardo Reis, Jorge Barbosa, Rodrigo Righi, Cristiano Costa, Rodolfo Antunes, Marcio Gomes, Clauter Gatti, Mariangela Vanzin, et al. 2020. Conversational agents in business: A systematic literature review and future research directions. Computer Science Review, 36:100239.
  7. 7.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  8. 8.Cristina Bicchieri. 2005. The grammar of society: The nature and dynamics of social norms. Cambridge University Press.
  9. 9.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow.
  10. 10.Petter Bae Brandtzaeg and Asbjørn Følstad. 2017. Why people use chatbots. In International conference on internet science, pages 377–392. Springer.
  11. 11.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  12. 12.Amanda Cercas Curry and Verena Rieser. 2018. #MeToo Alexa: How conversational systems respond to sexual harassment. In Proceedings of the Second ACL Workshop on Ethics in Natural Language Processing, pages 7–14, New Orleans, Louisiana, USA. Association for Computational Linguistics.
  13. 13.Veena Chattaraman, Wi-Suk Kwon, Juan E Gilbert, and Kassandra Ross. 2019. Should ai-based, conversational digital assistants employ social-or task-oriented interaction style? a task-competency and reciprocity perspective for older adults. Computers in Human Behavior, 90:315–330.
  14. 14.Kimberly E Culley and Poornima Madhavan. 2013. A note of caution regarding anthropomorphism in hci agents. Computers in Human Behavior, 29(3):577–579.
  15. 15.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  16. 16.Morteza Dehghani, Emmett Tomai, Kenneth D Forbus, and Matthew Klenk. 2008. An integrated reasoning approach to moral decision-making. In AAAI, pages 1280–1286.
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  18. 18.Djellel Eddine Difallah, Elena Filatova, and Panos Ipeirotis. 2018. Demographics and dynamics of mechanical turk workers. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, WSDM 2018, Marina Del Rey, CA, USA, February 5-9, 2018, pages 135–143. ACM.
  19. 19.Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2021. Anticipating safety issues in e2e conversational ai: Framework and tooling. ArXiv preprint, abs/2107.03451.
  20. 20.Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021. Moral stories: Situated reasoning about norms, intents, actions, and their consequences. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 698–718, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  21. 21.Jessica Ficler and Yoav Goldberg. 2017. Controlling linguistic style aspects in neural language generation. In Proceedings of the Workshop on Stylistic Variation, pages 94–104, Copenhagen, Denmark. Association for Computational Linguistics.
  22. 22.Fionn Delahunty. 2018. Reddit QA Corpus. https://github.com/FionnD/Reddit-QA-Corpus. Online; accessed XXX.
  23. 23.Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. Social chemistry 101: Learning to reason about social and moral norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 653–670, Online. Association for Computational Linguistics.
  24. 24.Jianfeng Gao, Michel Galley, and Lihong Li. 2018. Neural approaches to conversational AI. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR 2018, Ann Arbor, MI, USA, July 08-12, 2018, pages 1371–1374. ACM.
  25. 25.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2021. The pile: An 800gb dataset of diverse text for language modeling. ArXiv preprint, abs/2101.00027.
  26. 26.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  27. 27.Bernard Gert and Joshua Gert. 2002. The definition of morality.
  28. 28.Joseph K Goodman, Cynthia E Cryder, and Amar Cheema. 2013. Data collection in a flat world: The strengths and weaknesses of mechanical turk samples. Journal of Behavioral Decision Making, 26(3):213–224.
  29. 29.Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013. Moral foundations theory: The pragmatic validity of moral pluralism. In Advances in experimental social psychology, volume 47, pages 55–130. Elsevier.
  30. 30.Jesse Graham, Jonathan Haidt, and Brian A Nosek. 2008. The moral foundations questionnaire. MoralFoundations. org.
  31. 31.Jesse Graham, Jonathan Haidt, and Brian A Nosek. 2009. Liberals and conservatives rely on different sets of moral foundations. Journal of personality and social psychology, 96(5):1029.
  32. 32.Herbert P Grice. 1975. Logic and conversation. In Speech acts, pages 41–58. Brill.
  33. 33.Joshua Grossman, Zhiyuan Lin, Hao Sheng, Johnny Tian-Zheng Wei, Joseph J Williams, and Sharad Goel. 2019. Mathbot: Transforming online resources for learning math into conversational interactions. AAAI 2019 Story-Enabled Intelligence.
  34. 34.Suchin Gururangan, Ana Marasovic, Swabha ´ Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  35. 35.Jonathan Haidt. 2012. The righteous mind: Why good people are divided by politics and religion. Vintage.
  36. 36.Jonathan Haidt and Jesse Graham. 2007. When morality opposes justice: Conservatives have moral intuitions that liberals may not recognize. Social Justice Research, 20(1):98–116.
  37. 37.Jonathan Haidt, Silvia Helena Koller, and Maria G Dias. 1993. Affect, culture, and morality, or is it wrong to eat your dog? Journal of personality and social psychology, 65(4):613.
  38. 38.Richard Mervyn Hare, Richard Mervyn Hare, Richard Mervyn Hare Hare, and Richard M Hare. 1981. Moral thinking: Its levels, method, and point. Oxford: Clarendon Press; New York: Oxford University Press.
  39. 39.Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020. Aligning ai with shared human values. ArXiv preprint, abs/2008.02275.
  40. 40.Wilhelm Hofmann, Daniel C Wisneski, Mark J Brandt, and Linda J Skitka. 2014. Morality in everyday life. Science, 345(6202):1340–1343.
  41. 41.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  42. 42.Peng Hu, Yaobin Lu, et al. 2021. Dual humanness and trust in conversational ai: A person-centered approach. Computers in Human Behavior, 119:106727.
  43. 43.Minlie Huang, Xiaoyan Zhu, and Jianfeng Gao. 2020. Challenges in building intelligent open-domain dialog systems. ACM Transactions on Information Systems (TOIS), 38(3):1–32.
  44. 44.Ravi Iyer, Stephen J Read, and Jane Correia. 2010. Functional justice: Productivity and well-being goals define fairness. Available at SSRN 1691969.
  45. 45.Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Maxwell Forbes, Jon Borchardt, Jenny Liang, Oren Etzioni, Maarten Sap, and Yejin Choi. 2021. Delphi: Towards machine ethics and norms. ArXiv preprint, abs/2110.07574.
  46. 46.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. ArXiv preprint, abs/1909.05858.
  47. 47.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  48. 48.Liliana Laranjo, Adam G Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie YS Lau, et al. 2018. Conversational agents in healthcare: a systematic review. Journal of the American Medical Informatics Association, 25(9):1248–1258.
  49. 49.Sven Laumer, Christian Maier, and Fabian Tobias Gubler. 2019. Chatbot acceptance in healthcare: Explaining user adoption of conversational agents for disease diagnosis.
  50. 50.Peter Lee. Learning from tay’s introduction.
  51. 51.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pretraining for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  52. 52.Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016. Deep reinforcement learning for dialogue generation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1192–1202, Austin, Texas. Association for Computational Linguistics.
  53. 53.Q. Vera Liao, Muhammed Mas-ud Hussain, Praveen Chandar, Matthew Davis, Yasaman Khazaeni, Marco Patricio Crasso, Dakuo Wang, Michael J. Muller, N. Sadat Shami, and Werner Geyer. 2018. All work and no play? In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI 2018, Montreal, QC, Canada, April 21-26, 2018, page 3. ACM.
  54. 54.Chin-Yew Lin and Eduard Hovy. 2003. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, pages 150–157.
  55. 55.Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Lili Wang, and Soroush Vosoughi. 2021a. Mitigating political bias in language models through reinforced calibration. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 14857–14866.
  56. 56.Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang. 2021b. Towards emotional support dialog systems. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3469–3483, Online. Association for Computational Linguistics.
  57. 57.Nicholas Lourie, Ronan Le Bras, and Yejin Choi. 2021. Scruples: A corpus of community ethical judgments on 32, 000 real-life anecdotes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13470–13479.
  58. 58.Alexandra Sasha Luccioni and Joseph D Viviano. 2021. What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus. ArXiv preprint, abs/2105.02732.
  59. 59.Ewa Luger and Abigail Sellen. 2016. "like having a really bad pa": The gulf between user expectation and experience of conversational agents. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, San Jose, CA, USA, May 7-12, 2016, pages 5286–5297. ACM.
  60. 60.Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob N. Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel. 2019. A survey of reinforcement learning informed by natural language. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, pages 6309–6317. ijcai.org.
  61. 61.Roger C Mayer, James H Davis, and F David Schoorman. 1995. An integrative model of organizational trust. Academy of management review, 20(3):709–734.
  62. 62.D Harrison McKnight, Vivek Choudhury, and Charles Kacmar. 2002. Developing and validating trust measures for e-commerce: An integrative typology. Information systems research, 13(3):334–359.
  63. 63.Peter Meindl, Ravi Iyer, and Jesse Graham. 2019. Distributive justice beliefs are guided by whether people think the ultimate goal of society is well-being or power. Basic and applied social psychology, 41(6):359–385.
  64. 64.György Molnár and Zoltán Szüts. 2018. The role of chatbots in formal education. In 2018 IEEE 16th International Symposium on Intelligent Systems and Informatics (SISY), pages 000197–000202. IEEE.
  65. 65.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849, San Diego, California. Association for Computational Linguistics.
  66. 66.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA.
  67. 67.Xiangyu Peng, Siyan Li, Spencer Frazier, and Mark Riedl. 2020. Reducing non-normative text generation from language models. In Proceedings of the 13th International Conference on Natural Language Generation, pages 374–383, Dublin, Ireland. Association for Computational Linguistics.
  68. 68.Shrimai Prabhumoye, Brendon Boldt, Ruslan Salakhutdinov, and Alan W Black. 2021. Case study: Deontological ethics in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3784–3798, Online. Association for Computational Linguistics.
  69. 69.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  70. 70.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  71. 71.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  72. 72.Rezvaneh Rezapour, Saumil H. Shah, and Jana Diesner. 2019. Enhancing the measurement of social effects by capturing morality. In Proceedings of the Tenth Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 35–45, Minneapolis, USA. Association for Computational Linguistics.
  73. 73.Sarah T Roberts. 2016. Commercial content moderation: Digital laborers’ dirty work.
  74. 74.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, Online. Association for Computational Linguistics.
  75. 75.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5477–5490, Online. Association for Computational Linguistics.
  76. 76.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. ArXiv preprint, abs/2103.00453.
  77. 77.Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin Rothkopf, and Kristian Kersting. 2021. Language models have a moral dimension. ArXiv preprint, abs/2103.11790.
  78. 78.Anna-Maria Seeger, Jella Pfeiffer, and Armin Heinzl. 2017. When do we need a human? anthropomorphic design and trustworthiness of conversational agents. In Proceedings of the Sixteenth Annual Pre-ICIS Workshop on HCI Research in MIS, AISeL, Seoul, Korea, volume 10.
  79. 79.Kim Bartel Sheehan. 2018. Crowdsourcing research: data collection with amazon’s mechanical turk. Communication Monographs, 85(1):140–156.
  80. 80.Richard A Shweder. 1990. In defense of moral realism: Reply to gabennesch. Child Development, 61(6):2060–2067.
  81. 81.Eric Michael Smith, Diana Gonzalez-Rico, Emily Dinan, and Y-Lan Boureau. 2020. Controlling style in generated dialogue. ArXiv preprint, abs/2009.10855.
  82. 82.Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan. 2015. A neural network approach to context-sensitive generation of conversational responses. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 196–205, Denver, Colorado. Association for Computational Linguistics.
  83. 83.SSA. 2018. Popular names in 2018.
  84. 84.Constantine Stephanidis, Gavriel Salvendy, Margherita Antona, Jessie YC Chen, Jianming Dong, Vincent G Duffy, Xiaowen Fang, Cali Fidopiastis, Gino Fragomeni, Limin Paul Fu, et al. 2019. Seven hci grand challenges. International Journal of Human–Computer Interaction, 35(14):1229–1269.
  85. 85.Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2021. A word on machine ethics: A response to jiang et al.(2021). ArXiv preprint, abs/2111.04158.
  86. 86.Yi Tay, Donovan Ong, Jie Fu, Alvin Chan, Nancy Chen, Anh Tuan Luu, and Chris Pal. 2020. Would you rather? a new benchmark for learning machine alignment with cultural values and social preferences. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5369–5373, Online. Association for Computational Linguistics.
  87. 87.Aditya Nrusimha Vaidyam, Hannah Wisniewski, John David Halamka, Matcheri S Kashavan, and John Blake Torous. 2019. Chatbots and conversational agents in mental health: a review of the psychiatric landscape. The Canadian Journal of Psychiatry, 64(7):456–464.
  88. 88.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2153–2162, Hong Kong, China. Association for Computational Linguistics.
  89. 89.Weiquan Wang and Izak Benbasat. 2008. Attributions of trust in decision support technologies: A study of recommendation agents for e-commerce. Journal of Management Information Systems, 24(4):249–273.
  90. 90.Weiquan Wang and Izak Benbasat. 2016. Empirical assessment of alternative designs for enhancing different types of trusting beliefs in online recommendation agents. Journal of Management Information Systems, 33(3):744–775.
  91. 91.Theresa Wilson, Janyce Wiebe, and Paul Hoffmann. 2005. Recognizing contextual polarity in phrase-level sentiment analysis. In Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing, pages 347–354, Vancouver, British Columbia, Canada. Association for Computational Linguistics.
  92. 92.Marty J Wolf, Keith W Miller, and Frances S Grodzinsky. 2017. Why we should have seen that coming: comments on microsoft’s tay “experiment,” and wider implications. The ORBIT Journal, 1(2):1–12.
  93. 93.Bo Xiao and Izak Benbasat. 2007. E-commerce product recommendation agents: Use, characteristics, and impact. MIS quarterly, pages 137–209.
  94. 94.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2950–2968, Online. Association for Computational Linguistics.
  95. 95.Shanshan Yang and Chris Evans. 2019. Opportunities and challenges in using ai chatbots in higher education. In Proceedings of the 2019 3rd International Conference on Education and E-Learning, pages 79–83.
  96. 96.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 9051–9062.
  97. 97.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020a. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  98. 98.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020b. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278, Online. Association for Computational Linguistics.
  99. 99.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. ArXiv preprint, abs/1909.08593.
  100. 100.John Zoshak and Kristin Dew. 2021. Beyond kant and bentham: How ethical theories are being used in artificial moral agents. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–15.

Citation

MLA
Ziems, C., et al. “The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3755–73, https://doi.org/10.18653/v1/2022.acl-long.261.
APA
Ziems, C., Yu, J. A., Wang, Y.-C., Halevy, A. Y., & Yang, D. (2022). The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3755–3773. https://doi.org/10.18653/v1/2022.acl-long.261
Chicago
Ziems, C., J. A. Yu, Y.-C. Wang, A. Y. Halevy, and D. Yang. 2022. “The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3755–73. https://doi.org/10.18653/v1/2022.acl-long.261.
Harvard
Ziems, C. et al. (2022) “The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3755–3773. Available at: https://doi.org/10.18653/v1/2022.acl-long.261.
Vancouver
1. Ziems C, Yu JA, Wang Y-C, Halevy AY, Yang D (2022) The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3755–3773

BibTeX

@inproceedings{ziems-etal-2022-moral,
    title = "The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems",
    author = "Ziems, Caleb  and
      Yu, Jane A.  and
      Wang, Yi-Chia  and
      Halevy, Alon Y.  and
      Yang, Diyi",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.261/",
    doi = "10.18653/v1/2022.acl-long.261",
    pages = "3755--3773"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/