The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
Caleb ZiemsJane A. YuYi-Chia WangAlon Y. HalevyDiyi Yang
Presents a large-scale benchmark of prompt-reply pairs annotated with 99k conversational Rules of Thumb to systematically evaluate, explain, and improve how open-domain dialogue agents handle competing moral assumptions.
Conversational artificial intelligence systems are increasingly deployed across critical public domains, such as healthcare, education, and customer operations. However, these systems often generate insensitive, harmful, or inconsistent statements learned from broad internet data, directly undermining user trust and raising safety and reputational risks. Standard safeguards like simple word filtering or isolated safety scoring fail because conversational norms are highly contextual, subject to competing social values, and rarely universally agreed upon. To address this challenge, the article introduces a benchmark called the Moral Integrity Corpus (MIC), which provides a structured methodology to transparently evaluate, explain, and moderate the moral and social assumptions embedded in chatbot dialogues.
The research team compiled a diverse benchmark consisting of 38,000 unique human question and chatbot response pairs evaluated across three leading conversational architectures. A trained pool of 186 annotators developed 99,000 distinct natural-language "Rules of Thumb"—concise principles explaining why a given reply is acceptable or problematic—and attached 114,000 structured attribute profiles detailing moral dimensions, perceived severity, and global consensus. Using this dataset, the researchers fine-tuned generative transformer models to automatically draft Rules of Thumb for previously unseen dialogues and trained classification models to categorize the underlying social and ethical attributes of those rules.
The evaluation yielded several key findings regarding model capabilities and benchmark dynamics. First, generative transformer models successfully learned to articulate relevant moral rules, with the best-performing model achieving a high ROUGE-L similarity score of 53 and matching or exceeding human evaluators in fluency and structural well-formedness. Second, models evaluated via beam search achieved human-level relevance scores of 4.03 out of 5, outperforming simple retrieval techniques. Third, despite these high averages, modern generative models remain brittle, producing irrelevant or mismatched moral explanations nearly 28% of the time. Fourth, transfer experiments demonstrated that models trained on narrative text benchmarks fail when applied to open dialogue, proving that conversational settings present unique challenges such as leading questions and conversational pragmatics. Finally, attribute classifiers reliably predicted severity, consensus, and core moral foundations, outperforming human rater baselines on several categorical metrics.
These findings indicate that conversational AI safety cannot rely on static narrative rules or single universal verdicts. Instead, organizations must implement explainable, multi-perspective frameworks capable of interpreting subtle conversational nuances. By providing transparent rationales and alternative, revised answers, the benchmark provides a foundation to steer AI models via reinforcement learning, build nuanced safety classifiers, and design adaptable content moderation systems that respect diverse cultural perspectives.
Decision-makers and developers should leverage the corpus to train automated penalty and steering mechanisms rather than treating AI-generated ethical judgments as authoritative moral advice. Deployment should focus on hybrid moderation workflows where algorithmic rule generators provide transparent diagnostic signals to human operators. Because the underlying training data is limited to English-speaking contributors in the United States, leaders should exercise caution when deploying systems in cross-cultural settings and plan additional pilots to adapt the framework to broader global contexts.
No sufficiently relevant recommendations were found.
- Paper: ProsocialDialog: A Prosocial Backbone for Conversational Agents, Hyunwoo Kim et al. (2022). Building on rule-of-thumb-based moral analysis, ProsocialDialog turns norm-grounded dialogue data into systems that detect unsafe contexts and respond constructively.
