RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Samuel GehmanSuchin GururanganMaarten SapYejin ChoiNoah A. Smith
Introduces RealToxicityPrompts, a 100,000-prompt benchmark to evaluate toxic degeneration in language models, demonstrating that benign prompts can trigger severe toxicity and that current mitigation methods remain inadequate.
Pretrained neural language models increasingly serve as foundational components for modern conversational agents, automated writing assistants, and content generation tools. However, these systems frequently produce toxic, abusive, or discriminatory text, creating severe ethical, brand, and safety risks that hinder their deployment in public-facing applications.
The article evaluates the extent to which major language models generate toxic content when prompted with various types of text, and it assesses the effectiveness of several controllable text generation methods designed to mitigate this issue. In addition, the article examines the composition of the web corpora used to pretrain these models to identify the underlying sources of this toxic behavior.
To conduct this evaluation, the researchers created a benchmark dataset of 100,000 naturally occurring sentence prompts extracted from web text, paired with toxicity scores from a widely used commercial detector. They tested five prominent language models—including GPT-1, GPT-2, and GPT-3—measuring the likelihood and severity of toxic generation under unprompted conditions and when supplied with both toxic and non-toxic prompts. The study also evaluated data-based and decoding-based steering techniques, such as non-toxic fine-tuning, vocabulary manipulation, discriminator-guided generation, and basic word filtering, while auditing the training datasets for toxic content and unreliable sources.
The analysis revealed several critical findings. First, all evaluated models reliably generate high levels of toxicity even without any prompting, often reaching severe toxicity within 100 to 1,000 generations. Second, conditioning models on seemingly benign, non-toxic prompts still results in a toxicity probability near or above 50% across all models. Third, while advanced steering techniques—particularly domain-adaptive pretraining on non-toxic text and discriminator-guided decoding—reduce toxic outputs significantly more effectively than basic word blocklists, no tested method completely prevents toxic generation. Finally, pretraining corpora contain substantial abusive material, including 2.1% to 4.3% toxic documents, hundreds of thousands of texts linked from quarantined or banned internet forums, and significant material from low-reliability news sites.
These findings demonstrate that language models systematically acquire and retain toxicity directly from web-scraped pretraining data, and post-training interventions cannot currently guarantee safe outputs. For organizations deploying generative text systems, relying solely on simple profanity filters or standard steering algorithms leaves substantial compliance, reputational, and safety vulnerabilities.
To address these risks, the article recommends re-evaluating data collection strategies by implementing rigorous, transparent curation standards rather than relying on uncurated web scraping. Practitioners should prioritize data-intensive mitigations, such as domain-adaptive pretraining on curated non-toxic text, while combining them with advanced decoding controls. Furthermore, developers must adopt human-centered design frameworks to ensure data selection reflects appropriate safety norms without inadvertently censoring underrepresented dialects.
These conclusions are bounded by certain limitations, most notably the reliance on an automated toxicity classifier that exhibits known lexical and social biases. Additionally, because complete metadata was unavailable for proprietary pretraining corpora, the documented prevalence of toxic sources represents a conservative lower bound. Nonetheless, the evidence strongly supports high confidence in the finding that current models remain fundamentally prone to toxic degeneration without deeper structural interventions in training data curation.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). Its demonstration that language corpora encode human-like racial and gender biases provides the empirical foundation for interpreting toxic degeneration as a consequence of learned web-text associations.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). This later taxonomy broadens RealToxicityPrompts’ focused toxicity analysis into a systematic account of the wider ethical and social harms produced by language models.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). It continues the source’s evaluation of toxicity mitigation by testing whether RLHF can jointly improve helpfulness and harmlessness without sacrificing model capabilities.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). It extends the source’s concern with unsafe pretrained outputs into a concrete human-feedback pipeline for reducing toxicity, bias, and other harmful behaviors.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). It follows the source’s finding that simple safeguards are not failsafe by analyzing why safety training generalizes poorly and remains vulnerable to jailbreaks.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). It advances the source’s robustness testing from passive toxic degeneration to automated adversarial prompts that deliberately bypass alignment and induce harmful outputs.
