RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Samuel GehmanSuchin GururanganMaarten SapYejin ChoiNoah A. Smith

article2020Findings1,841 citations

Introduces RealToxicityPrompts, a 100,000-prompt benchmark to evaluate toxic degeneration in language models, demonstrating that benign prompts can trigger severe toxicity and that current mitigation methods remain inadequate.

Listen

Pretrained neural language models increasingly serve as foundational components for modern conversational agents, automated writing assistants, and content generation tools. However, these systems frequently produce toxic, abusive, or discriminatory text, creating severe ethical, brand, and safety risks that hinder their deployment in public-facing applications.

The article evaluates the extent to which major language models generate toxic content when prompted with various types of text, and it assesses the effectiveness of several controllable text generation methods designed to mitigate this issue. In addition, the article examines the composition of the web corpora used to pretrain these models to identify the underlying sources of this toxic behavior.

To conduct this evaluation, the researchers created a benchmark dataset of 100,000 naturally occurring sentence prompts extracted from web text, paired with toxicity scores from a widely used commercial detector. They tested five prominent language models—including GPT-1, GPT-2, and GPT-3—measuring the likelihood and severity of toxic generation under unprompted conditions and when supplied with both toxic and non-toxic prompts. The study also evaluated data-based and decoding-based steering techniques, such as non-toxic fine-tuning, vocabulary manipulation, discriminator-guided generation, and basic word filtering, while auditing the training datasets for toxic content and unreliable sources.

The analysis revealed several critical findings. First, all evaluated models reliably generate high levels of toxicity even without any prompting, often reaching severe toxicity within 100 to 1,000 generations. Second, conditioning models on seemingly benign, non-toxic prompts still results in a toxicity probability near or above 50% across all models. Third, while advanced steering techniques—particularly domain-adaptive pretraining on non-toxic text and discriminator-guided decoding—reduce toxic outputs significantly more effectively than basic word blocklists, no tested method completely prevents toxic generation. Finally, pretraining corpora contain substantial abusive material, including 2.1% to 4.3% toxic documents, hundreds of thousands of texts linked from quarantined or banned internet forums, and significant material from low-reliability news sites.

These findings demonstrate that language models systematically acquire and retain toxicity directly from web-scraped pretraining data, and post-training interventions cannot currently guarantee safe outputs. For organizations deploying generative text systems, relying solely on simple profanity filters or standard steering algorithms leaves substantial compliance, reputational, and safety vulnerabilities.

To address these risks, the article recommends re-evaluating data collection strategies by implementing rigorous, transparent curation standards rather than relying on uncurated web scraping. Practitioners should prioritize data-intensive mitigations, such as domain-adaptive pretraining on curated non-toxic text, while combining them with advanced decoding controls. Furthermore, developers must adopt human-centered design frameworks to ensure data selection reflects appropriate safety norms without inadvertently censoring underrepresented dialects.

These conclusions are bounded by certain limitations, most notably the reliance on an automated toxicity classifier that exhibits known lexical and social biases. Additionally, because complete metadata was unavailable for proprietary pretraining corpora, the documented prevalence of toxic sources represents a conservative lower bound. Nonetheless, the evidence strongly supports high confidence in the finding that current models remain fundamentally prone to toxic degeneration without deeper structural interventions in training data curation.

Cover for RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Abstract

Pretrained neural language models (LMs) are prone to generating racist, sexist, or otherwise toxic language which hinders their safe deployment. We investigate the extent to which pretrained LMs can be prompted to generate toxic language, and the effectiveness of controllable text generation algorithms at preventing such toxic degeneration. We create and release RealToxicityPrompts, a dataset of 100K naturally occurring, sentence-level prompts derived from a large corpus of English web text, paired with toxicity scores from a widely-used toxicity classifier. Using RealToxicityPrompts, we find that pretrained LMs can degenerate into toxic text even from seemingly innocuous prompts. We empirically assess several controllable generation methods, and find that while data- or compute-intensive methods (e.g., adaptive pretraining on non-toxic data) are more effective at steering away from toxicity than simpler solutions (e.g., banning "bad" words), no current method is failsafe against neural toxic degeneration. To pinpoint the potential cause of such persistent toxic degeneration, we analyze two web text corpora used to pretrain several LMs (including GPT-2; Radford et. al, 2019), and find a significant amount of offensive, factually unreliable, and otherwise toxic content. Our work provides a test bed for evaluating toxic generations by LMs and stresses the need for better data selection processes for pretraining.

Table of Contents

  • 1 Introduction
  • 2 Operationalizing Toxicity
  • 2.1 Perspective API Toxicity
  • 2.2 Biases in Toxic Language Detection
  • 3 Out-of-the-Box Generation Toxicity
  • 3.1 Unprompted Toxicity in Neural Models
  • 4 RealToxicityPrompts
  • 4.1 Prompt Creation and Selection
  • 4.2 Prompted Toxicity in Neural Models
  • 5 Detoxifying Generations
  • 5.1 Data-Based Detoxification
  • 5.2 Decoding-Based Detoxification
  • 5.3 Effect of Controllable Solutions on Generation Toxicity
  • 6 Analyzing Toxicity in Web Text
  • 6.1 Toxicity in Web Text
  • 6.2 Sources of Toxic Content in Web Text
  • 7 Discussion and Recommendations
  • 8 Related Work
  • 9 Conclusion
  • 10 Acknowledgments
  • References
  • A Creating RealToxicityPrompts
  • B Modeling Details
  • B.1 Out of the Box Models
  • B.2 Detoxification Data
  • B.3 Detoxification Procedure
  • B.4 Generation Procedure
  • B.5 Hyperparameters
  • B.6 Verifying Language Model Quality
  • B.7 Comparing GPT-2 to GPT-2-medium
  • C Lexical Cues and Racial Bias in Toxicity Detection
  • C.1 Racial Bias in Perspective API
  • C.2 Profanity and Identity Mentions
  • D Further Analyses of Corpora
  • D.1 All Perspective API Toxicity Attributes
  • D.2 Further Analyses of OpenWebText Corpus and OpenAI-WT
  • D.3 BooksCorpus
  • E Generation Examples

Knowls

  1. Knowl 1 — REALTOXICITYPROMPTS Dataset

    definition

    REALTOXICITYPROMPTS is a benchmark dataset comprising 100,000 naturally occurring English sentence prompts designed to measure and evaluate neural toxic degeneration in conditional language generation.

    The dataset is constructed from the OpenWebTextCorpus (OWTC), a web text collection scraped from Reddit outbound links with a karma score ≥3\ge 3. To ensure uniform coverage across the toxicity spectrum, 25,000 sentences are sampled from each of four equal-width Perspective API toxicity score ranges: [0.00,0.25)[0.00, 0.25), [0.25,0.50)[0.25, 0.50), [0.50,0.75)[0.50, 0.75), and [0.75,1.00][0.75, 1.00]. Each sentence is split in half into a prompt prefix and a continuation.

    Dataset composition and summary statistics:

    • Toxic Prompts (Toxicity≥0.5\text{Toxicity} \ge 0.5): 21,744 prompts
    • Non-Toxic Prompts (Toxicity<0.5\text{Toxicity} < 0.5): 77,272 prompts
    • Prompt Length: 11.7±4.211.7 \pm 4.2 tokens
    • Continuation Length: 12.0±4.212.0 \pm 4.2 tokens
    • Mean Prompt Toxicity: 0.29±0.270.29 \pm 0.27
    • Mean Continuation Toxicity: 0.38±0.310.38 \pm 0.31

    Prompt toxicity and continuation toxicity exhibit a slight negative correlation (r=−0.08r = -0.08, p≤0.001p \le 0.001), indicating that toxicity within web text sentences is typically localized to one half of the sentence.

  2. Knowl 2 — Metrics for Evaluating Neural Toxic Degeneration

    experimental setup

    Toxic degeneration in conditional text generation is evaluated using Perspective API's calibrated TOXICITY score, which estimates the probability that a reader finds a comment rude, disrespectful, or unreasonable ([0,1][0, 1]). A generated text span is defined as toxic if TOXICITY≥0.5\text{TOXICITY} \ge 0.5.

    For conditional evaluation on REALTOXICITYPROMPTS, a language model generates k=25k = 25 independent completions per prompt using nucleus sampling (p=0.9p = 0.9) up to a maximum length of 20 tokens. Performance is measured across two primary metrics:

    1. Expected Maximum Toxicity: The mean and standard deviation of the maximum toxicity score observed across k=25k = 25 generations for each prompt: E[max⁡i=1,…,kTOXICITY(yi)]\mathbb{E}\left[\max_{i=1,\dots,k} \text{TOXICITY}(y_i)\right] where yiy_i is the ii-th generated completion.

    2. Empirical Toxicity Probability: The empirical probability of the model generating at least one toxic completion (TOXICITY≥0.5\text{TOXICITY} \ge 0.5) among the k=25k = 25 completions: P(∃i∈{1,…,k} s.t. TOXICITY(yi)≥0.5)P\left(\exists i \in \{1,\dots,k\} \text{ s.t. } \text{TOXICITY}(y_i) \ge 0.5\right)

    For unprompted generation, the expected maximum toxicity over NN generations (N≤10,000N \le 10,000) is estimated by generating a pool of 10,000 unconditional samples and running 1,000 bootstrap iterations of sampling NN generations with replacement.

  3. Knowl 3 — Unprompted Toxic Degeneration in Autoregressive Language Models

    empirical result

    Pretrained autoregressive Transformers generate toxic text even without any textual prompt when initialized solely with start-of-sequence or domain control tokens (e.g., <|endoftext|> for GPT-2 and GPT-3, . for GPT-1, and [Links] or [Wiki] for CTRL).

    Bootstrap analysis of NN unprompted generations demonstrates:

    • Within N=100N = 100 unprompted generations, all evaluated models (GPT-1, GPT-2, GPT-3 175B DaVinci, CTRL, and CTRL-WIKI) achieve an expected maximum toxicity score exceeding 0.50.5. For example, GPT-2 generates an expected maximum toxicity of 0.650.65 within only 100 generations.
    • Within N=1,000N = 1,000 unprompted generations, GPT-1, GPT-2, GPT-3, and CTRL reach an expected maximum toxicity exceeding 0.90.9.
    • GPT-1 reaches higher toxicity with fewer generations than other models, corresponding to the high toxicity of its BookCorpus pretraining data.
    • GPT-2 and CTRL show similar toxicity trajectories due to web text pretraining data overlap; GPT-3 (175B) closely mirrors GPT-2's curve.
    • CTRL-WIKI (conditioned on the Wikipedia domain token) exhibits substantially lower unprompted toxicity than models conditioned on web text across all generation pool sizes NN.
  4. Knowl 4 — Prompted Toxicity in Pretrained Language Models

    data/table

    Conditioning language models on REALTOXICITYPROMPTS demonstrates that even non-toxic prompts frequently induce toxic completions across standard pretrained autoregressive models.

    The following table reports the Expected Maximum Toxicity (mean with standard deviation subscript) and Empirical Toxicity Probability over k=25k = 25 generations per prompt (nucleus sampling p=0.9p = 0.9, 20 max tokens):

    Expected Max Toxicity Toxicity Probability
    Model Toxic Prompts Non-Toxic Prompts Toxic Prompts Non-Toxic Prompts
    GPT-1 0.780.180.78_{0.18} 0.580.220.58_{0.22} 0.90 0.60
    GPT-2 0.750.190.75_{0.19} 0.510.220.51_{0.22} 0.88 0.48
    GPT-3 0.750.200.75_{0.20} 0.520.230.52_{0.23} 0.87 0.50
    CTRL 0.730.200.73_{0.20} 0.520.210.52_{0.21} 0.85 0.50
    CTRL-W 0.710.200.71_{0.20} 0.490.210.49_{0.21} 0.82 0.44

    Key empirical results:

    1. On toxic prompts, all models produce toxic continuations with probability between 0.820.82 (CTRL-WIKI) and 0.900.90 (GPT-1), and expected maximum toxicity ranging from 0.710.71 to 0.780.78.
    2. On non-toxic prompts, every model produces toxic text with high probability (0.440.44 to 0.600.60) and expected maximum toxicity near or above 0.500.50 (0.490.49 to 0.580.58).
    3. Even CTRL-WIKI (trained on Wikipedia text) exhibits prompted toxicity probabilities comparable to web-trained models (0.440.44 on non-toxic prompts, 0.820.82 on toxic prompts).
  5. Knowl 5 — Vocabulary Shifting for Decoding-Based Detoxification

    model/method

    Vocabulary Shifting (VOCAB-SHIFT) is a decoding-time detoxification method that modifies the unnormalized token logit distribution of a language model to systematically upweight non-toxic tokens without fine-tuning model parameters.

    Given a language model vocabulary VV and a balanced text corpus with toxicity annotations, a two-dimensional association matrix W∈R∣V∣×2W \in \mathbb{R}^{|V| \times 2} is learned, where the columns capture token associations with non-toxicity and toxicity. At decoding step tt, the modified logit vector z′∈R∣V∣\mathbf{z}' \in \mathbb{R}^{|V|} is computed as: z′=z+βWt\mathbf{z}' = \mathbf{z} + \beta W \mathbf{t} where:

    • z∈R∣V∣\mathbf{z} \in \mathbb{R}^{|V|} is the language model's original logit output vector over the vocabulary.
    • t=[1,0]⊤∈R2\mathbf{t} = [1, 0]^\top \in \mathbb{R}^2 is a one-hot control vector targeting the non-toxic class.
    • W∈R∣V∣×2W \in \mathbb{R}^{|V| \times 2} is the token toxicity association matrix.
    • β∈R\beta \in \mathbb{R} is the boosting strength hyperparameter (set to β=3\beta = 3).

    After logit modification, standard sampling (e.g., nucleus sampling p=0.9p = 0.9) is performed on the updated distribution.

  6. Knowl 6 — Data-Based and Decoding-Based Detoxification Frameworks

    model/method

    Detoxification strategies for autoregressive language models (evaluated on GPT-2) are divided into two paradigms:

    1. Data-Based Detoxification (requiring continued pretraining):

      • Domain-Adaptive Pretraining (DAPT Non-Toxic): Continued pretraining of GPT-2 on a filtered subset of approximately 150,000 non-toxic documents from OpenWebTextCorpus (OWTC). As a control, DAPT (Toxic) is trained on the complementary toxic subset.
      • Attribute Conditioning (ATCON): Prepending attribute tokens (<|toxic|> or <|nontoxic|>) to training documents during continued pretraining. During generation, the <|nontoxic|> token is prepended to the prompt to steer output.
    2. Decoding-Based Detoxification (inference-time parameter-frozen methods):

      • Vocabulary Shifting (VOCAB-SHIFT): Adding a learned linear term βWt\beta W \mathbf{t} (eta = 3) to the model's vocabulary logits to favor non-toxic tokens.
      • Word Filtering (WORD FILTER): Setting the generation probability of any word contained in a blocklist of profanities, slurs, and obscenities to zero.
      • Plug and Play Language Models (PPLM): Dynamically updating past and present hidden activation states at each generation step using backpropagated gradients from an external toxicity classifier discriminator.
  7. Knowl 7 — Comparative Performance of Detoxification Methods on GPT-2

    data/table

    Evaluating data-based and decoding-based detoxification methods on GPT-2 under REALTOXICITYPROMPTS reveals that while all methods mitigate toxicity, none completely eliminates toxic degeneration.

    The table below reports the Expected Maximum Toxicity and Toxicity Probability over 25 generations per prompt across Unprompted, Toxic Prompt, and Non-Toxic Prompt settings:

    Expected Max Toxicity Toxicity Probability
    Category Model Unprompted Toxic Non-Toxic Unprompted Toxic Non-Toxic
    Baseline GPT-2 0.440.170.44_{0.17} 0.750.190.75_{0.19} 0.510.220.51_{0.22} 0.33 0.88 0.48
    Data-based DAPT (Non-Toxic) 0.300.13\mathbf{0.30}_{0.13} 0.570.23\mathbf{0.57}_{0.23} 0.370.19\mathbf{0.37}_{0.19} 0.09\mathbf{0.09} 0.59\mathbf{0.59} 0.23\mathbf{0.23}
    DAPT (Toxic) 0.800.160.80_{0.16} 0.850.150.85_{0.15} 0.690.230.69_{0.23} 0.93 0.96 0.77
    ATCON 0.420.170.42_{0.17} 0.730.200.73_{0.20} 0.490.220.49_{0.22} 0.26 0.84 0.44
    Decoding-based VOCAB-SHIFT 0.430.180.43_{0.18} 0.700.210.70_{0.21} 0.460.220.46_{0.22} 0.31 0.80 0.39
    PPLM∗^* 0.280.11\mathbf{0.28}_{0.11} 0.520.26\mathbf{0.52}_{0.26} 0.320.19\mathbf{0.32}_{0.19} 0.05\mathbf{0.05} 0.49\mathbf{0.49} 0.17\mathbf{0.17}
    WORD FILTER 0.420.160.42_{0.16} 0.680.190.68_{0.19} 0.480.200.48_{0.20} 0.27 0.81 0.43

    (∗^*PPLM was evaluated on 10K prompts with 10 generations per prompt due to inference cost: 14 s/generation for PPLM vs. 0.2 s/generation for baseline GPT-2).

    Key comparisons:

    1. DAPT (Non-Toxic) is the most effective data-based method, reducing non-toxic prompt toxicity probability from 0.480.48 to 0.230.23 and unprompted toxicity probability from 0.330.33 to 0.090.09.
    2. PPLM achieves the lowest toxicity among decoding-based methods (toxicity probability of 0.170.17 on non-toxic prompts and 0.490.49 on toxic prompts), but incurs a 70×70\times generation latency increase.
    3. WORD FILTER and ATCON show modest reductions, lowering non-toxic prompt toxicity probability by only 0.050.05 and 0.040.04 respectively.
    4. Even the best steering methods still produce toxic completions for 17–23%17\text{--}23\% of non-toxic prompts and 49–59%49\text{--}59\% of toxic prompts.
  8. Knowl 8 — Toxicity Analysis of Web Pretraining Corpora (OpenAI WebText and OpenWebTextCorpus)

    empirical result

    Large-scale evaluation of the pretraining corpora OpenAI WebText (OPENAI-WT, 40 GB, ~8M documents) and OpenWebTextCorpus (OWTC, 38 GB, ~8M documents) using Perspective API shows that both corpora contain substantial quantities of toxic text:

    • OWTC Toxicity: 2.1%2.1\% of all documents have TOXICITY≥0.5\text{TOXICITY} \ge 0.5.
    • OPENAI-WT Toxicity: 4.3%4.3\% of all documents have TOXICITY≥0.5\text{TOXICITY} \ge 0.5.
    • Corpus Overlap: Locality-sensitive hashing (LSH) with MinHash indicates a 29%29\% document overlap between the two corpora (at least 2.3 million overlapping documents).
    • Impact of Filtering: Although OpenAI-WT used a blocklist of offensive and sexually explicit subreddits during dataset creation, it contains approximately twice the proportion of toxic documents compared to OWTC. Both percentages are consistent with web-wide abusive content rates (0.1–3%0.1\text{--}3\%).
  9. Knowl 9 — Toxicity Provenance from Low-Reliability News and Banned Subreddits

    empirical result

    Cross-referencing OpenWebTextCorpus (OWTC) and OpenAI WebText (OPENAI-WT) documents with domain factuality ratings and Reddit submission metadata reveals two primary drivers of corpus toxicity:

    1. News Source Factuality: There is a significant negative correlation between news site reliability and the proportion of toxic documents (Spearman ρ=−0.35\rho = -0.35). Low-reliability news sources contain a substantially higher fraction of toxic content than high-reliability sources. At least 12%12\% (272,000272,000 documents) of the overlapping news documents between OPENAI-WT and OWTC originate from low- or mixed-reliability news sites.
    2. Banned and Quarantined Subreddits: At least 3%3\% (212,000212,000 documents) of OWTC documents come from outbound links shared on Reddit communities that were subsequently quarantined or banned for abusive behavior. In the overlapping OPENAI-WT and OWTC subset, at least 63,00063,000 documents were submitted to quarantined or banned subreddits, including over 40,00040,000 documents from /r/The_Donald and over 4,0004,000 documents from /r/WhiteRights.
  10. Knowl 10 — Universal Trigger Prompts for Severe Toxic Degeneration

    empirical result

    Within REALTOXICITYPROMPTS, a specific subset of prompts acts as universal triggers for toxic generation across multiple distinct model architectures:

    • Multi-Model Universal Triggers: Exactly 327 prompts induce at least one generation with TOXICITY≥0.9\text{TOXICITY} \ge 0.9 across all evaluated models (GPT-1, GPT-2, GPT-3, CTRL, CTRL-WIKI, and all steered/detoxified GPT-2 variants).
    • Out-of-the-Box LM Triggers: Exactly 1,225 prompts induce TOXICITY≥0.9\text{TOXICITY} \ge 0.9 across all five unsteered baseline models (GPT-1, GPT-2, GPT-3, CTRL, CTRL-WIKI).

    Qualitative inspection indicates that while many of these challenging prompts contain explicit toxicity, several are seemingly benign sentence fragments featuring open quotation marks or beginnings of multiword expressions (e.g., "So, I'm starting to think she's full...", "So if you grab a woman by the..."). Additionally, at least 10%10\% of the 1,225 challenging prompts originate from unreliable news domains or banned/quarantined subreddits.

  11. Knowl 11 — Limitations of Automated Toxicity Scoring and Model Coverage

    limitation

    The methodology for measuring and mitigating toxic degeneration possesses three primary stated limitations:

    1. Toxicity Classifier Biases: Perspective API relies heavily on lexical surface cues (e.g., swearwords and slurs), which leads to false positives on benign minority dialect text (such as African American English) or identity mentions, while potentially failing to detect subtle social biases, implicit hate speech, and microaggressions.
    2. Architectural Scope: The evaluation is restricted to autoregressive Transformer language models (GPT-1, GPT-2, GPT-3, CTRL) and does not cover masked language models (e.g., BERT) or non-neural generation models.
    3. Lower-Bound Metadata Estimates: Because OpenAI-WT lacks original URL metadata and historical Reddit scraping coverage is incomplete, reported quantities of documents derived from unreliable news domains and banned subreddits represent lower-bound estimates.

Coverage note — No substantial contributed material was omitted; all key definitions, dataset creation steps, unprompted/prompted generation evaluations, detoxification techniques, pretraining corpus analyses, and stated limitations are covered.

References

  1. 1.Xavier Ferrer Aran, T. V. Nuenen, J. M. Such, and N. Criado. 2020. Discovering and categorising language biases in Reddit.
  2. 2.Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James Glass, and Preslav Nakov. 2018. Predicting factuality of reporting and bias of news media sources. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3528–3539, Brussels, Belgium. Association for Computational Linguistics.
  3. 3.Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: Allocative versus representational harms in machine learning. In SIGCIS.
  4. 4.Michael Barthel, Galen Stocking, Jesse Holcomb, and Amy Mitchell. 2016. Seven-in-Ten Reddit users get news on the site. Accessed: 2020-6-2.
  5. 5.Christine Basta, Marta R. Costa-jussà, and Noe Casas. 2019. Evaluating the underlying gender bias in contextualized word embeddings. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 33–39, Florence, Italy. Association for Computational Linguistics.
  6. 6.Emily M. Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  7. 7.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  8. 8.Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016. Demographic dialectal variation in social media: A case study of African-American English. In EMNLP.
  9. 9.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2016. Enriching word vectors with subword information. arXiv preprint arXiv:1607.04606.
  10. 10.Luke Breitfeller, Emily Ahn, David Jurgens, and Yulia Tsvetkov. 2019. Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1664–1674, Hong Kong, China. Association for Computational Linguistics.
  11. 11.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are Few-Shot learners.
  12. 12.Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Xiaodong Song. 2019. The secret sharer: Evaluating and testing unintended memorization in neural networks. In USENIX Security Symposium.
  13. 13.Mia Xu Chen, Benjamin N. Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Y. Lu, Jackie Tsay, Yinan Wang, Andrew M. Dai, Zhifeng Chen, Timothy Sohn, and Yonghui Wu. 2019. Gmail smart compose: Real-time assisted writing. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining.
  14. 14.Anna Chung. 2019. How automated tools discriminate against black language. Accessed: 2019-03-02.
  15. 15.Gloria Cowan and Désirée Khatchadourian. 2003. Empathy, ways of knowing, and interdependence as mediators of gender differences in attitudes toward hate speech and freedom of speech. Psychology of women quarterly, 27(4):300–308.
  16. 16.Cristian Danescu-Niculescu-Mizil, Moritz Sudhof, Dan Jurafsky, Jure Leskovec, and Christopher Potts. 2013. A computational approach to politeness with application to social factors. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 250–259, Sofia, Bulgaria. Association for Computational Linguistics.
  17. 17.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  18. 18.Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019. Racial bias in hate speech and abusive language detection datasets. In Proceedings of the Third Workshop on Abusive Language Online, pages 25–35, Florence, Italy. Association for Computational Linguistics.
  19. 19.Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman. 2020. Bringing the people back in: Contesting benchmark machine learning datasets. In ICML Workshop on Participatory Approaches to Machine Learning.
  20. 20.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  21. 21.Emily Dinan, A. Fan, Ledell Yu Wu, J. Weston, Douwe Kiela, and Adina Williams. 2020. Multi-dimensional gender bias classification.
  22. 22.Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019. Build it break it fix it for dialogue safety: Robustness from adversarial human attack. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4537–4546, Hong Kong, China. Association for Computational Linguistics.
  23. 23.Carl DiSalvo, Andrew Clement, and Volkmar Pipek. 2012. Communities: Participatory design for, with and by communities.
  24. 24.Lucas Dixon, John Li, Jeffrey Scott Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and mitigating unintended bias in text classification. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society.
  25. 25.Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019. Show your work: Improved reporting of experimental results. In EMNLP, pages 2185–2194, Hong Kong, China. Association for Computational Linguistics.
  26. 26.Jacob Eisenstein, Amr Ahmed, and Eric P. Xing. 2011. Sparse additive generative models of text. In Proceedings of the 28th International Conference on International Conference on Machine Learning, ICML’11, page 1041–1048, Madison, WI, USA. Omnipress.
  27. 27.Ethan Fast, Tina Vachovsky, and Michael S. Bernstein. 2016. Shirtless and dangerous: Quantifying linguistic signals of gender bias in an online fiction writing community.
  28. 28.Jessica Ficler and Yoav Goldberg. 2017. Controlling linguistic style aspects in neural language generation. In Proceedings of the Workshop on Stylistic Variation, pages 94–104, Copenhagen, Denmark. Association for Computational Linguistics.
  29. 29.Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018. Large scale crowdsourcing and characterization of Twitter abusive behavior. In ICWSM.
  30. 30.Batya Friedman, Peter H Kahn, and Alan Borning. 2008. Value sensitive design and information systems. The handbook of information and computer ethics, pages 69–101.
  31. 31.Sayan Ghosh, Mathieu Chollet, Eugene Laksana, Louis-Philippe Morency, and Stefan Scherer. 2017. Affect-LM: A neural language model for customizable affective text generation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 634–642, Vancouver, Canada. Association for Computational Linguistics.
  32. 32.Aaron Gokaslan and Vanya Cohen. 2019. Openwebtext corpus.
  33. 33.Jennifer Golbeck, Zahra Ashktorab, Rashad O. Banjo, Alexandra Berlinger, Siddharth Bhagwan, Cody Buntain, Paul Cheakalos, Alicia A. Geller, Quint Gergory, Rajesh Kumar Gnanasekaran, Raja Rajan Gunasekaran, Kelly M. Hoffman, Jenny Hottle, Vichita Jienjitlert, Shivika Khare, Ryan Lau, Marianna J. Martindale, Shalmali Naik, Heather L. Nixon, Piyush Ramachandran, Kristine M. Rogers, Lisa Rogers, Meghna Sardana Sarin, Gaurav Shahane, Jayanee Thanki, Priyanka Vengataraman, Zijian Wan, and Derek Michael Wu. 2017. A large labeled corpus for online harassment research. In Proceedings of the 2017 ACM on Web Science Conference, WebSci ’17, page 229–233, New York, NY, USA. Association for Computing Machinery.
  34. 34.Lisa Green. 2002. African American English: A Linguistic Introduction, 8.3.2002 edition edition. Cambridge University Press.
  35. 35.Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  36. 36.Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. 2018. Learning to write with cooperative discriminators. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1638–1649, Melbourne, Australia. Association for Computational Linguistics.
  37. 37.Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. International Conference on Learning Representations.
  38. 38.Matthew Honnibal and Ines Montani. 2017. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear.
  39. 39.Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl. 2020. Social biases in NLP models as barriers for persons with disabilities. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5491–5501, Online. Association for Computational Linguistics.
  40. 40.Abigail Z. Jacobs and Hanna M. Wallach. 2019. Measurement and fairness.
  41. 41.Eun Seo Jo and Timnit Gebru. 2020. Lessons from archives: Strategies for collecting sociocultural data in machine learning. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, page 306–316, New York, NY, USA. Association for Computing Machinery.
  42. 42.Mladen Karan and Jan Šnajder. 2019. Preemptive toxic language detection in Wikipedia comments using thread-level context. In Proceedings of the Third Workshop on Abusive Language Online, pages 129–134, Florence, Italy. Association for Computational Linguistics.
  43. 43.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. CTRL: A conditional Transformer language model for controllable generation.
  44. 44.Adam King. 2019. Talk to Transformer. Accessed 06-02-2020.
  45. 45.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526.
  46. 46.Pang Wei Koh and Percy Liang. 2017. Understanding black-box predictions via influence functions. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, page 1885–1894. JMLR.org.
  47. 47.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring bias in contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 166–172, Florence, Italy. Association for Computational Linguistics.
  48. 48.Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner. 1998. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324.
  49. 49.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach.
  50. 50.Xinyao Ma, Maarten Sap, Hannah Rashkin, and Yejin Choi. 2020. PowerTransformer: Unsupervised controllable revision for biased language correction. In EMNLP.
  51. 51.Adrienne Massanari. 2017. #gamergate and the fappening: How Reddit’s algorithm, governance, and culture support toxic technocultures. New Media & Society, 19(3):329–346.
  52. 52.Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On measuring social biases in sentence encoders. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 622–628, Minneapolis, Minnesota. Association for Computational Linguistics.
  53. 53.Kris McGuffie and Alex Newhouse. 2020. The radicalization risks of GPT-3 and advanced neural language models.
  54. 54.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, page 220–229, New York, NY, USA. Association for Computing Machinery.
  55. 55.Shruthi Mohan, Apala Guha, Michael Harris, Fred Popowich, Ashley Schuster, and Chris Priebe. 2017. The impact of toxic language on the health of Reddit communities. In Canadian Conference on AI.
  56. 56.Ji Ho Park and Pascale Fung. 2017. One-step and two-step classification for abusive language detection on Twitter. In Proceedings of the Workshop on Abusive Language Online.
  57. 57.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. dAlché Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc.
  58. 58.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  59. 59.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  60. 60.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text Transformer.
  61. 61.Ashwin Rajadesingan, Paul Resnick, and Ceren Budak. 2020. Quick, community-specific learning: How distinctive toxicity norms are maintained in political subreddits. Proceedings of the International AAAI Conference on Web and Social Media, 14(1):557–568.
  62. 62.Anand Rajaraman and Jeffrey David Ullman. 2011. Mining of massive datasets. Cambridge University Press.
  63. 63.Aja Romano. 2017. Reddit just banned one of its most toxic forums. but it won’t touch The_Donald. Accessed: 2020-02-23.
  64. 64.Björn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. 2017. Measuring the reliability of hate speech annotations: the case of the european refugee crisis. In NLP 4 CMC Workshop.
  65. 65.Elizabeth Sanders. 2002. From user-centered to participatory design approaches, pages 1–7.
  66. 66.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678, Florence, Italy. Association for Computational Linguistics.
  67. 67.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5477–5490, Online. Association for Computational Linguistics.
  68. 68.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  69. 69.Serge Sharoff. 2020. Know thy corpus! robust methods for digital curation of web corpora.
  70. 70.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3407–3412, Hong Kong, China. Association for Computational Linguistics.
  71. 71.Wessel Stoop, Florian Kunneman, Antal van den Bosch, and Ben Miller. 2019. Detecting harassment in real-time as conversations develop. In Proceedings of the Third Workshop on Abusive Language Online, pages 19–24, Florence, Italy. Association for Computational Linguistics.
  72. 72.Akhilesh Sudhakar, Bhargav Upadhyay, and Arjun Maheswaran. 2019. “transforming” delete, retrieve, generate approach for controlled text style transfer. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3269–3279, Hong Kong, China. Association for Computational Linguistics.
  73. 73.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6000–6010, Red Hook, NY, USA. Curran Associates Inc.
  74. 74.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal adversarial triggers for attacking and analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2153–2162, Hong Kong, China. Association for Computational Linguistics.
  75. 75.Alex Wang and Kyunghyun Cho. 2019. Bert has a mouth, and it must speak: Bert as a markov random field language model.
  76. 76.Zeerak Waseem. 2016. Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter. In Proceedings of the First Workshop on NLP and Computational Social Science, pages 138–142, Austin, Texas. Association for Computational Linguistics.
  77. 77.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2019. Neural text generation with unlikelihood training.
  78. 78.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, and Jamie Brew. 2019. HuggingFace’s Transformers: State-of-the-art natural language processing.
  79. 79.Bianca Zadrozny and Charles Elkan. 2002. Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’02, page 694–699, New York, NY, USA. Association for Computing Machinery.
  80. 80.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. In H. Wallach, H. Larochelle, A. Beygelzimer, F. dAlché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 9054–9065. Curran Associates, Inc.
  81. 81.Justine Zhang, Jonathan Chang, Cristian Danescu-Niculescu-Mizil, Lucas Dixon, Yiqing Hua, Dario Taraborelli, and Nithum Thain. 2018. Conversations gone awry: Detecting early signs of conversational failure. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1350–1361, Melbourne, Australia. Association for Computational Linguistics.
  82. 82.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019. Gender bias in contextualized word embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 629–634, Minneapolis, Minnesota. Association for Computational Linguistics.
  83. 83.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. 2015 IEEE International Conference on Computer Vision (ICCV).
  84. 84.Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences.

Citation

MLA
Gehman, S., et al. “RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models”. arXiv, 2020, http://arxiv.org/abs/2009.11462v2.
APA
Gehman, S., Gururangan, S., Sap, M., Choi, Y., & Smith, N. A. (2020). RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. arXiv. http://arxiv.org/abs/2009.11462v2
Chicago
Gehman, S., S. Gururangan, M. Sap, Y. Choi, and N. A. Smith. 2020. “RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models”. arXiv. http://arxiv.org/abs/2009.11462v2.
Harvard
Gehman, S. et al. (2020) “RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2009.11462v2.
Vancouver
1. Gehman S, Gururangan S, Sap M, Choi Y, Smith NA (2020) RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. arXiv

BibTeX

@article{gehman2020realtoxicityprompts,
  title = {RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models},
  author = {Gehman, Samuel and Gururangan, Suchin and Sap, Maarten and Choi, Yejin and Smith, Noah A.},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2009.11462v2},
  eprint = {2009.11462}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF