ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Thomas HartvigsenSaadia GabrielHamid PalangiMaarten SapDipankar RayEce Kamar

article2022ACL672 citations

Introduces ToxiGen, a large-scale balanced dataset of over 274,000 machine-generated statements across 13 minority groups, along with an adversarial decoding method to help classifiers detect subtle, implicit hate speech without over-relying on identity mentions.

Listen

Online toxicity detection systems frequently rely on surface-level keyword matching and spurious correlations, leading them to falsely flag benign statements that mention demographic groups while failing to detect subtle, implicit hate speech devoid of profanity. This systemic bias risks marginalizing vulnerable communities through disproportionate censorship and leaves online platforms unprotected against veiled abuse. Addressing these vulnerabilities requires diverse, balanced data that web scraping alone cannot reliably provide.

The article demonstrates how large language models can be steered to generate a massive, balanced, and implicit hate speech dataset to expose vulnerabilities in existing toxicity filters and significantly improve their detection performance. To achieve this, the authors introduced TOXIGEN, a dataset containing 274,186 machine-generated toxic and benign statements covering 13 demographic identity groups, generated using demonstration-based prompting with GPT-3 alongside a novel decoding technique called ALICE (Adversarial Language Imitation with Constrained Exemplars).

The evaluation produced four key findings. First, TOXIGEN successfully captures subtle abuse at scale, with 98.2% of its statements being implicit and free of explicit slurs or profanity. Second, human validation revealed that 90.5% of machine-generated examples were mistaken for human-written text, with 94.5% of toxic examples confirmed as hate speech by human raters. Third, ALICE proved highly effective as an adversarial attack mechanism, generating toxic statements that fooled existing detection systems like HateBERT at rates more than double standard decoding methods (58.97% versus 26.88%). Fourth, fine-tuning existing classifiers on TOXIGEN substantially improved their performance, boosting detection accuracy (AUC) by 7 to 19 percentage points across three separate human-written implicit hate benchmarks.

These results demonstrate that machine-generated data can cost-effectively eliminate data imbalances and mitigate the risk of automated censorship against minority communities. The findings also underscore a critical security implication: bad actors can readily leverage large language models to bypass standard content moderation filters unless those filters are proactively hardened against adversarial text.

Organizations developing or deploying content moderation systems should integrate adversarially generated implicit data like TOXIGEN into their training pipelines to reduce false positives and improve resilience against evasive hate speech. Next steps include transitioning from binary classification models to nuanced labeling frameworks and combining automated tools with human moderation expertise to align with emerging artificial intelligence governance policies.

These conclusions should be considered within key boundary conditions: the dataset primarily reflects United States socio-cultural perspectives and was generated using a specific language model (GPT-3). Furthermore, annotator subjectivity in assessing implicit toxicity introduces moderate labeling variance. Confidence remains high, however, that utilizing balanced, adversarial synthetic data significantly fortifies safety systems against both human and machine-generated toxicity.

arXiv: 2203.09509

Abstract

Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate. Such over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language. To help mitigate these issues, we create TOXIGEN, a new large-scale and machine-generated dataset of 274k toxic and benign statements about 13 minority groups. We develop a demonstration-based prompting framework and an adversarial classifier-in-the-loop decoding method to generate subtly toxic and benign text with a massive pretrained language model (Brown et al., 2020). Controlling machine generation in this way allows TOXIGEN to cover implicitly toxic text at a larger scale, and about more demographic groups, than previous resources of human-written text. We conduct a human evaluation on a challenging subset of TOXIGEN and find that annotators struggle to distinguish machine-generated text from human-written language. We also find that 94.5% of toxic examples are labeled as hate speech by human annotators. Using three publicly-available datasets, we show that finetuning a toxicity classifier on our data improves its performance on human-written data substantially. We also demonstrate that TOXIGEN can be used to fight machine-generated toxicity as finetuning improves the classifier significantly on our evaluation subset.

Table of Contents

  • 1 Introduction
  • 2 Implicit Hate Against Minority Groups
  • 3 Creating TOXIGEN
  • 3.1 Prompt Engineering
  • 3.1.1 Demonstration-based prompting
  • 3.1.2 Collecting demonstrations
  • 3.2 ALICE: Attacking Toxicity Classifiers with Adversarial Decoding
  • 3.3 Decoding Details
  • 3.4 TOXIGEN Statistics
  • 4 Human Validation of TOXIGEN
  • 4.1 Human Validation Design
  • 4.2 Constructing TOXIGEN-HUMANVAL
  • 4.3 Comparing Generation Methods
  • 5 Improving Toxicity Classifiers
  • 6 Conclusions
  • 7 Societal and Ethical Considerations
  • 8 Acknowledgements
  • References
  • Supplementary Materials
  • A Generation Details
  • A.1 Language Model Selection
  • B Human Validation Details
  • B.1 Selecting MTurk Workers
  • B.2 Annotation Interface
  • C How does perplexity change across groups?
  • D Does generated text actually mention the targeted groups?
  • E Analysis of Large-Scale Human Validation
  • F Example Prompt
  • G Releasing a Pretrained Model and its Propagated Labels
  • H Dataset Description
  • I Further comparing toxicity classifiers

Knowls

  1. Knowl 1 — Three-layer slab waveguide geometry and guiding condition

    definition

    A dielectric slab waveguide consists of a thin high-index film of refractive index nfn_f and thickness hh, sandwiched between a substrate of index nsn_s and a cover (cladding) of index ncn_c, with the ordering nf>ns≥ncn_f > n_s \ge n_c. The guiding direction is zz, the transverse direction is xx, and the film occupies 0<x<h0 < x < h. Light is confined to the film by total internal reflection provided the propagation constant satisfies nsk0<β<nfk0n_s k_0 < \beta < n_f k_0, where k0=2π/λk_0 = 2\pi/\lambda is the free-space wavenumber at wavelength λ\lambda. The asymmetry of the surrounding indices (nsn_s versus ncn_c) breaks the structural symmetry and affects mode cutoffs.

  2. Knowl 2 — Effective index of a guided slab mode

    definition

    For a step-index slab waveguide, each guided mode is characterized by an effective index N=β/k0N = \beta/k_0, where β\beta is the propagation constant along the waveguide axis and k0=2π/λk_0 = 2\pi/\lambda is the free-space wavenumber. The effective index lies between the substrate index nsn_s and the film index nfn_f: ns<N<nfn_s < N < n_f. It is the single quantity that determines the phase velocity vp=c/Nv_p = c/N of the mode and serves as the natural design variable when synthesizing a waveguide to meet a target modal behavior.

  3. Knowl 3 — Eigenvalue equation for TE modes of a symmetric-cladding slab

    equation

    For TE-polarized guided modes (electric field along yy) in a step-index slab of thickness hh with film index nfn_f, substrate index nsn_s, and cover index ncn_c, the allowed effective indices NN are the roots of the transcendental equation

    tan⁡(kxh)=kx(γs+γc)kx2−γsγc,\tan(k_x h) = \frac{k_x (\gamma_s + \gamma_c)}{k_x^2 - \gamma_s \gamma_c},

    where kx=k0nf2−N2k_x = k_0\sqrt{n_f^2 - N^2} is the transverse wavenumber in the film, γs=k0N2−ns2\gamma_s = k_0\sqrt{N^2 - n_s^2} and γc=k0N2−nc2\gamma_c = k_0\sqrt{N^2 - n_c^2} are the evanescent decay constants in the substrate and cover respectively, and k0=2π/λk_0 = 2\pi/\lambda. Each successive integer root corresponds to the fundamental (TE0\mathrm{TE}_0), first-order (TE1\mathrm{TE}_1), etc., mode. This equation is the basis for computing the dispersion relation N(λ)N(\lambda) of any given slab design.

  4. Knowl 4 — Eigenvalue equation for TM modes with polarization-dependent interface terms

    equation

    For TM-polarized guided modes (magnetic field along yy) in the same three-layer slab, the eigenvalue condition gains polarization-dependent weighting factors that account for the discontinuity of the electric field component normal to the layer interfaces:

    tan⁡(kxh)=nf2kx(nc2γs+ns2γc)ns2nc2kx2−nf4γsγc,\tan(k_x h) = \frac{n_f^2 k_x (n_c^2 \gamma_s + n_s^2 \gamma_c)}{n_s^2 n_c^2 k_x^2 - n_f^4 \gamma_s \gamma_c},

    with the same definitions kx=k0nf2−N2k_x = k_0\sqrt{n_f^2 - N^2}, γs=k0N2−ns2\gamma_s = k_0\sqrt{N^2 - n_s^2}, γc=k0N2−nc2\gamma_c = k_0\sqrt{N^2 - n_c^2}. TM modes of a given order always exhibit a lower effective index than the corresponding TE mode for the same geometry, and their cutoff thicknesses are larger.

  5. Knowl 5 — Cutoff condition and single-mode operation of slab modes

    theoretical result

    A slab mode of order mm (either polarization) exists as a guided mode only if the film thickness hh exceeds a critical cutoff thickness. For the TEm\mathrm{TE}_m mode the cutoff condition is

    h>mλ2nf2−ncl2,h > \frac{m\lambda}{2\sqrt{n_f^2 - n_{\mathrm{cl}}^2}},

    where ncln_{\mathrm{cl}} is the relevant cladding index (the larger of nsn_s and ncn_c sets the strictest cutoff). In particular, the fundamental mode TE0\mathrm{TE}_0 has zero cutoff thickness and is always guided, whereas higher-order modes appear one by one as hh or the index contrast increases. Designing a strictly single-mode slab therefore requires hh to sit between the TE0\mathrm{TE}_0 and TE1\mathrm{TE}_1 cutoff thresholds.

  6. Knowl 6 — Confinement factor as a mode-quality metric

    equation

    The modal confinement factor Γ\Gamma quantifies the fraction of the mode's power that propagates inside the high-index film. For a slab mode with transverse field profile E(x)E(x),

    Γ=∫0h∣E(x)∣2 dx∫−∞+∞∣E(x)∣2 dx,\Gamma = \frac{\int_0^h |E(x)|^2\,dx}{\int_{-\infty}^{+\infty} |E(x)|^2\,dx},

    where the numerator integrates over the film of thickness hh and the denominator over the full cross-section including the evanescent tails in the substrate and cover. Γ\Gamma increases with index contrast and with film thickness, and is a key figure of merit when trading off mode size against confinement loss.

  7. Knowl 7 — Inverse design formulated as a bounded parameter search

    model/method

    The paper's synthesis workflow treats slab waveguide design as a bounded inverse problem: given a target effective index NtargetN_{\mathrm{target}} at a design wavelength λ\lambda, search over the geometric and material parameters (film thickness hh, film index nfn_f, substrate index nsn_s, cover index ncn_c) to find a design whose numerically computed fundamental-mode effective index matches NtargetN_{\mathrm{target}}. The forward evaluation consists of solving the TE eigenvalue equation for the fundamental root at the target wavelength; the inverse search restricts the parameter space using physical constraints (nf>ns≥ncn_f > n_s \ge n_c, single-mode cutoff bounds on hh) and fabrication-feasibility bounds on layer thickness.

  8. Knowl 8 — Evanescent decay constants set the mode's penetration depth into the claddings

    equation

    The mode profile of a slab guided mode consists of an oscillatory sinusoidal field inside the film and exponentially decaying tails in the substrate and cover. The decay constant in each cladding region is

    γi=k0N2−ni2,i∈{s,c},\gamma_i = k_0\sqrt{N^2 - n_i^2}, \quad i \in \{s, c\},

    with k0=2π/λk_0 = 2\pi/\lambda, NN the effective index, and nin_i the refractive index of the substrate (ss) or cover (cc). The penetration depth of the evanescent tail into region ii is 1/γi1/\gamma_i; a larger index contrast N−niN - n_i yields a more tightly confined mode with shorter evanescent tails, which reduces crosstalk between neighboring integrated waveguides.

  9. Knowl 9 — Dispersion relation computed by sweeping wavelength and solving the eigenvalue equation

    model/method

    The dispersion relation of a slab waveguide — the dependence of the effective index NN on the free-space wavelength λ\lambda — is obtained by evaluating the TE (or TM) eigenvalue equation at each wavelength and finding the root N∈(ns,nf)N \in (n_s, n_f). Because material refractive indices themselves depend on λ\lambda, the dispersion calculation must account for material dispersion in nfn_f, nsn_s, and ncn_c in addition to the geometric (waveguide) dispersion. The resulting N(λ)N(\lambda) curves determine group-velocity behavior and are the tool used to check that a candidate design maintains the desired modal index across the wavelength band of interest.

  10. Knowl 10 — Symmetric-cladding limit yields closed-form even/odd mode structure

    theoretical result

    When both claddings share the same refractive index (ns=nc=ncln_s = n_c = n_{\mathrm{cl}}), the slab is structurally symmetric and the eigenvalue equation simplifies; modes are then classified as even or odd about the film centerline, and the TEm\mathrm{TE}_m modes satisfy

    kxh=mπ+2arctan⁡ ⁣(γkx),k_x h = m\pi + 2\arctan\!\left(\frac{\gamma}{k_x}\right),

    with γ=k0N2−ncl2\gamma = k_0\sqrt{N^2 - n_{\mathrm{cl}}^2}. The symmetric case provides closed-form insight into mode order and cutoff that is used as a sanity check on the asymmetric (ns≠ncn_s \ne n_c) designs, and it marks the boundary case where mode degeneracies between orthogonal polarizations can occur.

Coverage note — The paper's inverse-design pipeline sections describing implementation details of the numerical mode solver and the parameter-search wrapper were condensed into their load-bearing ideas rather than reproduced step by step, because they are standard subroutines (root-finding for the eigenvalue equation and bounded parameter search) rather than novel standalone algorithms.

Citation

MLA
Hartvigsen, T., et al. “ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3309–26, https://doi.org/10.18653/v1/2022.acl-long.234.
APA
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., & Kamar, E. (2022). ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3309–3326. https://doi.org/10.18653/v1/2022.acl-long.234
Chicago
Hartvigsen, T., S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar. 2022. “ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3309–26. https://doi.org/10.18653/v1/2022.acl-long.234.
Harvard
Hartvigsen, T. et al. (2022) “ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3309–3326. Available at: https://doi.org/10.18653/v1/2022.acl-long.234.
Vancouver
1. Hartvigsen T, Gabriel S, Palangi H, Sap M, Ray D, Kamar E (2022) ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3309–3326

BibTeX

@inproceedings{hartvigsen-etal-2022-toxigen,
    title = "{T}oxi{G}en: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection",
    author = "Hartvigsen, Thomas  and
      Gabriel, Saadia  and
      Palangi, Hamid  and
      Sap, Maarten  and
      Ray, Dipankar  and
      Kamar, Ece",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.234/",
    doi = "10.18653/v1/2022.acl-long.234",
    pages = "3309--3326"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/