Built independently by an author, for readers. Read the story and support ChapterPal

keyword

TOXIGEN dataset

The ToxiGen dataset is a large-scale, machine-generated benchmark designed to improve the detection of implicit hate speech and subtly toxic language in artificial intelligence systems. Comprising approximately 274,000 toxic and benign statements across 13 demographic groups, the dataset was produced using a massive pretrained language model guided by demonstration-based prompting and an adversarial classifier-in-the-loop decoding framework. By generating nuanced toxic examples alongside benign statements that mention demographic groups, the dataset aims to overcome common machine learning vulnerabilities, such as over-relying on explicit slurs or incorrectly flagging neutral discussions of marginalized identities as toxic. It serves as a standardized resource for evaluating and training toxicity classifiers to recognize subtle hostility, reduce false-positive rates on identity terms, and better detect machine-generated toxic text.

1 item

ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, Ece Kamar

OrganizationsAllen Institute for AICarnegie Mellon UniversityMassachusetts Institute of TechnologyMicrosoftUniversity of Washington

Why you should read this

Introduces ToxiGen, a large-scale balanced dataset of over 274,000 machine-generated statements across 13 minority groups, along with an adversarial decoding method to help classifiers detect subtle, implicit hate speech without over-relying on identity mentions.

Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate. Such over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language. To help mitigate these issues, we create TOXIGEN, a new large-scale and machine-generated dataset of 274k toxic and benign statements about 13 minority groups. We develop a demonstration-based prompting framework and an adversarial classifier-in-the-loop decoding method to generate subtly toxic and benign text with a massive pretrained language model (Brown et al., 2020). Controlling machine generation in this way allows TOXIGEN to cover implicitly toxic text at a larger scale, and about more demographic groups, than previous resources of human-written text. We conduct a human evaluation on a challenging subset of TOXIGEN and find that annotators struggle to distinguish machine-generated text from human-written language. We also find that 94.5% of toxic examples are labeled as hate speech by human annotators. Using three publicly-available datasets, we show that finetuning a toxicity classifier on our data improves its performance on human-written data substantially. We also demonstrate that TOXIGEN can be used to fight machine-generated toxicity as finetuning improves the classifier significantly on our evaluation subset.

Added

2026-09-26