StereoSet: Measuring stereotypical bias in pretrained language models
Moin NadeemAnna BethkeSiva Reddy
Introduces StereoSet, a large-scale benchmark to quantify and track stereotypical biases across gender, profession, race, and religion in popular pretrained language models.
Modern artificial intelligence models rely heavily on large-scale pretrained language representations to power widespread commercial natural language processing applications. However, because these systems are trained on massive datasets reflecting real-world human writing, they risk learning and amplifying harmful demographic stereotypes. Previous benchmarking efforts evaluated these behaviors mainly through artificial sentences, focused narrowly on masked models, or measured social bias in isolation without considering core language processing capabilities.
The article establishes a standardized benchmark to simultaneously measure stereotypical bias and language modeling performance across modern pretrained language models in natural language contexts.
The researchers developed StereoSet, an English dataset comprising 16,995 crowdsourced test instances across four key domains: gender, profession, race, and religion. StereoSet evaluates models through Context Association Tests at both the sentence level (intrasentence) and the discourse level (intersentence). Each test item presents a target group context alongside three completions: a stereotype, an anti-stereotype, and a meaningless distractor. The study evaluated popular model architectures—including BERT, RoBERTa, XLNet, and GPT-2—using three metrics: Language Modeling Score (ability to prefer meaningful over meaningless text, with an ideal score of 100), Stereotype Score (preference for stereotypes over anti-stereotypes, with an ideal score of 50 indicating neutrality), and an Idealized Context Association Test score that unifies accuracy and fairness into a single index.
The evaluation revealed several critical findings:
- Pretrained language models consistently demonstrate systemic stereotypical biases. Every tested model exhibited a Stereotype Score exceeding the neutral baseline of 50, with scores ranging from 50.5 (RoBERTa-base) to 60.0 (GPT-2-large), and reaching 62.3 for an ensemble model.
- Language modeling capability strongly and positively correlates with stereotypical bias (Spearman rank correlation of 0.87). As models improve at core linguistic prediction, their tendency to favor demographic stereotypes increases proportionally.
- Model scaling exacerbates bias. Within identical model families trained on the same data, larger parameter variants achieved higher language modeling scores but consistently displayed greater stereotypical bias than their smaller counterparts.
- Scoring mechanisms impact diagnostic performance across sequence lengths. Pseudo-likelihood scoring performed effectively for intrasentence evaluations (averaging 79.4 in language modeling) but degraded significantly on longer intersentence evaluations (averaging 75.98), where traditional likelihood scoring proved more reliable.
These results indicate that state-of-the-art language models internalize real-world societal biases directly from standard pretraining corpora. Deploying these systems without intervention introduces ethical, compliance, and reputational risks, as high-performing models are inherently more prone to generating or reinforcing discriminatory associations across user-facing platforms.
Organizations developing or deploying language models should avoid using raw linguistic accuracy as the sole deployment criterion. Technical teams should benchmark models using multidimensional metrics that penalize unfair preferences alongside accuracy. Because higher model capability currently entails greater bias risk, practitioners must actively decouple accuracy from bias by exploring new training objectives, curated corpora, and targeted debiasing techniques prior to production deployment.
Confidence in these findings is high for English-language systems evaluated within typical US cultural contexts, supported by strong validator agreement across thousands of natural examples. However, stakeholders should note key limitations: the dataset reflects cultural norms specific to US crowdworkers, and scoring long-form contexts remains sensitive to the choice of likelihood formulation. Furthermore, achieving a balanced score on StereoSet does not guarantee the complete absence of subtle or domain-specific biases.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). Introduces foundational techniques for quantifying human-like biases and stereotypes in learned vector representations of language, which StereoSet builds upon to evaluate contextual models.
- Paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al. (2016). Provides early definitions and geometric formulations of gender bias and occupational stereotypes in word embeddings that motivate modern benchmark evaluations in NLP.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). Introduces BERT, the core bidirectional Transformer architecture that StereoSet directly probes and evaluates for social and stereotypical bias.
- Paper: Mitigating Unwanted Biases with Adversarial Learning, Brian Hu Zhang et al. (2018). Establishes methodological paradigms for recognizing and mitigating demographic biases in NLP systems, providing essential conceptual grounding.
- Paper: Learning Fair Representations, Richard Zemel et al. (2013). Establishes fundamental concepts of demographic group fairness and representation that underpin the harm categories measured in StereoSet.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Presents a critical survey analyzing how bias benchmarks, including stereotype evaluations, are conceptualized and operationalized across the NLP literature.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Extends the evaluation of neural language model harms from stereotypical associations to open-ended toxic text generation and steering behaviors.
- Paper: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜, Emily M. Bender et al. (2021). Synthesizes broad societal, environmental, and bias risks associated with scaling pretrained models, contextualizing the empirical findings of benchmarks like StereoSet.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). Categorizes and expands upon the societal risks and discriminatory harms identified in pretrained language model benchmark evaluations.
- Paper: Scaling Language Models: Methods, Analysis & Insights from Training Gopher, Jack W. Rae et al. (2021). Evaluates how scaling language models to hundreds of billions of parameters affects downstream capabilities, toxicity, and demographic biases.
- Paper: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, Stella Biderman et al. (2023). Provides a controlled experimental suite to analyze the exact training dynamics and emergence of social biases across model scales.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). Surveys the holistic landscape of large language model evaluation, positioning bias and stereotype datasets within broader evaluation frameworks.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). Integrates diverse multi-task assessments, including social bias and fairness tasks, into a massive standardized benchmark for evaluating language models at scale.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Investigates how demographic biases and subtle cues undermine the faithfulness of reasoning in larger generative models.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). Standardizes adversarial red-teaming and safety refusal evaluation across advanced models, building on the necessity of measuring model harms.
