StereoSet: Measuring stereotypical bias in pretrained language models

Moin NadeemAnna BethkeSiva Reddy

article2020ACL1,483 citations

Introduces StereoSet, a large-scale benchmark to quantify and track stereotypical biases across gender, profession, race, and religion in popular pretrained language models.

Listen

Modern artificial intelligence models rely heavily on large-scale pretrained language representations to power widespread commercial natural language processing applications. However, because these systems are trained on massive datasets reflecting real-world human writing, they risk learning and amplifying harmful demographic stereotypes. Previous benchmarking efforts evaluated these behaviors mainly through artificial sentences, focused narrowly on masked models, or measured social bias in isolation without considering core language processing capabilities.

The article establishes a standardized benchmark to simultaneously measure stereotypical bias and language modeling performance across modern pretrained language models in natural language contexts.

The researchers developed StereoSet, an English dataset comprising 16,995 crowdsourced test instances across four key domains: gender, profession, race, and religion. StereoSet evaluates models through Context Association Tests at both the sentence level (intrasentence) and the discourse level (intersentence). Each test item presents a target group context alongside three completions: a stereotype, an anti-stereotype, and a meaningless distractor. The study evaluated popular model architectures—including BERT, RoBERTa, XLNet, and GPT-2—using three metrics: Language Modeling Score (ability to prefer meaningful over meaningless text, with an ideal score of 100), Stereotype Score (preference for stereotypes over anti-stereotypes, with an ideal score of 50 indicating neutrality), and an Idealized Context Association Test score that unifies accuracy and fairness into a single index.

The evaluation revealed several critical findings:

  1. Pretrained language models consistently demonstrate systemic stereotypical biases. Every tested model exhibited a Stereotype Score exceeding the neutral baseline of 50, with scores ranging from 50.5 (RoBERTa-base) to 60.0 (GPT-2-large), and reaching 62.3 for an ensemble model.
  2. Language modeling capability strongly and positively correlates with stereotypical bias (Spearman rank correlation of 0.87). As models improve at core linguistic prediction, their tendency to favor demographic stereotypes increases proportionally.
  3. Model scaling exacerbates bias. Within identical model families trained on the same data, larger parameter variants achieved higher language modeling scores but consistently displayed greater stereotypical bias than their smaller counterparts.
  4. Scoring mechanisms impact diagnostic performance across sequence lengths. Pseudo-likelihood scoring performed effectively for intrasentence evaluations (averaging 79.4 in language modeling) but degraded significantly on longer intersentence evaluations (averaging 75.98), where traditional likelihood scoring proved more reliable.

These results indicate that state-of-the-art language models internalize real-world societal biases directly from standard pretraining corpora. Deploying these systems without intervention introduces ethical, compliance, and reputational risks, as high-performing models are inherently more prone to generating or reinforcing discriminatory associations across user-facing platforms.

Organizations developing or deploying language models should avoid using raw linguistic accuracy as the sole deployment criterion. Technical teams should benchmark models using multidimensional metrics that penalize unfair preferences alongside accuracy. Because higher model capability currently entails greater bias risk, practitioners must actively decouple accuracy from bias by exploring new training objectives, curated corpora, and targeted debiasing techniques prior to production deployment.

Confidence in these findings is high for English-language systems evaluated within typical US cultural contexts, supported by strong validator agreement across thousands of natural examples. However, stakeholders should note key limitations: the dataset reflects cultural norms specific to US crowdworkers, and scoring long-form contexts remains sensitive to the choice of likelihood formulation. Furthermore, achieving a balanced score on StereoSet does not guarantee the complete absence of subtle or domain-specific biases.

arXiv: 2004.09456
Cover for StereoSet: Measuring stereotypical bias in pretrained language models

Abstract

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real world data, they are known to capture stereotypical biases. In order to assess the adverse effects of these models, it is important to quantify the bias captured in them. Existing literature on quantifying bias evaluates pretrained language models on a small set of artificially constructed bias-assessing sentences. We present StereoSet, a large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion. We evaluate popular models like BERT, GPT-2, RoBERTa, and XLNet on our dataset and show that these models exhibit strong stereotypical biases. We also present a leaderboard with a hidden test set to track the bias of future language models at this https URL

Table of Contents

  • 1 Introduction
  • 2 Task Formulation
  • 2.1 Intrasentence
  • 2.2 Intersentence
  • 3 Related Work
  • 3.1 Bias in word embeddings
  • 3.2 Bias in pretrained language models
  • 3.3 Measuring bias through extrinsic tasks
  • 4 Dataset Creation
  • 4.1 Target terms
  • 4.2 CATs collection
  • 4.3 CATs validation
  • 5 Dataset Analysis
  • 6 Experimental Setup
  • 6.1 Development and test sets
  • 6.2 Evaluation Metrics
  • 6.3 Baselines
  • 7 Main Experiments
  • 7.1 BERT
  • 7.2 RoBERTa
  • 7.3 XLNet
  • 7.4 GPT2
  • 8 Results and discussion
  • 9 Limitations
  • 10 Conclusion
  • References
  • A Appendix
  • A.1 Detailed Results
  • A.2 Mechanical Turk Task
  • A.3 Target Words
  • A.4 General Methods for Training a Next Sentence Prediction Head
  • A.5 Fine-Tuning BERT for Sentiment Analysis

Knowls

  1. Knowl 1 — Context Association Test Evaluation Framework

    model/method

    The Context Association Test (CAT) is an evaluation framework designed to jointly measure the stereotypical bias and language modeling ability of pretrained language models. Rather than evaluating bias in isolation or using artificial sentence templates, CAT provides target terms (representing social groups across gender, profession, race, and religion) with natural contexts and three associative options:

    1. Stereotypical association: An attribute or sentence reflecting a common societal stereotype about the target group.
    2. Anti-stereotypical association: An attribute or sentence that actively contradicts or combats the stereotype while remaining plausible.
    3. Meaningless association: An attribute or sentence that is completely unrelated or nonsensical in context, used to verify basic language modeling capability.

    CAT is structured into two complementary tasks:

    • Intrasentence CAT: Sentence-level evaluation where a fill-in-the-blank context describing a target group must be completed by one of the three candidate attribute words.
    • Intersentence CAT: Discourse-level evaluation where a context sentence introducing a target group is followed by one of three candidate continuation sentences.
  2. Knowl 2 — StereoSet Evaluation Metrics: LMS, SS, and ICAT

    equation

    StereoSet evaluates language models using three primary metrics based on ranking candidate associations:

    1. Language Modeling Score (lmslms): The percentage of instances in which a model assigns a higher probability/likelihood to a meaningful association (stereotypical or anti-stereotypical) over a meaningless association. For an ideal language model, lms=100%lms = 100\%; for a random baseline, lms=50%lms = 50\%.

    2. Stereotype Score (ssss): The percentage of instances in which a model prefers the stereotypical association over the anti-stereotypical association:

    • An ideal, unbiased model prefers neither and scores ss=50%ss = 50\%.
    • ss>50%ss > 50\% indicates a pro-stereotypical bias.
    • ss<50%ss < 50\% indicates an anti-stereotypical bias.
    1. Idealized CAT Score (icaticat): A unified score that penalizes models for both poor language modeling performance and deviations from fairness (ss=50ss = 50), defined as:

    icat=lms×min⁡(ss,100−ss)50icat = lms \times \frac{\min(ss, 100 - ss)}{50}

    Where:

    • lms∈[0,100]lms \in [0, 100] is the Language Modeling Score.
    • ss∈[0,100]ss \in [0, 100] is the Stereotype Score.
    • min⁡(ss,100−ss)50∈[0,1]\frac{\min(ss, 100 - ss)}{50} \in [0, 1] acts as a fairness factor that achieves its maximum of 11 when ss=50ss = 50 and its minimum of 00 when ss∈{0,100}ss \in \{0, 100\}.
    • An ideal model achieves icat=100icat = 100 (lms=100,ss=50lms = 100, ss = 50), a completely biased model achieves icat=0icat = 0 (ss∈{0,100}ss \in \{0, 100\}), and a random baseline achieves icat=50icat = 50 (lms=50,ss=50lms = 50, ss = 50).
  3. Knowl 3 — Likelihood and Pseudo-Likelihood Scoring Mechanisms for CAT

    model/method

    Because masked language models (MLMs) and autoregressive language models (ALMs) parameterize probability distributions differently, CAT evaluates model preferences using two scoring paradigms across sentence-level and discourse-level tests:

    • Likelihood-Based Scoring (llll):

      • Intrasentence: For MLMs, the score is the log probability of predicting the attribute token in the masked slot. For multi-token attributes, subwords are iteratively unmasked from left to right and the average per-subword log probability is computed. For ALMs, the probability of the full instantiated sentence is computed.
      • Intersentence: A Next Sentence Prediction (NSP) classification head is fine-tuned on Wikipedia consecutive vs. sampled negative sentence pairs (9.5M training examples) to score the log likelihood p(s∣c)p(s \mid c) of candidate sentence ss following context sentence cc.
    • Pseudo-Likelihood-Based Scoring (pllpll):

      • Intrasentence: Following pseudo-likelihood scoring for MLMs, the attribute token is never masked; instead, each context token is masked one by one to compute the pseudo-probability of the sentence conditioned on the attribute word.
      • Intersentence: For MLMs, tokens in the context sentence cc are iteratively masked conditioned on the attribute sentence ss. For ALMs, scoring uses the joint log probability ratio log⁡p(s∣c)p(s)\log \frac{p(s \mid c)}{p(s)} to measure the contextual association.
  4. Knowl 4 — Positive Correlation Between Language Modeling Ability and Stereotypical Bias

    empirical result

    Across evaluated pretrained language models (BERT, RoBERTa, XLNet, and GPT-2), there is a strong positive correlation between language modeling ability and stereotypical bias, with a Spearman rank correlation coefficient of ρ=0.87\rho = 0.87.

    As models improve at language modeling (achieving higher Language Modeling Scores lmslms), they exhibit systematically higher Stereotype Scores (ssss). An ensemble combining BERT-large, GPT2-medium, and GPT2-large achieves the highest language modeling score (lms=90.2lms = 90.2 on likelihood scoring) among all models, but also exhibits the highest level of stereotypical bias (ss=62.3ss = 62.3). This indicates that scaling and improving language modeling objectives on standard web corpora naturally increases the model's reflection of societal stereotypes.

  5. Knowl 5 — StereoSet Dataset Composition and Structure

    data/table

    StereoSet contains 16,995 English test triplets across 321 target terms covering four bias domains: gender, profession, race, and religion. Target terms were curated from Wikidata relation triples (profession P106, race P172, religion P140) and established gender term lists, filtered for frequency. Triplets were collected on Amazon Mechanical Turk from US crowdworkers and validated by 5 independent annotators, retaining instances with at least 3-way agreement (83% retention rate).

    Domain # Target Terms # CATs (triplets) Avg Len (# words)
    Intrasentence
    Gender 40 1,026 7.98
    Profession 120 3,208 8.30
    Race 149 3,996 7.63
    Religion 12 623 8.18
    Intrasentence Total 321 8,498 8.02
    Intersentence
    Gender 40 996 15.55
    Profession 120 3,269 16.05
    Race 149 3,989 14.98
    Religion 12 604 14.99
    Intersentence Total 321 8,497 15.39
    Overall 321 16,995 11.70

    The dataset is partitioned by target terms into a 25% development set and a disjoint 75% test set. There is no training split because StereoSet is designed strictly as an intrinsic bias benchmark for pretrained representations rather than a fine-tuning corpus.

  6. Knowl 6 — Performance of Pretrained Language Models on StereoSet

    data/table

    Evaluation of baselines and pretrained language models on the StereoSet test set under likelihood-based scoring:

    Model Language Model Score (lms) Stereotype Score (ss) Idealized CAT Score (icat)
    IDEALLM 100.0 50.0 100.0
    STEREOTYPEDLM – 100.0 0.0
    RANDOMLM 50.0 50.0 50.0
    SENTIMENTLM 65.1 60.8 51.1
    BERT-base 85.4 58.3 71.2
    BERT-large 85.8 59.2 69.9
    ROBERTA-base 68.2 50.5 67.5
    ROBERTA-large 75.8 54.8 68.5
    XLNET-base 67.7 54.1 62.1
    XLNET-large 78.2 54.0 72.0
    GPT2 83.6 56.4 73.0
    GPT2-medium 85.9 58.2 71.7
    GPT2-large 88.3 60.0 70.5
    ENSEMBLE 90.2 62.3 68.0

    All tested pretrained models exhibit stereotypical bias (ss>50.0ss > 50.0). GPT2-large achieves the highest individual language modeling capability (lms=88.3lms = 88.3) but exhibits the highest bias (ss=60.0ss = 60.0) among standalone models. RoBERTa-base shows the lowest stereotype score (ss=50.5ss = 50.5) but suffers from lower language modeling accuracy (lms=68.2lms = 68.2). Under pseudo-likelihood scoring, GPT2-large reaches lms=89.6,ss=62.7,icat=66.8lms = 89.6, ss = 62.7, icat = 66.8, and ENSEMBLE reaches lms=90.1,ss=62.2,icat=68.1lms = 90.1, ss = 62.2, icat = 68.1.

  7. Knowl 7 — Scaling Model Parameters Increases Stereotypical Bias

    empirical result

    Within every model architecture family trained on identical corpora, increasing the number of model parameters increases both the Language Modeling Score (lmslms) and the Stereotype Score (ssss):

    • BERT (Wikipedia + BookCorpus):
      • Base (110M params): lms=85.4,ss=58.3lms = 85.4, ss = 58.3
      • Large (340M params): lms=85.8,ss=59.2lms = 85.8, ss = 59.2
    • RoBERTa (160GB corpora):
      • Base (125M params): lms=68.2,ss=50.5lms = 68.2, ss = 50.5
      • Large (355M params): lms=75.8,ss=54.8lms = 75.8, ss = 54.8
    • GPT-2 (WebText / Reddit links, 40GB):
      • Small (117M params): lms=83.6,ss=56.4lms = 83.6, ss = 56.4
      • Medium (345M params): lms=85.9,ss=58.2lms = 85.9, ss = 58.2
      • Large (774M params): lms=88.3,ss=60.0lms = 88.3, ss = 60.0

    While larger capacity allows models to better represent natural text, it systematically exacerbates preference for stereotypical over anti-stereotypical associations.

  8. Knowl 8 — Length-Dependent Discrepancy Between Likelihood and Pseudo-Likelihood Scoring

    empirical result

    Comparing likelihood-based (llll) and pseudo-likelihood-based (pllpll) scoring across sentence lengths reveals a performance trade-off:

    • On intrasentence CATs (average length 8.02 words), pseudo-likelihood scoring outperforms likelihood scoring across language models (average lmspll=79.4lms_{pll} = 79.4 vs. average lmsll=75.7lms_{ll} = 75.7).
    • On intersentence CATs (longer discourse-level sequences, average length 15.39 words), pseudo-likelihood scoring degrades significantly (average lmspll=75.98lms_{pll} = 75.98 vs. average lmsll=78.82lms_{ll} = 78.82).

    Overall average language modeling scores across all test sets are narrow (average lmsll=79.88lms_{ll} = 79.88 vs. average lmspll=79.68lms_{pll} = 79.68), indicating that while pseudo-likelihood avoids penalizing infrequent attribute words in single-slot predictions, it struggles with variance and scoring quality as sequence lengths increase.

  9. Knowl 9 — Domain-Wise Bias Breakdown and Extremal Target Terms

    data/table

    Breakdown of the ENSEMBLE model across the four StereoSet domains and selected extremal (most and least biased) target terms using likelihood-based scoring:

    Domain / Target Term Language Model Score (lms) Stereotype Score (ss) Idealized CAT Score (icat)
    GENDER 92.4 63.9 66.7
    mother (most biased) 97.2 77.8 43.2
    grandfather (least biased) 96.2 52.8 90.8
    PROFESSION 88.8 62.6 66.5
    software developer (most biased) 94.0 75.9 45.4
    producer (least biased) 91.7 53.7 84.9
    RACE 91.2 61.8 69.7
    African (most biased) 91.8 74.5 46.7
    Crimean (least biased) 93.3 50.0 93.3
    RELIGION 93.5 63.8 67.7
    Bible (most biased) 85.0 66.0 57.8
    Muslim (least biased) 94.8 46.6 88.3

    Target terms that have frequent natural corpus presence and strongly established societal stereotypes (e.g., mother, software developer, African) yield high stereotype scores (ss≥74.5ss \ge 74.5). In contrast, terms lacking prevalent societal stereotypes (e.g., Crimean, producer) exhibit stereotype scores close to the neutral baseline (ss≈50.0ss \approx 50.0).

  10. Knowl 10 — Failure of Crowdsourced Neutral Continuations Due to Anchoring Bias

    limitation

    An initial design attempt to collect neutral associative contexts as a baseline for CAT failed due to human anchoring bias and poor inter-annotator agreement. When asked to provide contextually neutral continuations for target terms, crowdworkers frequently produced sentences that validation annotators classified as stereotypical or anti-stereotypical (e.g., 'The Ethiopian woman had never been so angry' characterizing the Angry Black Woman stereotype, or 'The chess player was bearded' invoking wisdom stereotypes).

    Because stereotypes are highly accessible in human cognition, crowdworkers struggled to remain neutral. As a result, explicit neutral options were abandoned in favor of meaningless associations to measure baseline language modeling capability without confounding bias measurements.

Coverage note — None was omitted; all key contributions, mathematical metric definitions, dataset statistics, experimental results across scoring methods and model scales, and qualitative findings are covered.

References

  1. 1.Vamsi Aribandi, Yi Tay, and Donald Metzler. 2021. How reliable are model diagnostics? In Proceedings of ACL Findings.
  2. 2.Emily M. Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  3. 3.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of FAccT.
  4. 4.Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam T. Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proceedings of Neural Information Processing Systems (NeurIPS), pages 4349–4357.
  5. 5.Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  6. 6.Patricia Hill Collins. 2004. Black sexual politics: African Americans, gender, and the new racism. Routledge.
  7. 7.Alexander M Czopp, Aaron C Kay, and Sapna Cheryan. 2015. Positive stereotypes are pervasive and powerful. Perspectives on Psychological Science, 10(4):451–463.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of North American Chapter of the Association for Computational Linguistics, pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Djellel Difallah, Elena Filatova, and Panos Ipeirotis. 2018. Demographics and dynamics of mechanical turk workers. In Proceedings of the ACM International Conference on Web Search and Data Mining, WSDM ’18, pages 135 – 143, New York, NY, USA. Association for Computing Machinery.
  10. 10.Hila Gonen and Yoav Goldberg. 2019. Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them.
  11. 11.Anthony G. Greenwald and Mahzarin R. Banaji. 1995. Implicit social cognition: attitudes, self-esteem, and stereotypes. Psychological review, 102(1):4.
  12. 12.Jeremy Howard and Sebastian Ruder. 2018. Universal Language Model Fine-tuning for Text Classification. In Proceedings of the Association for Computational Linguistics, pages 328–339, Melbourne, Australia. Association for Computational Linguistics.
  13. 13.Milos Jakubicek, Adam Kilgarriff, Vojtech Kovar, Pavel Rychly, and Vit Suchomel. 2013. The tenten corpus family. In Proceedings of the International Corpus Linguistics Conference CL.
  14. 14.Adam Kilgarriff. 2009. Simple maths for keywords. In Proceedings of the Corpus Linguistics Conference 2009 (CL2009), page 171.
  15. 15.Svetlana Kiritchenko and Saif Mohammad. 2018. Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems. In Proceedings of Joint Conference on Lexical and Computational Semantics, pages 43–53.
  16. 16.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring bias in contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 166–172, Florence, Italy. Association for Computational Linguistics.
  17. 17.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  18. 18.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the Association for Computational Linguistics, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  19. 19.Thomas Manzini, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019. Black is to criminal as caucasian is to police: Detecting and removing multiclass bias in word embeddings. In Proceedings of the North American Chapter of the Association for Computational Linguistics, pages 615–621, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On measuring social biases in sentence encoders. In Proceedings of the North American Chapter of the Association for Computational Linguistics, pages 622–628, Minneapolis, Minnesota. Association for Computational Linguistics.
  21. 21.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of Neural Information Processing Systems (NeurIPS), NIPS 13, pages 3111 – 3119, Red Hook, NY, USA. Curran Associates Inc.
  22. 22.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of Empirical Methods in Natural Language Processing (EMNLP), pages 1953–1967, Online. Association for Computational Linguistics.
  23. 23.Brian Nosek, Mahzarin Banaji, and Anthony Greenwald. 2002. Math = male, me = female, therefore math != me. Journal of personality and social psychology, 83:44–59.
  24. 24.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global vectors for word representation. In Proceedings of Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543.
  25. 25.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep Contextualized Word Representations. In Proceedings of the North American Chapter of the Association for Computational Linguistics), pages 2227–2237. Association for Computational Linguistics.
  26. 26.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog, 1(8).
  27. 27.Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender bias in coreference resolution. In Proceedings of North American Chapter of the Association for Computational Linguistics (NAACL), pages 8–14.
  28. 28.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked language model scoring. In Proceedings of the of the Association for Computational Linguistics, Online. Association for Computational Linguistics.
  29. 29.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. In Proceedings of the Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3407–3412, Hong Kong, China. Association for Computational Linguistics.
  30. 30.Amos Tversky and Daniel Kahneman. 1974. Judgment under uncertainty: Heuristics and biases. science, 185(4157):1124–1131.
  31. 31.Denny Vrandeˇci´c and Markus Krötzsch. 2014. Wikidata: A free collaborative knowledgebase. Commun. ACM, 57(10):78–85.
  32. 32.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d’e Buc, E. Fox, and R. Garnett, editors, Proceedings of Neural Information Processing Systems (NeurIPS), pages 5753–5763. Curran Associates, Inc.
  33. 33.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. In Proceedings of North American Chapter of the Association for Computational Linguistics, pages 15–20.
  34. 34.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), ICCV 15, pages 19 – 27, USA. IEEE Computer Society.

Citation

MLA
Nadeem, M., et al. “StereoSet: Measuring Stereotypical Bias in Pretrained Language Models”. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, pp. 5356–71, https://doi.org/10.18653/v1/2021.acl-long.416.
APA
Nadeem, M., Bethke, A., & Reddy, S. (2021). StereoSet: Measuring stereotypical bias in pretrained language models. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 5356–5371. https://doi.org/10.18653/v1/2021.acl-long.416
Chicago
Nadeem, M., A. Bethke, and S. Reddy. 2021. “StereoSet: Measuring Stereotypical Bias in Pretrained Language Models”. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 5356–71. https://doi.org/10.18653/v1/2021.acl-long.416.
Harvard
Nadeem, M., Bethke, A. and Reddy, S. (2021) “StereoSet: Measuring stereotypical bias in pretrained language models”, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5356–5371. Available at: https://doi.org/10.18653/v1/2021.acl-long.416.
Vancouver
1. Nadeem M, Bethke A, Reddy S (2021) StereoSet: Measuring stereotypical bias in pretrained language models. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, pp 5356–5371

BibTeX

@inproceedings{nadeem-etal-2021-stereoset,
    title = "{S}tereo{S}et: Measuring stereotypical bias in pretrained language models",
    author = "Nadeem, Moin  and
      Bethke, Anna  and
      Reddy, Siva",
    editor = "Zong, Chengqing  and
      Xia, Fei  and
      Li, Wenjie  and
      Navigli, Roberto",
    booktitle = "Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
    month = aug,
    year = "2021",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.acl-long.416/",
    doi = "10.18653/v1/2021.acl-long.416",
    pages = "5356--5371"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/