An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models
Nicholas MeadeElinor Poole-DayanSiva Reddy
Evaluates five prominent debiasing techniques across multiple language models and benchmarks, revealing that apparent reductions in social bias frequently stem from degraded language modeling capabilities rather than true mitigation.
Large pre-trained language models acquire widespread social stereotypes from the unmoderated internet text used during their training. While multiple debiasing methods have been developed to address these harmful associations, their broader effectiveness across non-gender domains and their side effects on core model capabilities remain poorly understood.
The article evaluates five prominent bias mitigation techniques across four major language models to determine which approach reduces bias most effectively and how these methods impact general language modeling ability and performance on downstream tasks.
The authors conducted a comprehensive empirical evaluation comparing Counterfactual Data Augmentation, increased Dropout regularization, Iterative Nullspace Projection, SentenceDebias, and Self-Debias across four architectures: BERT, ALBERT, RoBERTa, and GPT-2. The study evaluated gender, racial, and religious biases using three standard intrinsic benchmarks: the Sentence Encoder Association Test, StereoSet, and Crowdsourced Stereotype Pairs. In addition, the authors evaluated general language modeling fluency on WikiText-2 and downstream task performance on the standard General Language Understanding Evaluation benchmark using fine-tuned models.
The findings reveal several critical insights for model deployment. First, Self-Debias emerged as the most consistently effective debiasing technique, reducing stereotype scores across all bias domains while maintaining the model's underlying language generation capabilities. Second, most debiasing techniques performed substantially worse and less consistently when mitigating racial and religious biases compared to gender bias. Third, bias reductions on benchmarks like StereoSet and CrowS-Pairs were frequently accompanied by a degradation in language modeling fluency—for example, SentenceDebias doubled the perplexity error rate of GPT-2 relative to baseline. Finally, fine-tuning debiased models on specific downstream language understanding tasks showed virtually no degradation, indicating that fine-tuning enables models to preserve or relearn necessary task-specific representations.
These results demonstrate that reported reductions in intrinsic bias benchmarks can be misleading; lower stereotype scores often reflect an overall degradation in language modeling capability rather than genuine bias removal. For organizations deploying language models, adopting debiasing interventions without monitoring language modeling degradation introduces the risk of deploying impaired systems. However, downstream task performance remains robust across most debiased models.
Organizations seeking to reduce bias during text generation should prioritize self-debiasing and prompting techniques over invasive representation alterations. Practitioners should simultaneously evaluate standard language modeling fluency alongside any bias benchmark to ensure model quality is preserved. Further research is necessary before establishing definitive operational standards, specifically to extend debiasing to downstream tasks and evaluate extrinsic harms in real-world applications.
The conclusions should be interpreted with caution due to several constraints. The evaluation was limited to English-language models, relied on crowdsourced benchmarks that reflect North American cultural perspectives, operated under simplified binary gender definitions, and utilized intrinsic benchmarks that identify the presence of bias but cannot guarantee that a model is unbiased.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). Read StereoSet first to understand a key benchmark this survey uses to evaluate stereotypical bias alongside language-modeling ability.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). This critical survey frames the concepts and limitations of NLP bias measurement that motivate evaluating debiasing methods beyond benchmark scores.
- Paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al. (2016). Bolukbasi et al.’s embedding debiasing work supplies foundational mitigation ideas that help situate the later techniques examined in the survey.
- Paper: "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset, Eric Michael Smith et al. (2022). HOLISTICBIAS extends benchmark-based bias evaluation to a broader demographic taxonomy and conversational behaviors, testing what mitigation-focused assessments may miss.
- Paper: BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation, Tianxiang Sun et al. (2022). This study carries the question of pretrained-model bias into language-model evaluation metrics and tests practical mitigation strategies for that downstream setting.
- Paper: From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models, Shangbin Feng et al. (2023). This work extends bias analysis from intrinsic model benchmarks to the propagation of pretraining political biases into consequential downstream NLP tasks.
