An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models

Nicholas MeadeElinor Poole-DayanSiva Reddy

article2022ACL170 citations

Evaluates five prominent debiasing techniques across multiple language models and benchmarks, revealing that apparent reductions in social bias frequently stem from degraded language modeling capabilities rather than true mitigation.

Listen

Large pre-trained language models acquire widespread social stereotypes from the unmoderated internet text used during their training. While multiple debiasing methods have been developed to address these harmful associations, their broader effectiveness across non-gender domains and their side effects on core model capabilities remain poorly understood.

The article evaluates five prominent bias mitigation techniques across four major language models to determine which approach reduces bias most effectively and how these methods impact general language modeling ability and performance on downstream tasks.

The authors conducted a comprehensive empirical evaluation comparing Counterfactual Data Augmentation, increased Dropout regularization, Iterative Nullspace Projection, SentenceDebias, and Self-Debias across four architectures: BERT, ALBERT, RoBERTa, and GPT-2. The study evaluated gender, racial, and religious biases using three standard intrinsic benchmarks: the Sentence Encoder Association Test, StereoSet, and Crowdsourced Stereotype Pairs. In addition, the authors evaluated general language modeling fluency on WikiText-2 and downstream task performance on the standard General Language Understanding Evaluation benchmark using fine-tuned models.

The findings reveal several critical insights for model deployment. First, Self-Debias emerged as the most consistently effective debiasing technique, reducing stereotype scores across all bias domains while maintaining the model's underlying language generation capabilities. Second, most debiasing techniques performed substantially worse and less consistently when mitigating racial and religious biases compared to gender bias. Third, bias reductions on benchmarks like StereoSet and CrowS-Pairs were frequently accompanied by a degradation in language modeling fluency—for example, SentenceDebias doubled the perplexity error rate of GPT-2 relative to baseline. Finally, fine-tuning debiased models on specific downstream language understanding tasks showed virtually no degradation, indicating that fine-tuning enables models to preserve or relearn necessary task-specific representations.

These results demonstrate that reported reductions in intrinsic bias benchmarks can be misleading; lower stereotype scores often reflect an overall degradation in language modeling capability rather than genuine bias removal. For organizations deploying language models, adopting debiasing interventions without monitoring language modeling degradation introduces the risk of deploying impaired systems. However, downstream task performance remains robust across most debiased models.

Organizations seeking to reduce bias during text generation should prioritize self-debiasing and prompting techniques over invasive representation alterations. Practitioners should simultaneously evaluate standard language modeling fluency alongside any bias benchmark to ensure model quality is preserved. Further research is necessary before establishing definitive operational standards, specifically to extend debiasing to downstream tasks and evaluate extrinsic harms in real-world applications.

The conclusions should be interpreted with caution due to several constraints. The evaluation was limited to English-language models, relied on crowdsourced benchmarks that reflect North American cultural perspectives, operated under simplified binary gender definitions, and utilized intrinsic benchmarks that identify the presence of bias but cannot guarantee that a model is unbiased.

Cover for An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models

Abstract

Recent work has shown pre-trained language models capture social biases from the large amounts of text they are trained on. This has attracted attention to developing techniques that mitigate such biases. In this work, we perform an empirical survey of five recently proposed bias mitigation techniques: Counterfactual Data Augmentation (CDA), Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebias. We quantify the effectiveness of each technique using three intrinsic bias benchmarks while also measuring the impact of these techniques on a model’s language modeling ability, as well as its performance on downstream NLU tasks. We experimentally find that: (1) Self-Debias is the strongest debiasing technique, obtaining improved scores on all bias benchmarks; (2) Current debiasing techniques perform less consistently when mitigating non-gender biases; And (3) improvements on bias benchmarks such as StereoSet and CrowS-Pairs by using debiasing strategies are often accompanied by a decrease in language modeling ability, making it difficult to determine whether the bias mitigation was effective.

Table of Contents

  • 1 Introduction
  • 2 Techniques for Measuring Bias
  • 3 Debiasing Techniques
  • 4 Which Technique is Most Effective in Mitigating Bias?
  • 5 How Does Debiasing Impact Language Modeling?
  • 6 How Does Debiasing Impact Downstream Task Performance?
  • 7 Discussion and Limitations
  • 8 Conclusion
  • 9 Acknowledgements
  • 10 Further Ethical Considerations
  • A SEAT Test Specifications
  • SEAT-Religion-1
  • SEAT-Religion-1b
  • SEAT-Religion-2
  • SEAT-Religion-2b
  • B Bias Attribute Words
  • C Debiasing Details
  • C.1 CDA
  • C.2 Dropout
  • C.3 INLP
  • C.4 Self-Debias
  • C.5 SentenceDebias
  • D GLUE Details
  • E Additional Results

Knowls

  1. Knowl 1 — Empirical Superiority of Self-Debias Among Pre-trained Language Model Debiasing Techniques

    empirical result

    In a systematic evaluation of five debiasing techniques—Counterfactual Data Augmentation (CDA), Dropout regularization, Iterative Nullspace Projection (INLP), SentenceDebias, and Self-Debias—across four pre-trained language models (BERT-base-uncased, ALBERT-base-v2, RoBERTa-base, and GPT-2) and three social bias domains (gender, race, and religion), Self-Debias consistently achieves the strongest and most dependable reductions in stereotypical bias across intrinsic benchmarks (StereoSet and CrowS-Pairs). Unlike projection-based methods, Self-Debias introduces minimal degradation to language modeling capability. Furthermore, while all techniques mitigate gender bias with moderate consistency, their effectiveness decreases substantially and exhibits high variability when applied to non-gender biases (race and religion). In addition, on sentence-level embedding association tests (SEAT), aggressive projection methods (INLP and SentenceDebias) achieve greater reductions in effect size for masked language models than retraining methods (CDA and Dropout).

  2. Knowl 2 — Language Modeling Degradation as a Confounder in Stereotype Benchmark Scores

    empirical result

    Evaluation on intrinsic stereotype benchmarks such as StereoSet and CrowS-Pairs sets the ideal unbiased stereotype score at 50%50\% (representing an equal probability of preferring stereotypical versus anti-stereotypical sentence completions). However, a degenerate or random language model that assigns equal likelihoods to all continuations also attains a stereotype score of 50%50\% in expectation. Consequently, debiasing techniques that degrade general language modeling capacity—reflected by elevated perplexity on WikiText-2 and reduced StereoSet Language Modeling (LM) scores—can artificially produce improved (near-50%50\%) stereotype scores without actually removing representational bias. For example, SentenceDebias applied to GPT-2 for gender debiasing increases WikiText-2 perplexity from 30.15830.158 to 65.49365.493 while reducing the StereoSet stereotype score from 62.65%62.65\% to 56.05%56.05\%. Intrinsic stereotype metrics therefore cannot be reliably interpreted without simultaneously assessing language modeling performance.

  3. Knowl 3 — Preservation of Downstream NLU Task Performance After Representational Debiasing

    empirical result

    Fine-tuning debiased models (CDA, Dropout, INLP, and SentenceDebias) on the General Language Understanding Evaluation (GLUE) benchmark demonstrates that debiasing pre-trained representations does not significantly impair downstream task accuracy. On gender-debiased BERT base (baseline GLUE average: 77.7477.74), fine-tuned average GLUE scores remain virtually unchanged: 77.5277.52 for CDA, 76.2876.28 for Dropout, 76.7676.76 for INLP, and 77.8177.81 for SentenceDebias. For gender-debiased GPT-2 (baseline GLUE average: 73.0173.01), fine-tuned average scores are 74.2174.21 for CDA, 73.1673.16 for Dropout, 73.0673.06 for INLP, and 72.6372.63 for SentenceDebias. Downstream supervised fine-tuning allows models to retain and relearn task-essential information even if specific representational directions were altered or removed during pre-training or projection debiasing.

  4. Knowl 4 — Sentence Encoder Association Test Formulation

    equation

    The Sentence Encoder Association Test (SEAT) measures representation-level bias by extending the Word Embedding Association Test (WEAT) to sentence embeddings obtained from pre-trained language models via synthetic sentence templates.

    Let XX and YY denote two disjoint sets of target sentence representations, and let AA and BB denote two disjoint sets of attribute sentence representations. The SEAT test statistic s(X,Y,A,B)s(X, Y, A, B) is defined as: s(X,Y,A,B)=∑x∈Xs(x,A,B)−∑y∈Ys(y,A,B)s(X, Y, A, B) = \sum_{x \in X} s(x, A, B) - \sum_{y \in Y} s(y, A, B) where the differential association s(w,A,B)s(w, A, B) for a sentence vector ww is the difference between its mean cosine similarities with vectors in AA and BB: s(w,A,B)=1∣A∣∑a∈Acos⁡(w,a)−1∣B∣∑b∈Bcos⁡(w,b)s(w, A, B) = \frac{1}{|A|} \sum_{a \in A} \cos(w, a) - \frac{1}{|B|} \sum_{b \in B} \cos(w, b) with cos⁡(u,v)=u⋅v∥u∥2∥v∥2\cos(u, v) = \frac{u \cdot v}{\|u\|_2 \|v\|_2}.

    The effect size dd (Cohen's dd) measures the normalized separation between target distributions: d=μ({s(x,A,B)}x∈X)−μ({s(y,A,B)}y∈Y)σ({s(t,X,Y)}t∈A∪B)d = \frac{\mu(\{s(x, A, B)\}_{x \in X}) - \mu(\{s(y, A, B)\}_{y \in Y})}{\sigma(\{s(t, X, Y)\}_{t \in A \cup B})} where μ(⋅)\mu(\cdot) and σ(⋅)\sigma(\cdot) denote sample mean and standard deviation, respectively. An effect size dd closer to zero indicates a lower degree of bias in the representations.

  5. Knowl 5 — Intrinsic Stereotype Metrics: StereoSet and Masked Token CrowS-Pairs Scoring

    definition

    Intrinsic stereotypical bias benchmarks evaluate language models via the relative probability assigned to pairs or triplets of minimally contrastive sentences:

    1. StereoSet (Intrasentence): Evaluates a context sentence with three candidate completions: stereotypical (cstereoc_{\text{stereo}}), anti-stereotypical (cantic_{\text{anti}}), and unrelated (cunrelatedc_{\text{unrelated}}).

      • Stereotype Score (SS): The percentage of examples where the model prefers the stereotypical completion over the anti-stereotypical completion: SS=1N∑i=1NI(P(cstereo(i)∣context(i))>P(canti(i)∣context(i)))×100%SS = \frac{1}{N} \sum_{i=1}^N \mathbb{I}\left(P(c_{\text{stereo}}^{(i)} \mid \text{context}^{(i)}) > P(c_{\text{anti}}^{(i)} \mid \text{context}^{(i)})\right) \times 100\% An ideal unbiased model achieves SS=50%SS = 50\%.
      • Language Modeling Score (LM Score): The percentage of examples where the model prefers an associatively meaningful candidate (cstereoc_{\text{stereo}} or cantic_{\text{anti}}) over cunrelatedc_{\text{unrelated}}: LM=1N∑i=1NI(P(cmeaningful(i)∣context(i))>P(cunrelated(i)∣context(i)))×100%LM = \frac{1}{N} \sum_{i=1}^N \mathbb{I}\left(P(c_{\text{meaningful}}^{(i)} \mid \text{context}^{(i)}) > P(c_{\text{unrelated}}^{(i)} \mid \text{context}^{(i)})\right) \times 100\% An ideal model achieves LM=100%LM = 100\%.
    2. Crowdsourced Stereotype Pairs (CrowS-Pairs): Evaluates sentence pairs where the first reflects a stereotype and the second violates it. To eliminate pseudo-likelihood calibration artifacts, scoring is performed using masked token probabilities: for each differing token tt unique to a sentence, its masked probability P(t∣sentence∖t)P(t \mid \text{sentence}_{\setminus t}) is computed individually and averaged across multiple differing tokens. The stereotype score is the percentage of pairs where the model assigns higher average masked token probability to the stereotypical sentence than to the anti-stereotypical sentence (ideal score: 50%50\%).

  6. Knowl 6 — Representational Projection Debiasing: SentenceDebias and INLP

    model/method

    Linear projection debiasing removes demographic attribute information directly from sentence embeddings extracted by taking the mean token representation of the final hidden layer (last_hidden_state\text{last\_hidden\_state}):

    1. SentenceDebias: Estimates a bias subspace BB by identifying pairs of sentences generated via Counterfactual Data Augmentation (CDA) word substitutions from 2.5%2.5\% of English Wikipedia. Principal Component Analysis (PCA) is applied to the difference vectors of these sentence pairs. The first KK principal components define the orthonormal basis {vk}k=1K\{v_k\}_{k=1}^K of BB. A sentence representation hh is debiased by subtracting its orthogonal projection: hdebiased=h−∑k=1K(h⋅vk)vkh_{\text{debiased}} = h - \sum_{k=1}^K (h \cdot v_k) v_k

    2. Iterative Nullspace Projection (INLP): Extracts representations for 10,00010{,}000 sentences per bias class (e.g., male, female, neutral) from 2.5%2.5\% of English Wikipedia. A linear classifier is trained iteratively to predict the protected class from sentence representations. At each iteration ii with classifier weight matrix WiW_i, representations are projected onto the nullspace N(Wi)=I−Wi+Wi\mathcal{N}(W_i) = I - W_i^+ W_i. INLP applies this projection iteratively across mm sequential classifiers (m=80m = 80 for BERT, ALBERT, and RoBERTa; m=10m = 10 for GPT-2) to exhaustively remove linear information predictive of the protected attribute.

  7. Knowl 7 — Training-Based and Post-Hoc Debiasing: CDA, Dropout, and Self-Debias

    model/method

    Three non-projection debiasing methods are applied across masked and autoregressive language models:

    1. Counterfactual Data Augmentation (CDA): Rebalances text by swapping demographic attribute word pairs across a 10%10\% English Wikipedia corpus for gender, race, and religion. Models undergo an additional pre-training phase for 2,0002{,}000 steps with effective batch size 512512 (using 15%15\% masked language modeling for BERT, ALBERT, RoBERTa; autoregressive causal modeling for GPT-2).

    2. Dropout Regularization: Pre-trained models undergo an additional 2,0002{,}000 pre-training steps on 10%10\% English Wikipedia with batch size 512512 under elevated dropout probabilities to disrupt learned associative shortcuts:

      • BERT and RoBERTa: hidden_dropout_prob=0.20\text{hidden\_dropout\_prob} = 0.20, attention_probs_dropout_prob=0.15\text{attention\_probs\_dropout\_prob} = 0.15.
      • ALBERT: hidden_dropout_prob=0.05\text{hidden\_dropout\_prob} = 0.05, attention_probs_dropout_prob=0.05\text{attention\_probs\_dropout\_prob} = 0.05.
      • GPT-2: resid_p_dropout=0.15\text{resid\_p\_dropout} = 0.15, embd_dropout=0.15\text{embd\_dropout} = 0.15, attn_dropout=0.15\text{attn\_dropout} = 0.15.
    3. Self-Debias: A post-hoc generation algorithm that discourages biased outputs without updating model parameters. During generation, token likelihoods under a domain-specific toxic prompt prefix are computed and subtracted from non-prompted continuation probabilities. Prompts include:

      • Gender: "The following text discriminates against people because of their gender:"
      • Race: "The following text discriminates against people because of their race/color:"
      • Religion: "The following text discriminates against people because of their religion:" Because Self-Debias alters generation probabilities rather than internal representations or model weights, it cannot be applied to representation benchmarks like SEAT or downstream NLU classification tasks.
  8. Knowl 8 — StereoSet Benchmark Results for Debiased Language Models

    data/table

    StereoSet stereotype scores (ideal = 50.00%50.00\%) and language modeling scores (LM Score, ideal = 100.00%100.00\%) for BERT-base, ALBERT-base-v2, RoBERTa-base, and GPT-2 small across gender, race, and religion debiasing.

    Model Gender Race Religion
    Stereo (%) LM (%) Stereo (%) LM (%) Stereo (%) LM (%)
    BERT 60.28 84.17 57.03 84.17 59.70 84.17
    + CDA 59.61 83.08 56.73 83.41 58.37 83.24
    + DROPOUT 60.66 83.04 57.07 83.04 59.13 83.04
    + INLP 57.25 80.63 57.29 83.12 60.31 83.36
    + SELF-DEBIAS 59.34 84.09 54.30 84.24 57.26 84.23
    + SENTENCEDEBIAS 59.37 84.20 57.78 83.95 58.73 84.26
    ALBERT 59.93 89.77 57.51 89.77 60.32 89.77
    + CDA 55.85 77.11 53.15 79.09 58.70 75.85
    + DROPOUT 58.40 77.05 51.98 77.05 57.15 77.05
    + INLP 58.05 86.58 55.00 87.81 63.77 88.86
    + SELF-DEBIAS 61.52 89.54 55.94 89.63 59.83 89.59
    + SENTENCEDEBIAS 58.38 88.98 57.95 89.70 56.09 88.80
    RoBERTa 66.32 88.93 61.67 88.93 64.28 88.93
    + CDA 64.43 88.83 60.95 88.55 64.51 88.86
    + DROPOUT 66.26 88.81 60.41 88.81 62.08 88.81
    + INLP 60.82 88.23 58.26 88.96 60.34 88.11
    + SELF-DEBIAS 65.04 88.26 58.78 88.40 62.84 88.53
    + SENTENCEDEBIAS 62.77 88.94 62.72 88.32 63.91 88.70
    GPT-2 62.65 91.01 58.90 91.01 63.26 91.01
    + CDA 64.02 90.36 57.31 90.36 63.55 90.36
    + DROPOUT 63.35 90.40 57.50 90.40 64.17 90.40
    + INLP 60.17 91.62 58.96 91.06 63.95 91.17
    + SELF-DEBIAS 60.84 89.07 57.33 89.53 60.45 89.36
    + SENTENCEDEBIAS 56.05 87.43 56.43 91.38 59.62 90.53

    Self-Debias consistently reduces stereotype scores toward 50%50\% across all three bias domains with negligible loss in language modeling score (≤1.94%\le 1.94\% drop). In contrast, retraining ALBERT with CDA and Dropout causes severe LM score drops exceeding 1212 percentage points.

  9. Knowl 9 — CrowS-Pairs Benchmark Results for Debiased Language Models

    data/table

    CrowS-Pairs stereotype scores (%) across BERT, ALBERT, RoBERTa, and GPT-2 models for gender, race, and religion debiasing. An ideal unbiased score is 50.00%50.00\%.

    Model Gender Stereotype (%) Race Stereotype (%) Religion Stereotype (%)
    BERT 57.25 62.33 62.86
    + CDA 56.11 56.70 60.00
    + DROPOUT 55.34 59.03 55.24
    + INLP 51.15 67.96 60.95
    + SELF-DEBIAS 52.29 56.70 56.19
    + SENTENCEDEBIAS 52.29 62.72 63.81
    ALBERT 48.09 62.52 60.00
    + CDA 49.24 45.44 46.67
    + DROPOUT 51.53 48.54 42.86
    + INLP 47.33 55.34 57.14
    + SELF-DEBIAS 45.04 57.09 57.14
    + SENTENCEDEBIAS 47.33 62.14 25.71
    RoBERTa 60.15 63.57 60.00
    + CDA 56.32 63.76 59.05
    + DROPOUT 59.39 62.40 57.14
    + INLP 55.17 61.82 62.86
    + SELF-DEBIAS 57.09 62.40 51.43
    + SENTENCEDEBIAS 52.11 65.12 40.95
    GPT-2 56.87 59.69 62.86
    + CDA 56.87 60.66 51.43
    + DROPOUT 57.63 60.47 52.38
    + INLP 53.44 59.69 61.90
    + SELF-DEBIAS 56.11 53.29 58.10
    + SENTENCEDEBIAS 56.11 55.43 35.24

    Self-Debias consistently shifts stereotype scores toward 50%50\% across models and bias categories. Other methods display substantial volatility on smaller dataset slices; for example, on the 105105-example religion slice, SentenceDebias shifts GPT-2 from 62.86%62.86\% to 35.24%35.24\% and ALBERT to 25.71%25.71\%.

  10. Knowl 10 — SEAT Effect Sizes Across Debiased Models and Bias Types

    data/table

    Average absolute effect sizes across Sentence Encoder Association Test (SEAT) suites for gender (6 tests), race (7 tests), and religion (4 tests) across BERT, ALBERT, RoBERTa, and GPT-2 models. Effect sizes closer to 0.0000.000 indicate reduced association with stereotypical concept sets.

    Model Gender Avg. ∣d∣|d| Race Avg. ∣d∣|d| Religion Avg. ∣d∣|d|
    BERT 0.620 0.620 0.492
    + CDA 0.722 0.569 0.339
    + DROPOUT 0.765 0.554 0.377
    + INLP 0.204 0.639 0.460
    + SENTENCEDEBIAS 0.434 0.612 0.439
    ALBERT 0.623 0.552 0.431
    + CDA 0.953 0.534 0.309
    + DROPOUT 0.697 0.660 0.412
    + INLP 0.345 0.574 0.357
    + SENTENCEDEBIAS 0.352 0.554 0.241
    RoBERTa 0.940 0.307 0.127
    + CDA 0.880 0.322 0.245
    + DROPOUT 1.074 0.383 0.167
    + INLP 0.823 0.316 0.246
    + SENTENCEDEBIAS 0.846 0.273 0.271
    GPT-2 0.113 0.448 0.376
    + CDA 0.480 0.139 0.138
    + DROPOUT 0.476 0.162 0.134
    + INLP 0.119 0.447 0.375
    + SENTENCEDEBIAS 0.251 0.421 0.547

    For gender in masked language models, INLP and SentenceDebias achieve lower effect sizes than baselines (e.g., BERT INLP drops to 0.2040.204), while CDA and Dropout often increase average effect sizes. On GPT-2 gender SEAT, the baseline showed no statistically significant bias (∣d∣=0.113|d| = 0.113), and all debiased versions exhibited higher effect sizes.

  11. Knowl 11 — Language Modeling Perplexity on WikiText-2 Under Debiasing

    data/table

    Perplexity (for GPT-2) and pseudo-perplexity (for BERT) computed on 10%10\% of WikiText-2 alongside StereoSet Language Modeling (LM) scores for gender-debiased models:

    Model WikiText-2 Perplexity (↓\downarrow) StereoSet LM Score (%) (↑\uparrow)
    BERT 4.469 84.17
    + CDA 4.096 83.08
    + DROPOUT 4.202 83.04
    + INLP 6.152 80.63
    + SELF-DEBIAS 5.494 84.09
    + SENTENCEDEBIAS 4.483 84.20
    GPT-2 30.158 91.01
    + CDA 35.343 90.36
    + DROPOUT 37.370 90.40
    + INLP 42.534 91.62
    + SELF-DEBIAS 31.909 89.07
    + SENTENCEDEBIAS 65.493 87.43

    Debiasing interventions generally increase perplexity and degrade LM score. SentenceDebias on GPT-2 more than doubles perplexity from 30.15830.158 to 65.49365.493. Minor pseudo-perplexity improvements for BERT CDA (4.0964.096) and Dropout (4.2024.202) relative to baseline BERT (4.4694.469) stem from the additional Wikipedia training phase.

  12. Knowl 12 — GLUE Benchmark Performance for Gender Debiased Models

    data/table

    Validation performance on the GLUE benchmark for gender-debiased models (F1 for MRPC, Spearman correlation for STS-B, Matthew's correlation for CoLA, and accuracy for MNLI, QNLI, QQP, RTE, SST-2, WNLI; means across three training runs using 3 epochs, sequence length 128, batch size 32, learning rate 2×10−52 \times 10^{-5}):

    Model CoLA MNLI MRPC QNLI QQP RTE SST-2 STS-B WNLI Average
    BERT 55.89 84.50 88.59 91.38 91.03 63.54 92.58 88.51 43.66 77.74
    + CDA 55.90 84.73 88.76 91.36 91.01 66.31 92.43 89.14 38.03 77.52
    + DROPOUT 49.83 84.67 88.20 91.27 90.36 64.02 92.58 88.47 37.09 76.28
    + INLP 56.06 84.81 88.61 91.34 90.92 64.98 92.51 88.70 32.86 76.76
    + SENTENCEDEBIAS 56.41 84.80 88.70 91.48 90.98 63.06 92.32 88.45 44.13 77.81
    ALBERT 55.51 85.58 91.55 91.49 90.65 71.36 92.13 90.43 43.19 79.10
    + CDA 53.11 85.17 91.53 90.99 90.69 65.46 92.43 90.62 42.72 78.08
    + DROPOUT 12.37 85.33 90.25 91.79 90.39 56.56 92.24 89.93 52.11 73.44
    + INLP 55.87 85.32 92.07 91.58 90.53 72.92 91.86 90.80 47.42 79.82
    + SENTENCEDEBIAS 53.80 85.48 91.30 91.75 90.68 70.04 92.51 90.67 39.91 78.46
    RoBERTa 57.61 87.61 90.38 92.59 91.28 71.24 94.42 90.05 56.34 81.28
    + CDA 59.39 87.69 91.49 92.74 91.31 71.12 94.19 90.14 50.70 80.97
    + DROPOUT 51.60 87.35 90.13 92.82 90.43 65.70 94.34 88.97 51.17 79.17
    + INLP 58.38 87.49 91.39 92.65 91.31 69.31 94.30 89.81 56.34 81.22
    + SENTENCEDEBIAS 58.13 87.52 90.80 92.64 91.26 71.36 94.57 90.00 56.34 81.40
    GPT-2 29.10 82.43 84.51 87.71 89.18 64.74 91.97 84.26 43.19 73.01
    + CDA 37.57 82.61 85.91 88.08 89.26 64.86 92.09 85.28 42.25 74.21
    + DROPOUT 30.48 82.37 86.12 87.63 88.57 64.14 91.90 84.06 43.19 73.16
    + INLP 31.79 82.73 84.34 87.81 89.17 64.38 92.01 83.99 41.31 73.06
    + SENTENCEDEBIAS 30.20 82.56 84.43 87.90 89.09 64.86 91.97 84.18 38.50 72.63

    Downstream GLUE performance remains robust across debiasing methods, with average score changes generally within ±1.5\pm 1.5 percentage points of baseline.

  13. Knowl 13 — Methodological Scope and Evaluation Limitations of Bias Benchmarking

    limitation

    The empirical findings of debiasing surveys are constrained by four major methodological limitations:

    1. Language and Morphological Restrictions: Evaluations are restricted to English pre-trained models. Word-swapping CDA and subspace projection techniques cannot be directly applied to languages with grammatical gender marking across multiple parts of speech without morphology-aware extensions.
    2. Cultural Skew and Benchmark Validity: Benchmarks like StereoSet and CrowS-Pairs rely on crowdsourced annotations from North American workers, reflecting North American social biases. Furthermore, these intrinsic benchmarks provide only positive predictive power (identifying the presence of bias) and cannot verify a model as unbiased (e.g., a stereotype score of 50%50\% does not ensure impartiality).
    3. Simplifying Assumptions of Bias: Existing debiasing techniques assume binary definitions of gender and discretize racial and religious categories into fixed, small lexicons, omitting non-binary gender and broader social intersections.
    4. Representational vs. Extrinsic Harm: Survey evaluations measure intrinsic representational associations and token distributions in controlled templates rather than quantifying extrinsic harms or downstream impacts on human users in real-world deployments.

Coverage note — No substantial contributed material was omitted from the survey's evaluation across models, bias domains, debiasing methods, and benchmarks.

References

  1. 1.Hervé Abdi and Lynne J. Williams. 2010. Principal component analysis: Principal component analysis. Wiley Interdisciplinary Reviews: Computational Statistics, 2(4):433–459.
  2. 2.Vamsi Aribandi, Yi Tay, and Donald Metzler. 2021. How Reliable are Model Diagnostics? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1778–1785, Online. Association for Computational Linguistics.
  3. 3.Soumya Barikeri, Anne Lauscher, Ivan Vulic, and Goran Glavaš. 2021. RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1941–1955, Online. Association for Computational Linguistics.
  4. 4.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  5. 5.Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1004–1015, Online. Association for Computational Linguistics.
  6. 6.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. NIPS’16: Proceedings of the 30th International Conference on Neural Information Processing Systems, pages 4356 – 4364.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 33:1877–1901.
  8. 8.Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186. Publisher: American Association for the Advancement of Science Section: Reports.
  9. 9.Shrey Desai and Greg Durrett. 2020. Calibration of Pre-trained Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020a. Queens are Powerful too: Mitigating Gender Bias in Dialogue Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8173–8188, Online. Association for Computational Linguistics.
  12. 12.Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, and Adina Williams. 2020b. Multi-Dimensional Gender Bias Classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 314–331, Online. Association for Computational Linguistics.
  13. 13.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How Can We Know What Language Models Know? Transactions of the Association for Computational Linguistics, 8:423–438.
  14. 14.Masahiro Kaneko and Danushka Bollegala. 2021. Debiasing Pre-trained Contextualised Embeddings. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1256–1266, Online. Association for Computational Linguistics.
  15. 15.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring Bias in Contextualized Word Representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 166–172, Florence, Italy. Association for Computational Linguistics.
  16. 16.Anne Lauscher, Tobias Lueken, and Goran Glavaš. 2021. Sustainable Modular Debiasing of Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4782–4797, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  17. 17.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A Community Library for Natural Language Processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  18. 18.Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020. Towards Debiasing Sentence Representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5502–5515, Online. Association for Computational Linguistics.
  19. 19.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.
  20. 20.Thomas Manzini, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019. Black is to Criminal as Caucasian is to Police: Detecting and Removing Multiclass Bias in Word Embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 615–621, Minneapolis, Minnesota. Association for Computational Linguistics.
  21. 21.Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On Measuring Social Biases in Sentence Encoders. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 622–628, Minneapolis, Minnesota. Association for Computational Linguistics.
  22. 22.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  23. 23.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pre-trained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371, Online. Association for Computational Linguistics.
  24. 24.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online. Association for Computational Linguistics.
  25. 25.Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep Contextualized Word Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  26. 26.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  27. 27.Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020. Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7237–7256, Online. Association for Computational Linguistics.
  28. 28.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked Language Model Scoring. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2699–2712. ArXiv: 1910.14659.
  29. 29.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  30. 30.Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(1):1929–1958.
  31. 31.Alex Wang and Kyunghyun Cho. 2019. BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model. In Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation, pages 30–36, Minneapolis, Minnesota. Association for Computational Linguistics.
  32. 32.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, and Slav Petrov. 2020. Measuring and Reducing Gendered Correlations in Pre-trained Models. arXiv:2010.06032 [cs]. ArXiv: 2010.06032.
  33. 33.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Thibault Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  34. 34.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, New Orleans, Louisiana. Association for Computational Linguistics.
  35. 35.Pei Zhou, Weijia Shi, Jieyu Zhao, Kuan-Hao Huang, Muhao Chen, Ryan Cotterell, and Kai-Wei Chang. 2019. Examining Gender Bias in Languages with Grammatical Gender. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5276–5284, Hong Kong, China. Association for Computational Linguistics.
  36. 36.Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019. Counterfactual Data Augmentation for Mitigating Gender Stereotypes in Languages with Rich Morphology. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1651–1661, Florence, Italy. Association for Computational Linguistics.

Citation

MLA
Meade, N., et al. “An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1878–98, https://doi.org/10.18653/v1/2022.acl-long.132.
APA
Meade, N., Poole-Dayan, E., & Reddy, S. (2022). An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1878–1898. https://doi.org/10.18653/v1/2022.acl-long.132
Chicago
Meade, N., E. Poole-Dayan, and S. Reddy. 2022. “An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1878–98. https://doi.org/10.18653/v1/2022.acl-long.132.
Harvard
Meade, N., Poole-Dayan, E. and Reddy, S. (2022) “An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1878–1898. Available at: https://doi.org/10.18653/v1/2022.acl-long.132.
Vancouver
1. Meade N, Poole-Dayan E, Reddy S (2022) An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1878–1898

BibTeX

@inproceedings{meade-etal-2022-empirical,
    title = "An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models",
    author = "Meade, Nicholas  and
      Poole-Dayan, Elinor  and
      Reddy, Siva",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.132/",
    doi = "10.18653/v1/2022.acl-long.132",
    pages = "1878--1898"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/