Improving Gender Fairness of Pre-Trained Language Models without Catastrophic Forgetting
Zahra FatemiChen XingWenhao LiuCaiming Xiong
Presents GEnder Equality Prompt (GEEP), a prompt-tuning method that freezes base model parameters and trains dedicated profession embeddings on gender-neutral data to mitigate gender bias while preventing catastrophic forgetting on general NLP benchmarks.
Large pre-trained language models routinely absorb and amplify societal gender biases from web-scale data, often associating professions like "nurse" with women and "doctor" with men. While standard remediation methods train existing models on smaller, balanced datasets, this secondary training causes "catastrophic forgetting," where models lose general language knowledge and suffer major performance declines on downstream tasks.
The article demonstrates the extent of this forgetting problem and introduces a lightweight debiasing framework called GEnder Equality Prompt (GEEP) to reduce profession-based gender bias while preserving general model capabilities.
To evaluate debiasing methods, the authors constructed a 6.1-gigabyte gender-neutral dataset from Wikipedia by swapping gendered terms across profession-related sentences. Standard fine-tuning across all model parameters was compared against the proposed method, which freezes the original model and introduces newly initialized, trainable embeddings for 303 profession terms. Evaluations were conducted primarily on RoBERTa across standard fairness benchmarks—such as pronoun-profession resolution tasks—and the eight-task General Language Understanding Evaluation (GLUE) benchmark.
Updating all model parameters reduced overall language understanding on GLUE from 86.5 to 80.2 points, with steep drops of over 18 to 25 points on sensitive linguistic tasks. In contrast, the prompt-based method retained strong general capability, scoring 83.3 on GLUE while outperforming standard debiasing across fairness benchmarks. Specifically, it boosted profession pronoun resolution accuracy on the Winogender dataset to 64.5%, compared to 57.3% for standard debiasing and 50.9% for the baseline model. Furthermore, on direct pronoun prediction across 303 professions, the proposed method reduced gender bias scores by 70% compared to base BERT, and achieved near-peak fairness performance in one-fifth of the training iterations required by traditional retraining.
These findings indicate that updating entire neural networks on narrow debiasing data introduces severe operational risks by degrading broader linguistic competence. The proposed method resolves this trade-off by adding only 232,000 parameters (a 0.21% increase), providing a cost-effective, scalable way to debias language models rapidly without requiring full retraining.
Organizations seeking to improve artificial intelligence fairness should adopt parameter-efficient prompt tuning rather than full-model retraining. Future work should extend this approach beyond binary profession terms to address broader non-binary gender identities, intersectional fairness, and other demographic attributes before deploying models in high-stakes environments.
- Paper: An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models, Nicholas Meade et al. (2022). Its comparison of debiasing methods and their effects on GLUE performance establishes the fairness–utility trade-off that GEEP is designed to address.
- Paper: Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation, Tu Vu et al. (2022). Its comparison of prompt tuning with full-model fine-tuning shows how freezing a pretrained model can limit catastrophic forgetting, the central design choice behind GEEP.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). Its introduction of trainable continuous prompts for frozen language models provides the prompt-tuning mechanism that GEEP adapts for debiasing.
- Paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al. (2016). Its account of measuring and reducing gender associations in embeddings supplies foundational context for GEEP’s profession-bias target.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). Its StereoSet benchmark explains a standard way to evaluate stereotypical bias alongside language-modeling capability, a relevant concern in GEEP’s fairness–performance comparison.
- Paper: "Thinking" Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models, Shaz Furniturewala et al. (2024). It carries bias mitigation into structured, inference-time prompts, offering a follow-on route to reduce stereotypes when model parameters cannot be updated.
- Paper: MISGENDERED: Limits of Large Language Models in Understanding Pronouns, Tamanna Hossain et al. (2023). It extends profession-focused gender fairness evaluation toward singular they and neo-pronouns, directly probing an identity scope the source identifies as future work.
