Benchmarking Intersectional Biases in NLP

John LalorYi YangKendall SmithNicole ForsgrenAhmed Abbasi

article2022NAACL58 citations

Reveals that standard debiasing techniques fail to prevent amplified allocational harms across intersecting demographic dimensions in downstream natural language processing tasks despite preserving predictive accuracy.

Listen

As natural language processing (NLP) systems become widely deployed across high-stakes domains such as healthcare and hiring, ensuring their fairness is critical. Historically, most bias mitigation efforts have focused on representational harm—such as stereotypical associations in word representations—and examined only single demographic traits like gender. There has been limited evaluation of how debiasing methods prevent allocational harms, which occur when algorithms unfairly distribute opportunities or resources during real-world downstream prediction tasks, particularly across intersectional groups combining multiple demographic factors.

The article systematically evaluates the downstream predictive performance and fairness of multiple standard and debiased NLP models across several demographic intersections. The authors benchmark three prominent language representation models—BERT, RoBERTa, and GloVe—alongside four state-of-the-art debiasing techniques across ten sequence classification tasks derived from five datasets covering user-generated text from social media, web forums, and surveys. The evaluation incorporates up to five demographic variables: gender, race, age, income, and education, allowing for multi-group intersectional evaluations rather than isolated single-dimension assessments.

The analysis reveals several crucial findings. First, while existing debiasing strategies successfully preserve the predictive accuracy of language models, they fail to substantially alleviate bias in downstream tasks. Second, evaluating models on a single demographic dimension significantly masks their true unfairness: while single-dimension gender disparities remained modest (within 5% to 10%), disparate impact amplified dramatically across intersectional groups, worsening unfairness rates by 20% to 50% and increasing fairness violations by a factor of 3 to 10. Third, on sensitive prediction tasks such as assessing patient health literacy, numeracy, and personality traits, debiased models routinely produced predictions that violated legal and regulatory fairness thresholds, such as a three-to-one positive prediction disparity across demographic subgroups in numeracy. Finally, transformer-based models like BERT and RoBERTa consistently exhibited higher predictive accuracy and lower overall bias compared to older static embedding models like GloVe.

These findings indicate that current NLP debiasing practices provide a false sense of security. Organizations deploying language models based solely on standard single-variable debiasing tests face substantial compliance, legal, and operational risks of disparate impact when applying these tools in automated decision-making. Future debiasing research and organizational governance must shift focus from simply sanitizing upstream word representations to directly engineering fairer downstream classifiers. Organizations should also systematically collect multi-attribute demographic data to audit models for intersectional disparities prior to deployment in sensitive operational workflows.

Lalor et al (2022).pdf
Cover for Benchmarking Intersectional Biases in NLP

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Allocational and Representational Harms
  • 2.2 Intersectional Biases
  • 2.3 Debiasing
  • 3 Data, Models, Experiments
  • 3.1 Data
  • 3.2 Models and Debiasing Methods
  • 3.3 Evaluation
  • 4 Results and Discussion
  • 5 Conclusion
  • Acknowledgement
  • References
  • A Appendix: RoBERTa Results

Knowls

  1. Knowl 1 — Benchmark datasets cover ten prediction tasks and five demographic dimensions

    data/table

    The benchmark evaluates ten text-prediction tasks across five datasets, with sample sizes and demographic attributes as follows: Psychometrics (8,395 survey responses; anxiety, literacy, numeracy, and trust; gender, race, age, income, and education); Multilingual Twitter Corpus (83,078 tweets; hate-speech identification; gender, race, and age); Five Item Personality Inventory (6,805 survey responses; extraverted and stable traits; gender, race, age, income, and education); AskAPatient (20,000 forum posts; sentiment; gender and age); and Myers–Briggs Type Indicator (7,406 texts from 1,584 unique Reddit users; perceiving and thinking traits; gender and age). Psychometrics labels are participant-provided survey scores linked to free-text responses; the personality tasks use free text to predict personality traits. Demographics are self-reported in four datasets; the Twitter corpus uses inferred author demographics.

  2. Knowl 2 — Models, debiasing methods, and training design

    model/method

    The benchmark compares a GloVe-initialized word CNN with BERT-base-uncased and RoBERTa-base. The CNN has three parallel convolutional layers with kernel sizes 1, 2, and 3, each with 256 filters, ReLU activation, L2 regularization of 0.001, and global max pooling; it is trained for 35 epochs with batch size 32 and learning rate 10−410^{-4}. Baseline BERT and RoBERTa are fine-tuned for five epochs with batch size 32, learning rate 10−510^{-5}, and weight decay 0.01, retaining the checkpoint with lowest validation loss. Static GloVe embeddings are debiased with WordED, which learns a projection over 50 epochs. BERT and RoBERTa are debiased with ContextED using gender and stereotype word lists and News Commentary v15 sentences containing those words; all layers are debiased at token level, with debiasing loss weight 0.8, then fine-tuned for three epochs with batch size 32 and learning rate 5×10−55\times10^{-5}. Counterfactual data augmentation (CDA) and increased dropout are also tested as alternative BERT debiasing strategies. All datasets use five-fold cross-validation, with out-of-fold predictions combined for fairness calculations.

  3. Knowl 3 — Intersectional subgroups are constructed around gender

    experimental setup

    For intersectional audits, the benchmark uses gender as a reference demographic and enumerates protected groups for combinations of demographic dimensions that include gender. For example, a two-demographic audit can compare older women with the complementary younger-men group, or lower-education women with higher-education men. The same construction is extended to three- and four-demographic combinations; the five-demographic analysis considers all available dimensions together. The benchmark calculates disparate impact and fairness violation for these groups, using the corresponding complementary group as privileged. This procedure is intended to assess disjoint subgroup comparisons rather than combine results from separate single-demographic audits.

  4. Knowl 4 — Adjusted disparate impact accounts for observed label base rates

    equation

    Let A=0A=0 denote a protected group and A=1A=1 its privileged comparison group; let y∈{0,1}y\in\{0,1\} be the true label and y^∈{0,1}\hat y\in\{0,1\} the classifier prediction. Disparate impact (DI) is the ratio of positive-prediction rates, and adjusted disparate impact (ADI) divides this by the corresponding ratio of positive-label base rates:

    DI=P(y^=1∣A=0)P(y^=1∣A=1),DI∗=P(y=1∣A=0)P(y=1∣A=1),ADI=DIDI∗.\mathrm{DI}=\frac{P(\hat y=1\mid A=0)}{P(\hat y=1\mid A=1)},\qquad \mathrm{DI}^{*}=\frac{P(y=1\mid A=0)}{P(y=1\mid A=1)},\qquad \mathrm{ADI}=\frac{\mathrm{DI}}{\mathrm{DI}^{*}}.

    DI equal to 1 indicates equal positive-prediction rates across the two groups. The adjustment is intended to account for differences in observed label prevalence. The study applies additive smoothing when zero counts would otherwise make DI or ADI undefined.

  5. Knowl 5 — Fairness violation measures the worst subgroup true-positive-rate gap

    equation

    For a classifier evaluated on a dataset, let GfG_f be the set of demographic groups being audited, TPRg\mathrm{TPR}_g the true-positive rate among examples in group gg, and TPRD\mathrm{TPR}_D the true-positive rate over the full dataset. The benchmark's fairness violation is the largest absolute subgroup-to-dataset gap:

    FV=max⁡g∈Gf∣TPRg−TPRD∣.\mathrm{FV}=\max_{g\in G_f}\left|\mathrm{TPR}_g-\mathrm{TPR}_D\right|.

    The study uses the maximum rather than the average group gap to emphasize worst-case subgroup unfairness. The metric is compared with an acceptability parameter λ\lambda; no single value of λ\lambda is specified for the benchmark.

  6. Knowl 6 — Intersectional audits reveal wider disparities than single-demographic audits

    empirical result

    Across the evaluated tasks, fairness commonly looks closer to parity when only gender is considered than when gender is combined with other demographic dimensions. For BERT, gender-only ADI is generally within about 10% of 1 on most tasks; as additional demographics are included, the range of ADI values widens. Fairness-violation gaps also commonly grow with intersectional complexity, often by factors of 3 to 10 in the reported comparisons. The authors summarize intersectional unfairness rates as often 20–50% higher than single-demographic rates. These patterns show that a favorable single-axis fairness assessment does not establish fairness for demographic intersections.

  7. Knowl 7 — Embedding debiasing preserves predictive performance more reliably than it reduces downstream bias

    empirical result

    Across the benchmark, debiased models generally retain predictive performance close to their corresponding non-debiased models, while their downstream fairness improvements are small or inconsistent. Reducing bias for a target dimension such as gender does not reliably remove disparities for intersectional groups. The same broad increase in unfairness with additional demographic intersections appears for BERT when using ContextED, CDA, or Dropout on anxiety, literacy, numeracy, and trust tasks; RoBERTa with ContextED shows similar trends to BERT. Thus, the tested embedding-debiasing approaches preserve task performance comparatively well but do not reliably alleviate intersectional disparities in classification outcomes.

  8. Knowl 8 — Some tasks exhibit severe intersectional disparities

    empirical result

    The benchmark reports especially large disparities in some task-and-subgroup combinations. For BERT on numeracy prediction, the authors describe positive predictions for a five-demographic protected group as occurring at roughly a three-to-one ratio relative to its privileged comparison group; the reported five-way ADI is 3.23. For three-demographic hate-speech detection, positive predictions are reported as substantially less likely for the protected group than for the privileged group. These findings illustrate that the direction and magnitude of disparities depend on the task and subgroup, and that intersectional auditing can expose gaps obscured by single-demographic results.

  9. Knowl 9 — Bias and predictive performance vary by embedding family and task

    empirical result

    Across the benchmark as a whole, GloVe-based models tend to show more pronounced biases and lower predictive performance than the transformer-based models, leading the authors to characterize debiased transformer models as generally fairer and more predictive. This is a broad trend, not a guarantee for every task or subgroup. Anxiety and trust provide counterexamples to a uniform direction of bias: as more demographic dimensions are considered, GloVe tends to skew more unfairly against protected groups, whereas BERT remains closer to parity and skews slightly unfairly against privileged groups. The model comparison therefore does not support treating an embedding family as uniformly fair across tasks.

  10. Knowl 10 — The benchmark evaluates embedding debiasing, not classifier debiasing

    limitation

    The study explicitly limits its scope to debiasing input embeddings and does not evaluate methods that debias the downstream classifiers themselves. Its contextual embedding methods are targeted at gender bias, even though the fairness audits examine intersections involving gender, race, age, education, and income. The benchmark also combines different demographic measurement sources: four datasets provide self-reported demographics, while the Twitter corpus uses inferred demographics. Consequently, its findings characterize the tested embedding-debiasing methods and datasets rather than all classifier-level mitigation strategies or all demographic measurement settings.

Coverage note — The full model-by-task benchmark matrix and every plotted confidence interval are not reproduced; the knowls retain the main cross-task findings and selected task-specific results rather than the complete set of numerical comparisons.

References

  1. 1.Ahmed Abbasi, David Dobolyi, John P. Lalor, Richard Netemeyer, Kendall Smith, and Yi Yang. 2021. Constructing a psychometric testbed for fair natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  2. 2.Solon Barocas and Andrew D Selbst. 2016. Big data’s disparate impact. Calif. L. Rev., 104:671.
  3. 3.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, pages 610–623, New York, NY, USA. Association for Computing Machinery.
  4. 4.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  5. 5.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29:4349–4357.
  6. 6.Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency, pages 77–91. PMLR.
  7. 7.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  8. 8.Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019. Bias in bios: A case study of semantic representation bias in a high-stakes setting. In proceedings of the Conference on Fairness, Accountability, and Transparency, pages 120–128.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Sorelle A Friedler, Carlos Scheidegger, Suresh Venkatasubramanian, Sonam Choudhary, Evan P Hamilton, and Derek Roth. 2019. A comparative study of fairness-enhancing interventions in machine learning. In Proceedings of the conference on fairness, accountability, and transparency, pages 329–338.
  11. 11.Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences, 115(16):E3635–E3644.
  12. 12.Aparna Garimella, Akhash Amarnath, Kiran Kumar, Akash Pramod Yalla, N Anandhavelu, Niyati Chhaya, and Balaji Vasan Srinivasan. 2021. He is very intelligent, she is very beautiful? on mitigating social biases in language modelling and generation. In Findings of ACL, pages 4534–4545.
  13. 13.Matej Gjurkovic, Mladen Karan, Iva Vukojevi ´ c, Mi- ´ haela Bošnjak, and Jan Snajder. 2021. PANDORA talks: Personality and demographics on Reddit. In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media, pages 138–152, Online. Association for Computational Linguistics.
  14. 14.Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sanchez, Mugdha Pandya, and Adam Lopez. 2021. Intrinsic bias metrics do not correlate with application bias. In Proceedings of ACL.
  15. 15.Hila Gonen and Yoav Goldberg. 2019. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 609–614.
  16. 16.Yue Guo, Yi Yang, and Ahmed Abbasi. 2022. Auto-debias: Debiasing masked language models with automated biased prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL).
  17. 17.Xiaolei Huang, Linzi Xing, Franck Dernoncourt, and Michael Paul. 2020. Multilingual twitter corpus and baselines for evaluating demographic bias in hate speech recognition. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 1440–1448.
  18. 18.Masahiro Kaneko and Danushka Bollegala. 2019. Gender-preserving debiasing for pre-trained word embeddings. In Proceedings of ACL, pages 1641–1650.
  19. 19.Masahiro Kaneko and Danushka Bollegala. 2021. Debiasing pre-trained contextualised embeddings. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1256–1266.
  20. 20.Michael Kearns, Seth Neel, Aaron Roth, and Zhiwei Steven Wu. 2018. Preventing fairness gerrymandering: Auditing and learning for subgroup fairness. In International Conference on Machine Learning, pages 2564–2572. PMLR.
  21. 21.Nut Limsopatham and Nigel Collier. 2016. Normalising medical concepts in social media texts by learning semantic representation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1014–1023.
  22. 22.Haochen Liu, Wei Jin, Hamid Karimi, Zitao Liu, and Jiliang Tang. 2021. The authors matter: Understanding and mitigating implicit bias in deep text classification. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 74–85, Online. Association for Computational Linguistics.
  23. 23.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  24. 24.Chandler May, Alex Wang, Shikha Bordia, Samuel Bowman, and Rachel Rudinger. 2019. On measuring social biases in sentence encoders. In Proceedings of NAACL), pages 622–628.
  25. 25.Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR), 54(6):1–35.
  26. 26.Richard G Netemeyer, David G Dobolyi, Ahmed Abbasi, Gari Clifford, and Herman Taylor. 2020. Health literacy, health numeracy, and trust in doctor: effects on key patient health outcomes. Journal of Consumer Affairs, 54(1):3–42.
  27. 27.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543.
  28. 28.Flavien Prost, Nithum Thain, and Tolga Bolukbasi. 2019. Debiasing embeddings for reduced gender bias in text classification. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 69–75.
  29. 29.Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020. Null it out: Guarding protected attributes by iterative nullspace projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7237–7256.
  30. 30.Alexey Romanov, Maria De-Arteaga, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, Anna Rumshisky, and Adam Kalai. 2019. What’s in a name? reducing bias in bios without access to protected attributes. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4187–4195.
  31. 31.Deven Santosh Shah, H Andrew Schwartz, and Dirk Hovy. 2020. Predictive biases in natural language processing models: A conceptual framework and overview. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5248–5264.
  32. 32.Shivashankar Subramanian, Xudong Han, Timothy Baldwin, Trevor Cohn, and Lea Frermann. 2021. Evaluating debiasing techniques for intersectional biases. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2492–2498.
  33. 33.Harini Suresh and John V. Guttag. 2019. A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle.
  34. 34.Yi Chern Tan and L. Elisa Celis. 2019. Assessing social and intersectional biases in contextualized word representations. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  35. 35.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020. Measuring and reducing gendered correlations in pre-trained models. arXiv preprint arXiv:2010.06032.
  36. 36.Forest Yang, Mouhamadou Cisse, and Oluwasanmi O Koyejo. 2020. Fairness with overlapping groups; a probabilistic perspective. Advances in Neural Information Processing Systems, 33.
  37. 37.Chengxiang Zhai and John Lafferty. 2004. A study of smoothing methods for language models applied to information retrieval. ACM Transactions on Information Systems (TOIS), 22(2):179–214.
  38. 38.Jieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang, and Ahmed Hassan Awadallah. 2020. Gender bias in multilingual embeddings and cross-lingual transfer. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2896–2907.
  39. 39.Jieyu Zhao, Yichao Zhou, Zeyu Li, Wei Wang, and Kai-Wei Chang Chang. 2018. Learning gender-neutral word embeddings. In Proceedings of EMNLP.
  40. 40.Ran Zmigrod, Sabrina J Mielke, Hanna Wallach, and Ryan Cotterell. 2019. Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1651–1661.

Citation

MLA
Lalor, J. P., et al. “Benchmarking Intersectional Biases in NLP”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 3598–609, https://doi.org/10.18653/v1/2022.naacl-main.263.
APA
Lalor, J. P., Yang, Y., Smith, K., Forsgren, N., & Abbasi, A. (2022). Benchmarking Intersectional Biases in NLP. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3598–3609. https://doi.org/10.18653/v1/2022.naacl-main.263
Chicago
Lalor, J. P., Y. Yang, K. Smith, N. Forsgren, and A. Abbasi. 2022. “Benchmarking Intersectional Biases in NLP”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3598–3609. https://doi.org/10.18653/v1/2022.naacl-main.263.
Harvard
Lalor, J.P. et al. (2022) “Benchmarking Intersectional Biases in NLP”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 3598–3609. Available at: https://doi.org/10.18653/v1/2022.naacl-main.263.
Vancouver
1. Lalor JP, Yang Y, Smith K, Forsgren N, Abbasi A (2022) Benchmarking Intersectional Biases in NLP. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 3598–3609

BibTeX

@inproceedings{lalor-etal-2022-benchmarking,
    title = "Benchmarking Intersectional Biases in {NLP}",
    author = "Lalor, John  and
      Yang, Yi  and
      Smith, Kendall  and
      Forsgren, Nicole  and
      Abbasi, Ahmed",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.263/",
    doi = "10.18653/v1/2022.naacl-main.263",
    pages = "3598--3609"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/