Benchmarking Intersectional Biases in NLP
John LalorYi YangKendall SmithNicole ForsgrenAhmed Abbasi
Reveals that standard debiasing techniques fail to prevent amplified allocational harms across intersecting demographic dimensions in downstream natural language processing tasks despite preserving predictive accuracy.
As natural language processing (NLP) systems become widely deployed across high-stakes domains such as healthcare and hiring, ensuring their fairness is critical. Historically, most bias mitigation efforts have focused on representational harm—such as stereotypical associations in word representations—and examined only single demographic traits like gender. There has been limited evaluation of how debiasing methods prevent allocational harms, which occur when algorithms unfairly distribute opportunities or resources during real-world downstream prediction tasks, particularly across intersectional groups combining multiple demographic factors.
The article systematically evaluates the downstream predictive performance and fairness of multiple standard and debiased NLP models across several demographic intersections. The authors benchmark three prominent language representation models—BERT, RoBERTa, and GloVe—alongside four state-of-the-art debiasing techniques across ten sequence classification tasks derived from five datasets covering user-generated text from social media, web forums, and surveys. The evaluation incorporates up to five demographic variables: gender, race, age, income, and education, allowing for multi-group intersectional evaluations rather than isolated single-dimension assessments.
The analysis reveals several crucial findings. First, while existing debiasing strategies successfully preserve the predictive accuracy of language models, they fail to substantially alleviate bias in downstream tasks. Second, evaluating models on a single demographic dimension significantly masks their true unfairness: while single-dimension gender disparities remained modest (within 5% to 10%), disparate impact amplified dramatically across intersectional groups, worsening unfairness rates by 20% to 50% and increasing fairness violations by a factor of 3 to 10. Third, on sensitive prediction tasks such as assessing patient health literacy, numeracy, and personality traits, debiased models routinely produced predictions that violated legal and regulatory fairness thresholds, such as a three-to-one positive prediction disparity across demographic subgroups in numeracy. Finally, transformer-based models like BERT and RoBERTa consistently exhibited higher predictive accuracy and lower overall bias compared to older static embedding models like GloVe.
These findings indicate that current NLP debiasing practices provide a false sense of security. Organizations deploying language models based solely on standard single-variable debiasing tests face substantial compliance, legal, and operational risks of disparate impact when applying these tools in automated decision-making. Future debiasing research and organizational governance must shift focus from simply sanitizing upstream word representations to directly engineering fairer downstream classifiers. Organizations should also systematically collect multi-attribute demographic data to audit models for intersectional disparities prior to deployment in sensitive operational workflows.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Its distinction between representational and allocational harms supplies the conceptual framework for understanding why this benchmark tests downstream decisions rather than only stereotypical representations.
- Paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al. (2016). Its methods for measuring and reducing gender bias in word embeddings provide essential background for the debiasing techniques and representations compared in the benchmark.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). Its demonstration that embeddings reproduce human social biases establishes the upstream representation problem that the benchmark tests against downstream fairness outcomes.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). Its benchmark for measuring stereotypes in pretrained language models offers a useful contrast to this paper’s focus on whether such biases translate into unfair downstream predictions.
- Paper: Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, Joy Buolamwini et al. (2018). Its intersectional evaluation across gender and skin tone demonstrates why examining combined identities can reveal disparities hidden by aggregate or single-group results.
- Paper: Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning, Fan Zhou et al. (2023). Building on evidence that upstream debiasing can fail to protect downstream predictions, it integrates bias mitigation into fine-tuning to prevent bias from resurfacing in deployed tasks.
