Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models

Pieter DelobelleEwoenam Kwaku TokpoToon CaldersBettina Berendt

article2022NAACL104 citations

Demonstrates through empirical correlation analysis that existing intrinsic fairness metrics for language models are incompatible, overly sensitive to template and seed choices, and fail to predict extrinsic bias in downstream applications.

Listen

As artificial intelligence systems increasingly rely on pre-trained language models like BERT, measuring and mitigating social biases—such as gender stereotyping—has become a critical ethical and operational priority. However, the field currently lacks standardized evaluation practices. Numerous metrics exist to quantify internal model bias and downstream application unfairness, but practitioners struggle to determine whether these tools yield consistent, dependable evaluations.

The article systematically evaluates the compatibility and reliability of existing bias metrics for pre-trained language models. Its primary objective is to determine whether commonly used fairness measures agree with one another, how sensitive they are to arbitrary design choices, and whether internal bias in a base model accurately predicts unfair behavior in practical, downstream applications.

To evaluate these questions, the authors conducted a literature survey and experimental correlation analyses focusing on binary gender bias across professions. They tested five popular transformer-based language models, including standard, large, distilled, and multilingual variants of BERT, as well as RoBERTa. The analysis systematically varied sentence templates, seed word subsets (using 20 sampled groups of male- and female-stereotyped professions), and word representation techniques, while comparing internal model metrics against performance on real-world downstream tasks such as biography-based occupation prediction and pronoun coreference resolution.

The analysis revealed several critical findings regarding current fairness evaluations. First, internal bias metrics are highly unstable and deeply dependent on the specific sentence templates used. Even supposedly neutral, "semantically bleached" context sentences produced weak or negative correlations with each other (such as a negative 0.57 correlation between similar phrasing). Second, mathematical representations of words heavily distort results: extracting full sentence embeddings introduced severe contextual noise, whereas isolating the target word's specific representation provided much more stable measurements across templates. Third, internal bias metrics showed inconsistent or nonexistent relationships with external downstream unfairness. While certain targeted internal measures correlated well with profession classification bias, widely used benchmark datasets like CrowS-Pairs showed weak or conflicting relationships with downstream harms.

These findings indicate that relying on internal bias metrics provides a false sense of security. Because an internal score can fluctuate wildly based on minor, subjective choices—such as sentence framing or vector extraction methods—organizations risk approving models that appear unbiased in isolation but produce harmful, unfair outcomes when deployed. Furthermore, treating base model fairness scores as an insurance policy against downstream allocational harm is methodologically unsound.

Organizations evaluating language technologies should avoid sentence-level embedding metrics when measuring base representations, opting instead for metrics that measure target word tokens directly or methods that evaluate output probabilities without raw embedding comparisons. Most importantly, practitioners must focus validation resources primarily on extrinsic fairness testing within specific downstream applications, where actual allocation decisions and real-world impacts occur.

These conclusions should be interpreted in light of specific limitations. The experimental findings focus primarily on binary gender bias regarding English professions, and language-specific nuances (such as grammatical gender in German or Dutch) may introduce different dynamics. Nevertheless, there is high confidence in the core result: current internal fairness metrics are unreliable standalone indicators, and bias assessments must be anchored in application-specific testing.

arXiv: 2112.07447iPieter/biased-rulers
Cover for Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models

Abstract

An increasing awareness of biased patterns in natural language processing resources such as BERT has motivated many metrics to quantify ‘bias’ and ‘fairness’ in these resources. However, comparing the results of different metrics and the works that evaluate with such metrics remains difficult, if not outright impossible. We survey the literature on fairness metrics for pre-trained language models and experimentally evaluate compatibility, including both biases in language models and in their downstream tasks. We do this by combining traditional literature survey, correlation analysis and empirical evaluations. We find that many metrics are not compatible with each other and highly depend on (i) templates, (ii) attribute and target seeds and (iii) the choice of embeddings. We also see no tangible evidence of intrinsic bias relating to extrinsic bias. These results indicate that fairness or bias evaluation remains challenging for contextualized language models, among other reasons because these choices remain subjective. To improve future comparisons and fairness evaluations, we recommend to avoid embedding-based metrics and focus on fairness evaluations in downstream tasks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Fairness in word embeddings
  • 3 Measuring fairness in language models
  • 3.1 Intrinsic measures
  • 3.2 Extrinsic measures
  • 3.3 Measuring biases in other languages
  • 4 On the compatibility of measures
  • 4.1 Methodology
  • 4.2 Compatibility between templates
  • 4.3 Compatibility between representations
  • 4.4 Compatibility between metrics
  • 5 Code
  • 6 Discussion and ethical considerations
  • 7 Conclusion
  • Acknowledgements
  • References
  • A Templates
  • A.1 DisCo
  • A.2 SEAT
  • A.3 Vig et al. (2020)
  • A.4 BEC-Pro (English)
  • A.5 RobBERT (Dutch)
  • B Word lists for experiments
  • B.1 List of professions
  • B.2 List of target words
  • C Embedding methods
  • D Source code and datasets
  • E Evaluated templates

Knowls

  1. Knowl 1 — Extraction limitation: paper content unavailable

    limitation

    The source document provided for knowl extraction (file d950f863-da86-4ad6-8cc8-10b8527d3bd8.pdf) contained no readable text, so no methods, results, equations, algorithms, or data from the paper could be identified or summarized.

Coverage note — The PDF content was not accessible, so no contributed material from the paper could be extracted; all of the paper's contribution was necessarily omitted because nothing could be read.

References

  1. 1.Maria Antoniak and David Mimno. 2021. Bad seeds: Evaluating lexical methods for bias measurement. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 1889–1904. Association for Computational Linguistics.
  2. 2.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
  3. 3.Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2019. Fairness and Machine Learning. fairml-book.org. http://www.fairmlbook.org.
  4. 4.Marion Bartl, Malvina Nissim, and Albert Gatt. 2020. Unmasking contextual stereotypes: Measuring and mitigating BERT’s gender bias. In Proceedings of the Second Workshop on Gender Bias in Natural Language Processing, pages 1–16, Barcelona, Spain (Online). Association for Computational Linguistics.
  5. 5.Elisa Bassignana, Valerio Basile, and Viviana Patti. 2018. Hurtlex: A multilingual lexicon of words to hurt. In 5th Italian Conference on Computational Linguistics, CLiC-it 2018, volume 2253, pages 1–6. CEUR-WS.
  6. 6.Christine Basta, Marta R. Costa-jussa, and Noe Casas. 2019. Evaluating the underlying gender bias in contextualized word embeddings. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 33–39, Florence, Italy. Association for Computational Linguistics.
  7. 7.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA. Association for Computing Machinery.
  8. 8.Su Lin Blodgett, Solon Barocas, Hal Daume III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  9. 9.Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna M. Wallach. 2021. Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In ACL/IJCNLP.
  10. 10.Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, page 4356–4364, Red Hook, NY, USA. Curran Associates Inc.
  11. 11.Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  12. 12.Yang Trista Cao and Hal Daume III. 2020. Toward gender-inclusive coreference resolution. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4568–4595, Online. Association for Computational Linguistics.
  13. 13.Rodrigo Alejandro Chavez Mulsa and Gerasimos Spanakis. 2020. Evaluating bias in Dutch word embeddings. In Proceedings of the Second Workshop on Gender Bias in Natural Language Processing, pages 56–71, Barcelona, Spain (Online). Association for Computational Linguistics.
  14. 14.Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019. Bias in bios: A case study of semantic representation bias in a high-stakes setting. New York, NY, USA. Association for Computing Machinery.
  15. 15.Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2020. RobBERT: a Dutch RoBERTa-based Language Model. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3255–3265, Online. Association for Computational Linguistics.
  16. 16.Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2020. On measuring and mitigating biased inferences of word embeddings. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7659–7666.
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  18. 18.Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, and Adina Williams. 2020. Multi-dimensional gender bias classification. arXiv preprint arXiv:2005.00614.
  19. 19.Ismael Garrido-Munoz, Arturo Montejo-Raez, Fernando Martínez-Santiago, and L. Urena-Lopez. 2021. A survey on bias in deep nlp. Applied Sciences, 11:3184.
  20. 20.Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Munoz Sanchez, Mugdha Pandya, and Adam Lopez. 2021. Intrinsic bias metrics do not correlate with application bias. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1926–1940, Online. Association for Computational Linguistics.
  21. 21.Hila Gonen and Y. Goldberg. 2019. Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them. In NAACL.
  22. 22.Anthony G Greenwald, Debbie E McGhee, and Jordan LK Schwartz. 1998. Measuring individual differences in implicit cognition: the implicit association test. Journal of personality and social psychology, 74(6):1464.
  23. 23.Wei Guo and Aylin Caliskan. 2021. Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 122–133.
  24. 24.Masahiro Kaneko and Danushka Bollegala. 2021. Unmasking the mask–evaluating social biases in masked language models. arXiv preprint arXiv:2104.07496.
  25. 25.S. Kullback and R. A. Leibler. 1951. On Information and Sufficiency. The Annals of Mathematical Statistics, 22(1):79 – 86.
  26. 26.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring bias in contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 166–172, Florence, Italy. Association for Computational Linguistics.
  27. 27.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations.
  28. 28.Anne Lauscher, Tobias Luken, and Goran Glavaš. 2021. Sustainable modular debiasing of language models. arXiv preprint arXiv:2109.03646.
  29. 29.Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning.
  30. 30.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.
  31. 31.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692.
  32. 32.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  33. 33.Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On measuring social biases in sentence encoders. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 622–628, Minneapolis, Minnesota. Association for Computational Linguistics.
  34. 34.Tomas Mikolov, Kai Chen, G.s Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. Proceedings of Workshop at ICLR, 2013.
  35. 35.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. Stereoset: Measuring stereotypical bias in pre-trained language models. In ACL/IJCNLP.
  36. 36.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1953–1967, Online. Association for Computational Linguistics.
  37. 37.Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021. HONEST: Measuring hurtful sentence completion in language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2398–2406, Online. Association for Computational Linguistics.
  38. 38.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In EMNLP, volume 14, pages 1532–1543.
  39. 39.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations.
  40. 40.Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender bias in coreference resolution. arXiv preprint arXiv:1804.09301.
  41. 41.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Commun. ACM, 64(9):99–106.
  42. 42.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked language model scoring. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2699–2712, Online. Association for Computational Linguistics.
  43. 43.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. ArXiv, abs/1910.01108.
  44. 44.Joao Sedoc and Lyle Ungar. 2019. The role of protected class word lists in bias identification of contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 55–61, Florence, Italy. Association for Computational Linguistics.
  45. 45.Pamela Stone and Meg Lovejoy. 2004. Fast-track women and the “choice” to stay home. The ANNALS of the American Academy of Political and Social Science, 596(1):62–83.
  46. 46.Yi Chern Tan and L. Elisa Celis. 2019. Assessing social and intersectional biases in contextualized word representations. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  47. 47.Daniel de Vassimon Manela, David Errington, Thomas Fisher, Boris van Breugel, and Pasquale Minervini. 2021. Stereotype and skew: Quantifying gender bias in pre-trained and fine-tuned language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2232–2242, Online. Association for Computational Linguistics.
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.
  49. 49.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. In NeurIPS.
  50. 50.Ivan Vulic, Simon Baker, E. Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, T. Poibeau, Roi Reichart, and Anna Korhonen. 2020. Multi-simlex: A large-scale evaluation of multilingual and crosslingual lexical semantic similarity. Computational Linguistics, 46:847–897.
  51. 51.Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018. Mind the GAP: A balanced corpus of gendered ambiguous pronouns. Transactions of the Association for Computational Linguistics, 6:605–617.
  52. 52.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020. Measuring and reducing gendered correlations in pre-trained models. arXiv preprint arXiv:2010.06032.
  53. 53.Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. Ontonotes release 5.0 ldc2013t19. Linguistic Data Consortium, Philadelphia, PA, 23.
  54. 54.Jieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang, and Ahmed Hassan Awadallah. 2020. Gender bias in multilingual embeddings and cross-lingual transfer. arXiv preprint arXiv:2005.00699.
  55. 55.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019. Gender bias in contextualized word embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 629–634, Minneapolis, Minnesota. Association for Computational Linguistics.
  56. 56.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, New Orleans, Louisiana. Association for Computational Linguistics.

Citation

MLA
Delobelle, P., et al. “Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 1693–706, https://doi.org/10.18653/v1/2022.naacl-main.122.
APA
Delobelle, P., Tokpo, E., Calders, T., & Berendt, B. (2022). Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1693–1706. https://doi.org/10.18653/v1/2022.naacl-main.122
Chicago
Delobelle, P., E. Tokpo, T. Calders, and B. Berendt. 2022. “Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1693–1706. https://doi.org/10.18653/v1/2022.naacl-main.122.
Harvard
Delobelle, P. et al. (2022) “Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 1693–1706. Available at: https://doi.org/10.18653/v1/2022.naacl-main.122.
Vancouver
1. Delobelle P, Tokpo E, Calders T, Berendt B (2022) Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 1693–1706

BibTeX

@inproceedings{delobelle-etal-2022-measuring,
    title = "Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models",
    author = "Delobelle, Pieter  and
      Tokpo, Ewoenam  and
      Calders, Toon  and
      Berendt, Bettina",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.122/",
    doi = "10.18653/v1/2022.naacl-main.122",
    pages = "1693--1706"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/