Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models
Pieter DelobelleEwoenam Kwaku TokpoToon CaldersBettina Berendt
Demonstrates through empirical correlation analysis that existing intrinsic fairness metrics for language models are incompatible, overly sensitive to template and seed choices, and fail to predict extrinsic bias in downstream applications.
As artificial intelligence systems increasingly rely on pre-trained language models like BERT, measuring and mitigating social biases—such as gender stereotyping—has become a critical ethical and operational priority. However, the field currently lacks standardized evaluation practices. Numerous metrics exist to quantify internal model bias and downstream application unfairness, but practitioners struggle to determine whether these tools yield consistent, dependable evaluations.
The article systematically evaluates the compatibility and reliability of existing bias metrics for pre-trained language models. Its primary objective is to determine whether commonly used fairness measures agree with one another, how sensitive they are to arbitrary design choices, and whether internal bias in a base model accurately predicts unfair behavior in practical, downstream applications.
To evaluate these questions, the authors conducted a literature survey and experimental correlation analyses focusing on binary gender bias across professions. They tested five popular transformer-based language models, including standard, large, distilled, and multilingual variants of BERT, as well as RoBERTa. The analysis systematically varied sentence templates, seed word subsets (using 20 sampled groups of male- and female-stereotyped professions), and word representation techniques, while comparing internal model metrics against performance on real-world downstream tasks such as biography-based occupation prediction and pronoun coreference resolution.
The analysis revealed several critical findings regarding current fairness evaluations. First, internal bias metrics are highly unstable and deeply dependent on the specific sentence templates used. Even supposedly neutral, "semantically bleached" context sentences produced weak or negative correlations with each other (such as a negative 0.57 correlation between similar phrasing). Second, mathematical representations of words heavily distort results: extracting full sentence embeddings introduced severe contextual noise, whereas isolating the target word's specific representation provided much more stable measurements across templates. Third, internal bias metrics showed inconsistent or nonexistent relationships with external downstream unfairness. While certain targeted internal measures correlated well with profession classification bias, widely used benchmark datasets like CrowS-Pairs showed weak or conflicting relationships with downstream harms.
These findings indicate that relying on internal bias metrics provides a false sense of security. Because an internal score can fluctuate wildly based on minor, subjective choices—such as sentence framing or vector extraction methods—organizations risk approving models that appear unbiased in isolation but produce harmful, unfair outcomes when deployed. Furthermore, treating base model fairness scores as an insurance policy against downstream allocational harm is methodologically unsound.
Organizations evaluating language technologies should avoid sentence-level embedding metrics when measuring base representations, opting instead for metrics that measure target word tokens directly or methods that evaluate output probabilities without raw embedding comparisons. Most importantly, practitioners must focus validation resources primarily on extrinsic fairness testing within specific downstream applications, where actual allocation decisions and real-world impacts occur.
These conclusions should be interpreted in light of specific limitations. The experimental findings focus primarily on binary gender bias regarding English professions, and language-specific nuances (such as grammatical gender in German or Dutch) may introduce different dynamics. Nevertheless, there is high confidence in the core result: current internal fairness metrics are unreliable standalone indicators, and bias assessments must be anchored in application-specific testing.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). Its Word Embedding Association Test established an influential intrinsic-bias metric tradition that this study compares against later pretrained-language-model fairness measures.
- Paper: BERTScore: Evaluating Text Generation with BERT, Tianyi Zhang et al. (2019). This paper introduces BERTScore’s embedding-similarity approach, a central example of the embedding-based evaluation metrics whose fairness and comparability the source examines.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). StereoSet supplies a prominent intrinsic bias benchmark that helps contextualize the source’s comparison of bias measurements in pretrained language models.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Its critique of how NLP research defines and operationalizes bias clarifies the conceptual and methodological choices underlying the source’s survey of fairness metrics.
- Paper: From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models, Shangbin Feng et al. (2023). This study carries the source’s call to prioritize downstream fairness evaluations forward by tracing how pretraining biases propagate into consequential NLP tasks.
- Paper: BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation, Tianxiang Sun et al. (2022). It extends concerns about biased pretrained-model metrics to text-generation evaluation, showing how demographic bias can distort the scores used to assess systems.
