Built independently by an author, for readers. Read the story and support ChapterPal

keyword

intrinsic bias

Intrinsic bias refers to stereotypical, discriminatory, or unfair associations embedded directly within a machine learning model internal representations, parameters, or pre-trained probability distributions. In natural language processing, it reflects societal prejudices absorbed from large training corpora and stored within word embeddings or baseline token predictions. Unlike extrinsic bias, which evaluates performance disparities and unfair outcomes in specific downstream tasks or user-facing applications, intrinsic bias is measured independently of downstream fine-tuning by evaluating geometric distances between vector representations or assessing model preferences across sensitive demographic attributes and stereotypical associations.

3 items

Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models

Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models

Pieter Delobelle, Ewoenam Kwaku Tokpo, Toon Calders, Bettina Berendt

OrganizationsKU LeuvenTechnische Universität BerlinUniversity of AntwerpWeizenbaum Institute

Why you should read this

Demonstrates through empirical correlation analysis that existing intrinsic fairness metrics for language models are incompatible, overly sensitive to template and seed choices, and fail to predict extrinsic bias in downstream applications.

An increasing awareness of biased patterns in natural language processing resources such as BERT has motivated many metrics to quantify ‘bias’ and ‘fairness’ in these resources. However, comparing the results of different metrics and the works that evaluate with such metrics remains difficult, if not outright impossible. We survey the literature on fairness metrics for pre-trained language models and experimentally evaluate compatibility, including both biases in language models and in their downstream tasks. We do this by combining traditional literature survey, correlation analysis and empirical evaluations. We find that many metrics are not compatible with each other and highly depend on (i) templates, (ii) attribute and target seeds and (iii) the choice of embeddings. We also see no tangible evidence of intrinsic bias relating to extrinsic bias. These results indicate that fairness or bias evaluation remains challenging for contextualized language models, among other reasons because these choices remain subjective. To improve future comparisons and fairness evaluations, we recommend to avoid embedding-based metrics and focus on fairness evaluations in downstream tasks.

Added

2026-10-02

An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models

An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models

Nicholas Meade, Elinor Poole-Dayan, Siva Reddy

OrganizationsCIFARMcGill University & Mila

Why you should read this

Evaluates five prominent debiasing techniques across multiple language models and benchmarks, revealing that apparent reductions in social bias frequently stem from degraded language modeling capabilities rather than true mitigation.

Recent work has shown pre-trained language models capture social biases from the large amounts of text they are trained on. This has attracted attention to developing techniques that mitigate such biases. In this work, we perform an empirical survey of five recently proposed bias mitigation techniques: Counterfactual Data Augmentation (CDA), Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebias. We quantify the effectiveness of each technique using three intrinsic bias benchmarks while also measuring the impact of these techniques on a model’s language modeling ability, as well as its performance on downstream NLU tasks. We experimentally find that: (1) Self-Debias is the strongest debiasing technique, obtaining improved scores on all bias benchmarks; (2) Current debiasing techniques perform less consistently when mitigating non-gender biases; And (3) improvements on bias benchmarks such as StereoSet and CrowS-Pairs by using debiasing strategies are often accompanied by a decrease in language modeling ability, making it difficult to determine whether the bias mitigation was effective.

Added

2026-10-01

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation

Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang

OrganizationsFudan University

Why you should read this

Reveals that popular language model-based evaluation metrics like BERTScore perpetuate substantial social biases across demographic attributes, and introduces lightweight debiasing adapters to mitigate these unfair preferences without sacrificing evaluation accuracy.

WARNING: This paper contains examples that are offensive in nature. Automatic evaluation metrics are crucial to the development of generative systems. In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020), have been commonly adopted in various generation tasks. However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics. To that end, this work presents the first systematic study on the social bias in PLM-based metrics. We demonstrate that popular PLM-based metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status. In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs. In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation.

Added

2026-09-26