keyword
stereotypical bias
Stereotypical bias is a form of social and cognitive bias in which overgeneralized, simplified, and often unfair assumptions about specific demographic groups are systematically applied to individuals based on characteristics such as race, gender, religion, age, profession, or socioeconomic status. In artificial intelligence and natural language processing, this bias manifests when computational models learn societal prejudices embedded within human-generated training data and subsequently reproduce or amplify those associations in downstream tasks such as text generation, classification, and evaluation. Consequently, systems affected by stereotypical bias can generate disproportionately inaccurate, disparaging, or unequal representations and outcomes across different social groups.
3 items

"Thinking" Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models
Shaz Furniturewala, Surgan Jandial, Abhinav Java, Pragyan Banerjee, Simra Shahid, Sumit Bhatia, Kokil Jaidka
Why you should read this
Presents a System 2-inspired prompting framework that enables end users to significantly reduce social biases in closed-source and open-source language models without retraining weights or accessing internal probability distributions.
Existing debiasing techniques are typically training-based or require access to the model’s internals and output distributions, so they are inaccessible to end-users looking to adapt LLM outputs for their particular needs. In this study, we examine whether structured prompting techniques can offer opportunities for fair text generation. We evaluate a comprehensive end-user-focused iterative framework of debiasing that applies System 2 thinking processes for prompts to induce logical, reflective, and critical text generation, with single, multi-step, instruction, and role-based variants. By systematically evaluating many LLMs across many datasets and different prompting strategies, we show that the more complex System 2-based Implicative Prompts significantly improve over other techniques demonstrating lower mean bias in the outputs with competitive performance on the downstream tasks. Our work offers research directions for the design and the potential of end-user-focused evaluative frameworks for LLM use.
Added
2026-10-04

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation
Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang
Why you should read this
Reveals that popular language model-based evaluation metrics like BERTScore perpetuate substantial social biases across demographic attributes, and introduces lightweight debiasing adapters to mitigate these unfair preferences without sacrificing evaluation accuracy.
WARNING: This paper contains examples that are offensive in nature. Automatic evaluation metrics are crucial to the development of generative systems. In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020), have been commonly adopted in various generation tasks. However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics. To that end, this work presents the first systematic study on the social bias in PLM-based metrics. We demonstrate that popular PLM-based metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status. In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs. In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation.
Added
2026-09-26

StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, Siva Reddy
Why you should read this
Introduces StereoSet, a large-scale benchmark to quantify and track stereotypical biases across gender, profession, race, and religion in popular pretrained language models.
A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real world data, they are known to capture stereotypical biases. In order to assess the adverse effects of these models, it is important to quantify the bias captured in them. Existing literature on quantifying bias evaluates pretrained language models on a small set of artificially constructed bias-assessing sentences. We present StereoSet, a large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion. We evaluate popular models like BERT, GPT-2, RoBERTa, and XLNet on our dataset and show that these models exhibit strong stereotypical biases. We also present a leaderboard with a hidden test set to track the bias of future language models at this https URL
Added
2026-09-25
