keyword
social bias
Social bias is a systematic tendency to hold prejudiced attitudes, make generalized assumptions, or exhibit unfair preferences toward individuals or groups based on demographic characteristics such as race, gender, religion, age, physical appearance, sexual orientation, or socioeconomic status. Arising from cultural stereotypes, implicit associations, and historical inequalities, social bias leads people and societal institutions to evaluate individuals through the lens of preconceived group traits rather than individual merit. In technological and analytical contexts, social bias frequently manifests when automated systems and algorithms trained on human-generated data reproduce and amplify these societal prejudices, perpetuating harmful stereotypes and contributing to unfair outcomes for marginalized communities.
3 items

Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models
Yang Trista Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger, Linda Zou
Why you should read this
Adapts a social psychology framework to systematically quantify and evaluate how English language models reproduce human stereotypes across single and intersectional social groups using a novel sensitivity test validated against human judgments.
NLP models trained on text have been shown to reproduce human stereotypes, which can magnify harms to marginalized groups when systems are deployed at scale. We adapt the Agency-Belief-Communion (ABC) stereotype model of Koch et al. (2016) from social psychology as a framework for the systematic study and discovery of stereotypic group-trait associations in language models (LMs). We introduce the sensitivity test (SeT) for measuring stereotypical associations from language models. To evaluate SeT and other measures using the ABC model, we collect group-trait judgments from U.S.-based subjects to compare with English LM stereotypes. Finally, we extend this framework to measure LM stereotyping of intersectional identities.
Added
2026-10-03

French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English
Aurélie Névéol, Yoann Dupont, Julien Bezançon, Karën Fort
Why you should read this
Presents a culturally adapted French extension of the CrowS-Pairs dataset and revised English pairs to effectively measure and compare social biases across languages in masked language models.
Added
2026-10-03

BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation
Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang
Why you should read this
Reveals that popular language model-based evaluation metrics like BERTScore perpetuate substantial social biases across demographic attributes, and introduces lightweight debiasing adapters to mitigate these unfair preferences without sacrificing evaluation accuracy.
WARNING: This paper contains examples that are offensive in nature. Automatic evaluation metrics are crucial to the development of generative systems. In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020), have been commonly adopted in various generation tasks. However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics. To that end, this work presents the first systematic study on the social bias in PLM-based metrics. We demonstrate that popular PLM-based metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status. In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs. In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation.
Added
2026-09-26
