Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models
Yang Trista CaoAnna SotnikovaHal Daumé IIIRachel RudingerLinda Zou
Adapts a social psychology framework to systematically quantify and evaluate how English language models reproduce human stereotypes across single and intersectional social groups using a novel sensitivity test validated against human judgments.
Artificial intelligence language models trained on massive text corpora routinely reproduce and amplify human social stereotypes. When deployed in real-world systems like search engines or recruitment tools, these automated biases pose serious ethical, safety, and compliance risks by unfairly stereotyping marginalized communities. Prior efforts to detect these biases often relied either on rigid, ad hoc templates that fail to generalize to new groups or on crowdsourced natural text that lacks comprehensive theoretical coverage. The article addresses this challenge by establishing a standardized, theory-grounded methodology to evaluate how pre-trained language models capture societal stereotypes across diverse single and intersectional social identities.
The primary objective of the article is to adapt a proven social psychology framework to systematically measure group-trait associations in masked language models and assess how closely these model representations align with human stereotype judgments in the United States. To achieve this, the article evaluates two prominent transformer models, BERT and RoBERTa, across three stereotype dimensions: agency/socioeconomic success, conservative-progressive beliefs, and communion/warmth.
To conduct this evaluation, the study used 16 opposing trait pairs from the psychological Agency-Belief-Communion model and tested them against 25 broad social identity groups spanning gender, race, religion, age, and profession. Alongside existing evaluation methods, the article introduced a novel metric called the Sensitivity Test, which assesses the robustness of an association by calculating the minimal change required in a model's internal weights to make a specific trait its top prediction. The authors then conducted an IRB-approved crowdsourced study of 133 quality-vetted U.S. participants, asking them to evaluate how American society perceives these identity groups along the 16 trait dyads to serve as a baseline for model alignment.
The analysis yielded several key findings regarding how language models internalize stereotypes. First, language models show moderate alignment with human societal judgments; the RoBERTa model evaluated with the Sensitivity Test achieved the highest alignment, correctly matching human judgments on two out of three top-ranked group-trait associations (a precision of 65.3%). Second, RoBERTa consistently reflected human stereotypes more closely than BERT across all testing metrics. Third, when evaluating paired intersectional identities, the order of identity words had minimal impact on outputs, though the models tended to place slightly more weight on the second component word. Fourth, certain demographic categories strongly dominated others within paired identities: age and political stance exerted the strongest influence over model predictions, whereas race and nationality were predominantly overshadowed by other traits. Finally, while models could detect some compound stereotypes (such as associating male doctors with benevolence), they achieved only moderate success in capturing emergent intersectional stereotypes that do not exist in the isolated component identities.
These findings indicate that widely used language models inherently encode broad social stereotypes that mirror human societal biases, presenting significant risks if integrated into consumer-facing or decision-support software without mitigation. Because higher-capacity models like RoBERTa reflect human biases more strongly than earlier architectures, scaling up models will not naturally resolve bias issues. Organizations deploying language models should adopt theoretically grounded metrics like the Sensitivity Test to audit and benchmark stereotyping risks across intersecting identities before releasing downstream applications. However, leaders should note that the study evaluated abstract high-level traits within U.S. English cultural contexts, meaning further empirical work is required to determine exactly how these internal model associations translate into harmful behaviors in live operational environments.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). StereoSet establishes an influential benchmark for measuring stereotypes in pretrained language models, providing essential context for the source’s comparison with existing evaluation methods.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). Caliskan and colleagues show how language representations reproduce human social associations, grounding the source’s investigation of stereotypes encoded in model representations.
- Paper: From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models, Shangbin Feng et al. (2023). This study carries the source’s political-bias dimension into downstream hate-speech and misinformation tasks, tracing how model leanings can affect consequential applications.
- Paper: "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset, Eric Michael Smith et al. (2022). HOLISTICBIAS extends stereotype auditing to a much broader identity vocabulary and conversational behaviors, building on the source’s attention to diverse and intersecting identities.
