Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models

Yang Trista CaoAnna SotnikovaHal Daumé IIIRachel RudingerLinda Zou

article2022NAACL64 citations

Adapts a social psychology framework to systematically quantify and evaluate how English language models reproduce human stereotypes across single and intersectional social groups using a novel sensitivity test validated against human judgments.

Listen

Artificial intelligence language models trained on massive text corpora routinely reproduce and amplify human social stereotypes. When deployed in real-world systems like search engines or recruitment tools, these automated biases pose serious ethical, safety, and compliance risks by unfairly stereotyping marginalized communities. Prior efforts to detect these biases often relied either on rigid, ad hoc templates that fail to generalize to new groups or on crowdsourced natural text that lacks comprehensive theoretical coverage. The article addresses this challenge by establishing a standardized, theory-grounded methodology to evaluate how pre-trained language models capture societal stereotypes across diverse single and intersectional social identities.

The primary objective of the article is to adapt a proven social psychology framework to systematically measure group-trait associations in masked language models and assess how closely these model representations align with human stereotype judgments in the United States. To achieve this, the article evaluates two prominent transformer models, BERT and RoBERTa, across three stereotype dimensions: agency/socioeconomic success, conservative-progressive beliefs, and communion/warmth.

To conduct this evaluation, the study used 16 opposing trait pairs from the psychological Agency-Belief-Communion model and tested them against 25 broad social identity groups spanning gender, race, religion, age, and profession. Alongside existing evaluation methods, the article introduced a novel metric called the Sensitivity Test, which assesses the robustness of an association by calculating the minimal change required in a model's internal weights to make a specific trait its top prediction. The authors then conducted an IRB-approved crowdsourced study of 133 quality-vetted U.S. participants, asking them to evaluate how American society perceives these identity groups along the 16 trait dyads to serve as a baseline for model alignment.

The analysis yielded several key findings regarding how language models internalize stereotypes. First, language models show moderate alignment with human societal judgments; the RoBERTa model evaluated with the Sensitivity Test achieved the highest alignment, correctly matching human judgments on two out of three top-ranked group-trait associations (a precision of 65.3%). Second, RoBERTa consistently reflected human stereotypes more closely than BERT across all testing metrics. Third, when evaluating paired intersectional identities, the order of identity words had minimal impact on outputs, though the models tended to place slightly more weight on the second component word. Fourth, certain demographic categories strongly dominated others within paired identities: age and political stance exerted the strongest influence over model predictions, whereas race and nationality were predominantly overshadowed by other traits. Finally, while models could detect some compound stereotypes (such as associating male doctors with benevolence), they achieved only moderate success in capturing emergent intersectional stereotypes that do not exist in the isolated component identities.

These findings indicate that widely used language models inherently encode broad social stereotypes that mirror human societal biases, presenting significant risks if integrated into consumer-facing or decision-support software without mitigation. Because higher-capacity models like RoBERTa reflect human biases more strongly than earlier architectures, scaling up models will not naturally resolve bias issues. Organizations deploying language models should adopt theoretically grounded metrics like the Sensitivity Test to audit and benchmark stereotyping risks across intersecting identities before releasing downstream applications. However, leaders should note that the study evaluated abstract high-level traits within U.S. English cultural contexts, meaning further empirical work is required to determine exactly how these internal model associations translate into harmful behaviors in live operational environments.

Cover for Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models

Abstract

NLP models trained on text have been shown to reproduce human stereotypes, which can magnify harms to marginalized groups when systems are deployed at scale. We adapt the Agency-Belief-Communion (ABC) stereotype model of Koch et al. (2016) from social psychology as a framework for the systematic study and discovery of stereotypic group-trait associations in language models (LMs). We introduce the sensitivity test (SeT) for measuring stereotypical associations from language models. To evaluate SeT and other measures using the ABC model, we collect group-trait judgments from U.S.-based subjects to compare with English LM stereotypes. Finally, we extend this framework to measure LM stereotyping of intersectional identities.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 3 Measuring Stereotypes in LMs
  • 3.1 Measurements of Word Associations
  • 3.2 Implementation details
  • 4 Human Study
  • 5 Results
  • 5.1 Correlation on Individual Groups
  • 5.2 Intersectional Groups in LMs
  • 6 Limitations and Ethical Considerations
  • 7 Conclusion
  • Acknowledgments
  • References
  • A Traits
  • B Experiment Results with Single Groups
  • C Experiment Results of Intersectional Groups
  • D Human study setup
  • E Annotators demographics

Knowls

  1. Knowl 1 — SeT measures group–trait association by the model’s required parameter change

    model/method

    The Sensitivity Test (SeT) measures how much the output layer of a masked language model would need to change for a candidate trait to become its highest-probability vocabulary item. For group gg and trait token tt, let AA be the model’s final linear output matrix, and let hgh_g be the final hidden representation for a template containing group gg. Let hpriorh_{\mathrm{prior}} be the representation for the corresponding group-masked prior template. The logits are AhAh, and the minimum required change is

    Δ(A,h,t)=min⁡A′∥A′−A∥22subject to(A′h)t≥(A′h)t′+γ∀t′≠t,\Delta(A,h,t)=\min_{A'}\|A'-A\|_2^2\quad\text{subject to}\quad (A'h)_t\ge (A'h)_{t'}+\gamma\quad\forall t'\ne t,

    where A′A' is a candidate output matrix, t′t' ranges over other vocabulary items, and the fixed margin is γ=1\gamma=1. SeT is the prior-normalized log ratio log⁡ ⁣(Δ(A,hg,t)/Δ(A,hprior,t))\log\!\left(\Delta(A,h_g,t)/\Delta(A,h_{\mathrm{prior}},t)\right). Under this construction, a smaller required change in the group context relative to the prior indicates a stronger association. The optimization has no known closed-form solution in the paper and is solved iteratively with column squishing.

  2. Knowl 2 — The ABC framework turns stereotypes into a 16-dimensional group–trait profile

    definition

    The paper applies the Agency–Beliefs–Communion (ABC) stereotype framework to language-model group–trait associations. It represents a social group using 16 polar trait pairs: six in agency/socioeconomic success (powerful–powerless, high–low status, dominating–dominated, wealthy–poor, confident–unconfident, competitive–unassertive); four in beliefs (science-oriented–religious, alternative–conventional, liberal–conservative, modern–traditional); and six in communion (trustworthy–untrustworthy, sincere–dishonest, warm–cold, benevolent–threatening, likable–repellent, altruistic–egotistic). For each pair, the group score is the model’s score for one pole minus its score for the opposing pole; these pairwise differences form the group’s stereotype profile. The ABC traits are intended to provide broad, theory-grounded coverage of group stereotypes rather than fine-grained stereotype descriptions.

  3. Knowl 3 — The study compares three language-model association measurements

    model/method

    The paper evaluates SeT alongside increased log probability score (ILPS) and contextualized embedding association test (CEAT). For a group gg, trait token tt, and masked-word template, ILPS is the log probability of tt when the group is present divided by its probability in a prior template with the group masked: ILPS⁡(g,t)=log⁡ ⁣(P(t∣template(g))/P(t∣template([MASK])))\operatorname{ILPS}(g,t)=\log\!\left(P(t\mid\text{template}(g))/P(t\mid\text{template}([\mathrm{MASK}]))\right). The authors’ subword-aware variant, ILPS⋆^\star, computes a multi-subword trait probability by the chain rule. CEAT compares cosine similarity between contextualized group and trait embeddings: its per-group score is the difference between the mean cosine similarity to the two opposing trait sets, divided by the standard deviation of similarities across their union. CEAT embeddings are estimated from 1,000 randomly sampled Reddit sentences per word; the authors do not change CEAT to handle multi-subword words. For SeT, they score each subword separately and use the maximum subword score, which they characterize as a lower bound on the required parameter change.

  4. Knowl 4 — Experiments use ABC traits, curated U.S. groups, and varied templates

    experimental setup

    The language-model experiments use the pretrained large English masked language models BERT and RoBERTa. The authors manually compile social-group terms spanning gender/sexuality, race/ethnicity, religion, socioeconomic status, age, disability status, politics, and nationality. They score the 32 adjectives corresponding to the 16 ABC trait pairs. ILPS and SeT require templates, so the study varies grammatical and semantic forms, including singular versus plural, declarative versus interrogative, factual versus belief or social-expectation wording, and group-first versus trait-first forms. A pilot using Asian, Black, Hispanic, and immigrant groups selects at most two templates for each measurement–model pair by Kendall correlation against five human annotations per pilot group. RoBERTa’s best templates tend to include “That [group] is [trait],” while BERT’s tend to include “All [group] are [trait]” or “[Group] should be [trait].” The authors also test paired identities, omitting combinations judged logically impossible or grammatically awkward and using both component orders where possible.

  5. Knowl 5 — Human ratings provide comparison judgments for 25 groups

    experimental setup

    The human study asks U.S.-based participants recruited through Prolific how American society views each group, explicitly distinguishing that judgment from the participant’s personal belief. Participants rate each of the 16 opposing ABC trait pairs on a 0–100 slider, with the endpoints corresponding to the two poles. The study collects ratings for 25 social groups across five domains, with 20 quality-approved annotations per group. Each participant rates five groups. Three quality checks ask participants to recall the group and traits they just rated and to repeat a group rating; annotations are excluded for errors on either of the first two checks or less than 80% self-agreement on the repeated group. Of 247 participants, 133 passed the checks and their annotations were retained; all participants were paid. The retained participants represented 26 states, and more than 96% were aged 40 or younger.

  6. Knowl 6 — SeT with RoBERTa has the strongest overall rank alignment with human ratings

    empirical result

    Across the 25 human-rated groups and 16 trait pairs, the paper compares Kendall’s τ\tau and Precision at 3 (P@3) for each measurement–model combination. Kendall’s τ\tau assesses rank agreement between model and human group–trait scores. P@3 averages agreement on the model’s three highest-ranked traits at each polarity, with human judgments binarized at 50, and is then averaged over groups. In the order RoBERTa, BERT, the reported Kendall’s τ\tau values are: CEAT 0.019, 0.111; ILPS 0.169, 0.094; ILPS⋆^\star 0.175, 0.015; and SeT 0.199, 0.116. The corresponding P@3 values are: CEAT 0.500, 0.587; ILPS 0.620, 0.533; ILPS⋆^\star 0.653, 0.560; and SeT 0.653, 0.613. RoBERTa with SeT has the highest Kendall correlation, while RoBERTa with SeT and RoBERTa with ILPS⋆^\star tie for the highest P@3. The authors characterize the overall model–human alignment as moderate; the best reported Kendall correlation is 0.199.

  7. Knowl 7 — Intersectional analyses compare paired-group profiles with component identities

    model/method

    For intersectional experiments, the authors use RoBERTa with SeT, the best-aligned measurement–model pair in the individual-group comparison. They form paired groups from the curated single identities, retain both identity orders when possible, and compare each paired group’s vector of trait scores with those of its two component identities using Kendall’s τ\tau. This tests similarity to either component and sensitivity to identity order. To identify an emergent association for paired group g=(g1,g2)g=(g_1,g_2) and trait tt, they calculate S(g,t)−max⁡{S(g1,t),S(g2,t)}S(g,t)-\max\{S(g_1,t),S(g_2,t)\}, where SS is the language-model trait score; the reverse direction looks for paired-group scores below the minimum component score. The paper reserves “intersectional” for cases where the paired-group stereotype is more than the sum of component stereotypes.

  8. Knowl 8 — Paired-group scores show limited order effects and domain-specific dominance

    empirical result

    In the RoBERTa–SeT paired-identity analysis, the mean Kendall correlation between a paired group and its more similar component identity is 0.56. The mean correlation is 0.43 with the first component and 0.46 with the second; reversing identity order yields a mean correlation of 0.69 between the two orderings. The authors interpret these findings as showing that many paired-group profiles resemble one component, that order has little overall effect, and that the second component receives slightly more emphasis. To assess domain dominance, they compare average paired-group correlations with each component domain and count one domain as dominating another when the difference is at least 0.1. Age and political stance are reported as dominant domains; race/ethnicity and nationality are generally dominated domains. The authors note that the model’s tendency for race to be dominated contrasts with the human-behavior literature they discuss.

  9. Knowl 9 — Emergent-trait matches are promising but imperfect and include apparent artifacts

    empirical result

    Using the emergent-association score S(g,t)−max⁡{S(g1,t),S(g2,t)}S(g,t)-\max\{S(g_1,t),S(g_2,t)\}, the authors find examples such as Hispanic unemployed people scored as more egotistic (increase 0.0919), Democrat teenagers as more altruistic (0.0858), and male doctors as more benevolent (0.0819). Comparison against race/gender intersectional traits reported by Ghavami and Peplau yields precision 0.83 and recall 0.65, versus 0.72 precision and 0.50 recall for the paper’s random-guessing baseline. The authors therefore find some correspondence to human-study findings, but not reliable recovery of emergent stereotypes. They also report questionable patterns: many nationality-plus-mechanic pairs appear trustworthy or likable, while nationality-plus-autistic pairs appear egotistic. In these cases, the standalone mechanic or autistic score is low and the paired score rises toward an average level, so an apparent emergence need not represent a credible, distinctive intersectional stereotype.

  10. Knowl 10 — Interpretation is limited by defaulting, aggregation, and study scope

    limitation

    The authors caution that both human and model judgments may be affected by defaulting: an underspecified group such as men may be interpreted as cisgender, straight, and White, or language-model text may encode such defaults. The study measures stereotypes in English language models rather than deployed systems, so transfer to downstream system behavior is unclear; its scope is U.S. social stereotypes and English. The subjective human survey may still be affected by social-desirability concerns and by participants’ discomfort at speaking for how other people view groups, and collecting such judgments could itself reinforce stereotypes. Finally, averaging ratings into a single score suppresses differences among annotators and distributions: the same mean can represent widespread moderate judgments or sharply divided extreme judgments, and it cannot distinguish common weak stereotypes from rare strong ones.

Coverage note — The complete per-group rating tables, demographic subgroup comparisons, and full ranked list of 50 emergent associations are omitted because they are extensive appendix-level detail; representative findings and the central aggregate results are included.

References

  1. 1.Andrea Abele, Naomi Ellemers, Susan Fiske, Alex Koch, and Vincent Yzerbyt. 2020. Navigating the social world: Toward an integrated framework for evaluating self, individuals, and groups. Psychological review, 128.
  2. 2.Andrea Abele and Bogdan Wojciszke. 2014. Communal and agentic content a dual perspective model. Adv. Exp. Soc. Psychol., 50:198–255.
  3. 3.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  4. 4.Alina Beygelzimer, Sanjoy Dasgupta, and John Langford. 2008. Importance weighted active learning. CoRR, abs/0812.4952.
  5. 5.Victor Bittorf, Benjamin Recht, Christopher Ré, and Joel Tropp. 2012. Factoring nonnegative matrices with linear programs. Advances in Neural Information Processing Systems, 2.
  6. 6.Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1004–1015, Online. Association for Computational Linguistics.
  7. 7.Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. NeurIPS, pages 4349–4357.
  8. 8.Eduardo Bonilla-Silva. 1997. Rethinking racism: Toward a structural interpretation. American sociological review, pages 465–480.
  9. 9.Robert E. Botsch. 2011. Significance and Measures of Association.
  10. 10.Irene Browne and Joya Misra. 2003. The intersection of gender and race in the labor market. Annual review of sociology, 29(1):487–513.
  11. 11.J.S. Bruner, Brunswik E, L. Festinger, F. Heider, K.F. Muenzinger, C.E. Osgood, and D. Rapaport. 1957. Going beyond the information given. Contemporary approaches to cognition, pages 41–67.
  12. 12.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  13. 13.Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022. On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations.
  14. 14.Patricia Hill Collins. 2002. Black feminist thought: Knowledge, consciousness, and the politics of empowerment. routledge.
  15. 15.Combahee River Collective. 1977. A Black Feminist Statement. na.
  16. 16.Combahee River Collective. 1983. The combahee river collective statement. Home girls: A Black feminist anthology, 1:264–274.
  17. 17.Kimberlé Crenshaw. 1989. Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. u. Chi. Legal f., page 139.
  18. 18.Hal Daumé, III and Abhishek Kumar. 2017. Column squishing for multiclass updates (blog post).
  19. 19.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1).
  20. 20.Naomi Ellemers. 2017. Morality and the Regulation of Social Behavior: Groups as Moral Anchors.
  21. 21.Susan T. Fiske, Amy J. C. Cuddy, Peter Glick, and Jun Xu. 2002. A model of (often mixed) stereotype content: competence and warmth respectively follow from perceived status and competition. Journal of personality and social psychology, 82 6:878–902.
  22. 22.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé, and Kate Crawford. 2018. Datasheets for datasets.
  23. 23.Negin Ghavami and Letitia Anne Peplau. 2013. An intersectional analysis of gender and ethnic stereotypes: Testing three hypotheses. Psychology of Women Quarterly, 37(1):113–127.
  24. 24.Wei Guo and Aylin Caliskan. 2021. Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’21, page 122–133, New York, NY, USA. Association for Computing Machinery.
  25. 25.Timothy J. Hazen, Alexandra Olteanu, Gabriella Kazai, Fernando Diaz, and Michael Golebiewski. 2020. On the social and technical challenges of web search autosuggestion moderation. arXiv:2007.05039 [cs]. ArXiv: 2007.05039.
  26. 26.bell hooks. 1992. Yearning: Race, gender, and cultural politics. Hypatia, 7(2).
  27. 27.Dirk Hovy and Diyi Yang. 2021. The importance of modeling social factors of language: Theory and practice. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 588–602, Online. Association for Computational Linguistics.
  28. 28.Lynne M. Jackson. 2011. The psychology of prejudice: From attitudes to social action. American Psychological Association.
  29. 29.Deborah K King. 1988. Multiple jeopardy, multiple consciousness: The context of a black feminist ideology. Signs: Journal of women in culture and society, 14(1):42–72.
  30. 30.Alex Koch, Roland Imhoff, Ron Dotsch, Christian Unkelbach, and Hans Alves. 2016. The abc of stereotypes about groups: Agency/socioeconomic success, conservative-progressive beliefs, and communion. Journal of personality and social psychology, 110:675–709.
  31. 31.Alex Koch, Vincent Yzerbyt, Andrea Abele, Naomi Ellemers, and Susan T. Fiske. 2021. Social evaluation: Comparing models across interpersonal, intragroup, intergroup, several-group, and many-group contexts, volume 63, page 1–68. Elsevier.
  32. 32.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W. Black, and Yulia Tsvetkov. 2019. Measuring bias in contextualized word representations. CoRR, abs/1906.07337.
  33. 33.Carl A. Latkin, Catie Edwards, Melissa A. Davey-Rothwell, and Karin E. Tobin. 2017. The relationship between social desirability bias and self-reports of health, substance use, and social network factors among urban substance users in baltimore, maryland. Addictive Behaviors, 73:133–136.
  34. 34.Walter Lippmann. 1965. Public Opinion. New York :Free Press.
  35. 35.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  36. 36.Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On measuring social biases in sentence encoders. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 622–628, Minneapolis, Minnesota. Association for Computational Linguistics.
  37. 37.Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. Stereoset: Measuring stereotypical bias in pretrained language models. CoRR, abs/2004.09456.
  38. 38.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1953–1967, Online. Association for Computational Linguistics.
  39. 39.Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021. HONEST: Measuring hurtful sentence completion in language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2398–2406, Online. Association for Computational Linguistics.
  40. 40.Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. 2018. Comparing Bayesian Models of Annotation. Transactions of the Association for Computational Linguistics, 6:571–585.
  41. 41.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2019. Social bias frames: Reasoning about social and power implications of language. CoRR, abs/1911.03891.
  42. 42.Anna Sotnikova, Yang Trista Cao, Hal Daumé III, and Rachel Rudinger. 2021. Analyzing stereotypes in generative text inference tasks. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4052–4065, Online. Association for Computational Linguistics.
  43. 43.Dr. Charles Stangor. 2014. Principles of social psychology – 1st international edition. BCcampus.
  44. 44.Amy C Steinbugler, Julie E Press, and Janice Johnson Dias. 2006. Gender, race, and affirmative action: Operationalizing intersectionality in survey research. Gender & Society, 20(6):805–825.
  45. 45.Sheldon Stryker. 1980. Symbolic interactionism: a social structural version. Benjamin/Cummings Pub. Co.
  46. 46.S. Wheeler and Richard Petty. 2001. The effects of stereotype activation on behavior: A review of possible mechanisms. Psychological bulletin, 127:797–826.
  47. 47.I.J. Wod. 1985. Weight of evidence: A brief survey. Bayesian statistics, 2:249–270.
  48. 48.Vincent Y. Yzerbyt. 2018. The dimensional compensation model. Agency and Communion in Social Psychology.
  49. 49.Maxine Baca Zinn and Bonnie Thornton Dill. 1996. Theorizing difference from multiracial feminism. Feminist studies, 22(2):321–331.

Citation

MLA
Cao, Y. (Trista) ., et al. “Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 1276–95, https://doi.org/10.18653/v1/2022.naacl-main.92.
APA
Cao, Y. (Trista) ., Sotnikova, A., III, H. D., Rudinger, R., & Zou, L. (2022). Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1276–1295. https://doi.org/10.18653/v1/2022.naacl-main.92
Chicago
Cao, Y. (Trista) ., A. Sotnikova, H. D. III, R. Rudinger, and L. Zou. 2022. “Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1276–95. https://doi.org/10.18653/v1/2022.naacl-main.92.
Harvard
Cao, Y. (Trista) . et al. (2022) “Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 1276–1295. Available at: https://doi.org/10.18653/v1/2022.naacl-main.92.
Vancouver
1. Cao Y (Trista), Sotnikova A, III HD, Rudinger R, Zou L (2022) Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 1276–1295

BibTeX

@inproceedings{cao-etal-2022-theory,
    title = "Theory-Grounded Measurement of {U}.{S}. Social Stereotypes in {E}nglish Language Models",
    author = "Cao, Yang Trista  and
      Sotnikova, Anna  and
      Daum{\'e} III, Hal  and
      Rudinger, Rachel  and
      Zou, Linda",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.92/",
    doi = "10.18653/v1/2022.naacl-main.92",
    pages = "1276--1295"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/