MISGENDERED: Limits of Large Language Models in Understanding Pronouns
Tamanna HossainSunipa DevSameer Singh
Presents the MISGENDERED evaluation benchmark to reveal that popular large language models systematically fail at correctly using declared gender-neutral and neo-pronouns due to training data imbalances and rigid name associations.
Language models are increasingly embedded into customer-facing applications, virtual assistants, and document processing systems. While fairness research has historically focused on binary gender categories (male and female), societal use of language has expanded to encompass non-binary identities, including singular gender-neutral pronouns (they/them) and neo-pronouns (such as xe/xem or ze/zir). Misgendering individuals in automated communications poses serious risks to brand trust, inclusivity, and user safety, yet the extent to which modern artificial intelligence models can follow explicit pronoun instructions remains underexplored.
The article establishes a standardized framework, MISGENDERED, to evaluate how accurately popular language models respect and apply declared third-person personal pronouns. Specifically, it tests whether models can correctly predict the appropriate pronoun form in a sentence immediately after being given an individual's explicit or parenthetical pronoun preference.
The evaluation used a comprehensive dataset of 3.8 million instances generated across 50 sentence templates covering five grammatical forms: nominative, accusative, possessive-dependent, possessive-independent, and reflexive. The researchers populated these templates with 11 pronoun sets—spanning binary, singular neutral, and eight distinct neo-pronoun groups—and combined them with 500 popular first names (categorized as female, male, or unisex). The study evaluated masked and autoregressive models across six families (BART, T5, GPT-2, GPT-J, OPT, and BLOOM), ranging in size from 60 million to over 7 billion parameters, using a unified constrained decoding approach.
The findings show that standard out-of-the-box language models fail decisively at correctly using non-binary pronouns. While models correctly predict binary pronouns with an average accuracy of 75.3%, performance drops sharply to 31.0% for gender-neutral singular they/them and collapses to just 7.6% for neo-pronouns. Furthermore, increasing model scale does not reliably resolve the issue; while some model families exhibit selective improvements, others show stagnant or degrading performance. The root cause lies in training data imbalances and memorized associations: binary pronouns appear hundreds of times more frequently than neo-pronouns in standard training corpora, where neo-pronouns rarely appear in genuine pronoun contexts. Models also exhibit strong memorization biases, predicting correct pronouns at much higher rates when the individual's name aligns with conventional gender stereotypes and performing worst on neo-pronouns when paired with strongly gendered names.
These results demonstrate that deployed language models are inherently unreliable at processing explicitly declared non-binary pronouns, exposing organizations to compliance and reputational risks through automated misgendering. Although providing few-shot examples within prompts improves neo-pronoun accuracy up to 45.4%, the gains plateau after approximately six examples and are inefficient to implement universally across downstream tasks.
Organizations deploying automated language systems should avoid assuming out-of-the-box neutrality and instead implement targeted post-processing checks or rule-based guardrails when handling declared pronouns. Further research and engineering efforts must focus on debiasing training corpora and establishing active misgendering detection mechanisms, ideally developed in collaboration with affected non-binary and transgender communities.
Confidence in these findings is high for standard English language models within the tested configurations. However, leaders should note several limitations: the analysis relies on template-based evaluations, focuses strictly on English names and Western gender constructs, evaluates base models rather than fully tuned downstream products, and does not evaluate compound pronoun preferences (such as she/they). Careful auditing is advised when applying these insights to multi-lingual environments or more complex pronoun combinations.
- Paper: Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models, Pieter Delobelle et al. (2022). Its analysis of bias metrics—including pronoun coreference—provides useful context for why MISGENDERED uses a controlled, task-level evaluation rather than relying on a single generic bias score.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). This foundational demonstration that language representations absorb social associations from training text clarifies the data-driven bias mechanism MISGENDERED investigates.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Its account of representational harm and the limits of narrow bias measures frames the stakes and evaluation choices behind studying automated misgendering.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). Its benchmark for measuring stereotypical associations in pretrained models offers a useful foundation for understanding MISGENDERED’s controlled language-model evaluation.
No sufficiently relevant recommendations were found.
