MISGENDERED: Limits of Large Language Models in Understanding Pronouns

Tamanna HossainSunipa DevSameer Singh

article2023ACL66 citations

Presents the MISGENDERED evaluation benchmark to reveal that popular large language models systematically fail at correctly using declared gender-neutral and neo-pronouns due to training data imbalances and rigid name associations.

Listen

Language models are increasingly embedded into customer-facing applications, virtual assistants, and document processing systems. While fairness research has historically focused on binary gender categories (male and female), societal use of language has expanded to encompass non-binary identities, including singular gender-neutral pronouns (they/them) and neo-pronouns (such as xe/xem or ze/zir). Misgendering individuals in automated communications poses serious risks to brand trust, inclusivity, and user safety, yet the extent to which modern artificial intelligence models can follow explicit pronoun instructions remains underexplored.

The article establishes a standardized framework, MISGENDERED, to evaluate how accurately popular language models respect and apply declared third-person personal pronouns. Specifically, it tests whether models can correctly predict the appropriate pronoun form in a sentence immediately after being given an individual's explicit or parenthetical pronoun preference.

The evaluation used a comprehensive dataset of 3.8 million instances generated across 50 sentence templates covering five grammatical forms: nominative, accusative, possessive-dependent, possessive-independent, and reflexive. The researchers populated these templates with 11 pronoun sets—spanning binary, singular neutral, and eight distinct neo-pronoun groups—and combined them with 500 popular first names (categorized as female, male, or unisex). The study evaluated masked and autoregressive models across six families (BART, T5, GPT-2, GPT-J, OPT, and BLOOM), ranging in size from 60 million to over 7 billion parameters, using a unified constrained decoding approach.

The findings show that standard out-of-the-box language models fail decisively at correctly using non-binary pronouns. While models correctly predict binary pronouns with an average accuracy of 75.3%, performance drops sharply to 31.0% for gender-neutral singular they/them and collapses to just 7.6% for neo-pronouns. Furthermore, increasing model scale does not reliably resolve the issue; while some model families exhibit selective improvements, others show stagnant or degrading performance. The root cause lies in training data imbalances and memorized associations: binary pronouns appear hundreds of times more frequently than neo-pronouns in standard training corpora, where neo-pronouns rarely appear in genuine pronoun contexts. Models also exhibit strong memorization biases, predicting correct pronouns at much higher rates when the individual's name aligns with conventional gender stereotypes and performing worst on neo-pronouns when paired with strongly gendered names.

These results demonstrate that deployed language models are inherently unreliable at processing explicitly declared non-binary pronouns, exposing organizations to compliance and reputational risks through automated misgendering. Although providing few-shot examples within prompts improves neo-pronoun accuracy up to 45.4%, the gains plateau after approximately six examples and are inefficient to implement universally across downstream tasks.

Organizations deploying automated language systems should avoid assuming out-of-the-box neutrality and instead implement targeted post-processing checks or rule-based guardrails when handling declared pronouns. Further research and engineering efforts must focus on debiasing training corpora and establishing active misgendering detection mechanisms, ideally developed in collaboration with affected non-binary and transgender communities.

Confidence in these findings is high for standard English language models within the tested configurations. However, leaders should note several limitations: the analysis relies on template-based evaluations, focuses strictly on English names and Western gender constructs, evaluates base models rather than fully tuned downstream products, and does not evaluate compound pronoun preferences (such as she/they). Careful auditing is advised when applying these insights to multi-lingual environments or more complex pronoun combinations.

No sufficiently relevant recommendations were found.

Cover for MISGENDERED: Limits of Large Language Models in Understanding Pronouns

Abstract

Content Warning: This paper contains examples of misgendering and erasure that could be offensive and potentially triggering.

Gender bias in language technologies has been widely studied, but research has mostly been restricted to a binary paradigm of gender. It is essential also to consider non-binary gender identities, as excluding them can cause further harm to an already marginalized group. In this paper, we comprehensively evaluate popular language models for their ability to correctly use English gender-neutral pronouns (e.g., singular they, them) and neo-pronouns (e.g., ze, xe, thon) that are used by individuals whose gender identity is not represented by binary pronouns. We introduce MISGENDERED, a framework for evaluating large language models’ ability to correctly use preferred pronouns, consisting of (i) instances declaring an individual’s pronoun, followed by a sentence with a missing pronoun, and (ii) an experimental setup for evaluating masked and auto-regressive language models using a unified method. When prompted out-of-the-box, language models perform poorly at correctly predicting neo-pronouns (averaging 7.6% accuracy) and gender-neutral pronouns (averaging 31.0% accuracy). This inability to generalize results from a lack of representation of non-binary pronouns in training data and memorized associations. Few-shot adaptation with explicit examples in the prompt improves the performance but plateaus at only 45.4% for neo-pronouns. We release the full dataset, code, and demo at https://tamannahossainkay.github.io/misgendered/.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 MISGENDERED Framework
  • 3.1 Dataset Construction
  • 3.2 Evaluation Setup
  • 3.2.1 Constrained Decoding
  • 3.3 Experiments
  • 4 Results
  • 4.1 Direct Prompting
  • 4.2 Explaining Direct Prompting Results
  • 4.3 In-Context Learning
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgements
  • Limitations
  • Ethics Statement
  • References
  • A Templates
  • B Constrained Decoding Example
  • C Data and Code

Knowls

  1. Knowl 1 — Zero-shot models frequently misuse declared non-binary pronouns

    empirical result

    Across the paper’s zero-shot direct-prompting experiments, language models achieved average accuracy of 75.3% for binary pronouns, 31.0% for singular they, and 7.6% for neo-pronouns. These results measure whether a model filled a missing pronoun with the declared, correct pronoun, using the paper’s constrained candidate selection. Performance did not follow a consistent scaling rule across model families. For example, OPT’s accuracy for singular they rose from 21.2% for OPT-350M to 94.2% for OPT-6.7B, while BLOOM’s accuracy for neutral pronouns decreased slightly as model size grew.

  2. Knowl 2 — MISGENDERED dataset for evaluating declared-pronoun use

    model/method

    MISGENDERED evaluates whether a language model uses a person’s declared third-person pronouns when a sentence contains a pronoun blank. Each instance supplies a first name and pronoun declaration, followed by a manually written template with a missing pronoun. The dataset uses one pronoun group per instance, rather than combinations such as she/they, and contains 3.8 million instances.

    The five forms are nominative, accusative, possessive-dependent, possessive-independent, and reflexive. The 11 pronoun groups, listed in that form order, are: he/him/his/his/himself; she/her/her/hers/herself; they/them/their/theirs/themself; thon/thon/thons/thons/thonself; e/em/es/ems/emself; ae/aer/aer/aers/aerself; co/co/cos/cos/coself; vi/vir/vis/virs/virself; xe/xem/xyr/xyrs/xemself; ey/em/eir/eirs/emself; and ze/zir/zir/zirs/zirself.

    The dataset combines these pronouns with 300 sampled unisex U.S. names and the 100 most statistically female-associated and 100 most statistically male-associated names from U.S. Social Security data. It uses ten manually written templates for each pronoun form. Pronouns are declared either explicitly (for example, “Aamari’s pronouns are xe/xem”) or parenthetically after the person’s name. The resource is restricted to English and U.S. names; the pronoun list excludes several rare or less-used categories and is not intended to be exhaustive.

  3. Knowl 3 — Unified constrained decoding for masked and autoregressive models

    model/method

    MISGENDERED evaluates each instance by selecting the lowest-loss candidate among the 11 pronoun groups, restricted to the form required by the blank. Let xx be an evaluation sentence with a missing pronoun, ff its required pronoun form, PP the set of 11 pronoun groups, and pfp_f the form-ff pronoun belonging to candidate group pp. For each candidate, x(pf)x(p_f) is the model-specific input and y(pf)y(p_f) is its corresponding target. The prediction is

    p^=arg⁡min⁡p∈PL(x(pf),y(pf)),\hat{p}=\arg\min_{p\in P}\mathcal{L}\bigl(x(p_f),y(p_f)\bigr),

    where L\mathcal{L} is the evaluated model’s loss and p^\hat{p} is the predicted pronoun group. For autoregressive models, the candidate pronoun is inserted into the sentence in both input and target. For BART, the input uses its mask token and the target is the completed sentence. For T5, the missing span is represented by its sentinel token and the target contains that sentinel, the candidate pronoun, and the closing sentinel. This gives masked and autoregressive models a common candidate-ranking evaluation.

  4. Knowl 4 — Model families and prompting conditions evaluated

    experimental setup

    The direct-prompting evaluation covered autoregressive GPT-2 (124M, 355M, 774M, and 1.5B parameters), GPT-J (6.7B), BLOOM (560M, 1.1B, 3B, and 7.1B), and OPT (350M, 1.3B, 2.7B, and 6.7B), plus span-masked BART (140M and 400M) and T5 (60M, 220M, and 3B). The study varied declaration format (explicit or parenthetical) and declared pronoun forms (two through five forms). The combinations of forms were nominative and accusative; those two plus either possessive-independent or possessive-dependent; those two plus reflexive and either possessive form; and all five forms.

    In-context experiments used GPT-J-6B and OPT-6.7B with 2, 4, 6, 10, or 20 examples. They used explicit declarations of all five pronoun forms. Prompt examples were randomly sampled and excluded the templates, names, and pronouns of the evaluation instance.

  5. Knowl 5 — Accuracy varies across pronoun groups and grammatical forms

    empirical result

    In direct prompting, accuracy was 75.8% for she, 74.7% for he, and 31.0% for they. Among the tested neo-pronouns, accuracy was 18.5% for thon, 12.9% for xe, 9.2% for ey, 8.5% for ze, 6.2% for e, 2.2% for co, 2.0% for ae, and 1.1% for vi.

    Accuracy also depended on pronoun form. In the order nominative, accusative, reflexive, possessive-dependent, and possessive-independent, the results were: binary pronouns, 78.5%, 79.0%, 75.9%, 73.9%, and 60.0%; neutral pronouns, 18.1%, 27.2%, 11.4%, 40.1%, and 39.0%; neo-pronouns, 3.0%, 6.1%, 11.2%, 6.1%, and 12.2%. Thus, the best-performing form differed by category: nominative for binary pronouns, possessive-dependent for they, and possessive-independent for neo-pronouns.

  6. Knowl 6 — Declaration format and number of forms affect accuracy differently by pronoun type

    empirical result

    Direct-prompting accuracy differed between explicit and parenthetical pronoun declarations. For binary, neutral, and neo-pronouns respectively, explicit declarations yielded 68.8%, 24.1%, and 9.2%; parenthetical declarations yielded 81.6%, 37.8%, and 6.0%. Parenthetical declarations performed better for binary and neutral pronouns, while explicit declarations performed better for neo-pronouns.

    The number of declared forms also had category-dependent effects. For 2, 3, 4, and 5 declared forms, accuracy was respectively 74.9%, 75.0%, 75.4%, and 75.7% for binary pronouns; 32.5%, 31.8%, 30.2%, and 29.4% for neutral pronouns; and 4.6%, 6.6%, 9.2%, and 9.3% for neo-pronouns. More declared forms corresponded to higher neo-pronoun accuracy, a slight increase for binary pronouns, and lower accuracy for they.

  7. Knowl 7 — Pronoun predictions track name-gender associations

    empirical result

    Model accuracy varied with the gender association of the name, consistent with the paper’s interpretation that models may have memorized name–pronoun associations. For she, accuracy was 91.2% with female-associated names, 42.4% with male-associated names, and 81.7% with unisex names. For he, the corresponding accuracies were 32.5%, 91.8%, and 82.9%. For they, they were 23.7%, 24.9%, and 35.4%. The names with the strongest neo-pronoun performance were all classified as unisex, while the lowest-performing names were mostly gender-associated names. These results are based on the study’s U.S. name categories and do not establish the cause of the association.

  8. Knowl 8 — Pretraining-corpus counts show sparse representation of many neo-pronouns

    empirical result

    The authors counted corpus documents containing each pronoun token in C4, Open Web Text (OpenWT), and the Pile. The reported counts are shown below; the corpus measures document occurrences, not confirmed semantic use as a pronoun.

    • he: C4 552.7M; OpenWT 15.8M; Pile 161.9M. she: 348.0M; 5.5M; 68.0M. they: 769.3M; 13.5M; 180.4M.
    • thon: 2.1M; 5.5K; 83.4K. xe: 2.5M; 2.3K; 133.4K. ze: 1.8M; 3.3K; 177.2K.
    • co: 172.0M; 1.3M; 27.7M. e: 248.7M; 537.8K; 23.2M. ae: 5.4M; 7.9K; 412.2K. ey: 15.8M; 63.2K; 2.2M. vi: 12.9M; 45.2K; 2.2M.

    Many neo-pronoun tokens were much less represented than binary pronouns, although counts for short strings such as co and e can include non-pronoun uses. The authors also note that retrieved examples often did not use the token as a pronoun: for instance, co appeared in a Colorado address and e as a letter sequence; retrieved uses of they were often plural. The counts therefore support a possible data-scarcity explanation for poor generalization but do not demonstrate that scarcity causes the observed errors.

  9. Knowl 9 — Few-shot examples improve neo-pronoun accuracy up to six shots, then gains stall

    empirical result

    With explicit declarations of all five pronoun forms, in-context examples improved neo-pronoun accuracy for GPT-J-6B and OPT-6.7B through six shots, but the gain did not continue at larger shot counts. At 0, 2, 4, 6, 10, and 20 shots, GPT-J-6B accuracy was 6.7%, 30.4%, 39.7%, 45.4%, 24.8%, and 30.5%; OPT-6.7B accuracy was 11.9%, 31.7%, 33.7%, 38.8%, 23.9%, and 31.8%. The best neo-pronoun result was 45.4% for GPT-J-6B at six shots.

    For they, GPT-J-6B accuracy at those shot counts was 33.4%, 50.9%, 62.0%, 66.6%, 48.0%, and 51.1%. OPT-6.7B’s was 94.2%, 69.2%, 68.8%, 67.9%, 69.3%, and 68.6%, so adding examples reduced its accuracy relative to its zero-shot result. The experiments show that prompting can aid adaptation in some settings, not that it reliably mitigates bias in downstream use.

  10. Knowl 10 — Evaluation scope limits the conclusions about misgendering

    limitation

    MISGENDERED uses manually written templates, and measured accuracy may be sensitive to template choice; the authors caution that the results are not a definitive verdict on model misgendering. The study evaluates upstream language-model behavior rather than downstream applications. Its coverage is limited to English, a Western conception of gender, U.S. names and gender-assignment categories, and a selected, non-exhaustive set of pronouns. It does not account for name or gender changes or people who use multiple pronoun sets. Larger models were also omitted because of computational-budget constraints. Finally, the evaluation measures one specific form of misgendering; absence of errors under this setup would not demonstrate absence of misgendering or other gender-related harms.

Coverage note — The individual sentence templates and example excerpts from retrieved corpus documents are omitted because they are illustrative artifacts rather than additional findings; the dataset design, corpus-count evidence, results, and stated limitations are included.

References

  1. 1.Sarah Alnegheimish, Alicia Guo, and Yi Sun. 2022. Using natural sentence prompts for understanding biases in language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2824–2830, Seattle, United States. Association for Computational Linguistics.
  2. 2.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc.
  3. 3.Stephanie Brandl, Ruixiang Cui, and Anders Søgaard. 2022. How conservative are language models? adapting to the introduction of gender-neutral pronouns. arXiv preprint arXiv:2204.10281.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Yang Trista Cao and Hal Daumé III. 2020. Toward gender-inclusive coreference resolution. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4568–4595, Online. Association for Computational Linguistics.
  6. 6.Prafulla Kumar Choubey, Anna Currey, Prashant Mathur, and Georgiana Dinu. 2021. GFST: Gender-filtered self-training for more accurate gender in translation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1640–1654, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  7. 7.Sharyn Davies. 2007. Challenging gender norms: five genders among Bugis in Indonesia. Gale Cengage.
  8. 8.Pieter Delobelle, Ewoenam Tokpo, Toon Calders, and Bettina Berendt. 2022. Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1693–1706, Seattle, United States. Association for Computational Linguistics.
  9. 9.Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2021a. OSCaR: Orthogonal subspace correction and rectification of biases in word embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5034–5050, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  10. 10.Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. 2021b. Harms of gender exclusivity and challenges in non-binary representation in language technologies. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1968–1994, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  11. 11.EqualDex. 2022. Legal recognition of non-binary gender.
  12. 12.Andrew Flowers. 2015. The most common unisex names in america: Is yours one of them?
  13. 13.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  14. 14.Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. 2019. Openwebtext corpus.
  15. 15.Yue Guo, Yi Yang, and Ahmed Abbasi. 2022. Auto-debias: Debiasing masked language models with automated biased prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1012–1023, Dublin, Ireland. Association for Computational Linguistics.
  16. 16.Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models. Advances in neural information processing systems, 34:2611–2624.
  17. 17.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring bias in contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 166–172, Florence, Italy. Association for Computational Linguistics.
  18. 18.Anne Lauscher, Archie Crowley, and Dirk Hovy. 2022. Welcome to the modern world of pronouns: Identity-inclusive natural language processing beyond gender. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1221–1232, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  19. 19.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  20. 20.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  21. 21.Erin R Markman. 2011. Gender identity disorder, the gender binary, and transgender oppression: Implications for ethical social work. Smith College Studies in Social Work, 81(4):314–327.
  22. 22.NIH. What are sex & gender?
  23. 23.NIH. 2020. What are gender pronouns? why do they matter?
  24. 24.NIH. 2022. The importance of gender pronouns & their use in workplace communications.
  25. 25.National Library of Medicine NIH. 2021. Intersex.
  26. 26.Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. " i’m fully who i am": Towards centering transgender and non-binary voices to measure biases in open language generation. arXiv preprint arXiv:2305.09941.
  27. 27.Virginia Prince. 2005. Sex vs. gender. International Journal of Transgenderism, 8:29 – 32.
  28. 28.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  30. 30.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  31. 31.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  32. 32.Nikil Roashan Selvam, Sunipa Dev, Daniel Khashabi, Tushar Khot, and Kai-Wei Chang. 2022. The tail wagging the dog: Dataset construction biases of social bias benchmarks. arXiv preprint arXiv:2210.10040.
  33. 33.Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022. Quantifying social biases using templates is unreliable. arXiv preprint arXiv:2210.04337.
  34. 34.Social Security. 2022. Top names over the last 100 years.
  35. 35.Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019. Evaluating gender bias in machine translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1679–1684, Florence, Italy. Association for Computational Linguistics.
  36. 36.Them. 2021. How to affirm the people in your life who use multiple sets of pronouns.
  37. 37.U.S. Dept of State. 2022. X gender marker available on u.s. passports starting april 11. Press Statement.
  38. 38.Stanley R Vance Jr, Diane Ehrensaft, and Stephen M Rosenthal. 2014. Psychological and medical care of gender nonconforming youth. Pediatrics, 134(6):1184–1192.
  39. 39.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  40. 40.WHO. 2021. Gender and health.
  41. 41.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  42. 42.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, New Orleans, Louisiana. Association for Computational Linguistics.

Citation

MLA
Hossain, T., et al. “MISGENDERED: Limits of Large Language Models in Understanding Pronouns”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 5352–67, https://doi.org/10.18653/v1/2023.acl-long.293.
APA
Hossain, T., Dev, S., & Singh, S. (2023). MISGENDERED: Limits of Large Language Models in Understanding Pronouns. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5352–5367. https://doi.org/10.18653/v1/2023.acl-long.293
Chicago
Hossain, T., S. Dev, and S. Singh. 2023. “MISGENDERED: Limits of Large Language Models in Understanding Pronouns”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5352–67. https://doi.org/10.18653/v1/2023.acl-long.293.
Harvard
Hossain, T., Dev, S. and Singh, S. (2023) “MISGENDERED: Limits of Large Language Models in Understanding Pronouns”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5352–5367. Available at: https://doi.org/10.18653/v1/2023.acl-long.293.
Vancouver
1. Hossain T, Dev S, Singh S (2023) MISGENDERED: Limits of Large Language Models in Understanding Pronouns. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5352–5367

BibTeX

@inproceedings{hossain-etal-2023-misgendered,
    title = "{MISGENDERED}: Limits of Large Language Models in Understanding Pronouns",
    author = "Hossain, Tamanna  and
      Dev, Sunipa  and
      Singh, Sameer",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.293/",
    doi = "10.18653/v1/2023.acl-long.293",
    pages = "5352--5367"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/