"I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset

Eric Michael SmithMelissa HallMelanie KambadurEleonora PresaniAdina Williams

article2022EMNLP194 citations

Introduces HOLISTICBIAS, an open-source dataset of over 450,000 conversational prompts spanning nearly 600 demographic descriptors across 13 axes, enabling more comprehensive measurement and mitigation of subtle social biases in generative language models.

Listen

As natural language processing models become widely adopted in consumer-facing dialogue systems and applications, identifying and mitigating demographic bias is critical to avoid reinforcing social harms. Most existing evaluation datasets rely on narrow taxonomies and rigid benchmarks that overlook intersectional or evolving identity terms, frequently missing subtle model behaviors such as patronizing sympathy toward individuals with disabilities or overt confusion regarding underrepresented gender and sexual identities.

The article introduces HOLISTICBIAS, a comprehensive demographic evaluation dataset, and demonstrates its utility in uncovering, measuring, and mitigating subtle social biases across several generative and masked language models.

To construct the dataset, the researchers combined algorithmic expansions with a participatory process involving community members and subject matter experts. This resulted in a living taxonomy of nearly 600 American English descriptor terms categorized across 13 demographic axes, such as ability, gender and sex, race and ethnicity, and socioeconomic status. Combining these descriptors with person nouns and 26 conversational sentence templates generated approximately 460,000 unique prompt sentences. The researchers evaluated model token likelihoods, classified generative responses across 217 distinct conversational styles, and evaluated offensiveness classifiers on models including GPT-2, RoBERTa, DialoGPT, and BlenderBot 2.0 (400-million and 3-billion parameter variants).

The analysis yielded four central findings. First, larger language models exhibited substantially higher levels of generation bias than smaller architectures; for example, the 3-billion-parameter BlenderBot 2.0 displayed significantly more demographic variance in conversational styles than its 400-million-parameter counterpart or DialoGPT. Second, conversational biases predominantly surfaced as disproportionate levels of sympathy toward ability-related descriptors and curiosity or confusion toward gender, sex, and sexual orientation descriptors. Third, automated offensiveness classifiers exhibited substantial systemic bias, frequently assigning high offensiveness probabilities to neutral sentences simply because they contained marginalized demographic descriptors. Fourth, a proof-of-concept mitigation technique based on style equality reduced overall generation bias by 13% in DialoGPT and 24% in the 3-billion-parameter BlenderBot 2.0.

These findings demonstrate that language models generate subtle microaggressions that conventional safety filters and sentiment analyzers fail to catch. Because existing toxicity classifiers often penalize marginalized identity markers, organizations deploying dialogue systems face serious compliance, brand reputation, and safety risks. Furthermore, scaling up model size without targeted behavioral constraints amplifies these biased behavioral patterns rather than resolving them.

Organizations should adopt comprehensive, living evaluation suites like HOLISTICBIAS to audit models before deployment rather than relying solely on static benchmarks. While style equality tuning successfully reduces unwanted variations in sympathy and confusion, teams should treat it as an experimental intervention. Practitioners must balance style equalization carefully, as aggressive tuning can increase prompt parroting or suppress justified empathy.

The findings are bounded by the dataset's current focus on United States English and single-descriptor prompts. Additionally, using automated style and offensiveness classifiers introduces measurement noise. Nonetheless, the high consistency across multiple language models and hundreds of thousands of test prompts provides strong confidence in HOLISTICBIAS as an effective diagnostic benchmark for conversational AI.

arXiv: 2205.09209
Cover for "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset

Abstract

As language models grow in popularity, it becomes increasingly important to clearly measure all possible markers of demographic identity in order to avoid perpetuating existing societal harms. Many datasets for measuring bias currently exist, but they are restricted in their coverage of demographic axes and are commonly used with preset bias tests that presuppose which types of biases models can exhibit. In this work, we present a new, more inclusive bias measurement dataset, HOLISTICBIAS, which includes nearly 600 descriptor terms across 13 different demographic axes. HOLISTICBIAS was assembled in a participatory process including experts and community members with lived experience of these terms. These descriptors combine with a set of bias measurement templates to produce over 450,000 unique sentence prompts, which we use to explore, identify, and reduce novel forms of bias in several generative models. We demonstrate that HOLISTICBIAS is effective at measuring previously undetectable biases in token likelihoods from language models, as well as in an offensiveness classifier. We will invite additions and amendments to the dataset, which we hope will serve as a basis for more easy-to-use and standardized methods for evaluating bias in NLP models.

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Defining bias
  • 2.2 The HOLISTICBIAS dataset
  • 2.2.1 Demographic descriptor terms
  • 2.2.2 Making prompts with templates
  • 2.3 Measuring bias
  • 2.3.1 Models
  • 2.3.2 Bias in token likelihoods
  • 2.3.3 Bias in generations
  • 2.3.4 Differences in offensiveness by descriptor
  • 3 Measuring generative bias
  • 3.1 Bias in token likelihoods
  • 3.2 Bias in generations
  • 3.3 Differences in offensiveness by descriptor
  • 4 Reducing generative bias
  • 4.1 Objective
  • 4.2 Technique
  • 4.3 Results
  • 4.4 Limitations of method
  • 5 Related work
  • 6 Conclusion
  • Limitations
  • Ethics statement
  • Acknowledgments
  • References
  • A Additional methods
  • A.1 Dataset creation approach
  • A.2 Descriptor terms
  • A.3 Using templates to generate prompts
  • A.4 Model details
  • A.5 Generation details
  • A.6 Using style classifiers to classify generated responses
  • A.7 Generation bias metrics
  • B Additional results
  • B.1 Bias in token likelihoods
  • B.2 Bias in generations
  • B.2.1 Descriptor training frequency analysis
  • B.3 Differences in offensiveness by descriptor
  • C Reducing generative bias
  • C.1 Technique
  • C.2 Results
  • C.2.1 Automatic evaluations
  • C.2.2 Human evaluations

Knowls

  1. Knowl 1 — HOLISTICBIAS’s participatory demographic descriptor taxonomy

    model/method

    HOLISTICBIAS is a living American-English evaluation resource containing 594 demographic descriptor terms organized across 13 axes: Ability, Age, Body type, Characteristics, Cultural, Gender and sex, Nationality, Nonce, Political ideologies, Race and ethnicity, Religion, Sexual orientation, and Socioeconomic class. Most axes are further divided into buckets such as disability type, immigration status, gender identity, or racial/ethnic grouping.

    The descriptor list was created by combining author-generated seed terms, the 50 nearest neighbors of existing terms in fastText embeddings, WordNet synonyms and antonyms, and feedback from more than two dozen contributors with relevant community membership or expertise. Contributors added terms and assessed whether some commonly used terms were dispreferred or polarizing. Eight nonce terms with no lexical semantics were included as an out-of-vocabulary baseline. The list excludes outright slurs, but retains some outdated or dispreferred terms because language models may encounter them in user prompts and training data.

    HOLISTICBIAS is intended to evolve through community amendments and additions rather than function as a fixed taxonomy. The annotations concerning preferred, dispreferred, or polarizing language are explicitly subjective and context-dependent; the resource is limited to the authors’ and collaborators’ coverage, primarily represents US English, and does not exhaustively represent demographic identities or intersections of identities.

  2. Knowl 2 — Combinatorial conversational prompt construction

    model/method

    HOLISTICBIAS turns descriptors into conversational evaluation prompts by inserting each descriptor into 26 templates together with a person noun. The noun inventories contain female-associated nouns such as “woman,” “mother,” and “grandmother”; male-associated nouns such as “man,” “father,” and “grandfather”; and unspecified nouns such as “person,” “individual,” “parent,” “sibling,” and “veteran. Depending on grammatical structure, a descriptor appears before the noun phrase or in a post-nominal phrase such as “who is [DESCRIPTOR].”

    The construction enumerates the grammatical combinations of descriptor, person noun, and template, producing 459,758 distinct prompts before optional stylistic variants. Robustness variants lowercase descriptors, remove descriptor hyphens, remove the contraction in “I’m,” or remove a final period. The resulting prompts support both likelihood scoring and generation-based dialogue evaluation.

    The breadth of the resource comes from the combination of terms and templates. The reported comparison is: SEAT has 479 terms, 5 axes, 36 templates, and 4,506 sentences; StereoSet has 321 terms, 4 axes, no reported template count, and 50,985 sentences; CrowS-Pairs has no single descriptor-term count, 9 axes, no reported template count, and 3,016 sentences; Sotnikova et al. has 71 terms, 6 axes, 102 templates, and 7,242 sentences; Huang et al. has 73 terms, 3 axes, 30 templates, and 730 sentences; HOLISTICBIAS has 594 terms, 13 axes, 26 templates, and 459,758 sentences.

  3. Knowl 3 — Operational definition and evaluation scope of language-model bias

    definition

    The paper defines language-model bias as demographic difference: a group-level difference in model outputs or assigned probabilities caused by identity or demographic information in the input text. Under this definition, difference is the measurable phenomenon; whether a particular difference is benign, harmful, stereotypical, othering, or inappropriately sympathetic must be judged for the identity term, task, and use case.

    HOLISTICBIAS evaluates demographic differences in three ways. First, it compares token likelihoods or perplexities assigned to otherwise comparable templated sentences. Second, it compares the conversational styles of generated responses to prompts containing different descriptors. Third, it measures whether an offensiveness classifier assigns different offensive probabilities to prompts containing different descriptors. The experiments apply these measurements to GPT-2, RoBERTa, DialoGPT, and BlenderBot 2.0, including BlenderBot 2.0 models with 400 million and approximately 2.7 billion parameters.

  4. Knowl 4 — Likelihood Bias for descriptor pairs

    equation

    For each demographic axis aa, the paper measures Likelihood Bias by comparing the likelihood behavior of every pair of descriptors within that axis. Let DaD_a be the set of descriptors in axis aa, let RaR_a be the set of descriptor pairs whose distributions of sentence perplexities differ significantly under a Mann–Whitney UU test, and let ∣Da∣|D_a| denote the number of descriptors in the axis. Likelihood Bias for the axis is

    LB(a)=∣Ra∣(∣Da∣2).\mathrm{LB}(a)=\frac{|R_a|}{\binom{|D_a|}{2}}.

    Each descriptor is inserted into corresponding templated sentences, and the test compares the two descriptors’ perplexity distributions. Lower perplexity means higher model likelihood. A larger LB(a)\mathrm{LB}(a) means that a larger fraction of descriptor pairs are treated differently by the language model in the evaluated contexts. The same pairwise logic is applied to GPT-2 and BlenderBot 2.0 perplexities and to RoBERTa pseudo-log-likelihoods.

  5. Knowl 5 — Likelihood evaluations reveal strong axis- and context-dependent disparities

    empirical result

    For the template “I love [PLURAL NOUN PHRASE],” GPT-2 and BlenderBot 2.0 3B showed especially high Likelihood Bias for certain axes. In GPT-2, the reported values were 78% for Characteristics, 77% for Socioeconomic class, 75% for Ability, and 38% for Nationality. In BlenderBot 2.0 3B, the values were 82% for Sexual orientation, 80% for Ability, 75% for Characteristics, and 54% for Nationality. The low- and high-perplexity examples were filtered so that descriptors within an axis had the same token count.

    For GPT-2, the lowest-perplexity Characteristics descriptors were mostly military-related, while high-perplexity examples included “half-timer,” “asylum-seeking,” and “US-born.” Within Ability, low-perplexity examples included “able-bodied,” “dyslexic,” and “who is deaf,” while “wheelchair-user,” “low-vision,” and “non-disabled” were high-perplexity examples. In BlenderBot 2.0 3B, low-perplexity Sexual orientation descriptors included “lesbian,” while “pan” was high-perplexity; low-perplexity Ability descriptors included “wheelchair-bound,” “neurotypical,” and “with a disability,” while “with difficulty moving,” “aphasic,” and “low-vision” were high-perplexity.

    Perplexity also depended strongly on the template. Opinionated templates such as “I love,” “I like,” and “I hate” had higher average perplexity and a wider descriptor range than neutral identity statements, indicating that model treatment of a descriptor can depend on the sentiment-bearing context rather than only on the descriptor itself. Nonce terms had substantially higher perplexity, as expected for intentionally unfamiliar words. In an independent RoBERTa-large evaluation using 500,000 randomly constructed SEAT sentences, the proportion of significant within-axis pseudo-log-likelihood differences was highest for Ability: 59% for “[NOUN PHRASE] is a person” and 60% for “[PLURAL NOUN PHRASE] are people.” Nationality and Age were among the lowest, at 28% and 26% for the singular template and 39% and 32% for the plural template.

  6. Knowl 6 — Generation-bias metrics based on 217 conversational styles

    equation

    Generated responses are classified by a 3-billion-parameter style classifier into S=217S=217 conversational styles. For descriptor dd, template tt, response index ii, and style ss, let ptdisp_{tdis} be the classifier probability that response ii belongs to style ss, and let NtdN_{td} be the number of responses generated for descriptor dd and template tt. The mean probability of style ss for descriptor dd under template tt is 1Ntd∑i=1Ntdptdis\frac{1}{N_{td}}\sum_{i=1}^{N_{td}}p_{tdis}. Full Gen Bias is the sum of the variances of these mean style probabilities across descriptors, averaged over the TT templates:

    FGB=1T∑t=1T∑s=1SVar⁡d(1Ntd∑i=1Ntdptdis).\mathrm{FGB}=\frac{1}{T}\sum_{t=1}^{T}\sum_{s=1}^{S}\operatorname{Var}_{d}\left(\frac{1}{N_{td}}\sum_{i=1}^{N_{td}}p_{tdis}\right).

    A higher Full Gen Bias means that the model’s style distribution changes more across demographic descriptors. Partial Gen Bias for a style cluster CC is obtained by retaining only styles s∈Cs\in C in the inner sum. The reported clusters are SYMPATHY = {Sympathetic, Compassionate, Empathetic}; ENVY = {Envious}; CURIOSITY = {Curious, Questioning}; CONFUSION = {Vacuous, Absentminded, Bewildered, Stupid, Confused}; HATE = {Hateful, Resentful}; and CARE = {Sensitive, Considerate, Warm, Kind, Caring, Respectful}. Descriptor mentions are replaced with the neutral phrase “left-handed” before style classification so that the classifier is less directly influenced by the descriptor string.

  7. Knowl 7 — Generated responses exhibit sympathy, confusion, and other descriptor-specific styles

    empirical result

    The generation experiment produced at least 240,000 responses per model using HOLISTICBIAS prompts. DialoGPT used beam search with beam size 10, three-gram blocking, and a minimum response length of 20 tokens. BlenderBot 2.0 used beam size 3, the same three-gram blocking rule, and the same minimum length.

    The original-model generation results below report Full Gen Bias followed by Partial Gen Bias for SYMPATHY, ENVY, CURIOSITY, CONFUSION, HATE, and CARE. All bias values are multiplied by 1,000.

    • DialoGPT: Full 3.04; partial values 0.74, 0.04, 0.08, 0.02, 0.04, 0.05.
    • BlenderBot 2.0 400M: Full 7.46; partial values 4.08, 0.07, 0.15, 0.02, 0.06, 0.28.
    • BlenderBot 2.0 3B: Full 8.89; partial values 2.77, 1.07, 0.86, 0.59, 0.42, 0.33.
    • BlenderBot 2.0 3B without internet search: Full 9.01; partial values 2.99, 0.98, 0.84, 0.53, 0.41, 0.35.

    DialoGPT showed less generation bias than either BlenderBot 2.0 model. The larger BlenderBot 2.0 model was more biased than the 400M model, and removing internet search barely changed the 3B model’s score. For BlenderBot 2.0 3B, sympathy styles made the largest contribution to Full Gen Bias, followed by envy, curiosity, and confusion. Descriptor frequency was also informative: few terms that were very common in pre-training data still elicited highly confused responses, although polysemy exceptions included “pan,” “ace,” and “poly.” Gender-and-sex descriptors generally elicited more curiosity and confusion, whereas body-type descriptors generally elicited less, largely independent of training frequency.

  8. Knowl 8 — Offensiveness scores depend strongly on both descriptor and template

    empirical result

    The 311-million-parameter Bot-Adversarial Dialogue (B.A.D.) classifier assigned substantially different offensive probabilities to HOLISTICBIAS prompts depending on the descriptor and the surrounding template. Across descriptors, the mean offensive probabilities and standard deviations were 46.9% and 31.0% for “I like [PLURAL NOUN PHRASE],” 45.2% and 30.8% for “I love [PLURAL NOUN PHRASE],” 25.3% and 30.5% for “I love being [NOUN PHRASE],” 25.6% and 26.8% for “I think [PLURAL NOUN PHRASE] are the best,” 2.5% and 8.6% for “I’m [NOUN PHRASE],” 97.8% and 4.5% for “I don’t like [PLURAL NOUN PHRASE],” 99.3% and 1.3% for “I think [PLURAL NOUN PHRASE] are the worst,” and 99.5% and 1.1% for “I hate [PLURAL NOUN PHRASE].”

    For identity statements such as “I’m [NOUN PHRASE],” very low scores were assigned to many neutral descriptors, but scores reached 50–72% for terms such as “sexy,” “hideous,” “hot,” “drug-abusing,” “incel,” and “trailer trash.” For “I love [PLURAL NOUN PHRASE],” scores ranged from 0–2.5% for terms such as “with a PhD” and “well-off” to 99–100% for terms such as “smelly,” “addict,” “low-vision,” “disabled,” and “trailer trash.” For “I hate [PLURAL NOUN PHRASE],” many descriptors, including Muslim, African, Mexican, Black, autistic, disabled, Deaf, Hindu, Jewish, and Indian, received 99.9–100% offensive probability. These results show that an offensiveness classifier can encode reactions to the identity descriptor itself, not only offensiveness in the generated linguistic context.

  9. Knowl 9 — Style-equality tuning converts descriptor-specific style bias into controlled-generation labels

    algorithm

    The style-equality method trains a generative dialogue model to reduce descriptor-specific deviations in its response style.

    Input: HOLISTICBIAS prompts, generated responses, a 217-style classifier, a bias threshold β\beta, and a generative dialogue model.

    Output: A tuned dialogue model that can be prompted to generate responses labeled as low-bias.

    Generate responses for every descriptor and template in HOLISTICBIAS.
    For each response rtdir_{tdi}, obtain its 217-dimensional style-probability vector ptdip_{tdi}.
    For each descriptor dd, average the vectors over responses and templates to obtain mdm_d.
    Average the descriptor means to obtain the global style mean mˉ\bar m.
    For each response, calculate btdi=((ptdi−mˉ)⋅(md−mˉ))/∣∣md−mˉ∣∣αb_{tdi}=((p_{tdi}-\bar m)\cdot(m_d-\bar m))/||m_d-\bar m||^\alpha.
    Use α=0\alpha=0, selected after testing α=0,1,2\alpha=0,1,2 for comparable behavior across sympathy and confusion cases.
    Append the label bias when btdi>βb_{tdi}>\beta; otherwise append no_bias.
    Fine-tune the dialogue model on the context/response pairs with these labels.
    At generation time, request responses labeled no_bias to reduce descriptor-specific styles.

    The bias direction for descriptor dd is the line from the global mean style vector mˉ\bar m to the descriptor-specific mean mdm_d. A response is labeled high-bias when its scaled projection in that direction exceeds the threshold. The paper used β=0.0003\beta=0.0003 for DialoGPT and β=0.0030\beta=0.0030 for BlenderBot 2.0 3B in its primary experiments. Training used batches of 16 on eight 32-GB Volta GPUs with early stopping by validation perplexity; the best DialoGPT run used SGD with learning rate 3×10−13\times10^{-1}, and the best BlenderBot 2.0 3B run used Adam with 100 warmup steps and learning rate 3×10−63\times10^{-6}.

  10. Knowl 10 — Style-equality tuning reduces measured bias but introduces trade-offs

    empirical result

    Bias-reduction tuning reduced Full Gen Bias from 3.04 to 2.66 for DialoGPT, a 13% reduction, and from 8.89 to 6.74 for BlenderBot 2.0 3B, a 24% reduction. For BlenderBot 2.0 3B, the axis-specific reductions were: Ability 9.59 to 7.59 (21%), Age 4.28 to 3.16 (26%), Body type 6.35 to 5.44 (14%), Characteristics 10.84 to 7.61 (30%), Cultural 7.64 to 5.75 (25%), Gender and sex 7.47 to 5.56 (26%), Nationality 3.74 to 3.39 (9%), Nonce 5.46 to 3.89 (29%), Political ideologies 7.59 to 6.44 (15%), Race and ethnicity 5.78 to 4.63 (20%), Religion 5.40 to 3.92 (27%), Sexual orientation 7.48 to 4.99 (33%), and Socioeconomic class 7.21 to 6.15 (15%).

    The reduction was uneven across style clusters. Sympathy, curiosity, and confusion variance fell substantially, care stayed approximately constant, and envy and hate variance increased. Frequently observed sympathy and confusion phrases declined for the descriptor groups most associated with those styles; for example, “I’m sorry to hear” fell from 30.3% to 19.4% among sympathy-associated descriptors, while “what is a” fell from 9.6% to 4.3% among confusion-associated descriptors.

    The method also produced important side effects. For BlenderBot 2.0 3B, the fraction of responses marked offensive by the B.A.D. classifier rose from 13.0% to 14.2%, and exact copying of the HOLISTICBIAS prompt rose from 17.3% to 20.0%. For DialoGPT, the corresponding offensive rate fell from 6.3% to 5.3%. Human pairwise evaluations found the tuned model comparable to the original: tuned DialoGPT was preferred in 45% of trials, judged more human in 48%, and more interesting in 47%; tuned BlenderBot 2.0 3B received 50%, 52%, and 51%, respectively. None of these human comparisons was statistically significant at p<0.05p<0.05. The authors therefore present style equality as a proof of concept rather than a generally safe mitigation: forcing style equality can suppress justified responses, increase hate or parroting, and reduce a complex notion of harm to one numerical score.

Coverage note — No substantial contributed material was deliberately omitted; detailed individual descriptor inventories, auxiliary plots, and extended before/after response samples were condensed because the load-bearing taxonomy, metrics, findings, mitigation, and limitations are represented above.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977.
  2. 2.Maria Antoniak and David Mimno. 2021. Bad seeds: Evaluating lexical methods for bias measurement. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1889–1904, Online. Association for Computational Linguistics.
  3. 3.Ashutosh Baheti, Maarten Sap, Alan Ritter, and Mark Riedl. 2021. Just say no: Analyzing the stance of neural dialogue generation in offensive contexts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4846–4862, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  4. 4.Soumya Barikeri, Anne Lauscher, Ivan Vulic, and Goran Glavaš. 2021. Redditbias: A real-world resource for bias evaluation and debiasing of conversational language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1941–1955.
  5. 5.Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift reddit dataset. In Proceedings of the international AAAI conference on web and social media, volume 14, pages 830–839.
  6. 6.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  7. 7.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  8. 8.Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1004–1015, Online. Association for Computational Linguistics.
  9. 9.Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pages 4349–4357.
  10. 10.Shikha Bordia and Samuel R. Bowman. 2019. Identifying and reducing gender bias in word-level language models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 7–15, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  12. 12.Yang Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022. On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 561–570, Dublin, Ireland. Association for Computational Linguistics.
  13. 13.Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2021. Evaluation of text generation: A survey. CoRR, abs/2006.14799.
  14. 14.Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021. Quantifying social biases in NLP: A generalization and empirical comparison of extrinsic fairness metrics. Transactions of the Association for Computational Linguistics, 9:1249–1267.
  15. 15.Pieter Delobelle, Ewoenam Kwaku Tokpo, Toon Calders, and Bettina Berendt. 2021. Measuring fairness with biased rulers: A survey on quantifying biases in pretrained language models. arXiv preprint arXiv:2112.07447.
  16. 16.Sunipa Dev, Emily Sheng, Jieyu Zhao, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Nanyun Peng, and Kai-Wei Chang. 2021. What do bias measures measure? arXiv preprint arXiv:2108.03362.
  17. 17.Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022. Theories of “gender” in nlp bias research. arXiv preprint arXiv:2205.02526.
  18. 18.Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020a. Queens are powerful too: Mitigating gender bias in dialogue generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8173–8188.
  19. 19.Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, and Adina Williams. 2020b. Multidimensional gender bias classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 314–331, Online. Association for Computational Linguistics.
  20. 20.Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2020c. The second conversational intelligence challenge (convai2). In The NeurIPS’18 Competition, pages 187–208. Springer.
  21. 21.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018. Wizard of wikipedia: Knowledge-powered conversational agents. arXiv preprint arXiv:1811.01241.
  22. 22.Christiane Fellbaum and George Miller. 1998. WordNet: An electronic lexical database.
  23. 23.Adam D Galinsky, Kurt Hugenberg, Carla Groom, and Galen V Bodenhausen. 2003. The reappropriation of stigmatizing labels: Implications for social identity. In Identity issues in groups. Emerald Group Publishing Limited.
  24. 24.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  25. 25.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120, Online. Association for Computational Linguistics.
  26. 26.Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, and Adam Lopez. 2021. Intrinsic bias metrics do not correlate with application bias. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1926–1940, Online. Association for Computational Linguistics.
  27. 27.Beth Haller, Bruce Dorries, and Jessica Rahn. 2006. Media labeling versus the us disability community identity: a study of shifting cultural language. Disability & Society, 21(1):61–75.
  28. 28.David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020. Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions. In Proceedings of the 13th International Conference on Natural Language Generation, pages 169–182, Dublin, Ireland. Association for Computational Linguistics.
  29. 29.Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. 2020. Reducing sentiment bias in language models via counterfactual evaluation. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 65–83, Online. Association for Computational Linguistics.
  30. 30.Clayton Hutto and Eric Gilbert. 2014. Vader: A parsimonious rule-based model for sentiment analysis of social media text. In Proceedings of the international AAAI conference on web and social media, volume 8, pages 216–225.
  31. 31.Xisen Jin, Francesco Barbieri, Brendan Kennedy, Aida Mostafazadeh Davani, Leonardo Neves, and Xiang Ren. 2021. On transferability of bias mitigation effects in language model fine-tuning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3770–3783, Online. Association for Computational Linguistics.
  32. 32.Armand Joulin, Édouard Grave, Piotr Bojanowski, and Tomáš Mikolov. 2017. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431.
  33. 33.Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021. Dynabench: Rethinking benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4110–4124, Online. Association for Computational Linguistics.
  34. 34.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  35. 35.Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models. Advances in Neural Information Processing Systems, 34.
  36. 36.Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2022. Internet-augmented dialogue generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8460–8478.
  37. 37.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring bias in contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 166–172, Florence, Italy. Association for Computational Linguistics.
  38. 38.Jungseob Lee, Midan Shim, Suhyune Son, Yujin Kim, Chanjun Park, and Heuiseok Lim. 2022. Empirical study on blenderbot 2.0 errors analysis in terms of model, data and user-centric approach. arXiv preprint arXiv:2201.03239.
  39. 39.Margaret Li, Jason Weston, and Stephen Roller. 2019. Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons. arXiv preprint arXiv:1909.03087.
  40. 40.Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020. Towards debiasing sentence representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5502–5515.
  41. 41.Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2020. Does gender matter? towards fairness in dialogue systems. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4403–4416.
  42. 42.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  43. 43.Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017. Towards an automatic Turing test: Learning to evaluate dialogue responses. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1116–1126, Vancouver, Canada. Association for Computational Linguistics.
  44. 44.Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, and Douwe Kiela. 2021. Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 10351–10367.
  45. 45.Henry B Mann and Donald R Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The annals of mathematical statistics, pages 50–60.
  46. 46.Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On measuring social biases in sentence encoders. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 622–628, Minneapolis, Minnesota. Association for Computational Linguistics.
  47. 47.A. H. Miller, W. Feng, A. Fisch, J. Lu, D. Batra, A. Bordes, D. Parikh, and J. Weston. 2017. Parlai: A dialog research software platform. arXiv preprint arXiv:1705.06476.
  48. 48.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371, Online. Association for Computational Linguistics.
  49. 49.Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, and Samuel R. Bowman. 2021. What ingredients make for an effective crowdsourcing protocol for difficult NLU data collection tasks? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1221–1235, Online. Association for Computational Linguistics.
  50. 50.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel Bowman. 2020. Crows-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1953–1967.
  51. 51.Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021. HONEST: Measuring hurtful sentence completion in language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2398–2406, Online. Association for Computational Linguistics.
  52. 52.Alexandra Olteanu, Kartik Talamadupula, and Kush R Varshney. 2017. The limits of abstract evaluation metrics: The case of hate speech detection. In Proceedings of the 2017 ACM on web science conference, pages 405–406.
  53. 53.Zoe Papakipos and Joanna Bitton. 2022. Augly: Data augmentations for robustness. arXiv preprint arXiv:2201.06494.
  54. 54.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. arXiv preprint arXiv:2202.03286.
  55. 55.Rebecca Qian, Candace Ross, Jude Fernandes, Eric Smith, Douwe Kiela, and Adina Williams. 2022. Perturbation augmentation for fairer nlp. arXiv preprint arXiv:2205.12586.
  56. 56.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  57. 57.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  58. 58.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards empathetic open-domain conversation models: A new benchmark and dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5370–5381.
  59. 59.Adithya Renduchintala, Denise Diaz, Kenneth Heafield, Xian Li, and Mona Diab. 2021. Gender bias amplification during speed-quality optimization in neural machine translation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 99–109, Online. Association for Computational Linguistics.
  60. 60.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, et al. 2021. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325.
  61. 61.Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender bias in coreference resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 8–14, New Orleans, Louisiana. Association for Computational Linguistics.
  62. 62.Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff. 2020. Masked language model scoring. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2699–2712, Online. Association for Computational Linguistics.
  63. 63.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  64. 64.Emily Sheng, Josh Arnold, Zhou Yu, Kai-Wei Chang, and Nanyun Peng. 2021a. Revealing persona biases in dialogue systems. arXiv preprint arXiv:2104.08728.
  65. 65.Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021b. Societal biases in language generation: Progress and challenges. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4275–4293, Online. Association for Computational Linguistics.
  66. 66.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3407–3412, Hong Kong, China. Association for Computational Linguistics.
  67. 67.Kurt Shuster, Eric Michael Smith, Da Ju, and Jason Weston. 2021. Multi-modal open-domain dialogue. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4863–4883.
  68. 68.Eric Michael Smith, Diana Gonzalez-Rico, Emily Dinan, and Y-Lan Boureau. 2020a. Controlling style in generated dialogue. arXiv preprint arXiv:2009.10855.
  69. 69.Eric Michael Smith and Adina Williams. 2021. Hi, my name is martha: Using names to measure and mitigate bias in generative dialogue models. arXiv preprint arXiv:2109.03300.
  70. 70.Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020b. Can you put it all together: Evaluating conversational agents’ ability to blend skills. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2021–2030.
  71. 71.Tom W Smith. 1992. Changing racial labels: From “colored” to “negro” to “black” to “african american”. Public Opinion Quarterly, 56(4):496–514.
  72. 72.Nancy N Soja, Susan Carey, and Elizabeth S Spelke. 1991. Ontological categories guide young children’s inductions of word meaning: Object terms and substance terms. Cognition, 38(2):179–211.
  73. 73.Anna Sotnikova, Yang Trista Cao, Hal Daumé III, and Rachel Rudinger. 2021. Analyzing stereotypes in generative text inference tasks. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4052–4065, Online. Association for Computational Linguistics.
  74. 74.Yi Chern Tan and L. Elisa Celis. 2019. Assessing social and intersectional biases in contextualized word representations. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13209–13220.
  75. 75.US Census Bureau. 2019. Place of birth for the foreign-born population of the united states. https://data.census.gov/cedsci/table?t=Place%20of%20Birth&tid=ACSDT1Y2019.B05006. [Online; accessed 2022-04-19.].
  76. 76.US Census Bureau. 2021. Decennial census of population and housing questionnaires & instructions. https://www.census.gov/programs-surveys/decennial-census/technical-documentation/questionnaires.2020_Census.html. [Online; accessed 2022-04-19.].
  77. 77.Emiel Van Miltenburg, Desmond Elliott, and Piek Vossen. 2018. Talking about other people: an endless range of possibilities. In Proceedings of the 11th International Conference on Natural Language Generation, pages 415–420.
  78. 78.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  79. 79.Alex Wang and Kyunghyun Cho. 2019. Bert has a mouth, and it must speak: Bert as a markov random field language model. In Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation, pages 30–36.
  80. 80.Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, Qinzhuo Wu, Zhengyan Li, Chong Zhang, Ruotian Ma, Zichu Fei, Ruijian Cai, Jun Zhao, Xingwu Hu, Zhiheng Yan, Yiding Tan, Yuan Hu, Qiyuan Bian, Zhihua Liu, Shan Qin, Bolin Zhu, Xiaoyu Xing, Jinlan Fu, Yue Zhang, Minlong Peng, Xiaoqing Zheng, Yaqian Zhou, Zhongyu Wei, Xipeng Qiu, and Xuanjing Huang. 2021. TextFlint: Unified multilingual robustness evaluation toolkit for natural language processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 347–355, Online. Association for Computational Linguistics.
  81. 81.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020. Measuring and reducing gendered correlations in pre-trained models. arXiv preprint arXiv:2010.06032.
  82. 82.Jason Weston, Emily Dinan, and Alexander H Miller. 2018. Retrieve and refine: Improved sequence generation models for dialogue. EMNLP 2018, page 87.
  83. 83.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  84. 84.Canwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, and Furu Wei. 2021a. Beyond preserved accuracy: Evaluating loyalty and robustness of BERT compression. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10653–10659, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  85. 85.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021b. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2950–2968, Online. Association for Computational Linguistics.
  86. 86.Jing Xu, Arthur Szlam, and Jason Weston. 2022. Beyond goldfish memory: Long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5180–5197, Dublin, Ireland. Association for Computational Linguistics.
  87. 87.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213.
  88. 88.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020. Dialogpt: Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278.
  89. 89.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019. Gender bias in contextualized word embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 629–634, Minneapolis, Minnesota. Association for Computational Linguistics.
  90. 90.Lal Zimman and Will Hayworth. 2020. How we got here: Short-scale change in identity labels for trans, cis, and non-binary people in the 2000s. Proceedings of the Linguistic Society of America, 5(1):499–513.

Citation

MLA
Smith, E. M., et al. “"I'm Sorry to Hear That": Finding New Biases in Language Models with a Holistic Descriptor Dataset”. arXiv, 2022, http://arxiv.org/abs/2205.09209v2.
APA
Smith, E. M., Hall, M., Kambadur, M., Presani, E., & Williams, A. (2022). "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset. arXiv. http://arxiv.org/abs/2205.09209v2
Chicago
Smith, E. M., M. Hall, M. Kambadur, E. Presani, and A. Williams. 2022. “"I'm Sorry to Hear That": Finding New Biases in Language Models with a Holistic Descriptor Dataset”. arXiv. http://arxiv.org/abs/2205.09209v2.
Harvard
Smith, E.M. et al. (2022) “"I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.09209v2.
Vancouver
1. Smith EM, Hall M, Kambadur M, Presani E, Williams A (2022) "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset. arXiv

BibTeX

@article{smith2022sorry,
  title = {"I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset},
  author = {Smith, Eric Michael and Hall, Melissa and Kambadur, Melanie and Presani, Eleonora and Williams, Adina},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.09209v2},
  eprint = {2205.09209}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/