"I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset
Eric Michael SmithMelissa HallMelanie KambadurEleonora PresaniAdina Williams
Introduces HOLISTICBIAS, an open-source dataset of over 450,000 conversational prompts spanning nearly 600 demographic descriptors across 13 axes, enabling more comprehensive measurement and mitigation of subtle social biases in generative language models.
As natural language processing models become widely adopted in consumer-facing dialogue systems and applications, identifying and mitigating demographic bias is critical to avoid reinforcing social harms. Most existing evaluation datasets rely on narrow taxonomies and rigid benchmarks that overlook intersectional or evolving identity terms, frequently missing subtle model behaviors such as patronizing sympathy toward individuals with disabilities or overt confusion regarding underrepresented gender and sexual identities.
The article introduces HOLISTICBIAS, a comprehensive demographic evaluation dataset, and demonstrates its utility in uncovering, measuring, and mitigating subtle social biases across several generative and masked language models.
To construct the dataset, the researchers combined algorithmic expansions with a participatory process involving community members and subject matter experts. This resulted in a living taxonomy of nearly 600 American English descriptor terms categorized across 13 demographic axes, such as ability, gender and sex, race and ethnicity, and socioeconomic status. Combining these descriptors with person nouns and 26 conversational sentence templates generated approximately 460,000 unique prompt sentences. The researchers evaluated model token likelihoods, classified generative responses across 217 distinct conversational styles, and evaluated offensiveness classifiers on models including GPT-2, RoBERTa, DialoGPT, and BlenderBot 2.0 (400-million and 3-billion parameter variants).
The analysis yielded four central findings. First, larger language models exhibited substantially higher levels of generation bias than smaller architectures; for example, the 3-billion-parameter BlenderBot 2.0 displayed significantly more demographic variance in conversational styles than its 400-million-parameter counterpart or DialoGPT. Second, conversational biases predominantly surfaced as disproportionate levels of sympathy toward ability-related descriptors and curiosity or confusion toward gender, sex, and sexual orientation descriptors. Third, automated offensiveness classifiers exhibited substantial systemic bias, frequently assigning high offensiveness probabilities to neutral sentences simply because they contained marginalized demographic descriptors. Fourth, a proof-of-concept mitigation technique based on style equality reduced overall generation bias by 13% in DialoGPT and 24% in the 3-billion-parameter BlenderBot 2.0.
These findings demonstrate that language models generate subtle microaggressions that conventional safety filters and sentiment analyzers fail to catch. Because existing toxicity classifiers often penalize marginalized identity markers, organizations deploying dialogue systems face serious compliance, brand reputation, and safety risks. Furthermore, scaling up model size without targeted behavioral constraints amplifies these biased behavioral patterns rather than resolving them.
Organizations should adopt comprehensive, living evaluation suites like HOLISTICBIAS to audit models before deployment rather than relying solely on static benchmarks. While style equality tuning successfully reduces unwanted variations in sympathy and confusion, teams should treat it as an experimental intervention. Practitioners must balance style equalization carefully, as aggressive tuning can increase prompt parroting or suppress justified empathy.
The findings are bounded by the dataset's current focus on United States English and single-descriptor prompts. Additionally, using automated style and offensiveness classifiers introduces measurement noise. Nonetheless, the high consistency across multiple language models and hundreds of thousands of test prompts provides strong confidence in HOLISTICBIAS as an effective diagnostic benchmark for conversational AI.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). Provides the foundational conceptual framing and taxonomy of representational versus allocational harms necessary for analyzing social bias in natural language processing systems.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). Introduces standard benchmark methodologies for measuring stereotypical associations in pretrained language models that the source expands through fine-grained conversational descriptors.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Establishes standard prompt-based evaluations for toxic degeneration and automated safety classifier assessments across language models.
- Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). Pioneers template-based behavioral and linguistic capability testing in NLP, establishing the prompt-matrix paradigm adapted by the source.
- Paper: Model Cards for Model Reporting, Margaret Mitchell et al. (2019). Establishes standard reporting frameworks for disaggregating model performance and bias across specific demographic subgroups.
- Paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al. (2016). Provides seminal mathematical and conceptual formulations for quantifying and mitigating demographic bias in learned semantic representations.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). Demonstrates that distributional semantics in web-scale corpora naturally inherit pervasive human social and demographic associations.
- Paper: BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation, Tianxiang Sun et al. (2022). Extends the study of demographic disparities in text generation by evaluating and mitigating social bias directly within language model-based evaluation metrics.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). Broadens the evaluation of demographic misalignment from descriptor-based conversational styles to the political and value opinions reflected by language models.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). Examines how modern preference alignment techniques (such as RLHF and DPO) systematically alter global and demographic representation in generative models.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). Integrates fine-grained social bias and safety evaluations into a unified multidimensional benchmark framework for model trustworthiness.
- Paper: Automatically Auditing Large Language Models via Discrete Optimization, Erik Jones et al. (2023). Applies automated discrete optimization to systematically discover rare, emergent failure modes and biased behaviors beyond fixed template prompts.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Scales up the discovery of subtle toxicity and demographic harms by using language models to automatically generate diverse red-teaming test cases.
- Paper: OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models, Moreno D'Incà et al. (2024). Generalizes the identification of fine-grained, open-ended demographic biases from conversational text systems to text-to-image generative models.
