Built independently by an author, for readers. Read the story and support ChapterPal

keyword

linguistic discrimination

Linguistic discrimination is the unfair treatment or devaluation of individuals and communities based on their language, dialect, accent, or speech patterns. This bias typically occurs when a dominant or standardized linguistic variety is treated as inherently superior, correct, or prestigious, while nonstandard varieties, regional accents, and minority vernaculars are stigmatized, excluded, or deemed low quality. Rooted in prevailing language ideologies, linguistic discrimination reflects and reinforces existing socioeconomic, racial, and cultural hierarchies, creating systemic disadvantages across education, employment, the justice system, and automated language technologies.

3 items

Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Suchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade, Leroy Z. Wang, Zeyu Wang, Luke Zettlemoyer, Noah A. Smith

OrganizationsAllen Institute for AIEmory UniversityUniversity of MichiganUniversity of New MexicoUniversity of Washington

Why you should read this

Reveals that automated quality filters used to train large language models systematically favor text from wealthier, more educated, and urban demographics rather than reflecting objective standards like factuality or literary acclaim.

Language models increasingly rely on massive web crawls for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and news often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles—written by students from across the country—we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban zones (ZIP codes) are more likely to be classified as high quality. We also show that this quality measurement is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.

Added

2026-09-26

VALUE: Understanding Dialect Disparity in NLU

VALUE: Understanding Dialect Disparity in NLU

Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, Diyi Yang

OrganizationsGeorgia Institute of Technology

Why you should read this

Presents a dialect-specific evaluation benchmark and linguistically validated transformation rules for African American Vernacular English to expose and analyze performance disparities in modern language models.

English Natural Language Understanding (NLU) systems have achieved great performances and even outperformed humans on benchmarks like GLUE and SuperGLUE. However, these benchmarks contain only textbook Standard American English (SAE). Other dialects have been largely overlooked in the NLP community. This leads to biased and inequitable NLU systems that serve only a sub-population of speakers. To understand disparities in current models and to facilitate more dialect-competent NLU systems, we introduce the VernAcular Language Understanding Evaluation (VALUE) benchmark, a challenging variant of GLUE that we created with a set of lexical and morphosyntactic transformation rules. In this initial release (V.1), we construct rules for 11 features of African American Vernacular English (AAVE), and we recruit fluent AAVE speakers to validate each feature transformation via linguistic acceptability judgments in a participatory design manner. Experiments show that these new dialectal features can lead to a drop in model performance.

Added

2026-09-26

Language (Technology) is Power: A Critical Survey of “Bias” in NLP

Language (Technology) is Power: A Critical Survey of “Bias” in NLP

Su Lin Blodgett, Solon Barocas, Hal Daum'e, Hanna M. Wallach

OrganizationsCollege of Information and Computer SciencesCornell UniversityMicrosoftUniversity of MarylandUniversity of Massachusetts Amherst

Why you should read this

Reveals pervasive conceptual weaknesses across 146 natural language processing bias studies and delivers essential guidelines for aligning technical mitigation methods with normative reasoning and social power dynamics.

We survey 146 papers analyzing "bias" in NLP systems, finding that their motivations are often vague, inconsistent, and lacking in normative reasoning, despite the fact that analyzing "bias" is an inherently normative process. We further find that these papers' proposed quantitative techniques for measuring or mitigating "bias" are poorly matched to their motivations and do not engage with the relevant literature outside of NLP. Based on these findings, we describe the beginnings of a path forward by proposing three recommendations that should guide work analyzing "bias" in NLP systems. These recommendations rest on a greater recognition of the relationships between language and social hierarchies, encouraging researchers and practitioners to articulate their conceptualizations of "bias"---i.e., what kinds of system behaviors are harmful, in what ways, to whom, and why, as well as the normative reasoning underlying these statements---and to center work around the lived experiences of members of communities affected by NLP systems, while interrogating and reimagining the power relations between technologists and such communities.

Added

2026-09-18