Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection
Suchin GururanganDallas CardSarah K. DreierEmily K. GadeLeroy Z. WangZeyu WangLuke ZettlemoyerNoah A. Smith
Reveals that automated quality filters used to train large language models systematically favor text from wealthier, more educated, and urban demographics rather than reflecting objective standards like factuality or literary acclaim.
Modern artificial intelligence language models rely on vast amounts of web-scraped data to learn how to generate human-like text. Because raw web data frequently contains unwanted content such as spam, code, and hate speech, developers routinely apply automated "quality filters" to identify and retain only desirable text. These filters are commonly trained to favor text resembling established corpora like Wikipedia, published books, and mainstream news. However, treating these reference sources as neutral benchmarks of high quality implicitly establishes a value judgment—a sociolinguistic "language ideology"—that dictates which writing styles and author perspectives are considered valuable and worthy of inclusion in artificial intelligence systems.
The article evaluates the demographic, topical, and stylistic biases embedded within standard automated text filters by replicating the classifier used in the prominent GPT-3 model. It investigates whose language is systematically favored or excluded during data filtering, while also measuring how well the filter's definition of quality aligns with other recognized standards, including news factuality, standardized test scores, and prestigious literary awards.
To conduct this evaluation, the researchers curated a new dataset containing 910,000 articles published between 2010 and 2019 from 1,410 high school newspapers across all U.S. states. They linked each school to local demographic data from the U.S. Census and the National Center for Education Statistics, enabling them to evaluate quality scores against community wealth, educational attainment, urbanization, and school size. In addition, the researchers tested the filter against a dataset of factually reliable and unreliable news outlets, 12,100 standardized English proficiency essays, and Pulitzer Prize-winning literature across multiple genres.
The findings reveal that automated quality filters exhibit substantial demographic and stylistic disparities. First, the filter systematically assigns higher quality scores to school newspapers located in wealthier, more highly educated, and more urban areas, as well as to larger schools. Even when controlling for other variables, urban schools score significantly higher than rural ones. Second, the classifier shows strong topical and stylistic preferences: political and sports topics score up to 35 percentage points higher than everyday subjects like food, while longer documents without first- or second-person pronouns receive higher quality marks. Third, the quality filter shows no statistical difference in how it rates factually reliable versus unreliable news sources, readily classifying demonstrably false disinformation as high quality. Finally, the filter's judgments correlate weakly with human proficiency scores on standardized essays and heavily disfavor award-winning poetry and plays compared to non-fiction.
These results indicate that automated data filtering is not a neutral preprocessing step, but an active mechanism that reinforces societal disparities in artificial intelligence training data. By favoring the formal styles of privileged, urban demographics and deprioritizing vernacular or community-focused writing, existing filters risk building artificial intelligence systems that struggle to comprehend or respect diverse users. Furthermore, relying on stylistic proxies rather than true measures of reliability introduces substantial compliance, safety, and misinformation risks, because fluent yet entirely false text easily passes through existing data pipelines.
To mitigate these risks, organizations and developers should avoid assuming a single "general-purpose" corpus fits all needs. Instead, practitioners should explicitly document their inclusion and exclusion criteria, deliberately source text from underrepresented communities and across varied genres, and implement dedicated factual integrity checks rather than relying on stylistic filters. Decision-makers should recognize the article's analytical scope: while the demographic findings are based on a specific sample of U.S. school newspapers and aggregate geographic census data, the clear biases identified provide a strong foundation for revising current data curation practices.
- Paper: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜, Emily M. Bender et al. (2021). This paper establishes the foundational critique that massive web-crawled datasets and heuristic filtering encode structural demographic biases and overrepresent dominant viewpoints.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). This critical survey formalizes how conceptualizations of bias in NLP often overlook normative grounding and sociolinguistic realities, setting the conceptual framework for evaluating language ideologies.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). This work introduces the GPT-3 quality filtering methodology using curated reference sets like Wikipedia and books, which the source study explicitly replicates and audits.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). This analysis provides critical context on how digital resource inequality systematically excludes vernacular and low-resource linguistic varieties from mainstream NLP corpora.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). This seminal study demonstrates how automated representations derived from large web corpora inherit and reinforce cultural and demographic human biases.
- Paper: DsDm: Model-Aware Dataset Selection with Datamodels, Logan Engstrom et al. (2024). This work directly addresses the failure modes of heuristic similarity-to-curation filtering exposed by the source, proposing model-aware influence optimization instead.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). This study advances the critique of quality filtering by demonstrating that heuristic web filters fail to outperform scaled unfiltered data across training horizons.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). This benchmark systematically investigates the direct impact of text extraction and model-based filtering pipelines on downstream language model capability.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). This study examines the downstream sociodemographic skew in opinion representation that emerges when language models are aligned on filtered human preferences.
- Paper: From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models, Shangbin Feng et al. (2023). This paper traces how pretraining data choices and ideological skews propagate downstream into severe unfairness for moderation and fact-checking applications.
- Paper: Pretraining Language Models with Human Preferences, Tomasz Korbak et al. (2023). This work explores alternative pretraining objectives using human preference conditioning to overcome the severe data bottlenecks and biases introduced by dataset filtering.
