Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Suchin GururanganDallas CardSarah K. DreierEmily K. GadeLeroy Z. WangZeyu WangLuke ZettlemoyerNoah A. Smith

article2022EMNLP97 citations

Reveals that automated quality filters used to train large language models systematically favor text from wealthier, more educated, and urban demographics rather than reflecting objective standards like factuality or literary acclaim.

Listen

Modern artificial intelligence language models rely on vast amounts of web-scraped data to learn how to generate human-like text. Because raw web data frequently contains unwanted content such as spam, code, and hate speech, developers routinely apply automated "quality filters" to identify and retain only desirable text. These filters are commonly trained to favor text resembling established corpora like Wikipedia, published books, and mainstream news. However, treating these reference sources as neutral benchmarks of high quality implicitly establishes a value judgment—a sociolinguistic "language ideology"—that dictates which writing styles and author perspectives are considered valuable and worthy of inclusion in artificial intelligence systems.

The article evaluates the demographic, topical, and stylistic biases embedded within standard automated text filters by replicating the classifier used in the prominent GPT-3 model. It investigates whose language is systematically favored or excluded during data filtering, while also measuring how well the filter's definition of quality aligns with other recognized standards, including news factuality, standardized test scores, and prestigious literary awards.

To conduct this evaluation, the researchers curated a new dataset containing 910,000 articles published between 2010 and 2019 from 1,410 high school newspapers across all U.S. states. They linked each school to local demographic data from the U.S. Census and the National Center for Education Statistics, enabling them to evaluate quality scores against community wealth, educational attainment, urbanization, and school size. In addition, the researchers tested the filter against a dataset of factually reliable and unreliable news outlets, 12,100 standardized English proficiency essays, and Pulitzer Prize-winning literature across multiple genres.

The findings reveal that automated quality filters exhibit substantial demographic and stylistic disparities. First, the filter systematically assigns higher quality scores to school newspapers located in wealthier, more highly educated, and more urban areas, as well as to larger schools. Even when controlling for other variables, urban schools score significantly higher than rural ones. Second, the classifier shows strong topical and stylistic preferences: political and sports topics score up to 35 percentage points higher than everyday subjects like food, while longer documents without first- or second-person pronouns receive higher quality marks. Third, the quality filter shows no statistical difference in how it rates factually reliable versus unreliable news sources, readily classifying demonstrably false disinformation as high quality. Finally, the filter's judgments correlate weakly with human proficiency scores on standardized essays and heavily disfavor award-winning poetry and plays compared to non-fiction.

These results indicate that automated data filtering is not a neutral preprocessing step, but an active mechanism that reinforces societal disparities in artificial intelligence training data. By favoring the formal styles of privileged, urban demographics and deprioritizing vernacular or community-focused writing, existing filters risk building artificial intelligence systems that struggle to comprehend or respect diverse users. Furthermore, relying on stylistic proxies rather than true measures of reliability introduces substantial compliance, safety, and misinformation risks, because fluent yet entirely false text easily passes through existing data pipelines.

To mitigate these risks, organizations and developers should avoid assuming a single "general-purpose" corpus fits all needs. Instead, practitioners should explicitly document their inclusion and exclusion criteria, deliberately source text from underrepresented communities and across varied genres, and implement dedicated factual integrity checks rather than relying on stylistic filters. Decision-makers should recognize the article's analytical scope: while the demographic findings are based on a specific sample of U.S. school newspapers and aggregate geographic census data, the clear biases identified provide a strong foundation for revising current data curation practices.

arXiv: 2201.10474
Cover for Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Abstract

Language models increasingly rely on massive web crawls for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and news often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles—written by students from across the country—we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban zones (ZIP codes) are more likely to be classified as high quality. We also show that this quality measurement is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.

Table of Contents

  • 1 Introduction
  • 2 Motivation: Data Sources
  • 3 Measuring the Language Ideology of the GPT-3 Quality Filter
  • 3.1 U.S. SCHOOL NEWS
  • 3.2 The GPT-3 Quality Filter
  • 3.3 Document-Level Analysis
  • 3.4 Demographic Analysis
  • 4 Alignment with Other Notions of Quality
  • 4.1 Data
  • 4.2 Results
  • 5 Discussion
  • 6 Related Work
  • 7 Conclusion
  • Ethical Considerations
  • Acknowledgments
  • References
  • A Appendix
  • A.1 Language Model Training Corpora
  • A.2 OpenWebText URL distribution
  • A.3 Datasheet
  • A.4 Quality Filter Hyperparameters
  • A.5 High School News Scores
  • A.6 Example Articles
  • A.7 Topic Modeling
  • A.8 Demographic Features
  • A.9 Additional Regressions
  • A.10 Low Factuality News Considered High Quality
  • A.11 TOEFL Exam Responses

Knowls

  1. Knowl 1 — Demographic Disparities in GPT-3 Quality Filtering of U.S. School Newspapers

    empirical result

    A linear regression of the average predicted probability of high quality, P(high quality)P(\text{high quality}), assigned by the replicated GPT-3 quality filter to articles from 968 U.S. high schools against demographic and institutional variables reveals systematic demographic associations (R2=0.140R^2 = 0.140, adjusted R2=0.133R^2 = 0.133):

    Feature Coefficient
    Intercept 0.076
    % Rural −0.069∗∗∗-0.069^{***}
    % Adults ≥\ge Bachelor Deg. 0.059∗∗0.059^{**}
    log⁡2(Median Home Value)\log_2(\text{Median Home Value}) 0.010∗0.010^{*}
    log⁡2(Number of students)\log_2(\text{Number of students}) 0.006∗0.006^{*}
    log⁡2(Student:Teacher ratio)\log_2(\text{Student:Teacher ratio}) −0.007-0.007
    Is Public 0.015∗0.015^{*}
    Is Magnet 0.013
    Is Charter 0.033

    Where ∗p<0.05,∗∗p<0.01,∗∗∗p<0.001^{*}p < 0.05, ^{**}p < 0.01, ^{***}p < 0.001.

    School size, student-to-teacher ratio, and median home values are log-transformed (log⁡2\log_2), while percentage variables are represented as raw proportions ([0,1][0, 1]). Articles from schools located in more urban, highly educated, and higher-home-value ZIP codes, as well as larger schools and public schools, receive systematically higher quality scores. Holding other variables constant, a 14 percentage point increase in urban population or a 17 percentage point increase in college-educated adults corresponds to a 1 percentage point increase in average quality score. A doubling of median home value or a quadrupling of school size also increases the score by 1 percentage point.

  2. Knowl 2 — Independence of Text Quality Classification from News Factuality

    empirical result

    Evaluating the replicated GPT-3 quality filter on 9.9K articles from factually reliable news outlets and 7.7K articles from factually unreliable news outlets (categorized via MediaBiasFactCheck) demonstrates no statistically significant difference in the predicted quality score distribution between high-factuality and low-factuality sources (p=0.085p = 0.085, two-sample Kolmogorov-Smirnov test).

    Articles containing debunked conspiracy theories and medical disinformation can receive high quality probabilities if written in formal, journalistic prose. For instance, articles promoting COVID-19 conspiracy claims achieved scores such as P(high quality)=0.93P(\text{high quality}) = 0.93 and P(high quality)=0.98P(\text{high quality}) = 0.98, illustrating that the quality filter evaluates stylistic conformity with standard news and encyclopedic corpora rather than factual reliability.

  3. Knowl 3 — Topic and Stylistic Predictors of Document Quality Scores in Student Opinion Articles

    empirical result

    An ordinary least squares linear regression on N=10,000N = 10{,}000 opinion pieces from the U.S. high school news dataset demonstrates that quality scores assigned by the replicated GPT-3 filter are strongly governed by topic and stylistic markers (R2=0.336R^2 = 0.336, adjusted R2=0.336R^2 = 0.336). Topics were extracted using a 10-topic Latent Dirichlet Allocation (LDA) model, with Topic 0 (food, restaurant, eat) serving as the omitted baseline category:

    Feature Coefficient
    Intercept 0.471∗∗∗0.471^{***}
    Topic 5 (christmas, dress, holiday) −0.056∗∗∗-0.056^{***}
    Topic 2 (school, college, year) −0.037∗∗∗-0.037^{***}
    Topic 6 (student, school, class) −0.004-0.004
    Topic 1 (people, just, like) 0.0030.003
    Topic 7 (movie, film, movies) 0.062∗∗∗0.062^{***}
    Topic 3 (music, album, song) 0.113∗∗∗0.113^{***}
    Topic 4 (people, women, media) 0.197∗∗∗0.197^{***}
    Topic 9 (game, team, players) 0.246∗∗∗0.246^{***}
    Topic 8 (Trump, president, election) 0.346∗∗∗0.346^{***}
    Presence of first/second person pronoun −0.054∗∗∗-0.054^{***}
    Presence of third person pronoun 0.0240.024
    log⁡2(Number of tokens)\log_2(\text{Number of tokens}) 0.088∗∗∗0.088^{***}

    Where ∗∗∗p<0.001^{***}p < 0.001.

    Political discourse (Topic 8: +0.346+0.346) and sports (Topic 9: +0.246+0.246) yield the largest positive score shifts relative to food-related topics, whereas holiday and local school topics reduce the predicted quality score. Stylistically, inclusion of first- or second-person pronouns decreases P(high quality)P(\text{high quality}) by 5.4 percentage points, and doubling document token length increases the score by 8.8 percentage points.

  4. Knowl 4 — Prompt Dependence Over Human Proficiency Scores in Quality Filtering of TOEFL Essays

    empirical result

    Evaluation of 12.1K participant essays from the Test of English as a Foreign Language (TOEFL11) dataset reveals that the replicated GPT-3 quality filter correlates only weakly with official human-graded language proficiency scores (Pearson r=0.12,p<0.001r = 0.12, p < 0.001). A regression modeling the filter's quality score as a function of score tier (low, medium, high, with medium as reference) and prompt identifier explains substantial variance (R2=0.712R^2 = 0.712, adjusted R2=0.711R^2 = 0.711), where prompt features exhibit far larger effects than human proficiency ratings:

    Feature Coefficient
    Intercept 0.0631∗∗∗0.0631^{***}
    Low score −0.0414-0.0414
    High score 0.03390.0339
    Prompt 7 (ideas vs. facts) −0.0283∗∗∗-0.0283^{***}
    Prompt 6 (traveling with tour guide) −0.0204∗∗∗-0.0204^{***}
    Prompt 2 (enjoyment of life across age) 0.0068∗∗∗0.0068^{***}
    Prompt 8 (risk taking vs. known tasks) 0.0346∗∗∗0.0346^{***}
    Prompt 3 (youth time helping communities) 0.0880∗∗∗0.0880^{***}
    Prompt 5 (fewer cars in twenty years) 0.1470∗∗∗0.1470^{***}
    Prompt 4 (advertisements exaggerating products) 0.6745∗∗∗0.6745^{***}

    Where ∗∗∗p<0.001^{***}p < 0.001.

    Prompt 1 (broad knowledge vs specialization) served as the reference prompt category. Essays responding to Prompt 4 (advertisements exaggerating products) average P(high quality)=0.74P(\text{high quality}) = 0.74, while essays responding to prompts eliciting personal narratives, such as Prompt 6 (travel in a tour group) and Prompt 7 (learning ideas vs facts), average quality scores of 0.050.05 and 0.040.04, respectively.

  5. Knowl 5 — Implementation and Training Specification of the GPT-3 Binary Quality Filter

    model/method

    The GPT-3 text quality filter is implemented as a binary logistic regression classifier trained to discriminate between positive reference corpora and negative web crawl data.

    Training data consists of:

    • Positive class: 240 million whitespace-separated tokens, sampled equally (80 million tokens each) from English Wikipedia, Books3 (a large corpus of copyrighted eBooks), and OpenWebText.
    • Negative class: 240 million whitespace-separated tokens sampled from the September 2019 Common Crawl snapshot.

    Hyperparameters and optimization configuration determined via 100-trial random hyperparameter search:

    • Feature representation: Hashing vectorizer with whitespace tokenization, unigrams and bigrams (ngram range [1,2][1, 2]), no stopword removal.
    • Regularization: L1L_1 penalty with inverse regularization strength C=0.977778C = 0.977778.
    • Optimization solver: liblinear.
    • Stopping tolerance: tol=0.000816\text{tol} = 0.000816.
    • Random state seed: 44555.

    On a held-out test split, this replicated filter achieves an F1F_1 score of 90.4%90.4\% and a classification accuracy of 91.7%91.7\%. The classifier outputs a probability score P(high quality)∈[0,1]P(\text{high quality}) \in [0, 1] per document.

  6. Knowl 6 — Genre Disparities of Quality Classifier Scores on Pulitzer Prize-Winning Works

    empirical result

    Applying the replicated GPT-3 quality classifier to Pulitzer Prize-winning literary works drawn from the Books3 and Project Gutenberg corpora demonstrates that the quality filter exhibits strong genre disparity rather than a universal proxy for literary excellence.

    Among award-winning literature:

    • Nonfiction works obtain the highest quality scores, with a median P(high quality)P(\text{high quality}) exceeding 0.750.75.
    • Fiction works achieve median quality scores between 0.500.50 and 0.650.65, comparable to the median quality of self-published novels in BookCorpus.
    • Poetry and drama works receive significantly lower quality scores, with median P(high quality)P(\text{high quality}) falling between 0.250.25 and 0.400.40.

    These findings demonstrate that the classifier penalizes poetic and dramatic syntactic structures, dialogue formats, and line breaks that deviate from the prose norms of encyclopedias, newswire, and expository non-fiction.

  7. Knowl 7 — High Top-Level Domain Concentration in OpenWebText

    empirical result

    An empirical audit of OpenWebText—an open-source reproduction of the WebText corpus used to train GPT-2 and utilized as a positive reference dataset for the GPT-3 quality filter—shows extreme source concentration: out of 311,000 unique top-level domains, just 1% generate 75% of all documents in the corpus.

    The ten most frequent top-level URL domains and their document shares are:

    URL Domain # Docs % of Total Docs
    bbc.co.uk 116K 1.50%
    theguardian.com 115K 1.50%
    washingtonpost.com 89K 1.20%
    nytimes.com 88K 1.10%
    reuters.com 79K 1.10%
    huffingtonpost.com 72K 0.96%
    cnn.com 70K 0.93%
    cbc.ca 67K 0.89%
    dailymail.co.uk 58K 0.77%
    go.com 48K 0.63%

    The corpus is heavily dominated by mainstream Anglo-American news outlets, sports sites, and tech journalism reflecting the user demographics and submission dynamics of Reddit.

  8. Knowl 8 — The U.S. School News Dataset for Language Quality Filtering Evaluation

    experimental setup

    The U.S. School News corpus is constructed to evaluate demographic, stylistic, and topical variation in automated text quality filtering. It comprises approximately 910,000 English-language articles published between 2010 and 2019 by 1,410 U.S. high schools across 1,329 ZIP codes (covering 552 U.S. counties in all 50 U.S. states and Washington, D.C.).

    Curation details include:

    • Source: Articles scraped from student newspaper websites utilizing the SNOsites WordPress template (identifying 2,483 schools initially, yielding 1.95 million raw articles).
    • Exclusions: Removal of non-high schools, articles categorized under photo, video, or multimedia URL slugs, and schools with fewer than 100 articles to eliminate scraping artifacts.
    • Metadata Integration: School geographic locations were mapped to ZIP codes using the Google Maps Place API and merged with 2017–2018 school-level data from the National Center for Education Statistics (NCES) (student counts, student-to-teacher ratio, school type indicators) and 2020 U.S. Census data (median home value, adult educational attainment, percent rural population).
  9. Knowl 9 — Demographic Regression Sensitivity to Racial Composition and Political Vote Share

    empirical result

    In extended regressions of high school average quality scores (P(high quality)P(\text{high quality})) across N=968N = 968 schools:

    1. Adding school racial composition covariates (R2=0.152R^2 = 0.152, adjusted R2=0.142R^2 = 0.142) shows a statistically significant positive coefficient only for percentage of Asian students (β=0.081,p<0.01\beta = 0.081, p < 0.01), while coefficients for Black (β=−0.009\beta = -0.009) and Hispanic students (β=−0.020\beta = -0.020) are not statistically significant. The percentage of Asian students positively correlates with median home value (Pearson r=0.32,p<0.001r = 0.32, p < 0.001), absorbing the statistical significance of log⁡2(Median Home Value)\log_2(\text{Median Home Value}) (β=0.007\beta = 0.007, p≥0.05p \ge 0.05).

    2. Adding county-level 2016 Republican presidential vote share (R2=0.164R^2 = 0.164, adjusted R2=0.157R^2 = 0.157) produces a significant negative coefficient for GOP vote share (β=−0.114,p<0.001\beta = -0.114, p < 0.001). This political covariate absorbs the statistical significance of percent rural population (β=−0.021,p≥0.05\beta = -0.021, p \ge 0.05) and home values (β=0.003,p≥0.05\beta = 0.003, p \ge 0.05), acting as a strong proxy for rurality and geographic income stratification.

  10. Knowl 10 — Methodological and Ethical Limitations of School News Quality Filtering Audits

    limitation

    The study's findings are subject to several specific methodological and ethical constraints:

    • Dataset Non-Representativeness: The U.S. School News corpus is restricted to high schools with adequate internet infrastructure and active subscription to the SNOsites hosting platform, potentially underrepresenting both the least resourced schools and elite private schools with bespoke platforms.
    • Demographic Resolution: Demographic attributes are merged at the ZIP code and county level rather than directly measuring the socioeconomic background, race, or gender of individual student authors.
    • Lack of Proprietary Access: Proprietary training datasets and filter models for GPT-3 were unavailable; analyses rely on an open-source reproduction rather than the proprietary deployment artifact.
    • Domain and Geographic Specificity: The evaluation focuses exclusively on U.S. English high school journalism, which represents a niche genre reflecting U.S.-centric cultural and educational contexts.
    • Author Privacy: Due to authors being minors and the absence of explicit consent for training use, raw text from the corpus is withheld from release, making available only school-level demographic aggregates and quality scores.

Coverage note — All primary empirical analyses, dataset curation procedures, model replication parameters, and documented limitations have been included as knowls; omitted material is limited to secondary literature reviews and qualitative prose samples.

References

  1. 1.Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-Muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society.
  2. 2.Armen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi, Hu Xu, Gargi Ghosh, and Luke Zettlemoyer. 2021. Htlm: Hyper-text pre-training and prompting of language models. arXiv, abs/2107.06955.
  3. 3.Gabriel Arana. 2018. Decades of failure. Columbia Journalism Review.
  4. 4.Tyler Baldwin and Joyce Chai. 2011. Beyond normalization: Pragmatics of word form in text messages. In Proceedings of 5th International Joint Conference on Natural Language Processing.
  5. 5.Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James Glass, and Preslav Nakov. 2018. Predicting factuality of reporting and bias of news media sources. In Proceedings of EMNLP.
  6. 6.David Bamman and Noah A Smith. 2014. Unsupervised discovery of biographical structure from text. Transactions of the Association for Computational Linguistics, 2:363–376.
  7. 7.Jack Bandy and Nicholas Vincent. 2021. Addressing “documentation debt” in machine learning: A retrospective datasheet for BookCorpus. In NeurIPS.
  8. 8.Michael Barthel, Galen Stocking, Jesse Holcomb, and Amy Mitchell. 2016. Seven-in-Ten Reddit users get news on the site. [online; accessed: 2020-6-2].
  9. 9.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of FAccT.
  10. 10.Julian R. Betts, Kim S. Reuben, and Anne Danenberg. 2000. Equal Resources, Equal Outcomes? The Distribution of School Resources and Student Achievement in California. Public Policy Institute of California.
  11. 11.Steve Bien-Aimé. 2016. AP stylebook normalizes sports as a male space. Newspaper Research Journal, 37(1):44–57.
  12. 12.Daniel Blanchard, Joel R. Tetreault, Derrick Higgins, A. Cahill, and Martin Chodorow. 2013. TOEFL11: A corpus of non-native English. ETS Research Report Series, 2013:15.
  13. 13.David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003. Latent Dirichlet allocation. J. Mach. Learn. Res., 3:993–1022.
  14. 14.Su Lin Blodgett. 2021. Sociolinguistically Driven Approaches for Just Natural Language Processing. Ph.D. thesis, University of Massachusetts Amherst.
  15. 15.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  16. 16.Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016. Demographic dialectal variation in social media: A case study of African-American English. In Proceedings of EMNLP.
  17. 17.Sarah Brayne. 2017. Big data surveillance: The case of policing. American Sociological Review, 82(5):977–1008.
  18. 18.Julian Brooke, Adam Hammond, and Graeme Hirst. 2015. GutenTag: an NLP-driven tool for digital humanities research in the Project Gutenberg corpus. In Proceedings of the Fourth Workshop on Computational Linguistics for Literature.
  19. 19.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. arXiv, abs/2005.14165.
  20. 20.Janice Kai Chen, Ilena Peng, Jasen Lo, Trisha Ahmed, Simon J. Levien, and Devan Karp. 2021. Voices investigation: Few black, latinx students are editors of top college newspapers. AAJA Voices.
  21. 21.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pre-training text encoders as discriminators rather than generators. In ICLR.
  22. 22.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale.
  23. 23.Justin T. Craft, Kelly E. Wright, Rachel Elizabeth Weissler, and Robin M. Queen. 2020. Language and discrimination: Generating meaning, perceiving identities, and discriminating outcomes. Annual Review of Linguistics, 6(1):389–407.
  24. 24.Stephanie Decker. 2013. The silence of the archives: business history, post-colonialism and archival ethnography. Management & Organizational History, 8(2):155–173.
  25. 25.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL.
  26. 26.Robert DiNicola. 1994. Teaching journalistic style with the AP stylebook: Beyond fussy rules and dogma of ‘correctness’. The Journalism Educator, 49(2):64–70.
  27. 27.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021. Documenting the English Colossal Clean Crawled Corpus. arXiv, abs/2104.08758.
  28. 28.Penelope Eckert. 1989. Jocks and burnouts: Social categories and identity in the high school. Teachers college press.
  29. 29.Jacob Eisenstein. 2013. What to do about bad language on the internet. In Proceedings of NAACL, pages 359–369.
  30. 30.Jacob Eisenstein, Brendan O’Connor, Noah A. Smith, and Eric P. Xing. 2014. Diffusion of lexical change in social media. PLoS ONE, 9.
  31. 31.William Fedus, Barret Zoph, and Noam Shazeer. 2021. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.
  32. 32.Casey Fiesler, Nathan Beard, and Brian Keegan. 2020. No robots, spiders, or scrapers: Legal and ethical regulation of data collection methods in social media terms of service. In Proceedings of ICWSM.
  33. 33.Paula Froke, Anna Jo Bratton, Jeff McMillan, Pia Sarkar, Jerry Schwartz, and Raghuram Vadarevu. 2020. The Associated Press stylebook 2020-2022. The Associated Press.
  34. 34.Susan Gal. 2016. Sociolinguistic differentiation, page 113–136. Cambridge University Press.
  35. 35.Susan Gal and Judith T. Irvine. 1995. The boundaries of languages and disciplines: How ideologies construct difference. Social Research, 62(4):967–1001.
  36. 36.Leo Gao. 2021. An empirical exploration in quality filtering of text data. arXiv, abs/2109.00698.
  37. 37.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800Gb dataset of diverse text for language modeling. arXiv, abs/2101.00027.
  38. 38.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Communications of the ACM, 64(12):86–92.
  39. 39.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020.
  40. 40.Joyce Still Gibson. 1961. A study of the status of high school newspapers in the virginia public schools. Master’s thesis, University of Richmond.
  41. 41.Aaron Gokaslan and Vanya Cohen. 2019. OpenWebText corpus.
  42. 42.Eduardo Graells-Garrido, Mounia Lalmas, and Filippo Menczer. 2015. First women, second sex: Gender bias in Wikipedia. In Proceedings of the 26th ACM conference on hypertext & social media.
  43. 43.Rob Greenwald, Larry V. Hedges, and Richard D. Laine. 1996. The effect of school resources on student achievement. Review of Educational Research, 66(3):361–396.
  44. 44.Elizabeth Grieco. 2018. Newsroom employees are less diverse than U.S. workers overall. Pew Research Center. [online; accessed 2022-01-22].
  45. 45.Dirk Hovy and Diyi Yang. 2021. The importance of modeling social factors of language: Theory and practice. In Proceedings of NAACL.
  46. 46.Keira Huang. 2013. Wikipedia fails to bridge gender gap. South China Morning Post. [online; accessed 2022-01-11].
  47. 47.Eun Seo Jo and Timnit Gebru. 2020. Lessons from archives: Strategies for collecting sociocultural data in machine learning. Proceedings of FAccT.
  48. 48.Paresh Kharya and Ali Alvi. 2021. Using DeepSpeed and Megatron to train Megatron-Turing NLG 530b, the world’s largest and most powerful generative language model. [online; accessed 2022-01-20].
  49. 49.William Labov. 2006. The Social Stratification of English in New York City, 2 edition. Cambridge University Press.
  50. 50.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A lite BERT for self-supervised learning of language representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  51. 51.Lee & Low Books. 2020. Where is the diversity in publishing? The 2019 diversity baseline survey results. [online; accessed 2021-11-24].
  52. 52.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  53. 53.Stephanie Lindemann. 2005. Who speaks “broken English”? US undergraduates’ perceptions of nonnative English. International Journal of Applied Linguistics, 15(2):187–212.
  54. 54.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach. arXiv, abs/1907.11692.
  55. 55.Li Lucy and David Bamman. 2021. Characterizing English Variation across Social Media Communities with BERT. Transactions of the Association for Computational Linguistics, 9:538–556.
  56. 56.Jeff MacSwan. 2020. Academic English as standard language ideology: A renewed research agenda for asset-based language education. Language Teaching Research, 24(1):28–36.
  57. 57.Michael Mandiberg. 2020. Mapping Wikipedia. The Atlantic. [online; accessed 2021-11-24].
  58. 58.Sorin Adam Matei and Brian C. Britt. 2017. Structural Differentiation in Social Media. Springer International Publishing.
  59. 59.Kris McGuffie and Alex Newhouse. 2020. The radicalization risks of GPT-3 and advanced neural language models. arXiv, abs/2009.06807.
  60. 60.Zion Mengesha, Courtney Heldreth, Michal Lahav, Juliana Sublewski, and Elyse Tuennerman. 2021. “I don’t think these devices are very culturally sensitive.”—Impact of automated speech recognition errors on African Americans. Frontiers in Artificial Intelligence, 4:169.
  61. 61.Meta-wiki. 2018. Community insights/2018 report/contributors. [online; accessed 2012-11-24].
  62. 62.Amy Mitchell, Jeffrey Gottfried, Michael Barthel, and Nami Sumida. 2018. Can Americans tell factual from opinion statements in the news? Pew Research Center’s Journalism Project. [online; accessed 2022-01-22].
  63. 63.Dong Nguyen, Laura Rosseel, and Jack Grieve. 2021. On learning and representing social meaning in NLP: A sociolinguistic perspective. In Proceedings of NAACL.
  64. 64.Katherine Panciera, Aaron Halfaker, and Loren Terveen. 2009. Wikipedians are born, not made: A study of power editors on Wikipedia. In Proceedings of the ACM 2009 International Conference on Supporting Group Work.
  65. 65.Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011. Scikit-learn: Machine learning in python. Journal of Machine Learning Research, 12(85):2825–2830.
  66. 66.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  67. 67.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding with unsupervised learning. [online; accessed 2022-01-22].
  68. 68.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. [online; accessed 2022-01-22].
  69. 69.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv, abs/2112.11446.
  70. 70.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  71. 71.Sean F Reardon and Ann Owens. 2014. 60 years after Brown: Trends and consequences of school segregation. Annual Review of Sociology, 40:199–218.
  72. 72.John R. Rickford. 1985. Ethnicity as a sociolinguistic boundary. American Speech, 60(2):99–125.
  73. 73.John R. Rickford and Sharese King. 2016. Language and linguistics on trial: Hearing Rachel Jeantel (and other vernacular speakers) in the courtroom and beyond. Language, 92(4):948–988.
  74. 74.Jonathan Rosa and Nelson Flores. 2017. Unsettling race and language: Toward a raciolinguistic perspective. Language in Society, 46(5):621–647.
  75. 75.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2021. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. arXiv, abs/2111.07997.
  76. 76.Dante J. Scala and Kenneth M. Johnson. 2017. Political polarization along the rural-urban continuum? the geography of the presidential vote, 2000–2016. The ANNALS of the American Academy of Political and Social Science, 672(1):162–184.
  77. 77.Ari Schlesinger, Kenton P. O’Hara, and Alex S. Taylor. 2018. Let’s talk about race: Identity, chatbots, and AI. In Proceedings of CHI.
  78. 78.Anissa Tanweer, Emily Kalah Gade, PM Krafft, and Sarah K Dreier. 2021. Why the data revolution needs qualitative thinking. Harvard Data Science Review.
  79. 79.Michel-Rolph Trouillot. 1995. Silencing the past: Power and the production of history. Beacon Press.
  80. 80.Rob van der Goot, Alan Ramponi, Arkaitz Zubiaga, Barbara Plank, Benjamin Muller, Iñaki San Vicente Roncal, Nikola Ljubešić, Özlem Çetinoğlu, Rahmad Mahendra, Talha Çolakoğlu, Timothy Baldwin, Tommaso Caselli, and Wladimir Sidorenko. 2021. MultiLexNorm: A shared task on multilingual lexical normalization. In Proceedings of the Seventh Workshop on Noisy User-generated Text.
  81. 81.Marlon Vanegas, Juan Restrepo, Yurley Zapata, Giovany Rodríguez, Luis Cardona, and Cristian Muñoz. 2016. Linguistic discrimination in an English language teaching program: Voices of the invisible others. Íkala, Revista de Lenguaje y Cultura, 21.
  82. 82.Fred Vultee. 2012. A paleontology of style. Journalism Practice, 6(4):450–464.
  83. 83.Claudia Wagner, David Garcia, Mohsen Jadidi, and Markus Strohmaier. 2015. It’s a man’s Wikipedia? Assessing gender inequality in an online encyclopedia. In Proceedings of the AAAI conference on web and social media.
  84. 84.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of LREC.
  85. 85.Matthew L Williams, Pete Burnap, and Luke Sloan. 2017. Towards an ethical framework for publishing Twitter data in social research: Taking into account users’ views, online context and algorithmic estimation. Sociology, 51(6):1149–1168.
  86. 86.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  87. 87.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of ICCV.

Citation

MLA
Gururangan, S., et al. “Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 2562–80, https://doi.org/10.18653/v1/2022.emnlp-main.165.
APA
Gururangan, S., Card, D., Dreier, S., Gade, E., Wang, L., Wang, Z., Zettlemoyer, L., & Smith, N. A. (2022). Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2562–2580. https://doi.org/10.18653/v1/2022.emnlp-main.165
Chicago
Gururangan, S., D. Card, S. Dreier, et al. 2022. “Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2562–80. https://doi.org/10.18653/v1/2022.emnlp-main.165.
Harvard
Gururangan, S. et al. (2022) “Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2562–2580. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.165.
Vancouver
1. Gururangan S, Card D, Dreier S, Gade E, Wang L, Wang Z, Zettlemoyer L, Smith NA (2022) Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2562–2580

BibTeX

@inproceedings{gururangan-etal-2022-whose,
    title = "Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection",
    author = "Gururangan, Suchin  and
      Card, Dallas  and
      Dreier, Sarah  and
      Gade, Emily  and
      Wang, Leroy  and
      Wang, Zeyu  and
      Zettlemoyer, Luke  and
      Smith, Noah A.",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.165/",
    doi = "10.18653/v1/2022.emnlp-main.165",
    pages = "2562--2580"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/