Language (Technology) is Power: A Critical Survey of “Bias” in NLP
Su Lin BlodgettSolon BarocasHal Daum'eHanna M. Wallach
Reveals pervasive conceptual weaknesses across 146 natural language processing bias studies and delivers essential guidelines for aligning technical mitigation methods with normative reasoning and social power dynamics.
As automated language technologies become increasingly integrated into high-stakes applications such as hiring, moderation, and search, concerns regarding algorithmic fairness have surged. The article addresses the foundational problem that current technical research on "bias" in natural language processing (NLP) often lacks conceptual clarity and rigorous ethical grounding. Without a structured understanding of what makes a system's behavior harmful and to whom, organizations and technologists risk developing superficial technical fixes that fail to prevent real-world harms or protect vulnerable user populations.
The article evaluates how "bias" is conceptualized, motivated, and operationalized across the language processing literature. To achieve this, it presents a comprehensive survey of 146 research publications released prior to May 2020 that focus on analyzing, measuring, or mitigating social bias in written language technologies. The analysis categorizes these works using an established taxonomy of harms, specifically distinguishing between representational harms—such as stereotyping, denigration, and demographic performance disparities—and allocational harms, which involve the unfair distribution of resources or opportunities.
The analysis reveals three central shortcomings across the surveyed literature. First, research motivations are frequently ill-defined and lack normative reasoning: approximately 33% of analyzed works present multiple disparate motivations, 16% offer purely vague motivations or none at all, and 32% frame their work around technical model performance rather than ethical or normative justifications. Second, there is a severe mismatch between stated goals and the quantitative methods used to address them. While 21% of the publications cite allocational harms—such as unfair resume screening—to justify their work, only four papers actually measure or mitigate allocational outcomes directly, focusing instead on narrow token-level associations. Third, the literature largely fails to engage with established scholarship outside of computer science, such as sociolinguistics and critical race studies, resulting in inconsistent definitions of bias even among systems designed for the exact same task.
These findings indicate that existing technical benchmarks provide an incomplete and potentially misleading picture of system fairness. Focusing solely on convenient mathematical formulations without understanding how language reinforces societal inequalities creates significant organizational and compliance risks. For example, language models may inadvertently penalize non-standard dialects, such as African-American English, mischaracterizing benign communication as toxic or low quality. Relying on current debiasing techniques without deeper normative clarity leaves organizations vulnerable to deploying systems that actively perpetuate social marginalization.
To establish a more rigorous and effective path forward, the article outlines three key recommendations. First, research and development must ground technical evaluations in interdisciplinary literature that investigates the relationship between language and social power dynamics, treating representational harms as consequential in their own right. Second, researchers and practitioners must explicitly define their conceptualization of bias, clearly articulating what system behaviors are considered harmful, which groups are affected, and the specific ethical values underlying those determinations. Third, organizations must directly engage the lived experiences of affected communities through participatory methods, critically assessing power dynamics and evaluating whether certain automated systems should be built at all.
The conclusions of the article are based on a systematic evaluation of written text processing across standard academic and industry venues. While the findings are limited by the historical scope ending in early 2020 and do not evaluate spoken language systems, confidence in the diagnosed conceptual and methodological shortcomings remains exceptionally high, warranting deliberate caution among leaders relying on off-the-shelf debiasing benchmarks.
- Paper: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al. (2016). This paper establishes the foundational methods for quantifying and mitigating geometric gender bias in word embeddings, providing the primary technical paradigm evaluated and critiqued in the survey.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). This foundational work introduces the Word Embedding Association Test (WEAT) to demonstrate human-like prejudice in text embeddings, serving as a primary target of the survey's critique regarding vague conceptualizations of bias.
- Paper: Fairness and Abstraction in Sociotechnical Systems, Andrew D. Selbst et al. (2019). This paper identifies key abstraction traps in fair machine learning, supplying the sociotechnical critique of mathematical formalisms that directly motivates the survey's normative analysis.
- Paper: A Survey on Bias and Fairness in Machine Learning, Ninareh Mehrabi et al. (2019). This survey provides a comprehensive taxonomy of definitions and metrics for bias in machine learning, offering the broader fair-ML context that the survey analyzes within natural language processing.
- Paper: Big Data's Disparate Impact, Solon Barocas et al. (2016). This work establishes the legal and normative foundations of algorithmic disparate impact, supplying the interdisciplinary reasoning that the survey argues NLP literature urgently lacks.
- Paper: Model Cards for Model Reporting, Margaret Mitchell et al. (2019). This paper establishes standardized model documentation and disaggregated demographic reporting, laying the practical groundwork for interrogating model harms across affected communities.
- Paper: Fairness through awareness, Cynthia Dwork et al. (2012). This landmark paper formalizes individual fairness and task-specific similarity metrics, framing the algorithmic non-discrimination theory that underpins downstream bias research in language systems.
- Paper: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜, Emily M. Bender et al. (2021). This paper expands the survey's analysis of linguistic harms and social hierarchies to investigate the systemic data, power, and environmental risks inherent in scaling up massive language models.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). This work operationalizes the survey's call for normative clarity by creating a structured multidisciplinary taxonomy of 21 social and ethical harms arising across the lifecycle of language models.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). This survey provides a comprehensive landscape of modern large language model evaluations, putting into practice broader multidimensional auditing paradigms for ethics, bias, and reliability.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). This benchmark project implements extensive social bias and behavioral evaluations across diverse scales and tasks to test language models against human baselines.
