From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
Shangbin FengChan Young ParkYuhan LiuYulia Tsvetkov
Establishes a quantitative framework to trace how political leanings in pretraining corpora embed ideological biases into language models and systematically degrade the fairness of downstream hate speech and misinformation classifiers across diverse demographic groups.
Online public discourse and media sources contain substantial ideological polarization and social bias. Because modern large language models rely extensively on broad web text, news, and social discussion boards for training data, these systems risk absorbing underlying societal biases. This dynamic is especially critical today as organizations increasingly deploy language models in high-stakes automated moderation, risk assessment, and fact-checking applications where unfair decisions can cause real-world harm.
The article aims to empirically quantify the political leanings of pretrained language models across social and economic dimensions. Furthermore, it demonstrates how ideological biases originating in pretraining data propagate through language models and create unfair disparities in high-stakes downstream tasks, specifically hate speech and misinformation detection.
To conduct this evaluation, the researchers developed a probing framework based on the 62-statement Political Compass test to map language models onto two-dimensional political coordinates representing social and economic values. They analyzed 14 major foundation models and further trained models on six distinct partisan corpora comprising left-, center-, and right-leaning news and social media text. The resulting partisan models were fine-tuned on benchmark datasets containing hundreds of thousands of examples to measure both aggregated performance and group-specific outcomes across targeted demographic identities and partisan media sources.
The findings show that pretrained models exhibit distinct political leanings, with older encoder models leaning more socially conservative while newer text-generation models lean socially liberal. Across all evaluated systems, ideological bias was significantly more pronounced on social issues (averaging a shift magnitude of 2.97) than on economic issues (averaging 0.87). Additionally, pretraining on partisan text shifted model coordinates accordingly, with post-2017 text driving models further toward ideological extremes. Crucially, while overall task accuracy remained superficially stable, partisan models showed sharp subgroup disparities: left-leaning models excelled at detecting hate speech against minority groups (such as LGBTQ+ and Black communities) but missed hate speech targeting dominant groups, whereas right-leaning models showed the opposite pattern. Similarly, models were notably less effective at detecting misinformation from news sources that aligned with their own political leaning.
These findings indicate that standard aggregate performance metrics conceal severe fairness risks and operational blind spots in artificial intelligence systems. Even when pretraining data is filtered to remove overtly toxic language, subtle ideological imbalances still produce downstream models with stark double standards. In deployment, these systemic skews expose organizations to significant compliance, reputational, and safety risks, as content moderation tools may systematically under-protect certain demographic groups while misidentifying partisan viewpoints as misinformation.
To mitigate these issues, decision-makers should avoid relying on single foundation models for sensitive classification tasks. The article demonstrates that a partisan ensemble approach—combining predictions from multiple models pretrained on diverse ideological viewpoints—substantially improves fairness and performance, raising the balanced accuracy on hate speech detection from roughly 88.6% to 90.2% and misinformation detection to 90.9%. Alternatively, practitioners can apply strategic domain-specific pretraining tailored to counter known context-specific vulnerabilities, while recognizing the associated data-curation trade-offs.
The conclusions should be interpreted with awareness of certain limitations, including the Western-centric nature of the two-axis political compass and the sensitivity of language model probing to prompt variations. Nonetheless, the high confidence in these findings underscores that developers must actively audit foundation models for ideological bias before deploying them in critical social applications.
- Paper: Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al. (2016). This seminal study establishes how statistical language representations implicitly absorb human societal prejudices from web corpora, providing the foundational basis for auditing ideological bias in pretrained models.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). It introduces standard benchmark methodologies to quantify stereotypical bias in language models alongside task performance, which directly underpins the source paper's evaluation of representational disparities.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). This critical survey establishes the normative taxonomy of harms and conceptual frameworks for evaluating algorithmic bias in NLP that the source applies to political leanings and downstream tasks.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). It demonstrates how generative language models can reflect fine-grained political and demographic viewpoints, laying groundwork for probing models along ideological dimensions.
- Paper: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜, Emily M. Bender et al. (2021). This paper articulates the systemic risks of training massive language models on uncurated web datasets that overrepresent dominant viewpoints, framing the data-to-model bias trail investigated in the source.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). It formalizes benchmarks and auditing methodologies for measuring neural toxicity degeneration originating from pretraining corpora in large language models.
- Paper: “Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection, William Yang Wang (2017). It creates the standard benchmark dataset for political statement truthfulness, which directly informs downstream evaluations of political bias in automated misinformation detection.
- Paper: Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter, Zeerak Waseem et al. (2016). It establishes core classification frameworks and annotation guidelines for hate speech detection, a primary downstream task probed for partisan fairness disparities.
- Paper: Mitigating Unwanted Biases with Adversarial Learning, Brian Hu Zhang et al. (2018). This work develops fundamental adversarial learning mechanisms to mitigate unwanted demographic biases in machine learning representations while preserving task utility.
- Paper: Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey, Bonan Min et al. (2021). It surveys the pretraining-fine-tuning pipeline across Transformer language models, providing the broader architectural context for tracking how biases propagate across training stages.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). This work extends the investigation of model ideological leanings by measuring how language model response distributions systematically align with specific human demographic and political subgroups.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). It incorporates fairness and social bias auditing into an overarching, multi-dimensional trustworthiness evaluation framework for foundation models across diverse operational settings.
- Paper: Towards Understanding Sycophancy in Language Models, Mrinank Sharma et al. (2023). It explores how feedback alignment creates sycophancy where models actively tailor opinions to confirm user biases, building on findings about latent political perspectives in pretrained systems.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). This paper demonstrates how implicit input biases lead language models to generate unfaithful step-by-step explanations, continuing the study of operational blind spots in model reasoning.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). It investigates how downstream fine-tuning can rapidly undermine model alignment safeguards, complementing the source paper's findings on data-induced model vulnerabilities.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench provides a standardized automated red-teaming benchmark to systematically test safety refusal behaviors against malicious inputs and downstream harm risks.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). It evaluates the systematic biases and reliability limitations of using large language models as automated evaluators across downstream classification and judgment tasks.
- Paper: Large Language Models are Geographically Biased, Rohin Manvi et al. (2024). It broadens the empirical analysis of pretraining corpus skews by evaluating how foundation models exhibit systemic geospatial and cultural biases.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). It critiques conventional pretraining data filtering paradigms at scale, providing a direct counterpoint to discussions around curation trade-offs in mitigating pretraining bias.
