From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

Shangbin FengChan Young ParkYuhan LiuYulia Tsvetkov

article2023ACL413 citationsBest Paper Award

Establishes a quantitative framework to trace how political leanings in pretraining corpora embed ideological biases into language models and systematically degrade the fairness of downstream hate speech and misinformation classifiers across diverse demographic groups.

Listen

Online public discourse and media sources contain substantial ideological polarization and social bias. Because modern large language models rely extensively on broad web text, news, and social discussion boards for training data, these systems risk absorbing underlying societal biases. This dynamic is especially critical today as organizations increasingly deploy language models in high-stakes automated moderation, risk assessment, and fact-checking applications where unfair decisions can cause real-world harm.

The article aims to empirically quantify the political leanings of pretrained language models across social and economic dimensions. Furthermore, it demonstrates how ideological biases originating in pretraining data propagate through language models and create unfair disparities in high-stakes downstream tasks, specifically hate speech and misinformation detection.

To conduct this evaluation, the researchers developed a probing framework based on the 62-statement Political Compass test to map language models onto two-dimensional political coordinates representing social and economic values. They analyzed 14 major foundation models and further trained models on six distinct partisan corpora comprising left-, center-, and right-leaning news and social media text. The resulting partisan models were fine-tuned on benchmark datasets containing hundreds of thousands of examples to measure both aggregated performance and group-specific outcomes across targeted demographic identities and partisan media sources.

The findings show that pretrained models exhibit distinct political leanings, with older encoder models leaning more socially conservative while newer text-generation models lean socially liberal. Across all evaluated systems, ideological bias was significantly more pronounced on social issues (averaging a shift magnitude of 2.97) than on economic issues (averaging 0.87). Additionally, pretraining on partisan text shifted model coordinates accordingly, with post-2017 text driving models further toward ideological extremes. Crucially, while overall task accuracy remained superficially stable, partisan models showed sharp subgroup disparities: left-leaning models excelled at detecting hate speech against minority groups (such as LGBTQ+ and Black communities) but missed hate speech targeting dominant groups, whereas right-leaning models showed the opposite pattern. Similarly, models were notably less effective at detecting misinformation from news sources that aligned with their own political leaning.

These findings indicate that standard aggregate performance metrics conceal severe fairness risks and operational blind spots in artificial intelligence systems. Even when pretraining data is filtered to remove overtly toxic language, subtle ideological imbalances still produce downstream models with stark double standards. In deployment, these systemic skews expose organizations to significant compliance, reputational, and safety risks, as content moderation tools may systematically under-protect certain demographic groups while misidentifying partisan viewpoints as misinformation.

To mitigate these issues, decision-makers should avoid relying on single foundation models for sensitive classification tasks. The article demonstrates that a partisan ensemble approach—combining predictions from multiple models pretrained on diverse ideological viewpoints—substantially improves fairness and performance, raising the balanced accuracy on hate speech detection from roughly 88.6% to 90.2% and misinformation detection to 90.9%. Alternatively, practitioners can apply strategic domain-specific pretraining tailored to counter known context-specific vulnerabilities, while recognizing the associated data-curation trade-offs.

The conclusions should be interpreted with awareness of certain limitations, including the Western-centric nature of the two-axis political compass and the sensitivity of language model probing to prompt variations. Nonetheless, the high confidence in these findings underscores that developers must actively audit foundation models for ideological bias before deploying them in critical social applications.

arXiv: 2305.08283
Cover for From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Measuring the Political Leanings of LMs
  • 2.2 Measuring the Effect of LM’s Political Bias on Downstream Task Performance
  • 3 Experiment Settings
  • 4 Results and Analysis
  • 4.1 Political Bias of Language Models
  • 4.2 Political Leaning and Downstream Tasks
  • 5 Reducing the Effect of Political Bias
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Probing Language Models (cont.)
  • A.1 Encoder-Based LMs
  • A.2 Decoder-Based LMs
  • B Recall and Precision
  • C Experiment Details
  • D Stability Analysis
  • E Qualitative Analysis (cont.)
  • F Hyperparameter Settings
  • G Computational Resources
  • H Scientific Artifacts
  • ACL 2023 Responsible NLP Checklist
  • A For every submission:
  • B Did you use or create scientific artifacts?
  • C Did you run computational experiments?
  • D Did you use human annotators (e.g., crowdworkers) or research with human participants?

Knowls

  1. Knowl 1 — Two-Dimensional Political Leaning Probing Framework for Language Models

    model/method

    To evaluate the political leaning of language models without relying solely on individual political figures, models are probed using the 62 political propositions from the Political Compass test. The test maps a model's responses from the set {STRONG DISAGREE,DISAGREE,AGREE,STRONG AGREE}62\{\text{STRONG DISAGREE}, \text{DISAGREE}, \text{AGREE}, \text{STRONG AGREE}\}^{62} to a coordinate (ssoc,seco)∈[−10,10]×[−10,10](s_{\text{soc}}, s_{\text{eco}}) \in [-10, 10] \times [-10, 10], representing social values (ranging from −10-10 liberal/libertarian to +10+10 conservative/authoritarian) and economic values (ranging from −10-10 left to +10+10 right) through weighted summation.

    For encoder-only language models, probing uses mask-filling with the template: "Please respond to the following statement: [STATEMENT] I <MASK> with this statement." The model generates probability distributions for the mask token. Normalized probabilities over predefined positive tokens (e.g., agree, support, endorse) and negative tokens (e.g., disagree, refute, oppose) among the top 10 tokens are aggregated. If the positive sum exceeds the negative sum by a probability difference greater than 0.30.3, the response is mapped to STRONG AGREE\text{STRONG AGREE}; if below 0.30.3, it is mapped to AGREE\text{AGREE}. Disagreement is categorized analogously.

    For decoder-based and autoregressive language models, probing uses prompted generation with the template: "Please respond to the following statement: [STATEMENT] \n Your response:" across 10 random generation seeds. A BART-large model fine-tuned on MultiNLI serves as a zero-shot stance detector evaluating whether the generated completion entails agreement or disagreement with the proposition. Generations yielding stance detection confidence lower than 0.900.90 are filtered out as unclear, and the stance scores of the remaining responses are averaged.

  2. Knowl 2 — Disparities in Downstream Hate Speech Target Detection Across Partisan Language Models

    empirical result

    Language models pretrained on ideologically partisan corpora exhibit systematic fairness disparities when fine-tuned for hate speech classification across target demographic identity groups:

    • Left-leaning models (pretrained on left-leaning news or Reddit data) achieve higher accuracy in identifying hate speech directed at historically marginalized or minority demographic groups, including Black, Muslim, LGBTQ+, Asian, and Latinx communities.
    • Right-leaning models (pretrained on right-leaning news or Reddit data) achieve higher accuracy in identifying hate speech directed at dominant demographic groups, such as Men, White, and Christian targets.

    For example, on hate speech directed at LGBTQ+ targets, News-Left and Reddit-Left models achieve accuracies of 90.19%90.19\% and 89.96%89.96\%, respectively, compared to 88.91%88.91\% for News-Right and 88.43%88.43\% for Reddit-Right. Conversely, on hate speech directed at White targets, Reddit-Right and News-Right models achieve accuracies of 86.86%86.86\% and 86.35%86.35\%, whereas News-Left and Reddit-Left models achieve 86.22%86.22\% and 85.13%85.13\%.

  3. Knowl 3 — Partisan Asymmetries in Downstream Misinformation Detection

    empirical result

    When fine-tuned for fake news and misinformation detection on the PolitiFact dataset, language models pretrained on partisan corpora exhibit asymmetrical sensitivity depending on the political leaning of the news source:

    • Left-leaning language models detect misinformation originating from right-leaning media outlets (e.g., Fox News, Washington Examiner, Breitbart, Washington Times) with higher accuracy, but exhibit reduced sensitivity to misinformation from left-leaning outlets (e.g., CNN, The New York Times).
    • Right-leaning language models exhibit the reciprocal pattern, identifying falsehoods from left-leaning sources with higher accuracy while showing reduced sensitivity to misinformation from right-leaning sources.

    For instance, on statements originating from The New York Times, Reddit-Right and News-Right models achieve 86.71%86.71\% detection accuracy compared to 83.54%83.54\% for Reddit-Left. On statements from Fox News, the News-Left model achieves 93.10%93.10\% accuracy compared to 88.51%88.51\% for News-Right.

  4. Knowl 4 — Partisan Ensemble for Downstream Task Fairness and Performance Enhancement

    model/method

    To mitigate the political bias inherent in any single language model checkpoint, a partisan ensemble combines predictions from multiple language models that were independently pretrained on distinct partisan corpora (left-leaning and right-leaning across news and social media domains).

    When evaluated on downstream hate speech detection (across the Hate-Identity and Hate-Demographic benchmarks) and misinformation detection on PolitiFact, the partisan ensemble outperforms both the average single model and the best individual partisan model:

    Model Strategy Hate-Identity Hate-Demographic Misinformation
    BACC F1 BACC F1 BACC F1
    Average Uni-Model 88.58±0.288.58 \pm 0.2 81.01±0.781.01 \pm 0.7 89.83±0.489.83 \pm 0.4 83.35±0.583.35 \pm 0.5 87.24±1.287.24 \pm 1.2 86.54±1.486.54 \pm 1.4
    Best Uni-Model 88.78 81.77 90.19 83.82 88.61 88.15
    Partisan Ensemble 90.21 83.57 91.84 86.16 90.88 90.50

    BACC denotes balanced accuracy across classes. By ensembling predictions across diverse ideological representations, the framework achieves an increase of +1.43+1.43 to +2.27+2.27 in BACC over the best single model while mitigating target-specific and source-specific blind spots.

  5. Knowl 5 — Ideological Divergence and Axis-Dependent Bias in Pretrained Language Models

    empirical result

    Probing 14 pretrained language models (including BERT, RoBERTa, distilBERT, ALBERT, BART, GPT-2, GPT-3 variants, GPT-J, LLaMA, Alpaca, Codex, ChatGPT, and GPT-4) on the two-dimensional political compass demonstrates two primary patterns:

    1. Language model families occupy distinct ideological clusters: BERT and its variants lean more socially conservative/authoritarian, whereas the GPT series and instruction-tuned variants (e.g., ChatGPT, GPT-4, LLaMA) lean more socially liberal/libertarian.
    2. Pretrained models exhibit significantly stronger bias and variance on social issues (yy-axis) than on economic issues (xx-axis). The mean absolute score magnitude across evaluated models is 2.972.97 (standard deviation 1.291.29) on the social axis compared to 0.870.87 (standard deviation 0.840.84) on the economic axis.
  6. Knowl 6 — Domain-Specific Shifts in Language Model Political Coordinates Under Continued Pretraining

    empirical result

    Further pretraining base language models on partisan corpora systematically shifts their political compass coordinates in the direction of the training corpus's alignment (left-leaning data induces a liberal shift, while right-leaning data induces a conservative shift).

    Furthermore, the textual domain determines which ideological axis shifts most strongly:

    • Social media pretraining exerts a greater influence on social values: for RoBERTa, Reddit corpora induce an average coordinate shift of 1.601.60 along the social axis versus 0.610.61 along the economic axis.
    • News media pretraining exerts a greater influence on economic values: for RoBERTa, news corpora induce an average shift of 0.900.90 along the economic axis versus 0.640.64 along the social axis.
  7. Knowl 7 — Heightened Language Model Polarization Induced by Post-2017 Pretraining Corpora

    empirical result

    Partitioning partisan pretraining corpora into pre-Trump (pre-January 20, 2017) and post-Trump (post-January 20, 2017) periods reveals that language models further pretrained on post-2017 data shift further away from the ideological center (0,0)(0,0) compared to models trained on pre-2017 data.

    The resulting shifts in (seco,ssoc)(s_{\text{eco}}, s_{\text{soc}}) coordinates for RoBERTa are:

    • News Left: Δ=(−2.75,−1.24)\Delta = (-2.75, -1.24)
    • News Center: Δ=(−0.13,−1.03)\Delta = (-0.13, -1.03)
    • News Right: Δ=(1.63,1.03)\Delta = (1.63, 1.03)
    • Reddit Left: Δ=(0.75,−3.64)\Delta = (0.75, -3.64)
    • Reddit Center: Δ=(−0.50,−3.64)\Delta = (-0.50, -3.64)
    • Reddit Right: Δ=(−1.75,0.92)\Delta = (-1.75, 0.92)

    This indicates that models absorb the temporal escalation of socio-political polarization in post-2017 public discourse.

  8. Knowl 8 — Downstream Task Performance Across Partisan Pretrained RoBERTa Variants

    data/table

    Evaluating base RoBERTa and four partisan-pretrained RoBERTa variants on Hate Speech Detection (Hate-Identity and Hate-Demographic splits) and Misinformation Detection (PolitiFact) yields the following performance metrics:

    Model Hate-Identity Hate-Demographic Misinformation
    BACC F1 BACC F1 BACC F1
    RoBERTa 88.74±0.488.74 \pm 0.4 81.15±0.581.15 \pm 0.5 90.26±0.290.26 \pm 0.2 83.79±0.483.79 \pm 0.4 88.80±0.588.80 \pm 0.5 88.37±0.688.37 \pm 0.6
    RoBERTa-News-Left 88.75±0.288.75 \pm 0.2 81.44±0.281.44 \pm 0.2 90.19±0.490.19 \pm 0.4 83.53±0.883.53 \pm 0.8 88.61±0.488.61 \pm 0.4 88.15±0.588.15 \pm 0.5
    RoBERTa-Reddit-Left 88.78±0.388.78 \pm 0.3 81.77±0.3∗81.77 \pm 0.3^* 89.95±0.789.95 \pm 0.7 83.82±0.583.82 \pm 0.5 87.84±0.2∗87.84 \pm 0.2^* 87.25±0.2∗87.25 \pm 0.2^*
    RoBERTa-News-Right 88.45±0.388.45 \pm 0.3 80.66±0.6∗80.66 \pm 0.6^* 89.30±0.7∗89.30 \pm 0.7^* 82.76±0.182.76 \pm 0.1 86.51±0.4∗86.51 \pm 0.4^* 85.69±0.7∗85.69 \pm 0.7^*
    RoBERTa-Reddit-Right 88.34±0.2∗88.34 \pm 0.2^* 80.19±0.4∗80.19 \pm 0.4^* 89.87±0.789.87 \pm 0.7 83.28±0.4∗83.28 \pm 0.4^* 86.01±0.5∗86.01 \pm 0.5^* 85.05±0.6∗85.05 \pm 0.6^*

    BACC represents balanced accuracy across classes. Values represent mean ±\pm standard deviation across runs, with asterisks (∗*) denoting statistically significant differences compared to vanilla RoBERTa (p<0.05p < 0.05 via two-tailed tt-test). Left-leaning pretraining achieves higher overall task performance than right-leaning pretraining, and pretraining on right-leaning social media data (RoBERTa-Reddit-Right) consistently impairs overall performance.

  9. Knowl 9 — Partisan Pretraining Corpus Construction and Hate Speech Filtering

    experimental setup

    Controlled partisan pretraining corpora are constructed across two domains (news articles and social media posts) and three political leanings (left, center, right), producing six corpora of balanced sizes: {LEFT,CENTER,RIGHT}×{NEWS,REDDIT}\{\text{LEFT}, \text{CENTER}, \text{RIGHT}\} \times \{\text{NEWS}, \text{REDDIT}\}.

    • News data is collected from the POLITICS dataset and categorized by source bias ratings according to AllSides.
    • Social media data is collected from Reddit via the Pushshift API based on curated political subreddit lists, with non-political subreddits serving as the center corpus. Corpus post counts: Left contains 796,939 posts (average 44.50 tokens), Center contains 952,152 posts (average 34.67 tokens), and Right contains 934,452 posts (average 50.43 tokens).
    • To prevent injecting toxic content into the language models, all candidate pretraining corpora are filtered prior to training using a RoBERTa-based hate speech classifier fine-tuned on the TweetEval benchmark.
  10. Knowl 10 — Saturation of Ideological Polarization Under Scaled Partisan Pretraining

    empirical result

    Evaluating RoBERTa under extended continued pretraining with increasing epochs (from 10 to 50 epochs) and increasing corpus portions (from 20%20\% to 100%100\%) on partisan data does not produce hyperpartisan models reaching the extreme boundaries of the political compass ({−10,+10}\{-10, +10\}).

    While social value scores shift with initial training volume, the shift saturates rather than expanding indefinitely, and economic value scores remain near the center axis regardless of pretraining epoch count or corpus scale.

Coverage note — Omitted the stability sensitivity analysis across alternative paraphrases and prompts (Appendix D) and the full itemized list of 62 Political Compass test statements (Appendix Table 13), as these serve as validation details rather than core standalone scientific findings.

References

  1. 1.Alan Abramowitz and Jennifer McCoy. 2019. United states: Racial resentment, negative partisanship, and polarization in trump’s america. The ANNALS of the American Academy of Political and Social Science, 681(1):137–156.
  2. 2.Sohail Akhtar, Valerio Basile, and Viviana Patti. 2020. Modeling annotator perspective and polarized opinions to improve hate speech detection. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 8, pages 151–154.
  3. 3.Jacob Amedie. 2015. The impact of social media on society.
  4. 4.Aishwarya Anegundi, Konstantin Schulz, Christian Rauh, and Georg Rehm. 2022. Modelling cultural and socio-economic dimensions of political bias in German tweets. In Proceedings of the 18th Conference on Natural Language Processing (KONVENS 2022), pages 29–40, Potsdam, Germany. KONVENS 2022 Organizers.
  5. 5.Lisa P. Argyle, E. Busby, Nancy Fulda, Joshua Ronald Gubler, Christopher Michael Rytting, and David Wingate. 2022. Out of one, many: Using language models to simulate human samples. ArXiv, abs/2209.06899.
  6. 6.Eugene Bagdasaryan and Vitaly Shmatikov. 2022. Spinning language models: Risks of propaganda-as-a-service and countermeasures. In 2022 IEEE Symposium on Security and Privacy (SP), pages 1532–1532. IEEE Computer Society.
  7. 7.Pieter Ballon. 2014. Old and new issues in media economics. In The Palgrave handbook of European media policy, pages 70–95. Springer.
  8. 8.Yejin Bang, Nayeon Lee, Etsuko Ishii, Andrea Madotto, and Pascale Fung. 2021. Assessing political prudence of open-domain chatbots. In Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 548–555, Singapore and Online. Association for Computational Linguistics.
  9. 9.Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. Tweeteval: Unified benchmark and comparative evaluation for tweet classification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1644–1650.
  10. 10.John A Bargh. 1999. The cognitive monster: The case against the controllability of automatic stereotype effects.
  11. 11.Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift reddit dataset. In Proceedings of the international AAAI conference on web and social media, volume 14, pages 830–839.
  12. 12.Duncan Bell. 2014. What is liberalism? Political theory, 42(6):682–715.
  13. 13.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  14. 14.Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. O’Reilly Media, Inc.
  15. 15.Irene V Blair. 2002. The malleability of automatic stereotypes and prejudice. Personality and social psychology review, 6(3):242–261.
  16. 16.Charles Blattberg. 2001. Political philosophies and political ideologies. Public Affairs Quarterly, 15(3):193–217.
  17. 17.Su Lin Blodgett, Solon Barocas, Hal Daume III, and Hanna Wallach. 2020a. Language (technology) is power: A critical survey of “bias” in nlp. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476.
  18. 18.Su Lin Blodgett, Solon Barocas, Hal Daume III, and Hanna Wallach. 2020b. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  19. 19.Norberto Bobbio. 1996. Left and right: The significance of a political distinction. University of Chicago Press.
  20. 20.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in neural information processing systems, pages 4349–4357.
  21. 21.Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019. Nuanced metrics for measuring unintended bias with real data for text classification. In Companion proceedings of the 2019 world wide web conference, pages 491–500.
  22. 22.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  23. 23.Philip Bump. 2016. How likely are bernie sanders supporters to actually vote for donald trump? here are some clues. Washingtonpost. com.
  24. 24.Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency, pages 77–91. PMLR.
  25. 25.Michael A Cacciatore, Ashley A Anderson, Doo-Hun Choi, Dominique Brossard, Dietram A Scheufele, Xuan Liang, Peter J Ladwig, Michael Xenos, and Anthony Dudo. 2012. Coverage of emerging technologies: A comparison between print and online media. New media & society, 14(6):1039–1059.
  26. 26.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  27. 27.Yang Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022. On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 561–570.
  28. 28.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  29. 29.Elanor Colleoni, Alessandro Rozza, and Adam Arvidsson. 2014. Echo chamber or public sphere? predicting political orientation and measuring political homophily in twitter using big data. Journal of communication, 64(2):317–332.
  30. 30.Michael C Corballis and Ivan L Beale. 2020. The psychology of left and right. Routledge.
  31. 31.Jarret T Crawford, Mark J Brandt, Yoel Inbar, John R Chambers, and Matt Motyl. 2017. Social and economic ideologies differentially predict prejudice across the political spectrum, but social issues are most divisive. Journal of personality and social psychology, 112(3):383.
  32. 32.Aida Mostafazadeh Davani, Mark Dıaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  33. 33.Dorottya Demszky, Nikhil Garg, Rob Voigt, James Zou, Jesse Shapiro, Matthew Gentzkow, and Dan Jurafsky. 2019. Analyzing polarization in social media: Method and application to tweets on 21 mass shootings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2970–3005.
  34. 34.Patricia G Devine. 1989. Stereotypes and prejudice: Their automatic and controlled components. Journal of personality and social psychology, 56(1):5.
  35. 35.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  36. 36.Stanley Diamond and Eric Wolf. 2017. In search of the primitive: A critique of civilization. Routledge.
  37. 37.Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and mitigating unintended bias in text classification. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, pages 67–73.
  38. 38.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1286–1305, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  39. 39.Maeve Duggan. 2017. Online harassment 2017.
  40. 40.Denis Emelin, Ronan Le Bras, Jena D Hwang, Maxwell Forbes, and Yejin Choi. 2021. Moral stories: Situated reasoning about norms, intents, actions, and their consequences. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 698–718.
  41. 41.R. S. Enikolopov, Maria Petrova, and Ekaterina Zhuravskaya. 2019. Political effects of the internet and social media. Political Behavior: Cognition.
  42. 42.Hans Jurgen Eysenck. 1957. Sense and nonsense in psychology.
  43. 43.William Falcon and The PyTorch Lightning team. 2019. PyTorch Lightning.
  44. 44.Shangbin Feng, Zilong Chen, Wenqian Zhang, Qingyao Li, Qinghua Zheng, Xiaojun Chang, and Minnan Luo. 2021. Kgap: Knowledge graph augmented political perspective detection in news media. arXiv preprint arXiv:2108.03861.
  45. 45.Shangbin Feng, Zhaoxuan Tan, Zilong Chen, Ningnan Wang, Peisheng Yu, Qinghua Zheng, Xiaojun Chang, and Minnan Luo. 2022. PAR: Political actor representation learning with social context and expert knowledge. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  46. 46.Anjalie Field, Su Lin Blodgett, Zeerak Waseem, and Yulia Tsvetkov. 2021. A survey of race, racism, and anti-racism in nlp. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers).
  47. 47.Claudia Flores-Saviaga, Shangbin Feng, and Saiph Savage. 2022. Datavoidant: An ai system for addressing political data voids on social media. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2):1–29.
  48. 48.Daniel J Galvin. 2020. Party domination and base mobilization: Donald trump and republican party building in a polarized era. In The Forum, volume 18, pages 135–168. De Gruyter.
  49. 49.Kiran Garimella, Gianmarco De Francisci Morales, Aristides Gionis, and Michael Mathioudakis. 2018. Political discourse on social media: Echo chambers, gatekeepers, and the price of bipartisanship. In Proceedings of the 2018 world wide web conference, pages 913–922.
  50. 50.Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019. Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1161–1166, Hong Kong, China. Association for Computational Linguistics.
  51. 51.Sayan Ghosh, Dylan Baker, David Jurgens, and Vinodkumar Prabhakaran. 2021. Detecting cross-geographic biases in toxicity modeling on social media. In Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021), pages 313–328.
  52. 52.Allen Gindler. 2021. The theory of the political spectrum. Journal of Libertarian Studies, 24(2):24375.
  53. 53.Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Munoz Sanchez, Mugdha Pandya, and Adam Lopez. 2021. Intrinsic bias metrics do not correlate with application bias. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1926–1940.
  54. 54.Hila Gonen and Kellie Webster. 2020. Automatically identifying gender issues in machine translation using perturbations. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1991–1995, Online. Association for Computational Linguistics.
  55. 55.Lauren Guggenheim, S Mo Jang, Soo Young Bae, and W Russell Neuman. 2015. The dynamics of issue frame competition in traditional and social media. The ANNALS of the American Academy of Political and Social Science, 659(1):207–224.
  56. 56.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360.
  57. 57.Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning. Advances in neural information processing systems, 29.
  58. 58.Camille Harris, Matan Halevy, Ayanna Howard, Amy Bruckman, and Diyi Yang. 2022. Exploring the role of grammar and word choice in bias toward african american english (aae) in hate speech classification. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 789–798.
  59. 59.Charles R Harris, K Jarrod Millman, Stefan J Van Der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al. 2020. Array programming with numpy. Nature, 585(7825):357–362.
  60. 60.Alfred Hermida. 2016. Social media and the news. The SAGE handbook of digital journalism, pages 81–94.
  61. 61.Alfred Hermida, Fred Fletcher, Darryl Korell, and Donna Logan. 2012. Share, like, recommend: Decoding the social media news consumer. Journalism studies, 13(5-6):815–824.
  62. 62.David G Horrell. 2005. Paul among liberals and communitarians: models for christian ethics. Pacifica, 18(1):33–52.
  63. 63.Michael Hout and Christopher Maggio. 2021. Immigration, race & political polarization. Daedalus, 150(2):40–55.
  64. 64.Dirk Hovy and Anders Sogaard. 2015. Tagging performance correlates with author age. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 483–488, Beijing, China. Association for Computational Linguistics.
  65. 65.Kenneth Hudson. 1978. The language of modern politics. Springer.
  66. 66.Ben Hutchinson and Margaret Mitchell. 2019. 50 years of test (un) fairness: Lessons for machine learning. In Proceedings of the conference on fairness, accountability, and transparency, pages 49–58.
  67. 67.Hang Jiang, Doug Beeferman, Brandon Roy, and Deb Roy. 2022a. CommunityLM: Probing partisan worldviews from language models. In Proceedings of the 29th International Conference on Computational Linguistics, pages 6818–6826, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  68. 68.Hang Jiang, Doug Beeferman, Brandon Roy, and Deb Roy. 2022b. CommunityLM: Probing partisan worldviews from language models. In Proceedings of the 29th International Conference on Computational Linguistics, pages 6818–6826, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  69. 69.Xisen Jin, Francesco Barbieri, Brendan Kennedy, Aida Mostafazadeh Davani, Leonardo Neves, and Xiang Ren. 2021. On transferability of bias mitigation effects in language model fine-tuning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3770–3783.
  70. 70.Christopher D Johnston and Julie Wronski. 2015. Personality dispositions and political preferences across hard and easy issues. Political Psychology, 36(1):35–53.
  71. 71.Kenneth Joseph and Jonathan M. Morgan. 2020. When do word embeddings accurately reflect surveys on our beliefs about people? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.
  72. 72.Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2019. Learning the difference that makes a difference with counterfactually-augmented data. In International Conference on Learning Representations.
  73. 73.Sachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, and Yulia Tsvetkov. 2022. Language generation models can cause harm: So what can we do about it? an actionable survey. arXiv preprint arXiv:2210.07700.
  74. 74.Anna Sophie Kumpel, Veronika Karnowski, and Till Keyling. 2015. News sharing in social media: A review of current research on news sharing users, content, and networks. Social media+ society, 1(2):2056305115610141.
  75. 75.Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. Measuring bias in contextualized word representations. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing.
  76. 76.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations.
  77. 77.Anti-Defamation League. 2019. Online hate and harassment: The American experience.
  78. 78.Anti-Defamation League. 2021. The dangers of disinformation.
  79. 79.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Annual Meeting of the Association for Computational Linguistics.
  80. 80.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  81. 81.Chang Li and Dan Goldwasser. 2019. Encoding social information with graph convolutional networks forPolitical perspective detection in news media. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2594–2604, Florence, Italy. Association for Computational Linguistics.
  82. 82.Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. 2021. Contextualized perturbation for textual adversarial attack. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5053–5069, Online. Association for Computational Linguistics.
  83. 83.Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar. 2020. UNQOVERing stereotyping biases via underspecified questions. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3475–3489, Online. Association for Computational Linguistics.
  84. 84.Yizhi Li, Ge Zhang, Bohao Yang, Chenghua Lin, Anton Ragni, Shi Wang, and Jie Fu. 2022. Herb: Measuring hierarchical regional bias in pre-trained language models. In Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, pages 334–346.
  85. 85.Inna Lin, Lucille Njoo, Anjalie Field, Ashish Sharma, Katharina Reinecke, Tim Althoff, and Yulia Tsvetkov. 2022. Gendered mental health stigma in masked language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.
  86. 86.Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2020. Does gender matter? towards fairness in dialogue systems. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4403–4416, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  87. 87.Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Lili Wang, and Soroush Vosoughi. 2021. Mitigating political bias in language models through reinforced calibration. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 14857–14866.
  88. 88.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692.
  89. 89.Yujian Liu, Xinliang Frederick Zhang, David Wegsman, Nicholas Beauchamp, and Lu Wang. 2022a. POLITICS: Pretraining with same-story article comparison for ideology prediction and stance detection. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1354–1374, Seattle, United States. Association for Computational Linguistics.
  90. 90.Yujian Liu, Xinliang Frederick Zhang, David Wegsman, Nicholas Beauchamp, and Lu Wang. 2022b. POLITICS: Pretraining with same-story article comparison for ideology prediction and stance detection. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1354–1374, Seattle, United States. Association for Computational Linguistics.
  91. 91.Peter Mair. 2007. Left–right orientations.
  92. 92.Aibek Makazhanov and Davood Rafiei. 2013. Predicting political preference of twitter users. In Proceedings of the 2013 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, pages 298–305.
  93. 93.Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. Hatexplain: A benchmark dataset for explainable hate speech detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 14867–14875.
  94. 94.Brian Patrick Mitchell. 2007. Eight ways to run the country: A new and revealing look at left and right. Greenwood Publishing Group.
  95. 95.Eni Mustafaraj and Panagiotis Takis Metaxas. 2011. What edited retweets reveal about online political discourse. In Workshops at the Twenty-Fifth AAAI Conference on Artificial Intelligence.
  96. 96.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371, Online. Association for Computational Linguistics.
  97. 97.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel Bowman. 2020. Crows-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1953–1967.
  98. 98.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  99. 99.Ji Ho Park, Jamin Shin, and Pascale Fung. 2018. Reducing gender bias in abusive language detection. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2799–2804, Brussels, Belgium. Association for Computational Linguistics.
  100. 100.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  101. 101.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  102. 102.Fabio Petroni, Tim Rocktaschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473.
  103. 103.Daniel Preotiuc-Pietro, Ye Liu, Daniel Hopkins, and Lyle Ungar. 2017. Beyond binary labels: Political ideology prediction of Twitter users. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 729–740, Vancouver, Canada. Association for Computational Linguistics.
  104. 104.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  105. 105.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  106. 106.Lee Rainie, Aaron Smith, Kay Lehman Schlozman, Henry Brady, Sidney Verba, et al. 2012. Social media and political engagement. Pew Internet & American Life Project, 19(1):2–13.
  107. 107.Cameron Raymond, Isaac Waller, and Ashton Anderson. 2022. Measuring alignment of online grassroots political communities with political campaigns. In Proceedings of the International AAAI Conference on Web and Social Media, volume 16, pages 806–816.
  108. 108.Milton Rokeach. 1973. The nature of human values. Free press.
  109. 109.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
  110. 110.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  111. 111.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  112. 112.Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2022. On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning. arXiv preprint arXiv:2212.08061.
  113. 113.Qinlan Shen and Carolyn Rose. 2021. What sounds “right” to me? experiential factors in the perception of political ideology. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume.
  114. 114.Ryan Steed, Swetasudha Panda, Ari Kobren, and Michael Wick. 2022. Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language Models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  115. 115.Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019. Mitigating gender bias in natural language processing: Literature review. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  116. 116.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  117. 117.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  118. 118.Megan Trudell. 2016. Sanders, trump and the us working class. International Socialism.
  119. 119.Tom Utley. 2001. I’m v. right-wing, says the bbc, but it’s not that simple.
  120. 120.Sebastian Valenzuela, Yonghwan Kim, and Homero Gil de Zuniga. 2012. Social networks that matter: Exploring the role of political discussion for online political participation. International Journal of Public Opinion Research, 24:163–184.
  121. 121.Alcides Velasquez. 2012. Social media and online political discussion: The effect of cues and informational cascades on participation in online political communities. New Media & Society, 14(8):1286–1303.
  122. 122.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  123. 123.Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. Adversarial glue: A multi-task benchmark for robustness evaluation of language models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  124. 124.William Yang Wang. 2017. “liar, liar pants on fire”: A new benchmark dataset for fake news detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers).
  125. 125.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers).
  126. 126.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  127. 127.Michael Yoder, Lynnette Ng, David West Brown, and Kathleen Carley. 2022. How hate speech varies by target identity: A computational analysis. In Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL).
  128. 128.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR.
  129. 129.Wenqian Zhang, Shangbin Feng, Zilong Chen, Zhenyu Lei, Jundong Li, and Minnan Luo. 2022. KCD: Knowledge walks and textual cues enhanced political perspective detection in news media. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  130. 130.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers).
  131. 131.Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. 2015 IEEE International Conference on Computer Vision (ICCV), pages 19–27.

Citation

MLA
Feng, S., et al. “From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 11737–62, https://doi.org/10.18653/v1/2023.acl-long.656.
APA
Feng, S., Park, C. Y., Liu, Y., & Tsvetkov, Y. (2023). From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11737–11762. https://doi.org/10.18653/v1/2023.acl-long.656
Chicago
Feng, S., C. Y. Park, Y. Liu, and Y. Tsvetkov. 2023. “From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11737–62. https://doi.org/10.18653/v1/2023.acl-long.656.
Harvard
Feng, S. et al. (2023) “From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11737–11762. Available at: https://doi.org/10.18653/v1/2023.acl-long.656.
Vancouver
1. Feng S, Park CY, Liu Y, Tsvetkov Y (2023) From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11737–11762

BibTeX

@inproceedings{feng-etal-2023-pretraining,
    title = "From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair {NLP} Models",
    author = "Feng, Shangbin  and
      Park, Chan Young  and
      Liu, Yuhan  and
      Tsvetkov, Yulia",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.656/",
    doi = "10.18653/v1/2023.acl-long.656",
    pages = "11737--11762"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/