Unintended Impacts of LLM Alignment on Global Representation

Michael J. RyanWilliam Barr HeldDiyi Yang

article2024ACL64 citations

Reveals how standard alignment techniques like RLHF and DPO introduce substantial global representation biases by widening performance disparities across English dialects and skewing model viewpoints toward US perspectives, while simultaneously improving non-English multilingual capabilities.

Listen

Large language models are rapidly expanding to hundreds of millions of users worldwide after undergoing alignment procedures such as supervised fine-tuning and preference tuning to make them helpful and safe. However, alignment reflects specific human preferences that are rarely universal, raising concerns about which cultural norms and linguistic variations are prioritized. The article evaluates how standard alignment workflows impact large language models across three global dimensions: regional English dialects, multilingual performance, and country-level opinions.

The authors conducted a comprehensive evaluation across nine open-source language models, tracking performance changes across their base, supervised fine-tuning, and preference-tuned stages. The analysis tested models on task-oriented dialogue intent prediction across American, Indian, and Nigerian English dialects; evaluated reading comprehension and question-answering across nine typologically diverse languages; and measured alignment with national opinion surveys across seven diverse countries. The authors also developed a dataset of 554 country-specific subjective questions to probe the geographical biases learned by an open-source reward model.

The investigation produced four critical findings. First, while alignment improved intent recognition across all English dialects, it sharply widened performance disparities: the performance gap between American English and non-Western dialects grew from approximately 1% in base models to up to 17.1% after alignment. Second, fine-tuning unexpectedly boosted multilingual comprehension across most evaluated languages because the fine-tuning datasets unintentionally included roughly 13% non-English text. Third, alignment consistently shifted model responses closer to United States public opinion, increasing divergence from countries such as Jordan, China, and Nigeria by 2% to 5% while remaining aligned with Western nations. Finally, the evaluated reward model favored Western nations and ranked nearly all other countries below the United States, though this bias did not directly propagate into the fine-tuned language model because country-opinion topics were absent during preference training.

These results demonstrate that standard alignment practices inadvertently introduce regional disparities that favor Western and American norms, which may hinder the equitable global adoption of artificial intelligence. To address these disparities, developers should implement transparent reporting on annotator demographics, data curation, and domain selections throughout the alignment process. Organizations should also intentionally incorporate diverse multilingual and cross-dialect data into fine-tuning, as even modest inclusions generate substantial capability gains across languages without compromising primary language performance.

The findings are bounded by certain limitations, such as reliance on publicly released model checkpoints where some intermediate training data was unavailable, and evaluation across a focused set of downstream benchmarks. Nevertheless, the consistency of the results across distinct model families provides high confidence that alignment decisions heavily dictate global representation, making deliberate and inclusive data selection essential for future deployments.

arXiv: 2402.15018
Cover for Unintended Impacts of LLM Alignment on Global Representation

Abstract

Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference Optimization (DPO). Current evaluations of these procedures focus on benchmarks of instruction following, reasoning, and truthfulness. However, human preferences are not universal, and aligning to specific preference sets may have unintended effects. We explore how alignment impacts performance along three axes of global representation: English dialects, multilingualism, and opinions from and about countries worldwide. Our results show that current alignment procedures create disparities between English dialects and global opinions. We find alignment improves capabilities in several languages. We conclude by discussing design decisions that led to these unintended impacts and recommendations for more equitable preference tuning. We make our code and data publicly available on Github^1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Alignment Process
  • 4 Global Representation: English Dialects
  • 5 Global Representation: Languages
  • 6 Global Representation: Opinions
  • 6.1 Reward Model Probing
  • 7 Discussion and Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Dataset Examples
  • B Language Model Setting
  • C AskReddit Dataset
  • D Language Identification on SFT Data
  • E Ask Reddit Full Results
  • F Additional Models
  • F.1 Dialect
  • F.2 Languages
  • F.3 Opinions

Knowls

  1. Knowl 1 — Stage-wise evaluation of alignment across global representation

    experimental setup

    The study evaluates alignment by comparing language models at three checkpoints: a pretrained base model, a supervised fine-tuned (SFT) model, and a preference-tuned (PT) model. SFT trains on instruction–completion examples, whereas PT uses ranked preferences through RLHF/PPO, DPO, or RLAIF. The primary comparison contains nine open 7B models from the Llama 2 and Mistral v0.1 families: Llama 2, Llama 2 Chat, Tülu 2 SFT, Tülu 2 DPO, Mistral, Mistral SFT, Zephyr-7B-beta, OpenChat, and Starling LM. The released checkpoints use different training mixtures: Tülu and Zephyr use DPO, Llama Chat and Starling use PPO-based preference optimization, and Starling also uses AI-generated feedback. The study measures alignment effects on English dialect intent prediction, multilingual question answering and reading comprehension, and global opinion agreement. Generation experiments use greedy decoding, and all models are evaluated with 8-bit quantization on an NVIDIA A6000 GPU. Because the authors compare released checkpoints rather than retraining models with controlled interventions, the stage comparisons measure practical downstream effects rather than isolating a fully causal effect of one training decision.

  2. Knowl 2 — Alignment increases English-dialect disparity while improving intent prediction

    empirical result

    On the Multi-dialect Dataset of Dialogues, the models infer a hidden intent from task-oriented conversations in US American, Indian, and Nigerian English. A prediction counts as correct only when the model produces the intended answer and does not produce any distractor word. Before alignment, the three dialects have similar accuracy—approximately 5% for Llama-based models and 8% for Mistral-based models—but alignment improves US English more than the other dialects. For Mistral followed by its SFT checkpoint, accuracy increases by 15.2% for Indian English, 20.3% for Nigerian English, and 29.3% for US English. For Mistral followed by OpenChat, the corresponding gains are 20.3%, 27.9%, and 36.3%. Across the evaluated models, the dialect performance disparity grows from about 1% before alignment to as much as 17.1% after alignment. Preference tuning usually has smaller effects than SFT, although the Mistral-SFT-to-Zephyr transition significantly benefits US English; the OpenChat-to-Starling transition is not statistically significant, with its largest decrease occurring for Nigerian English.

  3. Knowl 3 — English-oriented alignment often improves multilingual comprehension

    empirical result

    The study evaluates nine languages—Arabic, Bengali, English, Finnish, Indonesian, Korean, Russian, Swahili, and Telugu—using two tasks. TyDiQA Gold Passage is evaluated in a 1-shot setting with greedy extraction of an answer span, scored by CFMScore, a trained answer-equivalence classifier based on F1 and related text features. Belebele is evaluated as multiple-choice reading comprehension using the model probabilities of answer letters (A)–(D). Despite the alignment datasets being intended mainly for English assistants, SFT and preference tuning produce significant reading-comprehension improvements in most tested languages, with no significant reading-comprehension decrease. Tülu and Starling improve TyDiQA in most languages, whereas Zephyr decreases significantly in six of the nine languages. Bengali is the only language that worsens for every primary model: the decreases are 12.7% for Llama Chat, 8.2% for Tülu, 9.7% for Zephyr, and 0.8% for Starling.

  4. Knowl 4 — Non-English examples in SFT data account for much of the multilingual gain

    empirical result

    Language identification of the Tülu SFT mixture shows that 13.1% of its examples are non-English, despite the intended emphasis on English. The largest identified non-English sources are ShareGPT with 141,970 samples, FlanV2 with 16,801 samples, and Open Assistant with 11,441 samples. Tülu contains 71 Bengali examples, corresponding to only 0.000058% of its data, while the UltraChat SFT data used for Zephyr contains zero Bengali examples and is 99.9% English overall. The authors associate Tülu’s multilingual SFT data with its improvements on most languages in both TyDiQA and Belebele, and associate the near-monolingual UltraChat data with weaker multilingual transfer. The OpenChat SFT model also shows substantial multilingual gains, plausibly because its training sources overlap with ShareGPT and Open Assistant. This analysis suggests that a relatively small multilingual component can improve several languages without sacrificing the largest English gains.

  5. Knowl 5 — Preference tuning shifts model opinions toward the United States

    empirical result

    The study measures agreement with country-level opinions using 245 questions from GlobalOpinionsQA that have representative answers from the USA, China, Jordan, Brazil, Nigeria, Germany, and Australia. For each question, the model produces probabilities over the answer choices. If AM(c)A_M(c) is the average agreement of model MM with country cc, agreement is computed as 11 minus the Jensen–Shannon divergence between the model’s answer distribution and the empirical distribution for country cc. The analysis uses the USA-over-country gap GM(c)=AM(USA)−AM(c)G_M(c)=A_M(\mathrm{USA})-A_M(c). Every primary alignment path increases GM(c)G_M(c) for Jordan, China, and Nigeria, meaning that the aligned model becomes relatively closer to US opinions than to opinions from the Middle East, Asia, and Sub-Saharan Africa. For the Llama-to-Llama-Chat path, the gap changes from 0.3% to 4.5% for Jordan, from 1.4% to 3.1% for China, and from −2.5% to 3.5% for Nigeria. The corresponding relative agreement with Germany and Australia does not change significantly. All evaluated models switch from agreeing more with Nigeria than with the USA before alignment to agreeing more with the USA after alignment.

  6. Knowl 6 — Construction of a country-opinion probing dataset

    data/table

    The authors create a dataset of 554 subjective country questions from r/AskReddit to probe opinions about countries rather than opinions originating from countries. They initially collect 957 questions using searches such as “Which Country,” “What Country,” “Best Country,” and “Worst Country,” then remove duplicates, factual questions, and questions that cannot be answered with a specific country. Two authors independently label each question as positive or negative according to whether the question asks for something favorable or unfavorable about a country; their Cohen’s κ\kappa is 0.963, with 10 disagreements resolved by discussion. ChatGPT-generated answer templates are manually validated, and the questions are divided into 11 categories.

    Could not parse LaTeX table

    The dataset is used analytically, not as training or alignment data, and is paired with 181 countries having populations above 250,000.

  7. Knowl 7 — Starling’s reward model strongly reflects US country preferences

    empirical result

    The Starling 7B Reward Model is evaluated on every combination of the 554 country-opinion questions and 181 candidate countries. For each question, the reward model scores the templated completion for each country, ranks the countries by score, reverses the ranking for negative questions so that a better rank always means a more favorable evaluation, and averages ranks across questions. Its country ranking has a Spearman correlation of 0.926 with Gallup’s 2017 US-citizen country preferences (p=1.78×10−9p=1.78\times10^{-9}) and 0.849 with the 2023 preferences (p=1.12×10−6p=1.12\times10^{-6}), using the 19 countries shared with the Gallup comparisons. The reward model rates 99.4% of countries other than the USA more negatively than the USA on average. Across all 181 countries, Morocco, the USA, Slovenia, and New Zealand receive the most favorable average rankings, while Western Sahara, North Korea, Turkmenistan, and the Central African Republic receive the least favorable rankings. The lowest-ranked regions are disproportionately in Central Africa and the Middle East; the bottom 60% of the reward-model ranking contains 72.7% of the world’s population.

  8. Knowl 8 — The Starling reward signal does not propagate to out-of-distribution country preferences

    empirical result

    To test whether the Starling reward model’s country preferences appear in the language model it tuned, the authors compute country rankings for all models using perplexity on the same 554 country-opinion prompts and compare rankings with Spearman correlation. Models within the same base-model family produce highly similar rankings: Llama, Llama Chat, Tülu SFT, and Tülu DPO cluster together, as do Mistral, Mistral SFT, Zephyr, OpenChat, and Starling LM, with within-family correlations reaching approximately 0.99. In contrast, the Starling reward model has substantially lower correlations with the language models, including only 0.59 with Starling LM; its correlations with the other models range from 0.51 to 0.67 for the Llama family and from 0.52 to 0.60 for the Mistral family. Thus, the Starling language model does not inherit the reward model’s strong US-like country ranking on these out-of-distribution questions. The authors infer that when opinionated country questions are absent from preference-tuning data, pretraining and SFT data dominate the model’s behavior and the reward signal does not substantially steer it in this domain.

  9. Knowl 9 — Additional Chinese-developed models reproduce the global-opinion and dialect patterns

    empirical result

    The authors validate the main findings with Qwen1.5-7B versus Qwen1.5-7B-Chat and Yi-6B versus Yi-6B-Chat. Qwen uses PPO-based RLHF, while Yi undergoes SFT without preference tuning; both model families were developed by Chinese organizations rather than the Western research groups responsible for the primary models. On MD3, Yi’s alignment gains are largest for Indian English at 7.41% and US English at 3.98%, compared with 2.39% for Nigerian English. Qwen-Chat initially appears to lose performance because it generates explanatory preambles that exceed a 10-token limit; when evaluated with its chat template and a 100-token limit, preference tuning improves intent accuracy and US English remains its strongest dialect. Qwen improves Belebele performance in nearly every tested language except Telugu, whereas Yi has no significant Belebele improvements and shows some TyDiQA decreases. On GlobalOpinionsQA, both families generally move farther from Jordanian, Chinese, and Nigerian opinions relative to US opinions while remaining close to their previous relative similarity with Germany and Australia. The same pattern therefore occurs even when alignment is performed by non-Western developers.

  10. Knowl 10 — Scope limitations prevent causal attribution to individual alignment decisions

    limitation

    The study uses nine publicly released model checkpoints rather than retraining models under controlled changes. Because the Llama 2 SFT checkpoint is unavailable, the separate effects of Llama 2 SFT and RLHF cannot be disentangled. More generally, comparisons across released checkpoints may reflect differences in architectures, datasets, annotators, prompts, or optimization procedures, so they do not establish that a single alignment component caused each observed disparity. The evaluations cover four downstream tasks—dialect intent prediction, extractive question answering, reading comprehension, and global opinion agreement—and therefore may miss effects on other capabilities or language varieties. The authors also note that improved multilingual performance is treated as positive under the assumption that deployment benefits speakers, although multilingual systems can also enable misinformation, surveillance, or targeted harassment; the subjective opinion datasets have no objectively correct answers and should be used for analysis rather than preference training.

Coverage note — The paper’s practitioner recommendations, broader governance discussion, and detailed prompt examples were not made separate knowls because they summarize the empirical findings or provide implementation context rather than adding independent load-bearing results.

References

  1. 1.AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai. 2024. Yi: Open foundation models by 01.ai.
  2. 2.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  3. 3.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Daniel Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback.
  4. 4.Michiel A. Bakker, Martin J Chadwick, Hannah Sheahan, Michael Henry Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matthew Botvinick, and Christopher Summerfield. 2022. Fine-tuning language models to find agreement among humans with diverse preferences. In Advances in Neural Information Processing Systems.
  5. 5.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2023. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. arXiv preprint arXiv:2308.16884.
  6. 6.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  7. 7.Rishi Bommasani, Kevin Klyman, Shayne Longpre, Sayash Kapoor, Nestor Maslej, Betty Xiong, Daniel Zhang, and Percy Liang. 2023. The foundation model transparency index.
  8. 8.Megan Brenan. 2023. Canada, britain favored most in u.s.; russia, n. korea least.
  9. 9.Yang Trista Cao, Anna Sotnikova, Jieyu Zhao, Linda X. Zou, Rachel Rudinger, and Hal Daume III au2. 2023. Multilingual large language models leak human stereotypes across language boundaries.
  10. 10.Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca.
  11. 11.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.
  12. 12.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics.
  13. 13.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge.
  14. 14.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  15. 15.Together Computer. 2023. Redpajama: an open dataset for training large language models.
  16. 16.Luigi Daniele and Suphavadeeprasit. 2023. Amplify-instruct: Synthetically generated diverse multi-turn conversations for effecient llm training. arXiv preprint arXiv:(comming soon).
  17. 17.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. Llm.int8(): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339.
  18. 18.Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023. Queer people are people first: Deconstructing sexual identity stereotypes in large language models.
  19. 19.Esin Durmus, Karina Nyugen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Daniel Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2023. Towards measuring the representation of subjective global opinions in language models.
  20. 20.Jacob Eisenstein, Vinodkumar Prabhakaran, Clara Rivera, Dora Demszky, and Devyani Sharma. 2023. Md3: The multi-dialect dataset of dialogues. In Interspeech.
  21. 21.Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. 2023. Gpts are gpts: An early look at the labor market impact potential of large language models. arXiv preprint arXiv:2303.10130.
  22. 22.Emilio Ferrara. 2023. Should chatgpt be biased? challenges and risks of bias in large language models. First Monday.
  23. 23.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  24. 24.Thomas Krendl Gilbert, Nathan Lambert, Sarah Dean, Tom Zick, and Aaron Snoswell. 2023. Reward reports for reinforcement learning.
  25. 25.Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023. The political ideology of conversational ai: Converging evidence on chatgpt’s pro-environmental, left-libertarian orientation.
  26. 26.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  27. 27.Tom Hosking, Phil Blunsom, and Max Bartolo. 2023. Human feedback is not gold standard.
  28. 28.Jing Huang and Diyi Yang. 2023. Culturally aware natural language inference. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 7591–7609, Singapore. Association for Computational Linguistics.
  29. 29.Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. Camels in a changing climate: Enhancing lm adaptation with tulu 2.
  30. 30.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  31. 31.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016. Bag of tricks for efficient text classification. arXiv preprint arXiv:1607.01759.
  32. 32.Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2023. Understanding the effects of rlhf on llm generalisation and diversity.
  33. 33.Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in large language models. In Proceedings of The ACM Collective Intelligence Conference, CI ’23, page 12–24, New York, NY, USA. Association for Computing Machinery.
  34. 34.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations – democratizing large language model alignment.
  35. 35.Nathan Lambert, Thomas Krendl Gilbert, and Tom Zick. 2023. The history and risks of reinforcement learning and human feedback.
  36. 36.Zongxia Li, Ishani Mondal, Yijun Liang, Huy Nghiem, and Jordan Boyd-Graber. 2024. Cfmatch: Aligning automated answer equivalence evaluation with expert judgments for open-domain question answering.
  37. 37.Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium". 2023. Openorca: An open dataset of gpt augmented flan reasoning traces. https://https://huggingface.co/Open-Orca/OpenOrca.
  38. 38.Andreas Liesenfeld, Alianda Lopez, and Mark Dingemanse. 2023. Opening up chatgpt: Tracking openness, transparency, and accountability in instruction-tuned text generators. In Proceedings of the 5th International Conference on Conversational User Interfaces, CUI ’23, New York, NY, USA. Association for Computing Machinery.
  39. 39.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  40. 40.Yong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao, Jianmeng Liu, Jipeng Zhang, Rui Pan, Haoxiang Wang, Wenbin Hu, Hanning Zhang, Hanze Dong, Renjie Pi, Han Zhao, Nan Jiang, Heng Ji, Yuan Yao, and Tong Zhang. 2024. Mitigating the alignment tax of rlhf.
  41. 41.Gabrielle Kaili-May Liu. 2023. Perspectives on the social impacts of reinforcement learning with human feedback.
  42. 42.Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, Xinyi Wu, Enrico Shippole, Kurt Bollacker, Tongshuang Wu, Luis Villa, Sandy Pentland, and Sara Hooker. 2023. The data provenance initiative: A large scale audit of dataset licensing & attribution in ai.
  43. 43.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency, pages 220–229.
  44. 44.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371, Online. Association for Computational Linguistics.
  45. 45.Shuyo Nakatani. 2010. Language detection library for java.
  46. 46.Tarek Naous, Michael J. Ryan, Alan Ritter, and Wei Xu. 2023. Having beer after prayer? measuring cultural bias in large language models.
  47. 47.Gabriel Nicholas and Aliya Bhatia. 2023. Lost in translation: Large language models in non-english content analysis.
  48. 48.OpenAI. 2023a. Gpt-4 technical report.
  49. 49.OpenAI. 2023b. Openai devday: Opening keynote.
  50. 50.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  51. 51.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with gpt-4.
  52. 52.Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. 2023. Discovering language model behaviors with model-written evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13387–13434, Toronto, Canada. Association for Computational Linguistics.
  53. 53.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model.
  54. 54.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints.
  55. 55.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106.
  56. 56.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect?
  57. 57.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
  58. 58.Omar Shaikh, Kristina Gligoric, Ashna Khetan, ´ Matthias Gerstgrasser, Diyi Yang, and Dan Jurafsky. 2023. Grounding or guesswork? large language models are presumptive grounders. arXiv preprint arXiv:2311.09144.
  59. 59.Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. 2023. A long way to go: Investigating length correlations in rlhf.
  60. 60.Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024. A roadmap to pluralistic alignment.
  61. 61.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei. 2023. Challenging BIG-bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13003–13051, Toronto, Canada. Association for Computational Linguistics.
  62. 62.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  63. 63.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation.
  64. 64.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  65. 65.Christoph Treude and Hideaki Hata. 2023. She elicits requirements and he tests: Software engineering gender bias in large language models.
  66. 66.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Shengyi Huang, Kashif Rasul, Alexander M. Rush, and Thomas Wolf. 2023a. The alignment handbook. https://github.com/huggingface/alignment-handbook.
  67. 67.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf. 2023b. Zephyr: Direct distillation of lm alignment.
  68. 68.Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. "kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters.
  69. 69.Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. 2023. Openchat: Advancing open-source language models with mixed-quality data. arXiv preprint arXiv:2309.11235.
  70. 70.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  71. 71.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244.
  72. 72.Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach. 2023. Low-resource languages jailbreak gpt-4.
  73. 73.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284.
  74. 74.Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023. Mammoth: Building math generalist models through hybrid instruction tuning. arXiv preprint arXiv:2309.05653.
  75. 75.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy. Association for Computational Linguistics.
  76. 76.Ge Zhang, Yemin Shi, Ruibo Liu, Ruibin Yuan, Yizhi Li, Siwei Dong, Yu Shu, Zhaoqun Li, Zekun Wang, Chenghua Lin, Wenhao Huang, and Jie Fu. 2023. Chinese open instruction generalist: A preliminary release.
  77. 77.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  78. 78.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models.
  79. 79.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. LIMA: Less is more for alignment. In Thirty-seventh Conference on Neural Information Processing Systems.
  80. 80.Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, and Jiantao Jiao. 2023. Starling-7b: Improving llm helpfulness & harmlessness with rlaif.
  81. 81.Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023. MultiVALUE: A framework for cross-dialectal English NLP. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 744–768, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Ryan, M. J., et al. “Unintended Impacts of LLM Alignment on Global Representation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 16121–40, https://doi.org/10.18653/v1/2024.acl-long.853.
APA
Ryan, M. J., Held, W., & Yang, D. (2024). Unintended Impacts of LLM Alignment on Global Representation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16121–16140. https://doi.org/10.18653/v1/2024.acl-long.853
Chicago
Ryan, M. J., W. Held, and D. Yang. 2024. “Unintended Impacts of LLM Alignment on Global Representation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16121–40. https://doi.org/10.18653/v1/2024.acl-long.853.
Harvard
Ryan, M.J., Held, W. and Yang, D. (2024) “Unintended Impacts of LLM Alignment on Global Representation”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 16121–16140. Available at: https://doi.org/10.18653/v1/2024.acl-long.853.
Vancouver
1. Ryan MJ, Held W, Yang D (2024) Unintended Impacts of LLM Alignment on Global Representation. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 16121–16140

BibTeX

@inproceedings{ryan-etal-2024-unintended,
    title = "Unintended Impacts of {LLM} Alignment on Global Representation",
    author = "Ryan, Michael J  and
      Held, William  and
      Yang, Diyi",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.853/",
    doi = "10.18653/v1/2024.acl-long.853",
    pages = "16121--16140"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/