Unintended Impacts of LLM Alignment on Global Representation
Michael J. RyanWilliam Barr HeldDiyi Yang
Reveals how standard alignment techniques like RLHF and DPO introduce substantial global representation biases by widening performance disparities across English dialects and skewing model viewpoints toward US perspectives, while simultaneously improving non-English multilingual capabilities.
Large language models are rapidly expanding to hundreds of millions of users worldwide after undergoing alignment procedures such as supervised fine-tuning and preference tuning to make them helpful and safe. However, alignment reflects specific human preferences that are rarely universal, raising concerns about which cultural norms and linguistic variations are prioritized. The article evaluates how standard alignment workflows impact large language models across three global dimensions: regional English dialects, multilingual performance, and country-level opinions.
The authors conducted a comprehensive evaluation across nine open-source language models, tracking performance changes across their base, supervised fine-tuning, and preference-tuned stages. The analysis tested models on task-oriented dialogue intent prediction across American, Indian, and Nigerian English dialects; evaluated reading comprehension and question-answering across nine typologically diverse languages; and measured alignment with national opinion surveys across seven diverse countries. The authors also developed a dataset of 554 country-specific subjective questions to probe the geographical biases learned by an open-source reward model.
The investigation produced four critical findings. First, while alignment improved intent recognition across all English dialects, it sharply widened performance disparities: the performance gap between American English and non-Western dialects grew from approximately 1% in base models to up to 17.1% after alignment. Second, fine-tuning unexpectedly boosted multilingual comprehension across most evaluated languages because the fine-tuning datasets unintentionally included roughly 13% non-English text. Third, alignment consistently shifted model responses closer to United States public opinion, increasing divergence from countries such as Jordan, China, and Nigeria by 2% to 5% while remaining aligned with Western nations. Finally, the evaluated reward model favored Western nations and ranked nearly all other countries below the United States, though this bias did not directly propagate into the fine-tuned language model because country-opinion topics were absent during preference training.
These results demonstrate that standard alignment practices inadvertently introduce regional disparities that favor Western and American norms, which may hinder the equitable global adoption of artificial intelligence. To address these disparities, developers should implement transparent reporting on annotator demographics, data curation, and domain selections throughout the alignment process. Organizations should also intentionally incorporate diverse multilingual and cross-dialect data into fine-tuning, as even modest inclusions generate substantial capability gains across languages without compromising primary language performance.
The findings are bounded by certain limitations, such as reliance on publicly released model checkpoints where some intermediate training data was unavailable, and evaluation across a focused set of downstream benchmarks. Nevertheless, the consistency of the results across distinct model families provides high confidence that alignment decisions heavily dictate global representation, making deliberate and inclusive data selection essential for future deployments.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). Santurkar et al. directly establish the core premise that human-feedback alignment skews language models toward specific demographic opinions rather than representative global views.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). Kirk et al. document how narrow annotator demographics and localized cultural assumptions distort feedback learning in large language models.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). Rafailov et al. introduce Direct Preference Optimization (DPO), one of the primary alignment mechanisms evaluated for unintended representational impacts.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Ouyang et al. establish standard Reinforcement Learning from Human Feedback (RLHF) pipelines whose subjective tuning leads to demographic and linguistic disparities.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Ahuja et al. provide essential benchmarking methodology for evaluating the disparities between English and non-English model performance across diverse language families.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). Bai et al. formulate helpfulness and harmlessness alignment objectives using crowdworker feedback that can inadvertently embed demographic biases.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). Argyle et al. formulate the foundational framework for probing whether language models can accurately represent diverse human subpopulations and opinion distributions.
- Paper: From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models, Shangbin Feng et al. (2023). Feng et al. examine how pre-existing data biases propagate into downstream models, contrasting pretraining bias with alignment-induced distribution shifts.
- Paper: Large Language Models are Geographically Biased, Rohin Manvi et al. (2024). Manvi et al. directly extend the study of global representational disparities by formalizing and measuring systemic geographic biases in language model predictions.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). Li et al. build on multilingual alignment analysis by introducing an internal representation metric to quantify cross-lingual capability gaps between high- and low-resource languages.
- Paper: Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment, Rui Yang et al. (2024). Yang et al. address the problem of aligning to monolithic preference sets by developing dynamic, multi-objective preference adjustment during deployment.
- Paper: KTO: Model Alignment as Prospect Theoretic Optimization, Kawin Ethayarajh et al. (2024). Ethayarajh et al. offer an alternative alignment objective based on behavioral utility theory that moves beyond standard pairwise preference tuning.
- Paper: Mitigating the Alignment Tax of RLHF, Yong Lin et al. (2024). Lin et al. examine and mitigate the broader trade-offs and performance degradations incurred when fine-tuning models with RLHF.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). Raheja and Pochhi synthesize how specific design axes in direct preference algorithms lead to unintended optimization behaviors like mode collapse and likelihood displacement.
