The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
AakankshaArash AhmadianBeyza ErmisSeraphina Goldfarb-TarrantJulia KreutzerMarzieh FadaeeSara Hooker
Presents the Aya Red-teaming dataset across eight languages and demonstrates that Direct Preference Optimization effectively reduces both universally recognized and culturally specific harms across multiple languages without degrading general model capabilities.
Artificial intelligence systems are increasingly deployed worldwide, yet existing safety alignment efforts remain heavily concentrated on English and Western-centric values. This narrow focus leaves non-English interactions vulnerable to harmful outputs and allows non-English prompts to bypass safety guardrails. The article addresses the urgent challenge of mitigating both universal harms and culturally specific, context-dependent harms across multiple languages without compromising general model capabilities.
The main objective of the article is to demonstrate the viability of multilingual safety alignment techniques, comparing supervised fine-tuning and offline direct preference optimization while assessing how safety interventions impact general generation quality across diverse languages.
To investigate this, the authors built a human-annotated dataset comprising roughly 900 adversarial prompts across eight languages (English, Hindi, French, Spanish, Russian, Arabic, Serbian, and Filipino), categorizing harms into universally recognized global harms and culturally nuanced local harms. Using these human-verified examples as seeds, synthetic data expansion was performed to create safety and general-purpose preference datasets. The researchers then evaluated an 8-billion-parameter multilingual model fine-tuned under varying safety data mixtures (0%, 15%, and 100%) and training recipes, validating automated evaluations against compensated human judgments across six core languages.
The key findings reveal that Direct Preference Optimization applied on top of a Supervised Fine-Tuned checkpoint achieves the strongest balance between safety and performance. This approach reduced harmful model generations by 54.7% while simultaneously achieving a 71.0% win rate over the base model on general open-ended generation benchmarks. In contrast, applying preference optimization directly to the raw instruction-tuned base model underperformed, resulting in higher harm rates (by roughly 8% to 10%) and lower general quality. The interventions consistently reduced harm across all evaluated languages by at least 32% to 79%, demonstrating particular benefit in underrepresented languages like Hindi and Arabic. Furthermore, the analysis showed significant positive cross-harm transfer: training exclusively on local harms strongly reduced global harms (by up to 77.8%), and models exposed only to global harms successfully mitigated local harms as well.
These findings prove that the conventional trade-off between safety and general utility is not inevitable in multilingual language models. Properly staged preference alignment enables models to adhere to safety standards globally without degrading user experience or capability across non-English markets. This has direct operational implications for international compliance, brand safety, and risk reduction, showing that safety guardrails can be deployed across languages without requiring disjointed, single-language systems.
Decision-makers and practitioners deploying multilingual models should adopt a staged alignment pipeline—first applying supervised fine-tuning on high-quality preferred responses before running preference optimization—rather than tuning base models directly. Training data strategies should also deliberately incorporate culturally specific, local safety examples, as they provide robust transfer across both universal and regional risk categories. Organizations should balance safety data mixtures (such as a 15% ratio) with general task data to avoid performance degradation on specialized tasks like translation.
The findings are supported by strong alignment between automated evaluators and human judges; however, certain limitations remain. The dataset covers eight languages and a defined set of harm categories, which cannot fully capture the entire dynamic and evolving landscape of global linguistic and cultural risks. Practitioners should maintain high confidence in the staged optimization methodology while recognizing the ongoing need to monitor and update safety data for emerging local nuances.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). This foundational work establishes the supervised-fine-tuning-then-preference-optimization pipeline that the source compares and recommends for multilingual safety alignment.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). Its RLHF experiments show how human preference training can balance helpfulness and harmlessness, the alignment trade-off that motivates the source’s multilingual evaluation.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). Its analysis of whose values human-feedback pipelines encode gives essential context for the source’s distinction between global harms and culturally local preferences.
- Paper: XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity, Dasol Choi et al. (2026). This later benchmark extends the source’s global-versus-local safety framing by testing country-grounded adversarial robustness and cultural sensitivity across more settings.
- Paper: TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages, Victor Akinode et al. (2026). This later benchmark carries the source’s multilingual safety agenda into African languages, testing how culturally grounded jailbreaks challenge guardrails in underrepresented linguistic contexts.
