The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm

AakankshaArash AhmadianBeyza ErmisSeraphina Goldfarb-TarrantJulia KreutzerMarzieh FadaeeSara Hooker

article2024EMNLP82 citations

Presents the Aya Red-teaming dataset across eight languages and demonstrates that Direct Preference Optimization effectively reduces both universally recognized and culturally specific harms across multiple languages without degrading general model capabilities.

Listen

Artificial intelligence systems are increasingly deployed worldwide, yet existing safety alignment efforts remain heavily concentrated on English and Western-centric values. This narrow focus leaves non-English interactions vulnerable to harmful outputs and allows non-English prompts to bypass safety guardrails. The article addresses the urgent challenge of mitigating both universal harms and culturally specific, context-dependent harms across multiple languages without compromising general model capabilities.

The main objective of the article is to demonstrate the viability of multilingual safety alignment techniques, comparing supervised fine-tuning and offline direct preference optimization while assessing how safety interventions impact general generation quality across diverse languages.

To investigate this, the authors built a human-annotated dataset comprising roughly 900 adversarial prompts across eight languages (English, Hindi, French, Spanish, Russian, Arabic, Serbian, and Filipino), categorizing harms into universally recognized global harms and culturally nuanced local harms. Using these human-verified examples as seeds, synthetic data expansion was performed to create safety and general-purpose preference datasets. The researchers then evaluated an 8-billion-parameter multilingual model fine-tuned under varying safety data mixtures (0%, 15%, and 100%) and training recipes, validating automated evaluations against compensated human judgments across six core languages.

The key findings reveal that Direct Preference Optimization applied on top of a Supervised Fine-Tuned checkpoint achieves the strongest balance between safety and performance. This approach reduced harmful model generations by 54.7% while simultaneously achieving a 71.0% win rate over the base model on general open-ended generation benchmarks. In contrast, applying preference optimization directly to the raw instruction-tuned base model underperformed, resulting in higher harm rates (by roughly 8% to 10%) and lower general quality. The interventions consistently reduced harm across all evaluated languages by at least 32% to 79%, demonstrating particular benefit in underrepresented languages like Hindi and Arabic. Furthermore, the analysis showed significant positive cross-harm transfer: training exclusively on local harms strongly reduced global harms (by up to 77.8%), and models exposed only to global harms successfully mitigated local harms as well.

These findings prove that the conventional trade-off between safety and general utility is not inevitable in multilingual language models. Properly staged preference alignment enables models to adhere to safety standards globally without degrading user experience or capability across non-English markets. This has direct operational implications for international compliance, brand safety, and risk reduction, showing that safety guardrails can be deployed across languages without requiring disjointed, single-language systems.

Decision-makers and practitioners deploying multilingual models should adopt a staged alignment pipeline—first applying supervised fine-tuning on high-quality preferred responses before running preference optimization—rather than tuning base models directly. Training data strategies should also deliberately incorporate culturally specific, local safety examples, as they provide robust transfer across both universal and regional risk categories. Organizations should balance safety data mixtures (such as a 15% ratio) with general task data to avoid performance degradation on specialized tasks like translation.

The findings are supported by strong alignment between automated evaluators and human judges; however, certain limitations remain. The dataset covers eight languages and a defined set of harm categories, which cannot fully capture the entire dynamic and evolving landscape of global linguistic and cultural risks. Practitioners should maintain high confidence in the staged optimization methodology while recognizing the ongoing need to monitor and update safety data for emerging local nuances.

arXiv: 2406.18682
Cover for The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm

Abstract

A key concern with the concept of alignment is the implicit question of alignment to what? AI systems are increasingly used across the world, yet safety alignment is often focused on homogeneous monolingual settings. Additionally, preference training and safety measures often overfit to harms common in Western-centric datasets. Here, we explore the viability of different alignment approaches when balancing dual objectives: addressing and optimizing for a non-homogeneous set of languages and cultural preferences while minimizing both global and local harms. We collect the first set of human-annotated red-teaming prompts¹ in different languages distinguishing between global and local harm, which serve as a laboratory for understanding the reliability of alignment techniques when faced with preference distributions that are non-stationary across geographies and languages. While this setting is seldom covered by the literature to date, which primarily centers on English harm mitigation, it captures real-world interactions with AI systems around the world. We establish a new precedent for state-of-the-art alignment techniques across 6 languages with minimal degradation in general performance. Our work provides important insights into cross-lingual transfer and novel optimization approaches to safeguard AI systems designed to serve global populations.

Table of Contents

  • 1 Introduction
  • 2 Building the Multilingual Aya Red-teaming Dataset
  • 2.1 Human Annotation
  • 2.2 Generating Preference Data for Safety
  • 2.3 Training data mixtures
  • 2.4 Training Methods
  • 3 Experimental setup
  • 3.1 Evaluation
  • 4 Results and Analyses
  • 4.1 Safety and Performance Trade-offs
  • 4.2 All Languages Win
  • 4.3 Mitigation Technique Matters
  • 4.4 Global vs. Local Harm
  • 4.4.1 Transferability of global harms to mitigate local harm ('global-only' ablation)
  • 4.4.2 Transferability of local harms to mitigate global harm ('local-only' ablation)
  • 4.5 LLM-as-evaluator Aligns With Human Judgement
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A Data Collection Process
  • A.1 Annotator Guidelines
  • A.2 Prompt Examples in the Guidelines
  • B Aya Red-teaming Dataset Details
  • C Training Setup Details
  • D Examples of Model generations
  • E LLMs as evaluators
  • F Additional Results and Analyses

Knowls

  1. Knowl 1 — Aya Red-teaming distinguishes globally shared from culturally local harms

    data/table

    The Aya Red-teaming dataset contains human-written harmful prompts in eight languages, collected from compensated native speakers. Annotators supplied English translations, selected one or more harm categories, and labeled each prompt as either global—widely recognized as harmful across contexts—or local—whose harmfulness depends on particular cultural, historical, or linguistic context. The dataset covers bullying and harassment; discrimination and injustice; graphic material; harms of representation, allocation, and quality of service; hate speech; non-consensual sexual content; profanity; self-harm; and violence, threats, and incitement.

    The reported prompt counts are: English, 987 total (569 global, 418 local); French, 813 (450, 363); Spanish, 782 (510, 272); Hindi, 915 (608, 307); Arabic, 900 (730, 170); Russian, 1,007 (747, 260); Serbian, 1,006 (764, 242); and Filipino, 1,009 (512, 497). Thus, the dataset contains roughly 900 prompts per language, with substantially different global/local proportions across languages.

  2. Knowl 2 — Multilingual synthetic preference data construction

    model/method

    The authors used the human red-teaming annotations as seeds for a larger safety preference corpus. For each language, they sampled 100 seed prompts—50 labeled global and 50 local—and used Command R+ to rephrase them and generate alternatives. Command R+ and the API version of the 35B Aya 23 model then generated responses to the expanded prompts; GPT-4 assigned a preference between the two responses. This produced multilingual safety prompts, model completions, and synthetic preference labels.

    For general-purpose preferences, the authors randomly selected 10,000 examples from the 61,135-example English UltraFeedback Binarized dataset. They translated these examples into the experimental languages with NLLB-3.3B, translated only the preferred responses, generated competing responses in each language with Command R+, and used GPT-4 to label the resulting response pairs. The authors refer to these corpora as the safety-only and general-purpose datasets, respectively.

  3. Knowl 3 — Safety and general-purpose training mixtures

    experimental setup

    The authors varied the fraction of safety preference data to study the balance between harm reduction and general capability. The 100% safety mixture uses 5,457 safety-only samples and no general-purpose samples. The 15% safety mixture contains 35,457 samples, with 15% drawn from the safety-only corpus and the remainder from the general-purpose corpus; this is the paper’s default mixture unless stated otherwise. The 0% safety mixture contains 60,000 general-purpose samples and no safety data, serving as the no-safety-training comparison.

  4. Knowl 4 — SFT and DPO training objectives and initializations

    model/method

    Training examples consist of a prompt xx, a preferred completion y+y^+, and a rejected completion y−y^-. SFT-Preferred trains on the preferred completion alone, using the conditional cross-entropy loss −log⁡πθ(y+∣x)-\log \pi_\theta(y^+\mid x), where πθ\pi_\theta is the model being trained. SFT-Random instead trains on a randomly selected completion from the preference pair, providing a comparison that does not consistently select the preferred response.

    The offline preference method is Direct Preference Optimization (DPO), using the preferred and rejected completions and a reference policy πref\pi_{\mathrm{ref}}:

    LDPO=−E(x,y+,y−)∼D[log⁡σ(β[log⁡πθ(y+∣x)πref(y+∣x)−log⁡πθ(y−∣x)πref(y−∣x)])],\mathcal{L}_{\mathrm{DPO}}=-\mathbb{E}_{(x,y^+,y^-)\sim D}\left[\log \sigma\left(\beta\left[\log\frac{\pi_\theta(y^+\mid x)}{\pi_{\mathrm{ref}}(y^+\mid x)}-\log\frac{\pi_\theta(y^-\mid x)}{\pi_{\mathrm{ref}}(y^-\mid x)}\right]\right)\right],

    where DD is the preference-pair dataset, σ\sigma is the logistic sigmoid, and β\beta is the DPO temperature parameter. The authors compare DPO initialized directly from the instruction-fine-tuned base model (DPO(IFT)) with DPO initialized from an SFT-Preferred checkpoint (DPO(SFT)).

  5. Knowl 5 — Model training and evaluation design

    experimental setup

    Experiments fine-tuned the 8B Aya 23 model in English, French, Spanish, Hindi, Russian, and Arabic. SFT ran to convergence with effective batch size 256, learning rate 3×10−53\times10^{-5}, constant learning-rate schedule, cosine warmup ratio 0.3, weight decay 0.1, and 512-token input context. DPO used effective batch size 128, AdamW, 1,024-token context with 512 tokens allocated to the prompt and 512 to the completion, and a sweep over learning rates 5×10−75\times10^{-7} and 5×10−85\times10^{-8} and β\beta values 0.1 and 0.5. The selected DPO configuration used learning rate 5×10−85\times10^{-8}, constant schedule, 150-step linear warmup, and β=0.1\beta=0.1.

    Safety was evaluated with GPT-4’s binary harmfulness judgment on both the original human-annotated prompts and an evaluation set made by translating English prompts into the other languages. General open-ended generation was evaluated on Multilingual Dolly-200 using GPT-4 pairwise comparisons against the base model; translation was evaluated in both directions on FLORES-200 using spBLEU. Reported aggregate performance covers the six experimental languages.

  6. Knowl 6 — DPO(SFT) gives the strongest overall capability–safety balance

    empirical result

    With the 15% safety mixture, every trained model reduced harmful generations relative to the base model, but the methods differed in the balance between safety and general performance. The values below are aggregated across the six experimental languages; harmful-generation percentages are lower-is-better, Dolly-200 win rates are higher-is-better, and FLORES values are spBLEU scores.

    • Base: harmful generations 31.32%; Dolly-200 win rate not reported; FLORES English-to-other 32.29 and other-to-English 37.82.
    • SFT-Random: 20.99%; Dolly-200 58.83%; FLORES 28.48 and 31.50.
    • SFT-Preferred: 13.59%; Dolly-200 67.42%; FLORES 29.47 and 32.50.
    • DPO(IFT): 22.36%; Dolly-200 48.00%; FLORES 31.48 and 33.74.
    • DPO(SFT): 14.19%; Dolly-200 71.00%; FLORES 25.69 and 29.48.

    SFT-Preferred achieved the lowest harmful-generation rate in this comparison, a 56.6% reduction from the base rate; DPO(SFT) reduced it by 54.7% while achieving the highest Dolly-200 win rate. DPO initialized directly from the instruction-fine-tuned model performed substantially worse: its 48.00% win rate was below the base-model comparison point, and its harmful-generation rate was 9.6 percentage points higher than DPO(SFT). The FLORES scores also show that the base model scored highest in both translation directions.

  7. Knowl 7 — Harm reduction varies by language

    empirical result

    On the original human-annotated safety evaluation, both SFT-Preferred and DPO(SFT) reduced harmful generations in all six experimental languages. The reported DPO(SFT) reductions were 72.4% for Hindi and 79.0% for Arabic, while French had the smallest reported reduction, 32.1%. The authors suggest that larger gains in some languages may reflect their underrepresentation in the base model’s training data, but present this as a hypothesis rather than an established explanation. In the same evaluation, DPO(SFT) reduced global harms more than local harms, while still reducing both.

  8. Knowl 8 — Mitigation transfers between global and local harms

    empirical result

    An ablation trained SFT and DPO(SFT) models on global-only, local-only, or combined global-and-local safety examples, using the 15% safety mixture. Each result below gives the percentage of harmful generations on the stated evaluation subset, followed by the relative reduction from the base model.

    • SFT, local evaluation: global-only training, 11.4% (56.7% reduction); local-only, 10.9% (58.6%); global plus local, 10.5% (60.1%).
    • DPO(SFT), local evaluation: global-only, 10.7% (59.3% reduction); local-only, 7.6% (71.1%); global plus local, 10.6% (59.7%).
    • SFT, global evaluation: global-only training, 12.7% (64.8% reduction); local-only, 11.0% (69.5%); global plus local, 12.6% (65.1%).
    • DPO(SFT), global evaluation: global-only training, 12.2% (66.2% reduction); local-only, 8.0% (77.8%); global plus local, 12.9% (64.3%).

    Global-only training therefore transferred to local-harm evaluation, and local-only training also reduced global harms. For SFT, adding local examples to global-only training improved local-evaluation reduction from 56.7% to 60.1%; for DPO(SFT), the corresponding change was small, from 59.3% to 59.7%. Local-only DPO(SFT) produced the largest reported reduction on global harms, 77.8%. These results support cross-harm transfer, while showing that the benefit of combining harm types depends on the training method and evaluation subset.

  9. Knowl 9 — GPT-4 harmfulness judgments were compared with human judgments

    empirical result

    To validate GPT-4 as a safety evaluator, native compensated annotators judged prompt–completion pairs drawn from 100 held-out global and local prompts across the six experimental languages. The table reports percentages of generations judged harmful by GPT-4 and by humans, followed by the reported agreement value. GPT-4 values include standard errors of the mean across 10 random samples.

    • Base: GPT-4 30.8% ±\pm 0.61; human 56.7%; agreement 66.8.
    • SFT: GPT-4 11.2% ±\pm 0.46; human 25.6%; agreement 81.8.
    • DPO(IFT): GPT-4 22.3% ±\pm 0.41; human 38.3%; agreement 75.3.
    • DPO(SFT): GPT-4 14.4% ±\pm 0.48; human 30.5%; agreement 79.7.

    The evaluator comparison preserves the same ordering across models for harmful-generation rates: the base model is highest and SFT is lowest, with DPO(SFT) between them. The reported agreement values are higher for each fine-tuned model than for the base model.

  10. Knowl 10 — Dataset coverage and temporal limitations

    limitation

    The authors note that the Aya Red-teaming dataset does not cover every form of harm: nuanced or context-specific cases may be underrepresented or absent. Its eight languages do not span the linguistic diversity required for truly global deployment, and its static contents cannot keep pace with evolving harmful content. Consequently, the experiments establish results for the studied languages and prompts rather than comprehensive multilingual safety coverage.

Coverage note — Omitted the paper’s individual prompt and model-generation examples because they illustrate the dataset and mitigation behavior but do not add separate generalizable methods or findings.

References

  1. 1.Amro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel, Kamalika Chaudhuri, and Ari S. Morcos. 2024. Effective pruning of web-scale datasets based on complexity of concept clusters.
  2. 2.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  3. 3.Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker. 2024. Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.
  4. 4.Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, Logesh Kumar Umapathi, Carolyn Jane Anderson, Yangtian Zi, Joel Lamy Poirier, Hailey Schoelkopf, Sergey Troshin, Dmitry Abulkhanov, Manuel Romero, Michael Lappert, Francesco De Toni, Bernardo García del Río, Qian Liu, Shamik Bose, Urvashi Bhattacharyya, Terry Yue Zhuo, Ian Yu, Paulo Villegas, Marco Zocca, Sourab Mangrulkar, David Lansky, Huu Nguyen, Danish Contractor, Luis Villa, Jia Li, Dzmitry Bahdanau, Yacine Jernite, Sean Hughes, Daniel Fried, Arjun Guha, Harm de Vries, and Leandro von Werra. 2023. Santacoder: don’t reach for the stars!
  5. 5.Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein. 2022. Probing pre-trained language models for cross-cultural differences in values. arXiv preprint arXiv:2203.13722.
  6. 6.Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Aidan Gomez, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. 2024. Aya 23: Open weight releases to further multilingual progress.
  7. 7.Edmond Awad, Sohan Dsouza, Azim Shariff, Iyad Rahwan, and Jean-François Bonnefon. 2020. Universals and variations in moral decisions made in 42 countries by 70,000 participants. Proceedings of the National Academy of Sciences, 117(5):2332–2337.
  8. 8.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  9. 9.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022b. Constitutional ai: Harmlessness from ai feedback.
  10. 10.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA. Association for Computing Machinery.
  11. 11.Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. 2023. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. arXiv preprint arXiv:2309.07875.
  12. 12.Meriem Boubdir, Edward Kim, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. 2023. Which prompts make the difference? data prioritization for efficient human llm evaluation.
  13. 13.R. A. Bradley and M. E. Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons.
  14. 14.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  15. 15.Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt. 2024. Are aligned neural networks adversarially aligned?
  16. 16.Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dylan Hadfield-Menell. 2023. Explore, establish, exploit: Red teaming language models from scratch. arXiv preprint arXiv:2306.09442.
  17. 17.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  18. 18.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  19. 19.Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2023. Safe rlhf: Safe reinforcement learning from human feedback.
  20. 20.Adrian de Wynter, Ishaan Watts, Nektar Ege Altıntoprak, Tua Wongsangaroonsri, Minghui Zhang, Noura Farra, Lena Baur, Samantha Claudet, Pavel Gajdusek, Can Gören, Qilong Gu, Anna Kaminska, Tomasz Kaminski, Ruby Kuo, Akiko Kyuba, Jongho Lee, Kartik Mathur, Petter Merok, Ivana Milovanovic, Nani Paananen, Vesa-Matti Paananen, ´ Anna Pavlenko, Bruno Pereira Vidal, Luciano Strika, Yueh Tsao, Davide Turcato, Oleksandr Vakhno, Judit Velcsov, Anna Vickers, Stéphanie Visser, Herdyan Widarmanto, Andrey Zaikin, and Si-Qing Chen. 2024. Rtp-lx: Can llms evaluate toxicity in multilingual scenarios?
  21. 21.Chengyuan Deng, Yiqun Duan, Xin Jin, Heng Chang, Yijun Tian, Han Liu, Henry Peng Zou, Yiqiao Jin, Yijia Xiao, Yichen Wang, Shenghao Wu, Zongxing Xie, Kuofeng Gao, Sihong He, Jun Zhuang, Lu Cheng, and Haohan Wang. 2024a. Deconstructing the ethics of large language models from long-standing issues to new-emerging dilemmas.
  22. 22.Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2024b. Multilingual jailbreak challenges in large language models.
  23. 23.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurahit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona-assigned language models.
  24. 24.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire.
  25. 25.Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md. Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen Ahmed. 2023. Bias and fairness in large language models: A survey. ArXiv, abs/2309.00770.
  26. 26.Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022a. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned.
  27. 27.Deep Ganguli et al. 2022b. Predictability and surprise in large generative models. arXiv preprint arXiv:2202.07785.
  28. 28.Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao. 2023. Mart: Improving llm safety with multi-round automatic red-teaming. arXiv preprint arXiv:2311.07689.
  29. 29.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models.
  30. 30.Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett. 2023. This prompt is measuring <mask>: Evaluating bias evaluation in language models.
  31. 31.Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, and Hannaneh Hajishirzi. 2024. Olmo: Accelerating the science of language models.
  32. 32.Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023. Textbooks are all you need.
  33. 33.Vipul Gupta, Pranav Narayanan Venkit, Shomir Wilson, and R. Passonneau. 2023. Sociodemographic bias in language models: A survey and forward path.
  34. 34.Katharina Hämmerl, Bjoern Deiseroth, Patrick Schramowski, Jindˇrich Libovický, Constantin Rothkopf, Alexander Fraser, and Kristian Kersting. 2023. Speaking multiple languages affects the moral bias of language models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 2137–2156, Toronto, Canada. Association for Computational Linguistics.
  35. 35.Kai Hartung, Aaricia Herygers, Shubham Kurlekar, Khabbab Zakaria, Taylan Volkan, Sören Gröttrup, and Munir Georges. 2023. Measuring sentiment bias in machine translation.
  36. 36.Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2023. Aligning ai with shared human values.
  37. 37.Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022. Challenges and strategies in cross-cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6997–7013, Dublin, Ireland. Association for Computational Linguistics.
  38. 38.Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao. 2024a. An empirical study of llm-as-a-judge for llm evaluation: Fine-tuned judge models are task-specific classifiers.
  39. 39.Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I. Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli. 2024b. Collective constitutional ai: Aligning a language model with public input. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24. ACM.
  40. 40.Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao. 2024. Ai alignment: A comprehensive survey.
  41. 41.Meng Ji, Meng Ji, Pierrette Bouillon, and Mark Seligman. 2023. Cultural and Linguistic Bias of Neural Machine Translation Technology, Studies in Natural Language Processing, page 100–128. Cambridge University Press.
  42. 42.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  43. 43.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts.
  44. 44.Pratik Joshi, Christain Barnes, Sebastin Santy, Simran Khanuja, Sanket Shah, Anirudh Srinivasan, Satwik Bhattamishra, Sunayana Sitaram, Monojit Choudhury, and Kalika Bali. 2019. Unsung challenges of building and deploying language technologies for low resource language communities. In Proceedings of the 16th International Conference on Natural Language Processing, pages 211–219, International Institute of Information Technology, Hyderabad, India. NLP Association of India.
  45. 45.Khyati Khandelwal, Manuel Tonneau, Andrew M. Bean, Hannah Rose Kirk, and Scott A. Hale. 2023. Casteist but not racist? quantifying disparities in large language model bias between india and the west.
  46. 46.Md Tawkat Islam Khondaker, Abdul Waheed, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023. Gptaraeval: A comprehensive evaluation of chatgpt on arabic nlp.
  47. 47.Tom Kocmi and Christian Federmann. 2023. Large language models are state-of-the-art evaluators of translation quality. In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, pages 193–203, Tampere, Finland. European Association for Machine Translation.
  48. 48.Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in large language models. In Proceedings of The ACM Collective Intelligence Conference, CI ’23. ACM.
  49. 49.Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2023. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback.
  50. 50.Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable agent alignment via reward modeling: a research direction.
  51. 51.Haoran Li, Yulin Chen, Jinglong Luo, Yan Kang, Xiaojin Zhang, Qi Hu, Chunkit Chan, and Yangqiu Song. 2023a. Privacy in large language models: Attacks, defenses and future directions. ArXiv, abs/2310.10383.
  52. 52.Jie Li, Yi Liu, Chongyang Liu, Ling Shi, Xiaoning Ren, Yaowen Zheng, Yang Liu, and Yinxing Xue. 2024a. A cross-language investigation into jailbreak attacks in large language models. arXiv preprint arXiv:2401.16765.
  53. 53.Junlong Li, Jinyuan Wang, Zhuosheng Zhang, and Hai Zhao. 2024b. Self-prompting large language models for zero-shot open-domain qa.
  54. 54.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2023b. Starcoder: may the source be with you!
  55. 55.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, Kailong Wang, and Yang Liu. 2023. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860.
  56. 56.Zixuan Liu, Xiaolin Sun, and Zizhan Zheng. 2024. Enhancing llm safety via constrained direct preference optimization. arXiv preprint arXiv:2403.02475.
  57. 57.Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, Xinyi Wu, Enrico Shippole, Kurt Bollacker, Tongshuang Wu, Luis Villa, Sandy Pentland, and Sara Hooker. 2023. The data provenance initiative: A large scale audit of dataset licensing attribution in ai.
  58. 58.Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-Béguelin. 2023. Analyzing leakage of personally identifiable information in language models.
  59. 59.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct.
  60. 60.Max Marion, Ahmet Üstün, Luiza Pozzobon, Alex Wang, Marzieh Fadaee, and Sara Hooker. 2023. When less is more: Investigating data pruning for pretraining llms at scale.
  61. 61.Mike Maxwell and Baden Hughes. 2006. Frontiers in linguistic annotation for lower-density languages. In Proceedings of the Workshop on Frontiers in Linguistically Annotated Corpora 2006, pages 29–37, Sydney, Australia. Association for Computational Linguistics.
  62. 62.Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. 2018. Ray: A distributed framework for emerging ai applications.
  63. 63.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir R. Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023. Crosslingual generalization through multitask finetuning. In Annual Meeting of the Association for Computational Linguistics.
  64. 64.Anjishnu Mukherjee, Chahat Raj, Ziwei Zhu, and Antonios Anastasopoulos. 2023. Global Voices, local biases: Socio-cultural prejudices across languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15828–15845, Singapore. Association for Computational Linguistics.
  65. 65.Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. 2023. Scalable extraction of training data from (production) language models.
  66. 66.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744.
  67. 67.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3419–3448. Association for Computational Linguistics.
  68. 68.Luiza Pozzobon, Patrick Lewis, Sara Hooker, and Beyza Ermis. 2024. From one to many: Expanding the scope of toxicity mitigation in language models.
  69. 69.Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson. 2024. Safety alignment should be made more than just a few tokens deep.
  70. 70.Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693.
  71. 71.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model.
  72. 72.Aida Ramezani and Yang Xu. 2023. Knowledge of cultural moral norms in large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 428–446, Toronto, Canada. Association for Computational Linguistics.
  73. 73.Shaina Raza, Oluwanifemi Bamgbose, Shardul Ghuge, and Deepak John Reji. 2024. Safe and responsible large language model development.
  74. 74.Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021. Gender bias in machine translation. Transactions of the Association for Computational Linguistics, 9:845–874.
  75. 75.Reva Schwartz, Apostol T. Vassilev, Kristen Greene, Lori A. Perine, Andrew Burt, and Patrick Hall. 2022. Towards a standard for identifying and managing bias in artificial intelligence.
  76. 76.Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Synthetic prompting: Generating chain-of-thought demonstrations for large language models.
  77. 77.Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. 2024. The language barrier: Dissecting safety challenges of llms in multilingual contexts. arXiv preprint arXiv:2401.13136.
  78. 78.Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation.
  79. 79.Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzeminski, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker. 2024. Aya dataset: An open-access collection for multilingual instruction tuning.
  80. 80.Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. “i’m sorry to hear that”: Finding new biases in language models with a holistic descriptor dataset. In Conference on Empirical Methods in Natural Language Processing.
  81. 81.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2022. Learning to summarize from human feedback.
  82. 82.Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn, and Aviral Kumar. 2024. Preference fine-tuning of llms should leverage suboptimal, on-policy data.
  83. 83.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  84. 84.Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy. 2024. Gemma: Open models based on gemini research and technology.
  85. 85.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation.
  86. 86.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  87. 87.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf. 2023. Zephyr: Direct distillation of lm alignment.
  88. 88.Eva Vanmassenhove, Dimitar Shterionov, and Matthew Gwilliam. 2021. Machine translationese: Effects of algorithmic bias on linguistic complexity in machine translation.
  89. 89.Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023. On evaluating and mitigating gender biases in multilingual settings.
  90. 90.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2021. Universal adversarial triggers for attacking and analyzing nlp.
  91. 91.Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023a. Is chatgpt a good nlg evaluator? a preliminary study.
  92. 92.Wenxuan Wang, Zhaopeng Tu, Chang Chen, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael R Lyu. 2023b. All languages matter: On the multilingual safety of large language models. arXiv preprint arXiv:2310.00905.
  93. 93.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023c. Self-instruct: Aligning language models with self-generated instructions.
  94. 94.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023a. Jailbroken: How does llm safety training fail?
  95. 95.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a. Finetuned language models are zero-shot learners.
  96. 96.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022b. Emergent abilities of large language models.
  97. 97.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023b. Chain-of-thought prompting elicits reasoning in large language models.
  98. 98.Chenxi Whitehouse, Monojit Choudhury, and Alham Fikri Aji. 2023. Llm-powered data augmentation for enhanced cross-lingual performance.
  99. 99.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions.
  100. 100.Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek. 2024. Llm jailbreak attack versus defense techniques–a comprehensive study. arXiv preprint arXiv:2402.13457.
  101. 101.Vithya Yogarajan, Gillian Dobbie, Te Taka Keegan, and Rostam Josef Neuwirth. 2023. Tackling bias in pre-trained language models: Current trends and under-represented societies. ArXiv, abs/2312.01509.
  102. 102.Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach. 2024. Low-resource languages jailbreak gpt-4.
  103. 103.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models.
  104. 104.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.
  105. 105.Xin Zhou, Yi Lu, Ruotian Ma, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. Making harmful behaviors unlearnable for large language models. arXiv preprint arXiv:2311.02105.
  106. 106.Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models.
  107. 107.Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. Aya model: An instruction finetuned open-access multilingual language model.

Citation

MLA
Aakanksha, et al. “The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 12027–49, https://doi.org/10.18653/v1/2024.emnlp-main.671.
APA
Aakanksha, Ahmadian, A., Ermis, B., Goldfarb-Tarrant, S., Kreutzer, J., Fadaee, M., & Hooker, S. (2024). The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 12027–12049. https://doi.org/10.18653/v1/2024.emnlp-main.671
Chicago
Aakanksha, A. Ahmadian, B. Ermis, et al. 2024. “The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 12027–49. https://doi.org/10.18653/v1/2024.emnlp-main.671.
Harvard
Aakanksha et al. (2024) “The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 12027–12049. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.671.
Vancouver
1. Aakanksha, Ahmadian A, Ermis B, Goldfarb-Tarrant S, Kreutzer J, Fadaee M, Hooker S (2024) The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 12027–12049

BibTeX

@inproceedings{aakanksha-etal-2024-multilingual,
    title = "The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm",
    author = "Aakanksha  and
      Ahmadian, Arash  and
      Ermis, Beyza  and
      Goldfarb-Tarrant, Seraphina  and
      Kreutzer, Julia  and
      Fadaee, Marzieh  and
      Hooker, Sara",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.671/",
    doi = "10.18653/v1/2024.emnlp-main.671",
    pages = "12027--12049"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/