Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model
Ahmet ÜstünViraat AryabumiZheng Xin YongWei-Yin KoDaniel D'souzaGbemileke OniludeNeel BhandariShivalika SinghHui-Lee OoiAmr Kayid
Presents Aya, an open-access instruction-finetuned language model spanning 101 languages—over half of which are lower-resourced—that consistently outperforms mT0 and BLOOMZ across both discriminative and generative benchmarks.
Recent advances in artificial intelligence and large language models have primarily benefited a small number of data-rich languages, most notably English. This leaves the vast majority of the world's languages underserved and exacerbates global linguistic inequality in technology access. Most instruction-following datasets and models remain heavily skewed toward English, limiting the ability of non-English speakers to interact with advanced systems in their native languages.
The article demonstrates the development and evaluation of Aya, an open-source, massively multilingual generative language model designed to follow instructions across 101 languages. The study evaluates how expanding language coverage, curating diverse multilingual instruction mixtures, and applying safety context distillation can confer broad instruction-following capabilities without sacrificing quality.
The authors built a 13-billion-parameter encoder-decoder model by fine-tuning mT5 on an expanded training mixture comprising 203 million data points across 101 languages, where English represents only 21.5% of the data. The dataset combines human-curated prompts from native speakers, pruned prompt templates, machine-translated instruction sets, and synthetic conversational data. To evaluate performance across 99 languages, the team deployed a comprehensive evaluation suite encompassing unseen discriminative tasks, academic benchmarks, generative tasks, human preference assessments, and simulated evaluations using large language models as judges.
The evaluation revealed several key findings. First, Aya consistently outperformed previous massively multilingual open-source baselines such as mT0 and BLOOMZ across most tasks, achieving relative performance gains of 13.1% on discriminative benchmarks and 11.7% on generative benchmarks over a comparable 101-language baseline. Second, in human preference evaluations across seven languages, native speakers preferred Aya's responses 77% of the time over baseline models, favoring its richer, more natural, and more detailed generations. Third, prioritizing translation-heavy datasets provided the best balance for open-ended conversational generation and machine translation, while human data pruning effectively increased average instruction length and data quality. Fourth, applying multilingual safety context distillation reduced harmful model outputs on adversarial prompts by 78% to 89% with only a minor performance drop of 2% to 3% on standard benchmarks.
These findings indicate that language inequality in artificial intelligence can be effectively addressed by curating balanced, diverse multilingual datasets rather than relying solely on English-centric cross-lingual transfer. Achieving strong multilingual instruction-following capabilities does not require sacrificing performance, provided that training data is appropriately filtered and weighted. Furthermore, the results demonstrate that safety interventions can be effectively scaled multilingually during instruction tuning rather than relying solely on post-hoc guardrails.
Organizations developing multilingual applications should adopt balanced data-weighting strategies and implement language-inclusive safety distillation. For practical deployment, practitioners should explore compression and quantization techniques to run 13-billion-parameter models efficiently on standard hardware. Further research should focus on extending instruction coverage to additional regional dialects, refining multilingual refusal phrasing to make responses less repetitive, and expanding human evaluations to a wider range of lower-resourced languages.
While the results demonstrate strong improvements, several limitations remain. The model covers 101 languages, representing only a fraction of the roughly 7,000 languages spoken worldwide, and dialectal variations or code-switching are not fully represented. Additionally, some translated datasets may reflect Western cultural biases, and safety evaluations remain less comprehensive for very low-resourced languages. Consequently, stakeholders should deploy the model with appropriate monitoring for domain-specific safety, cultural nuances, and potential grammatical errors in lower-resource settings.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). BLOOM provides the open-science multilingual foundational architecture and baseline (BLOOMZ) that Aya directly builds upon and outperforms across low-resource languages.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). mT5 established the multilingual text-to-text benchmark across 101 languages and served as the pre-trained base for mT0, which Aya benchmarks against and directly surpasses.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). XLM-R demonstrates the fundamentals and capacity trade-offs of massively multilingual cross-lingual representation learning at a scale of 100 languages.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). MEGA establishes standard benchmarking methodologies and highlights severe performance gaps across 70 languages in generative models, motivating Aya's evaluation design.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). No Language Left Behind sets the methodological standard for human-centered data curation and scaling machine translation across hundreds of low-resource languages.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). This paper offers the foundational taxonomy and resource-availability analysis of global languages that defines the low-resource inclusion problem Aya addresses.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Glot500 examines horizontal scaling of multilingual corpora and models to 500 under-resourced languages, directly contextualizing Aya's multi-language data-pruning and fine-tuning challenges.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). This work explores zero-shot cross-lingual transfer mechanisms across over 100 languages in pre-trained models, underpinning Aya's instruction-following transfer.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). MEGAVERSE expands on multilingual generative evaluation suites by benchmarking modern LLMs across 83 languages and diverse modalities.
- Paper: Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models, Tianyi Tang et al. (2024). This study delves into the mechanistic interpretability of multilingual LLMs by isolating language-specific neurons that govern cross-lingual capabilities.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). Language Ranker extends multilingual assessment methods by introducing an internal representation metric to quantify performance gaps between high- and low-resource languages.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). MMTEB scales multilingual evaluation even further by benchmarking instruction-tuned embeddings across more than 250 languages and hundreds of tasks.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Omnilingual MT pushes the boundaries of broad multilingual coverage by expanding LLM-driven translation to over 1,600 marginalized and low-resource languages.
