Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
Shivalika SinghFreddie VargusDaniel D'souzaBörje KarlssonAbinaya MahendiranWei-Yin KoHerumb ShandilyaJay PatelDeividas MataciunasLaura O'Mahony
Presents a massive open-access multilingual instruction-tuning resource comprising over 204,000 human-annotated examples across 65 languages and an expanded 513-million-instance collection spanning 114 languages to help close the linguistic resource gap in large language model development.
Instruction fine-tuning has driven major breakthroughs in large language models, enabling them to follow complex natural language prompts. However, existing datasets and fine-tuning resources remain predominantly English-centric, leaving low-resource languages underserved and causing models to exhibit cultural bias and security vulnerabilities. To address this linguistic inequality, the article details a global participatory research initiative aimed at building open-access, multilingual instruction-tuning resources.
The initiative developed three core resources using an open-science framework that engaged 2,997 collaborators across 119 countries. First, the human-annotated Aya Dataset was curated through a structured pipeline where fluent speakers generated original prompts and completions, re-annotated existing pairs, and reviewed data quality. Second, the Aya Collection was assembled by combining fluent-speaker prompt templates across 44 NLP datasets with machine translations of 19 high-quality datasets. Third, the Aya Evaluation Suite was created to benchmark multilingual open-ended generation, comprising human-written, machine-translated, and professionally post-edited prompts.
The project yielded several key findings regarding scale and data quality. The Aya Dataset gathered 204,114 human-curated instances spanning 65 languages, including 31 low-resource languages, while the broader Aya Collection expanded to 513 million instances covering 114 languages. Human re-annotation increased the average completion length across all sources by 25%, with Aya original submissions growing by 40%. A positive correlation of 0.27 emerged between character length and annotator approval scores, demonstrating that richer, full-sentence responses directly improve perceived data quality. Finally, Aya original annotations achieved an approval ratio of 0.81, significantly outperforming existing resources like xP3, which scored 0.50.
These findings demonstrate that collaborative, human-in-the-loop data curation can successfully expand language coverage without sacrificing quality. Releasing these datasets under a permissive Apache 2.0 license lowers development costs and risks for researchers aiming to build globally representative, safer models. Decision-makers and developers should leverage the Aya Collection and Dataset to expand multilingual support, while using the Aya Evaluation Suite and human evaluation rather than purely automated metrics to assess open-ended generation.
Several limitations remain. The 114 covered languages represent only a fraction of global linguistic diversity, and unwritten languages along with regional dialects are largely omitted. Additionally, contribution activity was uneven across languages, with a small number of active annotators generating the majority of instances in certain languages, introducing potential cultural and personal biases. While human review minimized toxic content, users should exercise appropriate caution regarding potential gaps in under-represented language varieties.
- Paper: Scaling Instruction-Finetuned Language Models, Hyung Won Chung et al. (2024). This foundational work establishes the scaling dynamics and methodology of instruction fine-tuning across diverse task mixtures, providing the direct conceptual framework that Aya adapts to the massively multilingual setting.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). It introduces the massively multilingual text-to-text transformer framework and pre-training recipe across 101 languages, establishing the baseline multilingual backbone and representation paradigm on which Aya's datasets and tuning build.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). It demonstrates how instruction tuning enables zero-shot task generalization in language models, defining the core instruction-following paradigm that the Aya dataset curates across non-English languages.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). It outlines the foundational approach for generating and templating instruction-tuning datasets, which informs Aya's pipeline for bootstrapping and augmenting multilingual instruction collections.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). It provides the operational blueprint for large-scale, participatory open-access research to build multilingual NLP resources across global communities, directly inspiring Aya’s collaborative methodology.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). It details large-scale human-centered collection and parallel evaluation practices across hundreds of low-resource languages, establishing key multilingual scaling foundations leveraged by the Aya initiative.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It benchmarks the persistent performance gap between English and non-English generative models, demonstrating the empirical motivation for Aya's multilingual instruction tuning collection.
- Paper: Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model, Ahmet Üstün et al. (2024). It directly trains and evaluates the Aya multilingual generative model using the human-curated and templated instruction mixtures created in the Aya Dataset initiative.
- Paper: The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants, Lucas Bandarkar et al. (2024). It provides a 122-language parallel reading comprehension benchmark that enables comprehensive downstream evaluation of the multilingual capabilities enabled by Aya's instruction tuning.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). It extends massive multilingual evaluation to representation and embedding tasks across 250+ languages, building upon the participatory community-driven paradigm exemplified by Aya.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). It establishes internal representation metrics to quantify LLM disparities between high- and low-resource languages, offering a diagnostic method to analyze models tuned on datasets like Aya.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). It benchmarks cross-lingual transfer and in-context learning against fine-tuning across 54 languages, presenting an evaluation framework to assess models trained on massive multilingual instruction datasets.
