Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Shivalika SinghFreddie VargusDaniel D'souzaBörje KarlssonAbinaya MahendiranWei-Yin KoHerumb ShandilyaJay PatelDeividas MataciunasLaura O'Mahony

article2024ACL234 citations

Presents a massive open-access multilingual instruction-tuning resource comprising over 204,000 human-annotated examples across 65 languages and an expanded 513-million-instance collection spanning 114 languages to help close the linguistic resource gap in large language model development.

Listen

Instruction fine-tuning has driven major breakthroughs in large language models, enabling them to follow complex natural language prompts. However, existing datasets and fine-tuning resources remain predominantly English-centric, leaving low-resource languages underserved and causing models to exhibit cultural bias and security vulnerabilities. To address this linguistic inequality, the article details a global participatory research initiative aimed at building open-access, multilingual instruction-tuning resources.

The initiative developed three core resources using an open-science framework that engaged 2,997 collaborators across 119 countries. First, the human-annotated Aya Dataset was curated through a structured pipeline where fluent speakers generated original prompts and completions, re-annotated existing pairs, and reviewed data quality. Second, the Aya Collection was assembled by combining fluent-speaker prompt templates across 44 NLP datasets with machine translations of 19 high-quality datasets. Third, the Aya Evaluation Suite was created to benchmark multilingual open-ended generation, comprising human-written, machine-translated, and professionally post-edited prompts.

The project yielded several key findings regarding scale and data quality. The Aya Dataset gathered 204,114 human-curated instances spanning 65 languages, including 31 low-resource languages, while the broader Aya Collection expanded to 513 million instances covering 114 languages. Human re-annotation increased the average completion length across all sources by 25%, with Aya original submissions growing by 40%. A positive correlation of 0.27 emerged between character length and annotator approval scores, demonstrating that richer, full-sentence responses directly improve perceived data quality. Finally, Aya original annotations achieved an approval ratio of 0.81, significantly outperforming existing resources like xP3, which scored 0.50.

These findings demonstrate that collaborative, human-in-the-loop data curation can successfully expand language coverage without sacrificing quality. Releasing these datasets under a permissive Apache 2.0 license lowers development costs and risks for researchers aiming to build globally representative, safer models. Decision-makers and developers should leverage the Aya Collection and Dataset to expand multilingual support, while using the Aya Evaluation Suite and human evaluation rather than purely automated metrics to assess open-ended generation.

Several limitations remain. The 114 covered languages represent only a fraction of global linguistic diversity, and unwritten languages along with regional dialects are largely omitted. Additionally, contribution activity was uneven across languages, with a small number of active annotators generating the majority of instances in certain languages, introducing potential cultural and personal biases. While human review minimized toxic content, users should exercise appropriate caution regarding potential gaps in under-represented language varieties.

arXiv: 2402.06619
  • Paper: Scaling Instruction-Finetuned Language Models, Hyung Won Chung et al. (2024). This foundational work establishes the scaling dynamics and methodology of instruction fine-tuning across diverse task mixtures, providing the direct conceptual framework that Aya adapts to the massively multilingual setting.
  • Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). It introduces the massively multilingual text-to-text transformer framework and pre-training recipe across 101 languages, establishing the baseline multilingual backbone and representation paradigm on which Aya's datasets and tuning build.
  • Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). It demonstrates how instruction tuning enables zero-shot task generalization in language models, defining the core instruction-following paradigm that the Aya dataset curates across non-English languages.
  • Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). It outlines the foundational approach for generating and templating instruction-tuning datasets, which informs Aya's pipeline for bootstrapping and augmenting multilingual instruction collections.
  • Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). It provides the operational blueprint for large-scale, participatory open-access research to build multilingual NLP resources across global communities, directly inspiring Aya’s collaborative methodology.
  • Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). It details large-scale human-centered collection and parallel evaluation practices across hundreds of low-resource languages, establishing key multilingual scaling foundations leveraged by the Aya initiative.
  • Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It benchmarks the persistent performance gap between English and non-English generative models, demonstrating the empirical motivation for Aya's multilingual instruction tuning collection.
Cover for Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Abstract

Datasets are foundational to many breakthroughs in modern artificial intelligence (AI). Many recent achievements in the space of natural language processing (NLP) can be attributed to the fine-tuning of pre-trained models on a diverse set of tasks that enables a large language model (LLM) to respond to instructions. Instruction fine-tuning (IFT) requires specifically constructed and annotated datasets. However, existing datasets are almost all in the English language. In this work, our primary goal is to bridge the language gap by building a human-curated instruction-following dataset spanning 65 languages. We worked with fluent speakers of languages from around the world to collect natural instances of instructions and completions. Furthermore, we create the most extensive multilingual collection to date, comprising 513 million instances through templating and augmenting existing datasets across 114 languages. In total, we contribute three key resources: we develop and open-source the Aya¹ Dataset, the Aya Collection, and the Aya Evaluation Suite. The Aya initiative also serves as a valuable case study in participatory research, involving collaborators from 119 countries. We see this as an important framework for future research collaborations that aim to bridge gaps in resources.

Table of Contents

  • 1 Introduction
  • 2 Aya Dataset
  • 2.1 Annotation Tasks
  • 2.2 Validating the quality of contributions
  • 2.3 Criteria for Inclusion in Aya Dataset
  • 2.4 Analysis of the Aya Dataset
  • 3 Aya Collection
  • 3.1 Templating Existing Datasets
  • 3.2 Automatic Translation
  • 4 Analysis of Aya Collection
  • 5 Aya Evaluation Suite
  • 6 Conclusion
  • 7 Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A A Participatory Approach to Research
  • B Related Work
  • B.1 Multilingual datasets
  • B.2 Instruction-tuning datasets
  • B.3 Multilingual Instruction-Tuning Datasets
  • B.4 Participatory Research in Machine Learning
  • C Aya Dataset: Additional Analysis
  • C.1 Contributors
  • C.2 Annotation Guidelines
  • C.3 Language Groupings
  • C.4 Length difference by language
  • C.5 Annotator Skew
  • C.5.1 Annotator Skew Across Languages
  • C.5.2 Annotator Skew Within a Language.
  • D Aya Collection: Additional Analysis
  • D.1 Translation Quality
  • D.2 Tasks Covered Across Templated and Translated Datasets
  • D.3 Prompt and Completion Lengths
  • D.4 Quality Assessment of All Different Data Sources
  • E Aya Evaluation Suite: Additional Analysis
  • E.1 Post-Editing the DOLLY-MACHINE-TRANSLATED Test Set
  • E.1.1 Annotators
  • E.1.2 Annotation Process
  • E.1.3 Instructions
  • E.1.4 Post-Editing Effort
  • F Additional Figures
  • G Language Representation via Community
  • G.1 Division by Regions
  • G.2 Language Ambassadors
  • G.2.1 Regional Leads
  • G.3 Communication
  • G.3.1 Platforms
  • G.3.2 Meetings
  • H Data Cards

Knowls

  1. Knowl 1 — The Aya Multilingual Instruction Dataset

    model/method

    The Aya Dataset is an open-source, human-curated multilingual instruction fine-tuning dataset consisting of 204,114 prompt-completion pairs spanning 65 typologically diverse languages (22 high-resource, 12 mid-resource, and 31 low-resource languages).

    The dataset aggregates original prompt-completion pairs authored by fluent speakers and curated re-annotations of existing datasets:

    Annotation Source Count
    Original Annotations 138,844
    Re-Annotations (xP3 datasets) 2,859
    Re-Annotations (Translated datasets) 7,757
    Re-Annotations (Templated datasets) 11,013
    Re-Annotations (Original Annotations) 43,641
    Aya Dataset Total 204,114

    To eliminate duplicates and ensure meaningful contributions, re-annotated pairs were retained only if the character-level Levenshtein edit distance dd between the original and edited versions satisfied d≥5d \ge 5. Furthermore, a language was included in the release only if it accumulated at least 50 validated contributions.

  2. Knowl 2 — The Aya Collection for Scaled Multilingual Instruction Tuning

    model/method

    The Aya Collection is an open-access multilingual instruction fine-tuning corpus containing 513,579,625 instances across 114 languages covering 11 fine-grained Natural Language Processing (NLP) task categories. The collection combines three components:

    1. Templated Datasets (44 datasets): Monolingual and multilingual datasets converted into instruction format by applying natural language templates created by fluent speakers via the PromptSource framework. Prompts include monolingual instructions or code-mixed formats where English instructions wrap non-English source data.
    2. Translated Datasets (19 datasets): 19 instruction datasets (including subsets of Flan 2022, Dolly-v2, SODA, Mintaka, and PAWS-Wiki) translated into 101 languages and dialects using the No Language Left Behind (NLLB 3.3B parameter) machine translation model.
    3. Aya Dataset: The 204,114 human-curated prompt-completion instances.

    The tasks are organized into three broad task types: Question Answering (QA), Natural Language Generation (summarization, translation, paraphrasing, dialogue, and text simplification), and Text Classification (sentiment analysis, information extraction, named entity recognition, event linking, natural language inference, and scientific document representation).

  3. Knowl 3 — The Aya Evaluation Suite for Multilingual Open-Ended Generation

    experimental setup

    The Aya Evaluation Suite is an open-access evaluation benchmark comprising 25,750 prompt instances covering 101 languages (114 dialects) tailored to evaluate the open-ended text generation, planning, and conversational capabilities of large language models (LLMs). It contains three complementary splits:

    1. AYA-HUMAN-ANNOTATED: 1,750 prompts (250 randomly chosen human-written prompts across 7 languages). Languages were selected for structural, geographical, script, and resource diversity from languages with at least 2,000 original annotations: English (high-resource, Latin script, Indo-European), Portuguese (high-resource, Latin script, Indo-European), Simplified Chinese (high-resource, Han script, Sino-Tibetan), Standard Arabic (high-resource, Arabic script, Afro-Asiatic), Turkish (high-resource, Latin script, Turkic), Telugu (low-resource, Telugu script, Dravidian), and Yoruba (low-resource, Latin script, Atlantic-Congo).
    2. DOLLY-MACHINE-TRANSLATED: 200 culturally universal English prompts filtered from 500 candidate Dolly prompts by multiple human reviewers to remove region-specific idioms or trivia, translated into 101 languages (114 dialects) using NLLB 3.3B.
    3. DOLLY-HUMAN-EDITED: 200 post-edited prompts across 6 languages (Arabic, French, Hindi, Russian, Serbian, Spanish) produced by professional annotators to eliminate machine translation artifacts and translationese while preserving semantic intent.
  4. Knowl 4 — Find-Fix-Verify Annotation and Quality Control Pipeline

    model/method

    The Aya data curation framework uses an open-science, find-fix-verify pipeline across three sequential annotation interfaces:

    1. Original Annotations (Find/Create): Fluent speakers submit original, culturally grounded prompt-completion pairs without utilizing language model outputs.
    2. Re-Annotations (Fix): Annotators inspect existing prompt-completion pairs from candidate sources, provide binary quality feedback (upvote or downvote), and edit prompts and completions to improve fluency, detail, and grammatical correctness.
    3. Annotation Feedback (Verify): Annotators review peer-submitted re-annotations and assess their quality.

    Each annotator receives an individual quality score (the Aya score), calculated as the average rating assigned to their contributions by other fluent annotators reviewing that language. This score populates a dynamic leaderboard to incentivize high-quality submissions.

  5. Knowl 5 — Perceived Quality Across Multilingual Curation Paradigms

    empirical result

    Binary peer feedback on prompt-completion pairs across 40 datasets (evaluating datasets with at least 20 feedback votes) measured data quality using the average approval ratio:

    Approval Ratio=T+T\text{Approval Ratio} = \frac{T_+}{T}

    where T+T_+ is the total number of upvotes and TT is the total number of votes for a given dataset.

    Human-authored and human-edited curation paradigms achieved higher approval ratios than templated or automatically translated corpora:

    • Aya Dataset (Original Annotations): Achieved the highest average approval ratio at 0.810.81.
    • Translated Datasets (NLLB-translated): Achieved an average approval ratio of 0.700.70.
    • Templated Datasets: Achieved an average approval ratio of 0.660.66.
    • xP3 Datasets: Achieved the lowest average approval ratio of 0.500.50.
  6. Knowl 6 — Post-Editing Effort and Translation Error Rates in Dolly Evaluation Prompts

    empirical result

    Professional post-editing of 200 machine-translated Dolly evaluation prompts across six target languages shows substantial variation in translation quality and error severity, quantified by the percentage of prompts requiring edits, the Human-targeted Translation Edit Rate (HTER, word-level edit distance where higher values indicate greater effort), and the Human-targeted Character F-Score (HChrF, character-level overlap where lower values indicate greater effort):

    Language % of Prompts Edited HTER HChrF
    Arabic 41.0% 10.78 92.74
    French 84.5% 5.56 96.81
    Hindi 60.0% 6.16 95.00
    Russian 86.5% 37.43 75.92
    Serbian 72.5% 9.06 92.79
    Spanish 75.5% 9.13 93.25

    In all six evaluated languages, at least 41.0% of machine-translated prompts required human edits. Russian required the greatest post-editing effort (37.4337.43 HTER and 75.9275.92 HChrF), whereas French prompts, though frequently edited (84.5%84.5\% of prompts), required only minor corrections (5.565.56 HTER and 96.8196.81 HChrF).

  7. Knowl 7 — Impact of Human Re-Annotation on Completion Length and Perceived Quality

    empirical result

    Human re-annotation systematically increases the length and detail of instruction completions:

    • Across all source categories (xP3, machine translations, templated datasets, and original submissions), completion length increased by an average of 25% following human re-annotation.
    • For original Aya submissions, re-annotated completions were 40% longer on average than their initial unedited lengths.
    • Combined prompt and completion character length exhibits a positive correlation with average human approval ratings, with a Pearson correlation coefficient of r=0.27r = 0.27. Longer, detailed completions containing complete grammatical sentences correlate with higher perceived quality ratings from fluent reviewers.
  8. Knowl 8 — Annotator Skew and Imbalance in Participatory Multilingual Annotation

    empirical result

    Participatory data collection involving 2,997 registered contributors from 119 countries exhibited heavy-tailed contribution distributions across and within languages:

    • Across languages: Total contributions ranged from 14,597 instances for Malagasy to 79 instances for Kurdish.
    • Within languages: The median number of active annotators per language was 15 (mean: 24.75). For 12 languages, the top 5 most active annotators produced 100% of all submitted instances. For Zulu and Sindhi, a single contributor accounted for 100% of all annotations and re-annotations (top-1 contributor ratio of 1.0).
    • The most evenly distributed contributions occurred in Malagasy, Tamil, Nepali, Hindi, English, and Portuguese. English had the highest number of unique annotators (130 contributors, of which 95 were non-native speakers annotating English as a second language).
  9. Knowl 9 — Cross-Lingual Variations in Prompt and Completion Lengths

    empirical result

    The ratio of prompt to completion lengths in the Aya Dataset varies substantially across languages:

    • In Japanese, completions are on average 31% shorter than prompts.
    • In Urdu and Yoruba, completions are substantially longer than prompts: average completion lengths exceed average prompt lengths by 1258% for Urdu and 2516% for Yoruba.
    • Comparing across languages, the average completion length in Yoruba is 1591% longer than the average prompt length in Japanese. Low-resource languages occupy both the upper and lower extremes of total length distributions.
  10. Knowl 10 — Linguistic, Dialectal, and Modal Coverage Limitations in Aya

    limitation

    The Aya Dataset and Aya Collection are subject to several structural limitations:

    • Exclusion of unwritten languages: The collection process requires written text, leaving out spoken languages without established writing systems (which constitute approximately half of the world's ~7,000 languages).
    • Dialectal omissions: Sub-regional spoken dialects (such as regional varieties of Malay spoken across Malaysia) are largely omitted, as submissions were standardized under parent language codes.
    • Monolingual isolation: The project enforced language isolation per entry to simplify classification and fine-tuning pipelines, failing to capture natural conversational code-switching.
    • Domain skew: In low-resource languages where existing digital corpora are predominantly news text (such as several African languages), templated datasets are heavily skewed toward journalistic style and vocabulary rather than colloquial language.

Coverage note — None was omitted; all key contributions—the Aya Dataset, the Aya Collection, the Aya Evaluation Suite, the participatory annotation pipeline, quality/length empirical analyses, and dataset limitations—are fully represented.

References

  1. 1.Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, and Benoı̂t Sagot. 2022. Towards a cleaner document-oriented multilingual crawled corpus. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 4344–4355, Marseille, France. European Language Resources Association.
  2. 2.Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2021. Ungoliant: An optimized pipeline for the generation of a very large-scale multilingual web corpus. In CMLC 2021-9th Workshop on Challenges in the Management of Large Corpora, Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-9) 2021. Limerick, 12 July 2021 (Online-Event), pages 1 – 9, Mannheim. Leibniz-Institut fur Deutsche Sprache.
  3. 3.Tilahun Abedissa, Ricardo Usbeck, and Yaregal Assabie. 2023. AmQA: Amharic Question Answering Dataset. arXiv preprint arXiv:2303.03290.
  4. 4.Gilles Adda, Sebastian Stuker, Martine Adda-Decker, Odette Ambouroue, Laurent Besacier, David Blachon, Hélène Bonneau-Maynard, Pierre Godard, Fatima Hamlaoui, Dmitry Idiatov, Guy-Noel Kouarata, Lori Lamel, Emmanuel-Moselly Makasso, Annie Rialland, Mark Van de Velde, François Yvon, and Sabine Zerbian. 2016. Breaking the unwritten language barrier: The bulb project. Procedia Computer Science, 81:8–14. SLTU-2016 5th Workshop on Spoken Language Technologies for Under-resourced languages 09-12 May 2016 Yogyakarta, Indonesia.
  5. 5.David Ifeoluwa Adelani, Marek Masiak, Israel Abebe Azime, Jesujoba Alabi, Atnafu Lambebo Tonja, Christine Mwase, Odunayo Ogundepo, Bonaventure F. P. Dossou, Akintunde Oladipo, Doreen Nixdorf, Chris Chinenye Emezue, sana al azzawi, Blessing Sibanda, Davis David, Lolwethu Ndolela, Jonathan Mukiibi, Tunde Ajayi, Tatiana Moteu, Brian Odhiambo, Abraham Owodunni, Nnaemeka Obiefuna, Muhidin Mohamed, Shamsuddeen Hassan Muhammad, Teshome Mulugeta Ababu, Saheed Abdullahi Salahudeen, Mesay Gemeda Yigezu, Tajuddeen Gwadabe, Idris Abdulmumin, Mahlet Taye, Oluwabusayo Awoyomi, Iyanuoluwa Shode, Tolulope Adelani, Habiba Abdulganiyu, Abdul-Hakeem Omotayo, Adetola Adeeko, Abeeb Afolabi, Anuoluwapo Aremu, Olanrewaju Samuel, Clemencia Siro, Wangari Kimotho, Onyekachi Ogbu, Chinedu Mbonu, Chiamaka Chukwuneke, Samuel Fanijo, Jessica Ojo, Oyinkansola Awosan, Tadesse Kebede, Toadoum Sari Sakayo, Pamela Nyatsine, Freedmore Sidume, Oreen Yousuf, Mardiyyah Oduwole, Tshinu Tshinu, Ussen Kimanuka, Thina Diko, Siyanda Nxakama, Sinodos Nigusse, Abdulmejid Johar, Shafie Mohamed, Fuad Mire Hassan, Moges Ahmed Mehamed, Evrard Ngabire, Jules Jules, Ivan Ssenkungu, and Pontus Stenetorp. 2023. MasakhaNEWS: News Topic Classification for African Languages.
  6. 6.Asif Agha. 2006. Language and Social Relations. Studies in the Social and Cultural Foundations of Language. Cambridge University Press.
  7. 7.Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019. Massively multilingual neural machine translation.
  8. 8.AhmadMustafa. 2023a. Urdu-Instruct-News-Article-Generation. https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Article-Generation. Accessed: 2023-11-28.
  9. 9.AhmadMustafa. 2023b. Urdu-Instruct-News-Category-Classification. https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Category-Classification. Accessed: 2023-11-28.
  10. 10.AhmadMustafa. 2023c. Urdu-Instruct-News-Headline-Generation. https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Headline-Generation. Accessed: 2023-11-28.
  11. 11.AI Tamil Nadu. 2023a. Tamil stories. https://huggingface.co/datasets/aitamilnadu/tamil_stories. Accessed: 2023-12-15.
  12. 12.AI Tamil Nadu. 2023b. Thirukkural Instruct. https://huggingface.co/datasets/aitamilnadu/thirukkural_instruct. Accessed: 2023-11-30.
  13. 13.Christopher Akiki, Giada Pistilli, Margot Mieskes, Matthias Gallé, Thomas Wolf, Suzana Ilic, and Yacine Jernite. 2022. Bigscience: A case study in the social construction of a multilingual large language model.
  14. 14.Waseem AlShikh, Manhal Daaboul, Kirk Goddard, Brock Imel, Kiran Kamble, Parikshith Kulkarni, and Melisa Russak. 2023. Becoming self-instruct: introducing early stopping criteria for minimal instruct tuning. arXiv preprint arXiv:2307.03692.
  15. 15.Hossein Amirkhani, Mohammad AzariJafari, Soroush Faridan-Jahromi, Zeinab Kouhkan, Zohreh Pourjafari, and Azadeh Amirak. 2023. Farstail: A Persian natural language inference dataset. Soft Computing.
  16. 16.Lauri Andress, Tristen Hall, Sheila Davis, Judith Levine, Kimberly Cripps, and Dominique Guinn. 2020. Addressing power dynamics in community-engaged research partnerships. Journal of Patient-Reported Outcomes 4: 24.
  17. 17.Md Adnan Arefeen, Biplob Debnath, and Srimat Chakradhar. 2023. Leancontext: Cost-efficient domain-specific question answering using llms. arXiv preprint arXiv:2309.00841.
  18. 18.Kailash Awati and Simon Buckingham Shum. 2015. Big data metaphors we live by. Towards Data Science.
  19. 19.Paul Azunre, Salomey Osei, Salomey Afua Addo, Lawrence Asamoah Adu-Gyamfi, Stephen Moore, Bernard Adabankah, Bernard Opoku, Clara Asare-Nyarko, Samuel Nyarko, Cynthia Amoaba, Esther Dansoa Appiah, Felix Akwerh, Richard Nii Lante Lawson, Joel Budu, Emmanuel Debrah, Nana Adowaa Boateng, Wisdom Ofori, Edwin Buabeng-Munkoh, Franklin Adjei, Isaac. K. E. Ampomah, Joseph Otoo., Reindorf Nartey Borkor, Standylove Birago Mensah, Lucien Mensah, Mark Amoako Marcel, Anokye Acheampong Amponsah, and James Ben Hayfron-Acquah. 2021. Nlp for ghanaian languages. ArXiv, abs/2103.15475.
  20. 20.Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al. 2022. Promptsource: An integrated development environment and repository for natural language prompts. arXiv preprint arXiv:2202.01279.
  21. 21.Loı̈c Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023. Seamlessm4t-massively multilingual & multimodal machine translation. arXiv preprint arXiv:2308.11596.
  22. 22.Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. 2020. Beat the ai: Investigating adversarial human annotation for reading comprehension. Transactions of the Association for Computational Linguistics, 8:662–678.
  23. 23.Michael S. Bernstein, Greg Little, Robert C. Miller, Björn Hartmann, Mark S. Ackerman, David R. Karger, David Crowell, and Katrina Panovich. 2015. Soylent: a word processor with a crowd inside. Commun. ACM, 58(8):85–94.
  24. 24.Steven Bird. 2022. Local languages, third spaces, and other high-resource scenarios. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7817–7829, Dublin, Ireland. Association for Computational Linguistics.
  25. 25.Abeba Birhane, William Isaac, Vinodkumar Prabhakaran, Mark Diaz, Madeleine Clare Elish, Iason Gabriel, and Shakir Mohamed. 2022. Power to the people? opportunities and challenges for participatory ai. In Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO ’22. ACM.
  26. 26.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence.
  27. 27.Verena Blaschke, Hinrich Schuetze, and Barbara Plank. 2023. A survey of corpora for Germanic low-resource languages and dialects. In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pages 392–414, Tórshavn, Faroe Islands. University of Tartu Library.
  28. 28.Jan A. Botha, Manaal Faruqui, John Alex, Jason Baldridge, and Dipanjan Das. 2018. Learning to split and rephrase from Wikipedia edit history. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 732–737, Brussels, Belgium. Association for Computational Linguistics.
  29. 29.Meriem Boubdir, Edward Kim, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. 2023. Which prompts make the difference? data prioritization for efficient human llm evaluation.
  30. 30.Anne Bowser, Caren Cooper, Alex De Sherbinin, Andrea Wiggins, Peter Brenton, Tyng-Ruey Chuang, Elaine Faustman, Mordechai Haklay, and Metis Meloche. 2020. Still in need of norms: the state of the data in citizen science. Citizen Science: Theory and Practice, 5(1).
  31. 31.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  32. 32.Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Winata, Bryan Wilie, Fajri Koto, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, et al. 2023. Nusacrowd: Open source initiative for indonesian nlp resources. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13745–13818.
  33. 33.Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8875–8898, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  34. 34.Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. 2023. Poisoning web-scale training datasets is practical.
  35. 35.Joe Castaldo. 2023. “AI chatbots fall short in dozens of languages. A non-profit project aims to fix that”. Globe and Mail.
  36. 36.Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, et al. 2023a. Alpagasus: Training a better alpaca with fewer data. arXiv preprint arXiv:2307.08701.
  37. 37.Pinzhen Chen, Shaoxiong Ji, Nikolay Bogoychev, Barry Haddow, and Kenneth Heafield. 2023b. Monolingual or multilingual instruction tuning: Which makes a better alpaca.
  38. 38.Sanxing Chen, Yongqiang Chen, and Börje F. Karlsson. 2023c. Dataset and baseline system for multi-lingual extraction and normalization of temporal and numerical expressions. arXiv preprint arXiv:2303.18103.
  39. 39.Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, and David Ha. 2018. Deep learning for classical japanese literature. ArXiv, abs/1812.01718.
  40. 40.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics.
  41. 41.Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. Xnli: Evaluating cross-lingual sentence representations. arXiv preprint arXiv:1809.05053.
  42. 42.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM.
  43. 43.ConseggioLigure. 2023a. Lij news instruct ita-lij. https://huggingface.co/datasets/ConseggioLigure/lijnews-instruct-ita-lij. Accessed: 2023-11-28.
  44. 44.ConseggioLigure. 2023b. Lij news instruct lij-ita. https://huggingface.co/datasets/ConseggioLigure/lijnews-instruct-lij-ita. Accessed: 2023-11-28.
  45. 45.ConseggioLigure. 2023c. Seed instruct eng-lij. https://huggingface.co/datasets/ConseggioLigure/seed-instruct-eng-lij. Accessed: 2023-11-28.
  46. 46.ConseggioLigure. 2023d. Seed instruct lij-eng. https://huggingface.co/datasets/ConseggioLigure/seed-instruct-lij-eng. Accessed: 2023-11-28.
  47. 47.Eric Corbett, Emily Denton, and Sheena Erete. 2023. Power and public participation in ai. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO ’23, New York, NY, USA. Association for Computing Machinery.
  48. 48.Kate Crawford. 2021. Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press, New Haven, Connecticut.
  49. 49.Raj Dabre, Chenhui Chu, and Anoop Kunchukuttan. 2020. A survey of multilingual neural machine translation. ACM Comput. Surv., 53(5).
  50. 50.Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The participatory turn in ai design: Theoretical foundations and the current state of practice. Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization.
  51. 51.Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2023. Multilingual jailbreak challenges in large language models. ArXiv, abs/2310.06474.
  52. 52.desik98. 2023. Telugu Riddles. https://huggingface.co/datasets/desik98/TeluguRiddles. Accessed: 2023-11-30.
  53. 53.Kaustubh D Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Shrivastava, Samson Tan, et al. 2021. Nl-augmenter: A framework for task-sensitive natural language augmentation. arXiv preprint arXiv:2112.02721.
  54. 54.Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2023. Towards leaving no Indic language behind: Building monolingual corpora, benchmark and models for Indic languages. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12402–12426, Toronto, Canada. Association for Computational Linguistics.
  55. 55.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021. Documenting the english colossal clean crawled corpus. CoRR, abs/2104.08758.
  56. 56.A Seza Doğruöz, Sunayana Sitaram, and Zheng-Xin Yong. 2023. Representativeness as a forgotten lesson for multilingual and code-switched data collection and preparation. arXiv preprint arXiv:2310.20470.
  57. 57.Oliver Falck, Stephan Heblich, Alfred Lameli, and Jens Südekum. 2012. Dialects, cultural identity, and economic exchange. Journal of urban economics, 72(2-3):225–239.
  58. 58.Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021. Beyond english-centric multilingual machine translation. Journal of Machine Learning Research, 22(107):1–48.
  59. 59.Mehrdad Farahani, Mohammad Gharachorloo, and Mohammad Manthouri. 2021. Leveraging parsbert and pretrained mt5 for persian abstractive text summarization. In 2021 26th International Computer Conference, Computer Society of Iran (CSICC), pages 1–6. IEEE.
  60. 60.∀, Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. 2020. Participatory research for low-resourced machine translation: A case study in African languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online. Association for Computational Linguistics.
  61. 61.ganeshjcs. 2023a. Hindi Article Summarization. https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization. Accessed: 2023-11-28.
  62. 62.ganeshjcs. 2023b. Hindi Headline Article Generation. https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation. Accessed: 2023-11-28.
  63. 63.Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina Mcmillan-major, Anna Shvets, Ashish Upadhyay, and Bernd Bohnet. 2022. GEMv2: Multilingual NLG benchmarking in a single line of code. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 266–281, Abu Dhabi, UAE. Association for Computational Linguistics.
  64. 64.Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2023. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. Journal of Artificial Intelligence Research, 77:103–166.
  65. 65.Eldho Ittan George. 2023a. Aya Indic Sentiment. https://huggingface.co/datasets/el2e10/aya-indicsentiment. Accessed: 2023-11-28.
  66. 66.Eldho Ittan George. 2023b. Aya Paraphrase. https://huggingface.co/datasets/el2e10/aya-paraphrase. Accessed: 2023-11-28.
  67. 67.Charles Goodwin. 2017. Co-Operative Action. Learning in Doing: Social, Cognitive and Computational Perspectives. Cambridge University Press.
  68. 68.Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021a. Larger-Scale Transformers for Multilingual Masked Language Modeling. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021). Association for Computational Linguistics.
  69. 69.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzman, and Angela Fan. 2021b. The flores-101 evaluation benchmark for low-resource and multilingual machine translation.
  70. 70.Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, and Hannaneh Hajishirzi. 2024. OLMo: Accelerating the Science of Language Models. arXiv preprint.
  71. 71.Adriana Guevara-Rukoz, Isin Demirsahin, Fei He, Shan-Hui Cathy Chu, Supheakmungkol Sarin, Knot Pipatsrisawat, Alexander Gutkin, Alena Butryna, and Oddur Kjartansson. 2020. Crowdsourcing Latin American Spanish for low-resource text-to-speech. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6504–6513, Marseille, France. European Language Resources Association.
  72. 72.Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023. Textbooks are all you need.
  73. 73.Himanshu Gupta, Kevin Scaria, Ujjwala Anantheswaran, Shreyas Verma, Mihir Parmar, Saurabh Arjun Sawant, Chitta Baral, and Swaroop Mishra. 2023. Targen: Targeted data generation with large language models.
  74. 74.Margot Hanley, Apoorv Khandelwal, Hadar Averbuch-Elor, Noah Snavely, and Helen Nissenbaum. 2020. An ethical highlighter for people-centric dataset creation. arXiv preprint arXiv:2011.13583.
  75. 75.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL-sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  76. 76.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in neural information processing systems, pages 1693–1701.
  77. 77.hghader1. 2023. FarsTail-Instruct-LLM. https://huggingface.co/datasets/hghader1/FarsTail-Instruct-LLM. Accessed: 2023-11-28.
  78. 78.Oskar Holmström and Ehsan Doostmohammadi. 2023. Making instruction finetuning accessible to non-English languages: A case study on Swedish models. In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pages 634–642, Tórshavn, Faroe Islands. University of Tartu Library.
  79. 79.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2023. Unnatural instructions: Tuning language models with (almost) no human labor. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14409–14428, Toronto, Canada. Association for Computational Linguistics.
  80. 80.Dirk Hovy and Shrimai Prabhumoye. 2021. Five sources of bias in natural language processing. Language and Linguistics Compass, 15(8):e12432.
  81. 81.Pei-Yun Hsueh, Prem Melville, and Vikas Sindhwani. 2009. Data quality from crowdsourcing: a study of annotation selection criteria. In Proceedings of the NAACL HLT 2009 workshop on active learning for natural language processing, pages 27–35.
  82. 82.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models.
  83. 83.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In International Conference on Machine Learning, pages 4411–4421. PMLR.
  84. 84.Kuan-Hao Huang, I-Hung Hsu, Premkumar Natarajan, Kai-Wei Chang, and Nanyun Peng. 2022. Multilingual generative language models for zero-shot cross-lingual event argument extraction.
  85. 85.Khalid Hussain, Nimra Mughal, Irfan Ali, Saif Hassan, and Sher Muhammad Daudpota. 2021. Urdu News Dataset 1M. Technical report, Mendeley Data, V3.
  86. 86.Mika Hämäläinen. 2021. Endangered Languages are not Low-Resourced!, page 1–11. University of Helsinki.
  87. 87.Iftitahu. 2023a. Indonesian Instruct Stories. https://huggingface.co/datasets/Iftitahu/indonesian_instruct_stories. Accessed: 2023-11-28.
  88. 88.Iftitahu. 2023b. Javanese Instruct Stories. https://huggingface.co/datasets/Iftitahu/javanese_instruct_stories. Accessed: 2023-11-28.
  89. 89.Iftitahu. 2023c. Sudanese Instruct Stories. https://huggingface.co/datasets/Iftitahu/sundanese_instruct_stories. Accessed: 2023-11-28.
  90. 90.Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, et al. 2022. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017.
  91. 91.jjzha. 2023. IMDB Dutch Instruct. https://huggingface.co/datasets/jjzha/imdb-dutch-instruct. Accessed: 2023-11-28.
  92. 92.Joseph Cheung. 2023. GuanacoDataset (Revision 8cf0d29) .
  93. 93.Pratik Joshi, Christain Barnes, Sebastin Santy, Simran Khanuja, Sanket Shah, Anirudh Srinivasan, Satwik Bhattamishra, Sunayana Sitaram, Monojit Choudhury, and Kalika Bali. 2019. Unsung challenges of building and deploying language technologies for low resource language communities. In Proceedings of the 16th International Conference on Natural Language Processing, pages 211–219, International Institute of Information Technology, Hyderabad, India. NLP Association of India.
  94. 94.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  95. 95.Jenna Kanerva, Filip Ginter, Li-Hsin Chang, Iiro Rastas, Valtteri Skantsi, Jemina Kilpeläinen, Hanna-Mari Kupari, Jenna Saarni, Maija Sevón, and Otto Tarkka. 2021. Finnish paraphrase corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pages 288–298, Reykjavik, Iceland (Online). Linköping University Electronic Press, Sweden.
  96. 96.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. Unifiedqa: Crossing format boundaries with a single qa system. arXiv preprint arXiv:2005.00700.
  97. 97.Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, and Yejin Choi. 2022. Soda: Million-scale dialogue distillation with social commonsense contextualization. ArXiv, abs/2212.10465.
  98. 98.Seungone Kim, Se June Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo. 2023. The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning. arXiv preprint arXiv:2305.14045.
  99. 99.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al. 2023. Openassistant conversations–democratizing large language model alignment. arXiv preprint arXiv:2304.07327.
  100. 100.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoı̂t Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  101. 101.Venni V. Krishna. 2020. Open science and its enemies: Challenges for a sustainable science–society social contract. Journal of Open Innovation: Technology, Market, and Complexity, 6(3):61.
  102. 102.Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, et al. 2023. Madlad-400: A multilingual and document-level large audited dataset. arXiv preprint arXiv:2309.04662.
  103. 103.Anoop Kunchukuttan, Siddharth Jain, and Rahul Kejriwal. 2021. A large-scale evaluation of neural machine transliteration for Indic languages. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3469–3475, Online. Association for Computational Linguistics.
  104. 104.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  105. 105.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations – democratizing large language model alignment.
  106. 106.Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4034–4048, Online. Association for Computational Linguistics.
  107. 107.Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A Rossi, and Thien Huu Nguyen. 2023. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback. arXiv preprint arXiv:2307.16039.
  108. 108.Guillaume Lample and Alexis Conneau. 2019. Cross-lingual language model pretraining. arXiv preprint arXiv:1901.07291.
  109. 109.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022. The bigscience roots corpus: A 1.6 tb composite multilingual dataset. Advances in Neural Information Processing Systems, 35:31809–31826.
  110. 110.Nayeon Lee, Chani Jung, and Alice Oh. 2023. Hate speech classifiers are culturally insensitive. In Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP), pages 35–46, Dubrovnik, Croatia. Association for Computational Linguistics.
  111. 111.Heather Lent, Kelechi Ogueji, Miryam de Lhoneux, Orevaoghene Ahia, and Anders Søgaard. 2022. What a creole wants, what a creole needs. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6439–6449, Marseille, France. European Language Resources Association.
  112. 112.Wei Qi Leong, Jian Gang Ngui, Yosephine Susanto, Hamsawardhini Rengarajan, Kengatharaiyer Sarveswaran, and William Chandra Tjhi. 2023. Bhasa: A holistic southeast asian linguistic and cultural evaluation suite for large language models.
  113. 113.Vladimir I Levenshtein et al. 1966. Binary codes capable of correcting deletions, insertions, and reversals. In Soviet physics doklady, volume 10, pages 707–710. Soviet Union.
  114. 114.Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. Mlqa: Evaluating cross-lingual extractive question answering.
  115. 115.Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. 2023a. Bactrian-x: A multilingual replicable instruction-following model with low-rank adaptation. arXiv preprint arXiv:2305.15011.
  116. 116.Haoran Li, Yulin Chen, Jinglong Luo, Yan Kang, Xiaojin Zhang, Qi Hu, Chunkit Chan, and Yangqiu Song. 2023b. Privacy in large language models: Attacks, defenses and future directions. ArXiv, abs/2310.10383.
  117. 117.Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, Lingpeng Kong, and Qi Liu. 2023c. M3it: A large-scale dataset towards multimodal multilingual instruction tuning.
  118. 118.Constantine Lignos, Nolan Holley, Chester Palen-Michel, and Jonne Sälevä. 2022. Toward more meaningful resources for lower-resourced languages. In Findings of the Association for Computational Linguistics: ACL 2022, pages 523–532, Dublin, Ireland. Association for Computational Linguistics.
  119. 119.Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, and Xiang Ren. 2021. Common sense beyond English: Evaluating and improving multilingual language models for commonsense reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1274–1287, Online. Association for Computational Linguistics.
  120. 120.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022. Few-shot learning with multilingual generative language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9019–9052, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  121. 121.Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott. 2021. Visually grounded reasoning across languages and cultures. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10467–10485, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  122. 122.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023a. The flan collection: Designing data and methods for effective instruction tuning. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  123. 123.Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al. 2023b. A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity. arXiv preprint arXiv:2305.13169.
  124. 124.Yanni Alexander Loukissas. 2019. All Data Are Local: Thinking Critically in a Data-Driven Society. MIT Press, Cambridge, Massachusetts.
  125. 125.Lalita Lowphansirikul, Charin Polpanumas, Attapol T Rutherford, and Sarana Nutanong. 2022. A large English–Thai parallel corpus from the web and machine-generated text. Language Resources and Evaluation, 56(2):477–499.
  126. 126.Alexandra Luccioni and Joseph Viviano. 2021. What’s in the box? an analysis of undesirable content in the Common Crawl corpus. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 182–189, Online. Association for Computational Linguistics.
  127. 127.Lucia Specia, Nicola Cancedda, and Marc Dymetman. 2010. A Dataset for Assessing Machine Translation Evaluation Metrics. International Conference on Language Resources and Evaluation.
  128. 128.Nils Lukas, A. Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-B’eguelin. 2023. Analyzing leakage of personally identifiable information in language models. 2023 IEEE Symposium on Security and Privacy (SP), pages 346–363.
  129. 129.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct.
  130. 130.Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, et al. 2023. Fingpt: Large generative models for a small language. arXiv preprint arXiv:2311.05640.
  131. 131.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  132. 132.Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, and Francisco Guzmán. 2023. Small data, big impact: Leveraging minimal data for effective machine translation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2740–2756, Toronto, Canada. Association for Computational Linguistics.
  133. 133.Max Marion, Ahmet Üstün, Luiza Pozzobon, Alex Wang, Marzieh Fadaee, and Sara Hooker. 2023. When less is more: Investigating data pruning for pretraining llms at scale.
  134. 134.Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa, Hila Gonen, Joseph Marvin Imperial, Börje F. Karlsson, Peiqin Lin, Nikola Ljubešić, LJ Miranda, Barbara Plank, Arij Riabi, and Yuval Pinter. 2023. Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark. arXiv preprint arXiv:2311.09122.
  135. 135.Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018. The natural language decathlon: Multitask learning as question answering. arXiv preprint arXiv:1806.08730.
  136. 136.Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR), 54(6):1–35.
  137. 137.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021. Metaicl: Learning to learn in context. arXiv preprint arXiv:2110.15943.
  138. 138.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In ACL.
  139. 139.Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre. 2023a. Octopack: Instruction tuning code large language models. arXiv preprint arXiv:2308.07124.
  140. 140.Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023b. Scaling data-constrained language models. arXiv preprint arXiv:2305.16264.
  141. 141.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023c. Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15991–16111, Toronto, Canada. Association for Computational Linguistics.
  142. 142.Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, and Stephen Arthur. 2023. AfriSenti: A Twitter sentiment analysis benchmark for African languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13968–13981, Singapore. Association for Computational Linguistics.
  143. 143.Carol Myers-Scotton. 2017. Code-switching. The handbook of sociolinguistics, pages 217–237.
  144. 144.Gabriel Nakamura, Bruno Soares, Valério Pillar, José Diniz-Filho, and Leandro Duarte. 2023. Three pathways to better recognize the expertise of global south researchers. npj Biodiversity.
  145. 145.Tarek Naous, Michael Joseph Ryan, and Wei Xu. 2023. Having beer after prayer? measuring cultural bias in large language models. ArXiv, abs/2305.14456.
  146. 146.Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. 2023. Scalable extraction of training data from (production) language models.
  147. 147.Huu Nguyen, Sameer Suri, Ken Tsui, and Christoph Schuhmann. 2023. The open instruction generalist (oig) dataset. https://laion.ai/blog/oig-dataset/.
  148. 148.Jinjie Ni, Fuzhao Xue, Kabir Jain, Mahir Hitesh Shah, Zangwei Zheng, and Yang You. 2023. Instruction in the wild: A user-based instruction dataset. https://github.com/XueFuzhao/InstructionWild.
  149. 149.Joel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis, and Daniel E Ho. 2023. Multilegalpile: A 689gb multilingual legal corpus. arXiv preprint arXiv:2306.02069.
  150. 150.NLLB-Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation.
  151. 151.Gianluca Nogara, Francesco Pierri, Stefano Cresci, Luca Luceri, Petter Törnberg, and Silvia Giordano. 2023. Toxic bias: Perspective api misreads german as more toxic.
  152. 152.Odunayo Ogundepo, Tajuddeen Gwadabe, Clara Rivera, Jonathan Clark, Sebastian Ruder, David Adelani, Bonaventure Dossou, Abdou Diop, Claytone Sikasote, Gilles Hacheme, Happy Buzaaba, Ignatius Ezeani, Rooweither Mabuya, Salomey Osei, Chris Emezue, Albert Kahira, Shamsuddeen Muhammad, Akintunde Oladipo, Abraham Owodunni, Atnafu Tonja, Iyanuoluwa Shode, Akari Asai, Anuoluwapo Aremu, Ayodele Awokoya, Bernard Opoku, Chiamaka Chukwuneke, Christine Mwase, Clemencia Siro, Stephen Arthur, Tunde Ajayi, Verrah Otiende, Andre Rubungo, Boyd Sinkala, Daniel Ajisafe, Emeka Onwuegbuzia, Falalu Lawan, Ibrahim Ahmad, Jesujoba Alabi, Chinedu Mbonu, Mofetoluwa Adeyemi, Mofya Phiri, Orevaoghene Ahia, Ruqayya Iro, and Sonia Adhiambo. 2023. Afriqa: Cross-lingual open-retrieval question answering for african languages. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 14957–14972, Singapore. Association for Computational Linguistics.
  153. 153.Pedro Javier Ortiz Su’arez, Benoit Sagot, and Laurent Romary. 2019. Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures. In 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7), Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-7) 2019. Cardiff, 22nd July 2019, pages 9 – 16, Mannheim. Leibniz-Institut f"ur Deutsche Sprache.
  154. 154.osyvokon. 2023. UA-GEC instruction tuning . https://huggingface.co/datasets/osyvokon/ua_gec_instruction_tuning. Accessed: 2023-11-28.
  155. 155.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The lambada dataset: Word prediction requiring a broad discourse context.
  156. 156.Michael Park, Erin Leahey, and Russell J. Funk. 2023. Papers and patents are becoming less disruptive over time. Nature, 613:138–144.
  157. 157.Kenny Peng, Arunesh Mathur, and Arvind Narayanan. 2021. Mitigating dataset harms requires stewardship: Lessons from 1000 papers. arXiv preprint arXiv:2108.02922.
  158. 158.Laura Perez-Beltrachini and Mirella Lapata. 2021. Models and datasets for cross-lingual summarisation. In Proceedings of The 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic.
  159. 159.Martin Petty. 2023. Explainer: Why is Myanmar’s military holding an election? Accessed on Jan. 17, 2024.
  160. 160.McKevitt C Pinel C, Prainsack B. 2020. Caring for data: Value creation in a data-intensive research laboratory. Social Studies of Science, 50(2):175–197.
  161. 161.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020. Xcopa: A multilingual dataset for causal common-sense reasoning. arXiv preprint arXiv:2005.00333.
  162. 162.Maja Popović. 2015. chrF: character n-gram F-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.
  163. 163.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  164. 164.Adithya Pratapa, Rishubh Gupta, and Teruko Mitamura. 2022. Multilingual event linking to Wikidata. In Proceedings of the Workshop on Multilingual Information Access (MIA), pages 37–58, Seattle, USA. Association for Computational Linguistics.
  165. 165.Cornelius Puschmann and Jean Burgess. 2014. Big data, big questions| metaphors of big data. International Journal of Communication, 8(0).
  166. 166.Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. 2022. Data cards: Purposeful and transparent dataset documentation for responsible ai. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, page 1776–1826, New York, NY, USA. Association for Computing Machinery.
  167. 167.PyThaiNLP. 2023a. scb_mt_2020_en2th_prompt. https://huggingface.co/datasets/pythainlp/scb_mt_2020_en2th_prompt. Accessed: 2023-11-29.
  168. 168.PyThaiNLP. 2023b. scb_mt_2020_th2en_prompt. https://huggingface.co/datasets/pythainlp/scb_mt_2020_th2en_prompt. Accessed: 2023-11-29.
  169. 169.PyThaiNLP. 2023c. Thai-Pos-prompt. https://huggingface.co/datasets/pythainlp/Thai-Pos-prompt. Accessed: 2023-11-29.
  170. 170.PyThaiNLP. 2023d. thai-wiktionary-prompt. https://huggingface.co/datasets/pythainlp/thai-wiktionary-prompt. Accessed: 2023-11-29.
  171. 171.PyThaiNLP. 2023e. thai_usembassy_en2th_prompt. https://huggingface.co/datasets/pythainlp/thai_usembassy_en2th_prompt. Accessed: 2023-11-29.
  172. 172.PyThaiNLP. 2023f. thai_usembassy_th2en_prompt. https://huggingface.co/datasets/pythainlp/thai_usembassy_th2en_prompt. Accessed: 2023-11-29.
  173. 173.Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266.
  174. 174.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084.
  175. 175.Reuters. 2023. Explainer: What is happening between Armenia and Azerbaijan over Nagorno-Karabakh? Accessed on Jan. 17, 2024.
  176. 176.Raf Van Rooy. 2021. Language or Dialect? The History of a Conceptual Pair. Oxford University Press.
  177. 177.Sebastian Ruder, Jonathan Clark, Alexander Gutkin, Mihir Kale, Min Ma, Massimo Nicosia, Shruti Rijhwani, Parker Riley, Jean-Michel Sarr, Xinyi Wang, John Wieting, Nitish Gupta, Anna Katanova, Christo Kirov, Dana Dickinson, Brian Roark, Bidisha Samanta, Connie Tao, David Adelani, Vera Axelrod, Isaac Caswell, Colin Cherry, Dan Garrette, Reeve Ingle, Melvin Johnson, Dmitry Panteleev, and Partha Talukdar. 2023. XTREME-UP: A user-centric scarce-data benchmark for under-represented languages. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 1856–1884, Singapore. Association for Computational Linguistics.
  178. 178.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, et al. 2021. Xtreme-r: Towards more challenging and nuanced multilingual evaluation. arXiv preprint arXiv:2104.07412.
  179. 179.Jon Saad-Falcon, Joe Barrow, Alexa Siu, Ani Nenkova, Ryan A Rossi, and Franck Dernoncourt. 2023. Pdf-triage: Question answering over long, structured documents. arXiv preprint arXiv:2309.08872.
  180. 180.Marta Sabou, Kalina Bontcheva, and Arno Scharl. 2012. Crowdsourcing research opportunities: Lessons from natural language processing. In Proceedings of the 12th International Conference on Knowledge Management and Knowledge Technologies, i-KNOW ’12, New York, NY, USA. Association for Computing Machinery.
  181. 181.Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, New York, NY, USA. Association for Computing Machinery.
  182. 182.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  183. 183.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022a. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  184. 184.Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Bideman, Hady Elsahar, Niklas Muennighoff, Jason Phang, et al. 2022b. What language model to train if you have one million gpu hours? arXiv preprint arXiv:2210.15424.
  185. 185.Nick Seaver. 2021. Care and scale: Decorrelative ethics in algorithmic recommendation. Cultural Anthropology, 36(3):509–537.
  186. 186.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. CoRR, abs/1704.04368.
  187. 187.Priyanka Sen, Alham Fikri Aji, and Amir Saffari. 2022. Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1604–1619, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  188. 188.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  189. 189.Shafagh. 2023a. Aya Persian Instruction pn Summary. https://huggingface.co/datasets/Shafagh/aya_persian_instruction_pn-summary. Accessed: 2023-11-28.
  190. 190.Shafagh. 2023b. Aya Persian Instruction pn Summary Title. https://huggingface.co/datasets/Shafagh/aya_persian_instruction_pn-summary-title. Accessed: 2023-11-28.
  191. 191.Jack Sidnell and N. J. Enfield. 2012. Language diversity and social action: A third locus of linguistic relativity. Current Anthropology, 53(3):302–333.
  192. 192.Kathleen Siminyu, Godson Kalipe, Davor Orlic, Jade Z. Abbott, Vukosi Marivate, Sackey Freshia, Prateek Sibal, Bhanu Neupane, David Ifeoluwa Adelani, Amelia V. Taylor, Jamiil Toure Ali, Kevin Degila, Momboladji Balogoun, Thierno Ibrahima Diop, Davis David, Chayma Fourati, Hatem Haddad, and Malek Naski. 2021. AI4D - african language program. In 2nd AfricaNLP Workshop Proceedings, AfricaNLP@EACL 2021, Virtual Event, April 19, 2021.
  193. 193.Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. 2022. SciRepEval: A Multi-Format Benchmark for Scientific Document Representations. ArXiv, abs/2211.13308.
  194. 194.Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Pete Walsh, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. 2024. Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. arXiv preprint.
  195. 195.Lucia Specia and Atefeh Farzindar. 2010. Estimating machine translation post-editing effort with HTER. In Proceedings of the Second Joint EM+/CNGL Workshop: Bringing MT to the User: Research on Integrating MT in the Translation Industry, pages 33–43, Denver, Colorado, USA. Association for Machine Translation in the Americas.
  196. 196.Bernard Spolsky. 2018. Language policy in french colonies and after independence. Current Issues in Language Planning, 19(3):231–315.
  197. 197.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  198. 198.Vivek Srivastava and Mayank Singh. 2021. Challenges and considerations with code-mixed nlp for multilingual societies.
  199. 199.Luke Stark and Anna Lauren Hoffmann. 2019. Data is the new what? popular metaphors & professional ethics in emerging data culture. Journal of Cultural Analytics, 4(1).
  200. 200.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  201. 201.Stephanie Strassel and Jennifer Tracey. 2016. LORELEI language packs: Data, tools, and resources for technology development in low resource languages. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3273–3280, Portorož, Slovenia. European Language Resources Association (ELRA).
  202. 202.Leon Strømberg-Derczynski, Manuel Ciosici, Rebekah Baglini, Morten H. Christiansen, Jacob Aarup Dalsgaard, Riccardo Fusaroli, Peter Juel Henrichsen, Rasmus Hvingelby, Andreas Kirkedal, Alex Speed Kjeldsen, Claus Ladefoged, Finn Arup Nielsen, Jens Madsen, Malte Lau Petersen, Jonathan Hvithamar Rystrøm, and Daniel Varab. 2021. The Danish Gigaword corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pages 413–421, Reykjavik, Iceland (Online). Linköping University Electronic Press, Sweden.
  203. 203.SuryaKrishna02. 2023a. Aya Telugu Food Recipes. https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-food-recipes. Accessed: 2023-11-28.
  204. 204.SuryaKrishna02. 2023b. Aya Telugu Jokes. https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-jokes. Accessed: 2023-11-28.
  205. 205.SuryaKrishna02. 2023c. Aya Telugu News Articles. https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles. Accessed: 2023-11-28.
  206. 206.SuryaKrishna02. 2023d. Aya Telugu Paraphrase. https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-paraphrase. Accessed: 2023-11-28.
  207. 207.SuryaKrishna02. 2023e. Aya Telugu Poems. https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-poems. Accessed: 2023-11-28.
  208. 208.Masahiro Suzuki, Masanori Hirano, and Hiroki Sakaji. 2023. From base to conversational: Japanese instruction dataset and tuning large language models. arXiv preprint arXiv:2309.03412.
  209. 209.syntaxshill. 2023. Arpa aya. https://huggingface.co/datasets/syntaxshill/arpa-aya. Accessed: 2023-11-28.
  210. 210.Oleksiy Syvokon, Olena Nahorna, Pavlo Kuchmiichuk, and Nastasiia Osidach. 2023. UA-GEC: Grammatical error correction and fluency corpus for the Ukrainian language. In Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP), pages 96–102, Dubrovnik, Croatia. Association for Computational Linguistics.
  211. 211.TahmidH. 2023. Annotated News Summary. https://huggingface.co/datasets/TahmidH/annotated_news_summary. Accessed: 2023-11-28.
  212. 212.Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2021. Multilingual translation from denoising pre-training. Findings of the Association for Computational Linguistics: ACL-IJCNLP.
  213. 213.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  214. 214.Tellarin.ai. 2023a. LLM Japanese Dataset Vanilla Aya Format. https://huggingface.co/datasets/tellarin-ai/llm-japanese-dataset-vanilla-aya-format. Accessed: 2023-11-28.
  215. 215.Tellarin.ai. 2023b. NTX LLM Instructions. https://huggingface.co/datasets/tellarin-ai/ntx_llm_instructions. Accessed: 2023-11-28.
  216. 216.theblackcat102. 2023. Joke explaination. https://huggingface.co/datasets/theblackcat102/joke_explaination. Accessed: 2023-11-29.
  217. 217.Jörg Tiedemann. 2012. Parallel data, tools and interfaces in opus. In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12), Istanbul, Turkey. European Language Resources Association (ELRA).
  218. 218.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  219. 219.TurkuNLP. 2023. Turku Paraphrase Corpus. https://huggingface.co/datasets/TurkuNLP/turku_paraphrase_corpus. Accessed: 2023-11-28.
  220. 220.Universal NER. 2023. UNER LLM Instructions. https://huggingface.co/datasets/universalner/uner_llm_instructions. Accessed: 2023-11-28.
  221. 221.Gorka Urbizu, Iñaki San Vicente, Xabier Saralegi, and Ander Corral. 2023. Not enough data to pre-train your language model? MT to the rescue! In Findings of the Association for Computational Linguistics: ACL 2023, pages 3826–3836, Toronto, Canada. Association for Computational Linguistics.
  222. 222.Cécile B. Vigouroux. 2013. Francophonie. Annual Review of Anthropology, 42(1):379–397.
  223. 223.Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning.
  224. 224.Guangyu Wang, Guoxing Yang, Zongxin Du, Longjun Fan, and Xiaohu Li. 2023. Clinicalgpt: Large language models finetuned with diverse medical data and comprehensive evaluation.
  225. 225.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022a. Self-instruct: Aligning language model with self generated instructions. ArXiv preprint, abs/2212.10560.
  226. 226.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Maitreya Patel, Kuntal Kumar Pal, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Shailaja Keyur Sampat, Savan Doshi, Siddhartha Mishra, Sujan Reddy, Sumanta Patro, Tanay Dixit, Xudong Shen, Chitta Baral, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi, and Daniel Khashabi. 2022b. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks.
  227. 227.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022c. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. arXiv preprint arXiv:2204.07705.
  228. 228.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022d. Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5085–5109, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  229. 229.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  230. 230.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022b. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  231. 231.Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, et al. 2023. Polylm: An open source polyglot large language model. arXiv preprint arXiv:2307.06018.
  232. 232.Richard E West, Timothy Newby, Zui Cheng, Alyssa Erickson, and Kyle Clements. 2020. Acknowledging all learning: Alternative, micro, and open credentials. Handbook of Research in Educational Communications and Technology: Learning Design, pages 593–613.
  233. 233.Chenxi Whitehouse, Monojit Choudhury, and Alham Fikri Aji. 2023. Llm-powered data augmentation for enhanced crosslingual performance. arXiv preprint arXiv:2305.14288.
  234. 234.W Bruce Willis. 1998. The Adinkra dictionary: A visual primer on the language of Adinkra. Pyramid Complex.
  235. 235.Genta Winata, Alham Fikri Aji, Zheng Xin Yong, and Thamar Solorio. 2023a. The decades progress on code-switching research in NLP: A systematic survey on trends and challenges. In Findings of the Association for Computational Linguistics: ACL 2023, pages 2936–2978, Toronto, Canada. Association for Computational Linguistics.
  236. 236.Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, and Sebastian Ruder. 2023b. NusaX: Multilingual parallel sentiment dataset for 10 Indonesian local languages. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 815–834, Dubrovnik, Croatia. Association for Computational Linguistics.
  237. 237.Peter Wittenburg. 2021. Open Science and Data Science. Data Intelligence, 3(1):95–105.
  238. 238.Sam Witteveen and Martin Andrews. 2019. Paraphrasing with large language models. arXiv preprint arXiv:1911.09661.
  239. 239.Walt Wolfram. 1997. Issues in dialect obsolescence: An introduction. American speech, 72(1):3–11.
  240. 240.Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862.
  241. 241.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I Wang, et al. 2022. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. arXiv preprint arXiv:2201.05966.
  242. 242.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023a. Wizardlm: Empowering large language models to follow complex instructions.
  243. 243.Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen. 2023b. Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models.
  244. 244.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  245. 245.Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li. 2023. Dataset pruning: Reducing training data by examining generalization influence.
  246. 246.Yi Yang, Wen-tau Yih, and Christopher Meek. 2015. WikiQA: A challenge dataset for open-domain question answering. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2013–2018, Lisbon, Portugal. Association for Computational Linguistics.
  247. 247.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP).
  248. 248.Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021. CrossFit: A few-shot learning challenge for cross-task generalization in NLP. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7163–7189, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  249. 249.Zheng Xin Yong, Cristina Menghini, and Stephen Bach. 2023a. Low-resource languages jailbreak GPT-4. In Socially Responsible Language Modelling Research.
  250. 250.Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, and Vassilina Nikoulina. 2023b. BLOOM+1: Adding language support to BLOOM for zero-shot prompting. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11682–11703, Toronto, Canada. Association for Computational Linguistics.
  251. 251.Zheng Xin Yong, Ruochen Zhang, Jessica Zosa Forde, Skyler Wang, Arjun Subramonian, Holy Lovenia, Samuel Cahyawijaya, Genta Indra Winata, Lintang Sutawika, Jan Christian Blaise Cruz, et al. 2023c. Prompting multilingual large language models to generate code-mixed texts: The case of south east asian languages. In Sixth Workshop on Computational Approaches to Linguistic Code-Switching.
  252. 252.Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2023. Large language model as attributed training data generator: A tale of diversity and bias.
  253. 253.Ann Yuan, Daphne Ippolito, Vitaly Nikolaev, Chris Callison-Burch, Andy Coenen, and Sebastian Gehrmann. 2021. Synthbio: A case study in human-ai collaborative curation of text datasets. CoRR, abs/2111.06467.
  254. 254.Marcos Zampieri, Preslav Nakov, and Yves Scherrer. 2020. Natural language processing for similar languages, varieties, and dialects: A survey. Natural Language Engineering, 26(6):595–612.
  255. 255.Ge Zhang, Yemin Shi, Ruibo Liu, Ruibin Yuan, Yizhi Li, Siwei Dong, Yu Shu, Zhaoqun Li, Zekun Wang, Chenghua Lin, Wenhao Huang, and Jie Fu. 2023a. Chinese open instruction generalist: A preliminary release.
  256. 256.Shaolei Zhang, Qingkai Fang, Zhuocheng Zhang, Zhengrui Ma, Yan Zhou, Langlin Huang, Mengyu Bu, Shangtong Gui, Yunji Chen, Xilin Chen, and Yang Feng. 2023b. Bayling: Bridging cross-lingual alignment and instruction following through interactive translation for large language models.
  257. 257.Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2023c. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models.
  258. 258.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. Paws: Paraphrase adversaries from word scrambling.
  259. 259.Zhirui Zhang, Shujie Liu, Mu Li, Ming Zhou, and Enhong Chen. 2018. Joint training for neural machine translation models with monolingual data. CoRR, abs/1803.00353.
  260. 260.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206.
  261. 261.Terry Yue Zhuo, Armel Zebaze, Nitchakarn Suppattarachai, Leandro von Werra, Harm de Vries, Qian Liu, and Niklas Muennighoff. 2024. Astraios: Parameter-efficient instruction tuning code large language models. arXiv preprint arXiv:2401.00788.
  262. 262.Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016. Transfer learning for low-resource neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1568–1575, Austin, Texas. Association for Computational Linguistics.

Citation

MLA
Singh, S., et al. “Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning”. arXiv, 2024, http://arxiv.org/abs/2402.06619v1.
APA
Singh, S., Vargus, F., Dsouza, D., Karlsson, B. F., Mahendiran, A., Ko, W.-Y., Shandilya, H., Patel, J., Mataciunas, D., OMahony, L., Zhang, M., Hettiarachchi, R., Wilson, J., Machado, M., Moura, L. S., Krzemiński, D., Fadaei, H., Ergün, I., Okoh, I., … Hooker, S. (2024). Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning. arXiv. http://arxiv.org/abs/2402.06619v1
Chicago
Singh, S., F. Vargus, D. Dsouza, et al. 2024. “Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning”. arXiv. http://arxiv.org/abs/2402.06619v1.
Harvard
Singh, S. et al. (2024) “Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.06619v1.
Vancouver
1. Singh S, Vargus F, Dsouza D, et al (2024) Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning. arXiv

BibTeX

@article{singh2024aya,
  title = {Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning},
  author = {Singh, Shivalika and Vargus, Freddie and Dsouza, Daniel and Karlsson, Börje F. and Mahendiran, Abinaya and Ko, Wei-Yin and Shandilya, Herumb and Patel, Jay and Mataciunas, Deividas and OMahony, Laura and Zhang, Mike and Hettiarachchi, Ramith and Wilson, Joseph and Machado, Marina and Moura, Luisa Souza and Krzemiński, Dominik and Fadaei, Hakimeh and Ergün, Irem and Okoh, Ifeoma and Alaagib, Aisha and Mudannayake, Oshan and Alyafeai, Zaid and Chien, Vu Minh and Ruder, Sebastian and Guthikonda, Surya and Alghamdi, Emad A. and Gehrmann, Sebastian and Muennighoff, Niklas and Bartolo, Max and Kreutzer, Julia and Üstün, Ahmet and Fadaee, Marzieh and Hooker, Sara},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.06619v1},
  eprint = {2402.06619}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/