IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

David Ifeoluwa AdelaniJessica OjoIsrael Abebe AzimeJian Yun ZhuangJesujoba Oluwadara AlabiXuanli HeMillicent OchiengSara HookerAndiswa BukulaEn-Shiun Annie Lee

article2025NAACL59 citationsOutstanding Paper Award

Introduces IrokoBench, a human-translated benchmark across 17 African languages that evaluates 16 open and proprietary language models on natural language inference, mathematical reasoning, and question answering to measure performance disparities against high-resource languages.

Listen

Large language models have advanced rapidly in solving complex, knowledge-intensive tasks, but their development and evaluation remain heavily concentrated on a few high-resource languages such as English. While African languages represent hundreds of millions of speakers, prior evaluation efforts have largely focused on basic text classification or relied on noisy machine translations. Consequently, organizations and developers lack a reliable understanding of how state-of-the-art language models handle advanced reasoning and multi-step problem solving across diverse African languages.

The article introduces and evaluates IrokoBench, a high-quality, human-translated benchmark designed to rigorously assess language models on complex tasks across 17 typologically diverse African languages. The benchmark covers natural language inference through AfriXNLI, multi-choice question answering through AfriMMLU, and grade-school mathematical reasoning via AfriMGSM. The authors evaluated 10 open-weight and six commercial proprietary models under direct in-language prompting, few-shot setups, and a translate-test approach where prompts are automatically translated into English prior to processing.

The evaluation reveals a steep performance drop between high-resource languages and African languages, with an average performance gap of roughly 45% across models. Proprietary systems significantly outperformed open models; the top-performing open model, Gemma 2 27B, achieved only 63% of the performance of the leading commercial model, GPT-4o. Furthermore, the volume of available web training data strongly correlated with accuracy: languages with less than 50 million characters of online text (such as Ewe, Lingala, and Wolof) saw severe performance degradation, whereas well-resourced languages like Swahili performed significantly better. Across task types, mathematical reasoning proved the most challenging, while the translate-test pipeline substantially improved reasoning scores for English-centric open models by up to 21 percentage points.

These findings indicate that current language models cannot be deployed reliably in native African languages for complex reasoning workflows without explicit adaptation. Relying on users to translate prompts into English introduces user-experience friction and potential failure points, even though it currently mitigates reasoning deficiencies in open models. The results also show that dedicated, smaller Africa-centric models pre-trained on regional data can match or outperform much larger generalist models on specific linguistic tasks.

Decision-makers and practitioners aiming to deploy language technology in these regions should invest in Africa-centric model pre-training and quality instruction tuning rather than relying solely on off-the-shelf global models. When using existing open-weight architectures for complex reasoning, teams should consider translating inputs to English as an interim mitigation strategy while building native-language support. Future work must expand human-annotated datasets to underrepresented language families and domain-specific education topics to ensure equitable AI access across the continent.

No sufficiently relevant recommendations were found.

Cover for IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

Abstract

Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (\eg African languages) are often evaluated only on basic text classification tasks due to the lack of appropriate or comprehensive benchmarks outside of high-resource languages. In this paper, we introduce IrokoBench -- a human-translated benchmark dataset for 17 typologically-diverse low-resource African languages covering three tasks: natural language inference~(AfriXNLI), mathematical reasoning~(AfriMGSM), and multi-choice knowledge-based question answering~(AfriMMLU). We use IrokoBench to evaluate zero-shot, few-shot, and translate-test settings~(where test sets are translated into English) across 10 open and six proprietary LLMs. Our evaluation reveals a significant performance gap between high-resource languages~(such as English and French) and low-resource African languages. We observe a significant performance gap between open and proprietary models, with the highest performing open model, Gemma 2 27B only at 63% of the best-performing proprietary model GPT-4o performance. In addition, machine translating the test set to English before evaluation helped to close the gap for larger models that are English-centric, such as Gemma 2 27B and LLaMa 3.1 70B. These findings suggest that more efforts are needed to develop and adapt LLMs for African languages.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 IrokoBench
  • 3.1 Languages covered by IrokoBench
  • 3.2 Tasks covered by IrokoBench
  • 3.3 Data collection process
  • 3.4 LLMs used for evaluation
  • 3.5 Evaluation Settings
  • 4 Results
  • 4.1 Overall Results
  • 4.2 Task-specific results
  • 4.3 Few-shot results, Cross-Lingual Transfer, Sensitivity to Prompt Templates
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 Language covered
  • A.2 Prompts Template and Evaluation Tool
  • A.3 AfriCOMET metric scores for XNLI translation
  • A.4 Task-specific results for all models
  • A.5 Comparison between in-language and translate-test results
  • A.6 Cross-lingual transfer results for XNLI

Knowls

  1. Knowl 1 — IrokoBench’s languages, tasks, and dataset sizes

    definition

    IrokoBench is a human-translated benchmark for evaluating language models on reasoning and knowledge-intensive tasks in African languages. It covers 17 African languages from West, East, Southern, and Central Africa, across the Mande, Afro-Asiatic, and Niger-Congo families, including eight Bantu languages; English and French are also included. Swahili is included in all three benchmarks.

    The benchmark contains three datasets:

    • AfriXNLI tests whether a premise entails, contradicts, or is neutral toward a hypothesis. Its translated subset covers 15 African languages other than Swahili, with 450 development and 600 test examples per language, sampled equally from 10 domains. The balanced three-class task is evaluated by accuracy.
    • AfriMMLU is multiple-choice question answering across five subjects selected as relatively internationally applicable and easier to translate: elementary mathematics, high-school geography, high-school microeconomics, international law, and global facts. It contains 608 translated question-answer pairs, including 25 training, 83 development, and 500 test examples; the test set has 100 questions per subject. It is evaluated by option accuracy.
    • AfriMGSM evaluates grade-school mathematical word-problem solving. It has 8 training examples for few-shot use and 250 test problems. The dataset overview reports 17 African languages excluding Swahili and including Vai. It is evaluated by exact match of the answer.
  2. Knowl 2 — Human translation and quality control

    model/method

    IrokoBench’s African-language data were produced through professional human translation coordinated by language leads. The translation work took about two months. Most translators worked from English; Ewe, Lingala, and Wolof were translated from French, with French versions of MMLU professionally prepared as an intermediate source. Language coordinators reviewed translations and corrected errors, and translators were paid after this review.

    The authors also assessed translation quality with AfriCOMET quality-estimation scores. Scores were generally between 0.7 and 1.0 for about 13 language pairs, but the metric gave lower averages for Lingala, Twi, and Wolof. They caution that these scores are not reliable for those languages because the AfriCOMET encoder was not pretrained on them; Twi’s scores had also shown weak correlation with human judgments.

  3. Knowl 3 — LLM evaluation protocol

    experimental setup

    The evaluation compares 10 open-weight models—including multilingual encoder-decoder models and decoder-only models—with six proprietary models: four OpenAI GPT variants, Gemini-1.5-Pro, and Claude Opus. The main settings are zero-shot evaluation in the benchmark language and translate-test evaluation, in which NLLB-200 (3.3B) translates the test input into English before model evaluation. Few-shot tests are restricted to the three selected best-performing models and use 5 examples for AfriXNLI and AfriMMLU and 8 for AfriMGSM. AfriMGSM uses chain-of-thought prompting in both language settings.

    Models are tested with five prompt templates per task. Open models use log-likelihood scoring for the classification and multiple-choice tasks, while closed models are evaluated through answer verbalizers; AfriMGSM uses generated answers and exact match. Accuracy is used for AfriXNLI and AfriMMLU.

  4. Knowl 4 — Large gaps between African-language and high-resource-language results

    empirical result

    In zero-shot evaluation, the best African-language average across the three tasks is GPT-4o’s 59.0, compared with 86.9 for English and 78.1 for French. Thus, even the strongest evaluated model has gaps of 27.9 points against English and 19.1 points against French. The averages are task averages over African languages, excluding English, French, and Vai.

    Proprietary models generally outperform open models. GPT-4o’s African-language average of 59.0 is 21.9 points higher than the best open model, Gemma 2 27B, at 37.1; the open model reaches about 63% of GPT-4o’s score. Performance also tracks language-data availability: Swahili tends to do better than other African languages, while Ewe, Lingala, Luganda, Twi, and Wolof—which the paper reports each have fewer than 50 million characters of web data—are among the weakest.

  5. Knowl 5 — Translate-test helps English-centric models, but not uniformly

    empirical result

    Translating test inputs into English often improves results for open models, but the benefit depends on the model and task. On AfriMGSM, Gemma 2 27B rises from 28.5 exact match in-language to 46.1 after translation, and LLaMa 3.1 70B rises from 24.6 to 45.6. GPT-4o instead falls from 52.6 to 42.6, so Gemma 2 27B’s translated-test score is 3.5 points above GPT-4o’s in that setting.

    On AfriMMLU, translate-test improves Gemma 2 27B from 39.9 to 48.8 and LLaMa 3.1 70B from 39.4 to 51.3, while GPT-4o scores higher in-language than after translation, 60.0 versus 54.1. The authors report that the open models improved on 15 of 16 African languages for AfriMMLU. In the overall task averages, Command-R and LLaMa 3.1 8B gain 19.8 and 13.3 points, respectively, with translate-test; GPT-4o and GPT-4-Turbo are among the exceptions that perform better in-language.

  6. Knowl 6 — Task and subject difficulty varies substantially

    empirical result

    AfriMGSM is the most difficult task overall, followed by AfriMMLU and AfriXNLI. In the zero-shot in-language averages, GPT-4o scores 64.3 on AfriXNLI, 60.0 on AfriMMLU, and 52.6 exact match on AfriMGSM; Gemma 2 27B scores 42.8, 39.9, and 28.5, respectively. The open-versus-proprietary difference is especially large on AfriMMLU and AfriMGSM: GPT-4o leads Gemma 2 27B by 20.1 and 24.1 points, respectively. For AfriXNLI, the smaller multilingual models mT0-XXL-MT and Aya-101 outperform the larger Gemma 2 27B and LLaMa 3.1 70B models.

    Within GPT-4o’s AfriMMLU results, elementary mathematics has the highest African-language average, 75.1, and international law averages 66.9. Global facts is lowest at 48.0; high-school geography and microeconomics average 55.6 and 54.3. The paper notes that multiple-choice math questions may be easier than free-form math reasoning, and that global-facts difficulty also affects English and French.

  7. Knowl 7 — Few-shot examples help selectively

    empirical result

    In-language few-shot tests used 5 examples for AfriXNLI and AfriMMLU and 8 examples for AfriMGSM. For the classification tasks, examples improved Gemma 2 27B by 13.2 points on AfriXNLI and 4.9 on AfriMMLU, and improved LLaMa 3.1 70B by 10.7 and 7.0 points. GPT-4o did not improve on these classification tasks with additional examples.

    For AfriMGSM, GPT-4o’s exact match rises from 49.8 zero-shot to 56.1 with 8 examples, a 6.3-point increase. Gemma 2 27B falls from 27.0 to 14.4, and LLaMa 3.1 70B falls from 23.2 to 14.9. The reported pattern is therefore task- and model-dependent: few-shot examples help the open models on classification, but not on non-English mathematical reasoning.

  8. Knowl 8 — English-supervised NLI transfer favors Africa-centric encoders

    empirical result

    For AfriXNLI, the authors fine-tuned multilingual masked language models on 400,000 English NLI training examples and evaluated transfer to other languages. AfroXLMR-76L-large, which was pretrained on all languages in AfriXNLI, achieves a reported average score of 65.7, compared with 62.7 for GPT-4o in the comparison. Other encoder averages are 61.4 for AfroXLMR-large, 54.9 for AfroXLMR-base, 53.1 for Serengeti, and 49.6 for XLM-R-large. The authors attribute the Africa-centric models’ advantage in part to their pretraining coverage of the evaluation languages; GPT-4o is competitive with these encoders but does not exceed AfroXLMR-76L-large on average.

  9. Knowl 9 — Prompt choice affects measured performance

    empirical result

    The authors test five prompt templates for each task. Across those templates, the reported standard deviations of African-language averages for Aya-101, Gemma 2 27B, and GPT-4o are, respectively: 2.2, 2.9, and 6.0 points on AfriXNLI; 0.1, 0.4, and 8.7 on AfriMMLU; and 0.1, 1.9, and 1.3 on AfriMGSM. Prompt sensitivity is therefore particularly notable for GPT-4o on AfriMMLU. The best AfriXNLI template also differs by model: Aya-101 performs best with a prompt naming the language, Gemma 2 27B with a simpler premise-and-question prompt, and GPT-4o with a detailed task description. The authors find AfriMGSM generally less sensitive to prompt wording.

  10. Knowl 10 — Benchmark limitations

    limitation

    IrokoBench uses translations of source-language material rather than examples originally authored in African languages, so its data may exhibit translationese; native-language generation would avoid that issue, although parallel translation enables comparison across languages on corresponding inputs. The benchmark also covers only three African language families. Nilo-Saharan, Austronesian, and Khoisan languages are absent, which the authors attribute partly to difficulty recruiting professional translators and limited translation budget.

Coverage note — The extensive per-language and per-prompt result tables are omitted because the knowls retain the main comparisons and reported patterns without reproducing exhaustive breakdowns.

References

  1. 1.Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, and Alcides Alcoba Inciarte. 2023. SERENGETI: Massively multilingual language models for Africa. In Findings of the Association for Computational Linguistics: ACL 2023, pages 1498–1537, Toronto, Canada. Association for Computational Linguistics.
  2. 2.David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memd￾jokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, and Sam Manthalu. 2022a. A few thousand translations go a long way! leveraging pre-trained models for African news translation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3053–3070, Seattle, United States. Association for Computational Linguistics.
  3. 3.David Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba Alabi, Yanke Mao, Haonan Gao, and En-Shiun Lee. 2024. SIB-200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 226–245, St. Julian’s, Malta. Association for Computational Linguistics.
  4. 4.David Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwadabe, Tosin Adewumi, Orevaoghene Ahia, and Joyce Nakatumba-Nabende. 2022b. MasakhaNER 2.0: Africa-centric transfer learning for named entity recognition. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4488–4508, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  5. 5.David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021. MasakhaNER: Named entity recognition for African languages. Transactions of the Association for Computational Linguistics, 9:1116–1131.
  6. 6.David Ifeoluwa Adelani, Marek Masiak, Israel Abebe Azime, Jesujoba Alabi, Atnafu Lambebo Tonja, Christine Mwase, Odunayo Ogundepo, Bonaventure F. P. Dossou, Akintunde Oladipo, Doreen Nixdorf, Chris Chinenye Emezue, Sana Al-azzawi, Blessing Sibanda, Davis David, Lolwethu Ndolela, Jonathan Mukiibi, Tunde Ajayi, Tatiana Moteu, Brian Odhiambo, Abraham Owodunni, Nnaemeka Obiefuna, Muhidin Mohamed, Shamsuddeen Hassan Muhammad, Teshome Mulugeta Ababu, Saheed Abdullahi Salahudeen, Mesay Gemeda Yigezu, Tajuddeen Gwadabe, Idris Abdulmumin, Mahlet Taye, Oluwabusayo Awoyomi, Iyanuoluwa Shode, Tolulope Adelani, Habiba Abdulganiyu, Abdul-Hakeem Omotayo, Adetola Adeeko, Abeeb Afolabi, Anuoluwapo Aremu, Olanrewaju Samuel, Clemencia Siro, Wangari Kimotho, Onyekachi Ogbu, Chinedu Mbonu, Chiamaka Chukwuneke, Samuel Fanijo, Jessica Ojo, Oyinkansola Awosan, Tadesse Kebede, Toadoum Sari Sakayo, Pamela Nyatsine, Freedmore Sidume, Oreen Yousuf, Mardiyyah Oduwole, Kanda Tshinu, Ussen Kimanuka, Thina Diko, Siyanda Nxakama, Sinodos Nigusse, Abdulmejid Johar, Shafie Mohamed, Fuad Mire Hassan, Moges Ahmed Mehamed, Evrard Ngabire, Jules Jules, Ivan Ssenkungu, and Pontus Stenetorp. 2023. MasakhaNEWS: News topic classification for African languages. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 144–159, Nusa Dua, Bali. Association for Computational Linguistics.
  7. 7.Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2023a. MEGA: Multilingual evaluation of generative AI. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4232–4267, Singapore. Association for Computational Linguistics.
  8. 8.Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023b. Megaverse: Benchmarking large language models across languages, modalities, models and tasks. ArXiv, abs/2311.07463.
  9. 9.Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. 2022. Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning. In Proceedings of the 29th International Conference on Computational Linguistics, pages 4336–4349, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  10. 10.Anthropic. 2024. Claude — anthropic.com. https://www.anthropic.com/claude. [Accessed 01-06-2024].
  11. 11.Anuoluwapo Aremu, Jesujoba O. Alabi, and David Ifeoluwa Adelani. 2023. Yorc: Yoruba reading comprehension dataset. Preprint, arXiv:2308.09768.
  12. 12.Israel Abebe Azime, Mitiku Yohannes Fuge, Atnafu Lambebo Tonja, Tadesse Destaw Belay, Aman Kassahun Wassie, Eyasu Shiferaw Jada, Yonas Chanie, Walelign Tewabe Sewunetie, and Seid Muhie Yimam. 2024. Enhancing amharic-llama: Integrating task specific and generative datasets. arXiv preprint arXiv:2402.08015.
  13. 13.Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2023. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. Preprint, arXiv:2308.16884.
  14. 14.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 675–718, Nusa Dua, Bali. Association for Computational Linguistics.
  15. 15.Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A. Wang, Genta Indra Winata, François Yvon, and Andy Zou. 2024. Lessons from the trenches on reproducible evaluation of language models. Preprint, arXiv:2405.14782.
  16. 16.BigScience-Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel ´ Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, and et al. 2023. Bloom: A 176b-parameter open-access multilingual language model. Preprint, arXiv:2211.05100.
  17. 17.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  18. 18.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. Preprint, arXiv:2005.14165.
  19. 19.Cohere. 2024. Command R — cohere.com. https://cohere.com/command. [Accessed 01-06-2024].
  20. 20.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  21. 21.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  22. 22.Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, Jeff Wang, and NLLB Team. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018):841–846.
  23. 23.Cheikh M. Bamba Dione, David Ifeoluwa Adelani, Peter Nabende, Jesujoba Alabi, Thapelo Sindane, Happy Buzaaba, Shamsuddeen Hassan Muhammad, Chris Chinenye Emezue, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jonathan Mukiibi, Blessing Sibanda, Bonaventure F. P. Dossou, Andiswa Bukula, Rooweither Mabuya, Allahsera Auguste Tapo, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Fatoumata Ouoba Kabore, Amelia Taylor, Godson Kalipe, Tebogo Macucwa, Vukosi Marivate, Tajuddeen Gwadabe, Mboning Tchiaze Elvis, Ikechukwu Onyenwe, Gratien Atindogbe, Tolulope Adelani, Idris Akinade, Olanrewaju Samuel, Marien Nahimana, Théogène Musabeyezu, Emile Niyomutabazi, Ester Chimhenga, Kudzai Gotosa, Patrick Mizha, Apelete Agbolo, Seydou Traore, Chinedu Uchechukwu, Aliyu Yusuf, Muhammad Abdullahi, and Dietrich Klakow. 2023. MasakhaPOS: Part-of-speech tagging for typologically diverse African languages. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10883–10900, Toronto, Canada. Association for Computational Linguistics.
  24. 24.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, and et al. 2024. The llama 3 herd of models. ArXiv, abs/2407.21783.
  25. 25.Emilio Ferrara. 2023. Should chatgpt be biased? challenges and risks of bias in large language models. arXiv preprint arXiv:2304.03738.
  26. 26.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  27. 27.Gemini-Team, Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry, Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, and et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. Preprint, arXiv:2403.05530.
  28. 28.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538.
  29. 29.Kai Hartung, Aaricia Herygers, Shubham Kurlekar, Khabbab Zakaria, Taylan Volkan, Sören Gröttrup, and Munir Georges. 2023. Measuring sentiment bias in machine translation. Preprint, arXiv:2306.07152.
  30. 30.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).
  31. 31.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021b. Measuring mathematical problem solving with the MATH dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  32. 32.Amr Hendy, Mohamed Gomaa Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a comprehensive evaluation. ArXiv, abs/2302.09210.
  33. 33.Meng Ji, Meng Ji, Pierrette Bouillon, and Mark Seligman. 2023. Cultural and Linguistic Bias of Neural Machine Translation Technology, page 100–128. Studies in Natural Language Processing. Cambridge University Press.
  34. 34.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts. Preprint, arXiv:2401.04088.
  35. 35.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  36. 36.Eric Khiu, Hasti Toossi, David Anugraha, Jinyu Liu, Jiaxu Li, Juan Flores, Leandro Roman, A. Seza Dogruöz, and En-Shiun Lee. 2024. ˘ Predicting machine translation performance on low-resource languages: The role of domain similarity. In Findings of the Association for Computational Linguistics: EACL 2024, pages 1474–1486, St. Julian’s, Malta. Association for Computational Linguistics.
  37. 37.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  38. 38.Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. Madlad-400: A multilingual and document-level large audited dataset. Preprint, arXiv:2309.04662.
  39. 39.Viet Lai, Chien Nguyen, Nghia Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan Rossi, and Thien Nguyen. 2023a. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 318–327, Singapore. Association for Computational Linguistics.
  40. 40.Viet Dac Lai, Nghia Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023b. ChatGPT beyond English: Towards a comprehensive evaluation of large language models in multilingual learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 13171–13189, Singapore. Association for Computational Linguistics.
  41. 41.En-Shiun Lee, Sarubi Thillainathan, Shravan Nayak, Surangika Ranathunga, David Adelani, Ruisi Su, and Arya McCarthy. 2022. Pre-trained multilingual sequence-to-sequence models: A hope for low-resource language translation? In Findings of the Association for Computational Linguistics: ACL 2022, pages 58–67, Dublin, Ireland. Association for Computational Linguistics.
  42. 42.Yinquan Lu, Wenhao Zhu, Lei Li, Yu Qiao, and Fei Yuan. 2024. Llamax: Scaling linguistic horizons of llm by enhancing translation capabilities beyond 100 languages. arXiv preprint arXiv:2407.05975.
  43. 43.Alexandra Luccioni and Joseph Viviano. 2021. What‘s in the box? an analysis of undesirable content in the Common Crawl corpus. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 182–189, Online. Association for Computational Linguistics.
  44. 44.Chunlan Ma, Ayyoob ImaniGooghari, Haotian Ye, Ehsaneddin Asgari, and Hinrich Schütze. 2023. Taxi1500: A multilingual dataset for text classification in 1500 languages. Preprint, arXiv:2305.08487.
  45. 45.Meta. 2024. Introducing Meta Llama 3: The most capable openly available LLM to date — ai.meta.com. https://ai.meta.com/blog/meta-llama-3/. [Accessed 01-06-2024].
  46. 46.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023. Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15991–16111, Toronto, Canada. Association for Computational Linguistics.
  47. 47.Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, Samuel Rutunda, Tadesse Destaw Belay, Wendimu Messelle, Hailu Balcha, Sisay Chala, Hagos Gebremichael, Bernard Opoku, and Stephen Arthur. 2023. AfriSenti: A Twitter sentiment analysis benchmark for African languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13968–13981, Singapore. Association for Computational Linguistics.
  48. 48.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4885–4901, Online. Association for Computational Linguistics.
  49. 49.Jessica Ojo, Kelechi Ogueji, Pontus Stenetorp, and David I Adelani. 2023. How good are large language models on african languages? arXiv preprint arXiv:2311.07978.
  50. 50.OpenAI. 2024. Introducing ChatGPT. https://openai.com/index/chatgpt/. [Accessed 01-06-2024].
  51. 51.OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, and et al. 2024. Gpt-4 technical report. Preprint, arXiv:2303.08774.
  52. 52.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  53. 53.Danti Pudjiati, Ninuk Lustyantie, Ifan Iskandar, and Tira Nur Fitria. 2022. Post-editing of machine translation: Creating a better translation of cultural specific terms. Language Circle: Journal of Language and Literature, 17(1):61–73.
  54. 54.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  55. 55.Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy P. Lillicrap, Jean-Baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M. Dai, Katie Millican, Ethan Dyer, Mia Glaese, Thibault Sottiaux, and et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. ArXiv, abs/2403.05530.
  56. 56.Gemma Team Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, L’eonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ram’e, Johan Ferret, Peter Liu, Pouya Dehghani Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, and et al. 2024. Gemma 2: Improving open language models at a practical size. ArXiv, abs/2408.00118.
  57. 57.Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021. Gender bias in machine translation. Preprint, arXiv:2104.06001.
  58. 58.Timo Schick and Hinrich Schütze. 2021. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2339–2352.
  59. 59.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2023. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations.
  60. 60.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations.
  61. 61.Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzeminski, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker. 2024. Aya dataset: An open-access collection for multilingual instruction tuning. Preprint, arXiv:2402.06619.
  62. 62.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  63. 63.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, and et al. 2023. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288.
  64. 64.A. Ustun, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. Aya model: An instruction finetuned open-access multilingual language model. ArXiv, abs/2402.07827.
  65. 65.Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. Aya model: An instruction finetuned open-access multilingual language model. Preprint, arXiv:2402.07827.
  66. 66.Eva Vanmassenhove, Dimitar Shterionov, and Matthew Gwilliam. 2021. Machine translationese: Effects of algorithmic bias on linguistic complexity in machine translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2203–2213, Online. Association for Computational Linguistics.
  67. 67.Jiayi Wang, David Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluwapo Aremu, Jessica Ojo, Shamsuddeen Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Mohamed, Hassan Ayinde, Oluwabusayo Awoyomi, Lama Alkhaled, Sana Al-azzawi, Naome Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Toadoum Sari Sakayo, Lyse Naomi Wamba, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Iro, Saheed Abdullahi, Stephen Moore, Bernard Opoku, Zainab Akinjobi, Abeeb Afolabi, Nnaemeka Obiefuna, Onyekachi Ogbu, Sam Ochieng’, Verrah Otiende, Chinedu Mbonu, Yao Lu, and Pontus Stenetorp. 2024. AfriMTE and AfriCOMET: Enhancing COMET to embrace under-resourced African languages. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 5997–6023, Mexico City, Mexico. Association for Computational Linguistics.
  68. 68.Jun Wang, Benjamin Rubinstein, and Trevor Cohn. 2022. Measuring and mitigating name biases in neural machine translation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2576–2590, Dublin, Ireland. Association for Computational Linguistics.
  69. 69.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.

Citation

MLA
Adelani, D. I., et al. “IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 2732–57, https://doi.org/10.18653/v1/2025.naacl-long.139.
APA
Adelani, D. I., Ojo, J., Azime, I. A., Zhuang, J. Y., Alabi, J., He, X., Ochieng, M., Hooker, S., Bukula, A., Lee, E.-S. A., Chukwuneke, C. I., Buzaaba, H., Sibanda, B. K., Kalipe, G. K., Mukiibi, J., Kabenamualu, S. K., Yuehgoh, F., Setaka, M., Ndolela, L., … Stenetorp, P. (2025). IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2732–2757. https://doi.org/10.18653/v1/2025.naacl-long.139
Chicago
Adelani, D. I., J. Ojo, I. A. Azime, et al. 2025. “IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2732–57. https://doi.org/10.18653/v1/2025.naacl-long.139.
Harvard
Adelani, D.I. et al. (2025) “IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2732–2757. Available at: https://doi.org/10.18653/v1/2025.naacl-long.139.
Vancouver
1. Adelani DI, Ojo J, Azime IA, et al (2025) IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 2732–2757

BibTeX

@inproceedings{adelani-etal-2025-irokobench,
    title = "{I}roko{B}ench: A New Benchmark for {A}frican Languages in the Age of Large Language Models",
    author = "Adelani, David Ifeoluwa  and
      Ojo, Jessica  and
      Azime, Israel Abebe  and
      Zhuang, Jian Yun  and
      Alabi, Jesujoba Oluwadara  and
      He, Xuanli  and
      Ochieng, Millicent  and
      Hooker, Sara  and
      Bukula, Andiswa  and
      Lee, En-Shiun Annie  and
      Chukwuneke, Chiamaka Ijeoma  and
      Buzaaba, Happy  and
      Sibanda, Blessing Kudzaishe  and
      Kalipe, Godson Koffi  and
      Mukiibi, Jonathan  and
      Kabongo Kabenamualu, Salomon  and
      Yuehgoh, Foutse  and
      Setaka, Mmasibidi  and
      Ndolela, Lolwethu  and
      Odu, Nkiruka  and
      Mabuya, Rooweither  and
      Osei, Salomey  and
      Muhammad, Shamsuddeen Hassan  and
      Samb, Sokhar  and
      Guge, Tadesse Kebede  and
      Sherman, Tombekai Vangoni  and
      Stenetorp, Pontus",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.139/",
    doi = "10.18653/v1/2025.naacl-long.139",
    pages = "2732--2757",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/