IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models
David Ifeoluwa AdelaniJessica OjoIsrael Abebe AzimeJian Yun ZhuangJesujoba Oluwadara AlabiXuanli HeMillicent OchiengSara HookerAndiswa BukulaEn-Shiun Annie Lee
Introduces IrokoBench, a human-translated benchmark across 17 African languages that evaluates 16 open and proprietary language models on natural language inference, mathematical reasoning, and question answering to measure performance disparities against high-resource languages.
Large language models have advanced rapidly in solving complex, knowledge-intensive tasks, but their development and evaluation remain heavily concentrated on a few high-resource languages such as English. While African languages represent hundreds of millions of speakers, prior evaluation efforts have largely focused on basic text classification or relied on noisy machine translations. Consequently, organizations and developers lack a reliable understanding of how state-of-the-art language models handle advanced reasoning and multi-step problem solving across diverse African languages.
The article introduces and evaluates IrokoBench, a high-quality, human-translated benchmark designed to rigorously assess language models on complex tasks across 17 typologically diverse African languages. The benchmark covers natural language inference through AfriXNLI, multi-choice question answering through AfriMMLU, and grade-school mathematical reasoning via AfriMGSM. The authors evaluated 10 open-weight and six commercial proprietary models under direct in-language prompting, few-shot setups, and a translate-test approach where prompts are automatically translated into English prior to processing.
The evaluation reveals a steep performance drop between high-resource languages and African languages, with an average performance gap of roughly 45% across models. Proprietary systems significantly outperformed open models; the top-performing open model, Gemma 2 27B, achieved only 63% of the performance of the leading commercial model, GPT-4o. Furthermore, the volume of available web training data strongly correlated with accuracy: languages with less than 50 million characters of online text (such as Ewe, Lingala, and Wolof) saw severe performance degradation, whereas well-resourced languages like Swahili performed significantly better. Across task types, mathematical reasoning proved the most challenging, while the translate-test pipeline substantially improved reasoning scores for English-centric open models by up to 21 percentage points.
These findings indicate that current language models cannot be deployed reliably in native African languages for complex reasoning workflows without explicit adaptation. Relying on users to translate prompts into English introduces user-experience friction and potential failure points, even though it currently mitigates reasoning deficiencies in open models. The results also show that dedicated, smaller Africa-centric models pre-trained on regional data can match or outperform much larger generalist models on specific linguistic tasks.
Decision-makers and practitioners aiming to deploy language technology in these regions should invest in Africa-centric model pre-training and quality instruction tuning rather than relying solely on off-the-shelf global models. When using existing open-weight architectures for complex reasoning, teams should consider translating inputs to English as an interim mitigation strategy while building native-language support. Future work must expand human-annotated datasets to underrepresented language families and domain-specific education topics to ensure equitable AI access across the continent.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2021). Its MMLU benchmark is the foundation for AfriMMLU, so reading it first clarifies the source’s multilingual adaptation of broad academic-knowledge evaluation.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). XNLI supplies the natural-language-inference benchmark lineage behind AfriXNLI, making its task and cross-lingual evaluation setup essential context for IrokoBench.
No sufficiently relevant recommendations were found.
