The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
Lucas BandarkarDavis LiangBenjamin MullerMikel ArtetxeSatya Narayan ShuklaDonald HusaNaman GoyalAbhinandan KrishnanLuke ZettlemoyerMadian Khabsa
Introduces a fully parallel reading comprehension benchmark across 122 language variants to evaluate and directly compare language models across high-, medium-, and low-resource languages.
As artificial intelligence applications expand globally, natural language processing systems face growing demands to support diverse global populations. However, the lack of high-quality, parallel evaluation datasets across a broad range of languages severely limits the assessment and development of technologies for non-English speakers, particularly in medium- and low-resource languages. Most existing language understanding benchmarks cover fewer than 30 languages, leaving critical blind spots regarding whether state-of-the-art models truly generalize across global linguistic boundaries.
To address this critical evaluation gap, the article introduces BELEBELE, a parallel multiple-choice reading comprehension benchmark designed to evaluate language models across 122 language variants. The primary objective is to measure, compare, and establish baseline performance for both multilingual masked language models and English-centric large language models across diverse language families and resource levels.
The benchmark consists of 900 curated four-option multiple-choice questions derived from 488 passages sourced from the FLORES-200 machine translation corpus, spanning 27 language families and 29 scripts. Professional human translators created the dataset without machine translation, ensuring full parallelism across all 122 language variants, including romanized variants of five Indo-Aryan languages. The authors designed questions to measure text comprehension while resisting simple pattern-matching shortcuts, confirmed via iterative quality checks and statistical featurization. The evaluation assessed multiple masked language models (such as XLM-V and INFOXLM) and prominent large language models (including LLAMA variants, FALCON, and GPT-3.5-Turbo) across few-shot, zero-shot, and fine-tuning configurations.
The analysis revealed several critical findings regarding language model capabilities. First, masked language models trained on balanced multilingual data with extensive vocabularies vastly outperform English-centric large language models on language coverage. For instance, XLM-V achieved above 50% accuracy on over 76% of evaluated languages when fine-tuned, whereas GPT-3.5-Turbo exceeded that threshold on only 43.4% of languages and LLAMA 2 (70B) on 38.5%. Second, English-centric models excel on top-tier languages—such as LLAMA 2 (70B) reaching 90.9% accuracy on English—but drop off sharply on medium- and low-resource languages. Third, cascading translation provides a substantial operational benefit for large language models: translating non-English inputs into English before inference (Translate-Test) raised LLAMA-2-CHAT (70B) average zero-shot accuracy from 44.0% to 57.1% across 91 languages, improving performance in 68 languages. Finally, human evaluation achieved 97.6% accuracy on English questions, substantially higher than all evaluated models, demonstrating that the benchmark presents a demanding, meaningful challenge.
These findings have major strategic implications for deploying language technologies across international markets. Relying solely on raw, English-centric foundation models introduces severe operational and performance risks in non-English contexts, risking total failure in low-resource environments. Vocabulary size and pretraining data distribution represent foundational bottlenecks; models with larger, tailored tokenizers such as XLM-V retain linguistic capabilities much deeper into the long tail of global languages. In addition, the success of translating inputs into English demonstrates that pipeline architectures combining specialized machine translation with large language models offer a cost-effective alternative to training massive native models for each language.
Organizations and developers should adopt specific, evidence-backed practices based on these results. When serving global user bases, teams should deploy architectures that leverage machine translation pipelines or specialized multilingual models rather than relying on direct zero-shot prompting of English-centric systems. Model builders should also allocate larger vocabulary budgets and include balanced multilingual corpora during pretraining. Looking ahead, future research should develop complementary benchmarks that measure nuanced cultural factors—such as local formality and societal values—which pure parallel translations do not capture.
The conclusions should be interpreted in light of certain limitations. The benchmark is derived from English source text and parallel translations, which introduces subtle translation artifacts and excludes localized cultural context. Furthermore, lack of transparency regarding the exact training data compositions for proprietary models like GPT-3.5-Turbo limits full experimental reproducibility. Nevertheless, the rigorous multi-stage human validation and statistical quality checks support high confidence in the benchmark’s comparative evaluations across all 122 language variants.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). This paper establishes XLM-R, the foundational massively multilingual masked language model architecture and pretraining methodology directly evaluated and analyzed in Belebele.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). This work introduces the foundational parallel multi-language evaluation paradigm (XNLI) that Belebele scales from sentence classification to machine reading comprehension across 122 language variants.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). This benchmark investigates multilingual evaluation and generative model capabilities across typologically diverse languages, directly motivating the need for Belebele's massively expanded language coverage.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). This paper introduces mT5 and the mC4 corpus, establishing essential baseline concepts for massively multilingual text representation and evaluation.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). This paper presents BLOOM, providing key insights into the training dynamics and cross-lingual evaluation of large multilingual open-access language models.
- Paper: Language Contamination Helps Explains the Cross-lingual Capabilities of English Pretrained Models, Terra Blevins et al. (2022). This work analyzes how language contamination in pretraining corpora drives cross-lingual transfer, helping contextualize Belebele's findings on English-centric LLM performance across languages.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). This foundational study introduces cross-lingual language model pretraining and cross-lingual NLU benchmarks that underpin multilingual model evaluation.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). This paper explores scaling language models and evaluation corpora to hundreds of low-resource languages, addressing the exact coverage gaps Belebele targets.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). This paper provides foundational probing into zero-shot cross-lingual transfer in masked language models across diverse scripts and typologies.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). This paper presents Llama 2, representative of the English-centric large language models evaluated in Belebele's cross-lingual comprehension studies.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). MEGAVERSE expands on Belebele's evaluation methodology by benchmarking both proprietary and open-source models across 83 languages, modalities, and tasks while accounting for test data contamination.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). BUFFET builds upon Belebele's findings comparing smaller multilingual models with large LLMs by benchmarking few-shot cross-lingual transfer and in-context learning across 54 languages.
- Paper: Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model, Ahmet Üstün et al. (2024). Aya extends massively multilingual evaluation and instruction-following across 101 languages, directly addressing the low-resource language comprehension challenges highlighted by Belebele.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). Language Ranker extends cross-lingual evaluation by introducing internal representation alignment metrics to quantify model proficiency across 94 high- and low-resource languages.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). MMTEB extends massive multilingual evaluation to over 250 languages across embedding and retrieval tasks, continuing the expansion of global language benchmarks.
- Paper: Do Llamas Work in English? On the Latent Language of Multilingual Transformers, Chris Wendler et al. (2024). This paper investigates the mechanistic interpretability behind the cross-lingual transfer mechanisms in multilingual LLMs that Belebele empirically evaluates.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). This work explores how post-training alignment impacts multilingual comprehension and regional variations across diverse global languages.
- Paper: Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?, Nishant Balepur et al. (2024). This paper analyzes multiple-choice evaluation dynamics and potential dataset artifacts in large language models, providing critical diagnostic depth for datasets structured like Belebele.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Omnilingual MT dramatically scales global language coverage to over 1,600 languages, extending the boundaries of multilingual NLP far beyond 122 language variants.
