MEGA: Multilingual Evaluation of Generative AI
Kabir AhujaHarshita DiddeeRishav HadaMillicent OchiengKrithika RameshPrachi JainAkshay Uttama NambiTanuja GanuSameer SegalMohamed Ahmed
Introduces MEGA, a comprehensive multilingual benchmark spanning 16 standard tasks across 70 typologically diverse languages to systematically compare generative large language models against prior state-of-the-art systems and identify critical performance gaps in low-resource settings.
Large Language Models (LLMs) such as GPT-3.5 and GPT-4 have demonstrated remarkable capabilities in English-language tasks, but their effectiveness across other global languages remains poorly understood. Because generative artificial intelligence is increasingly deployed in global products and public-facing services, uneven language performance risks creating severe digital inequalities. The article introduces the Multilingual Evaluation of Generative AI (MEGA) benchmark to comprehensively evaluate generative LLMs across diverse languages and standard natural language processing tasks, comparing them against specialized, fine-tuned baseline models.
The researchers assessed four major generative models (text-davinci-003, GPT-3.5-Turbo, GPT-4, and BLOOMZ) alongside strong fine-tuned baselines across 16 standard datasets covering 70 typologically diverse languages and 21 language families. The evaluation spanned five core task categories: text classification, question answering, sequence labeling, text summarization, and responsible AI metrics (gender bias and toxicity). The team tested multiple prompting approaches, including native monolingual prompting, zero-shot cross-lingual prompting with English examples, and translate-test workflows where non-English inputs are automatically translated into English prior to generation.
The findings show a substantial performance gap between English and non-English languages, which is especially severe for low-resource languages that use non-Latin scripts. While GPT-4 generally outperformed GPT-3.5 and BLOOMZ, fine-tuned models such as TULRv6 consistently surpassed generative models on structured tasks like classification, sequence labeling, and extractive question answering. For lower-resource languages like Burmese, Tamil, and Telugu, translating the input into English first (translate-test) boosted GPT-3.5-Turbo's accuracy by more than 30% relative to prompting in the native language, though a large gap compared to English baseline performance remained. In addition, tokenizer inefficiencies caused low-resource non-Latin languages to require up to ten times more tokens per word than English, which strongly correlated with lower task accuracy and dramatically higher processing costs.
These results demonstrate that organizations cannot assume off-the-shelf generative models will perform reliably or equitably across all global user bases. Relying on zero-shot prompting in low-resource native languages introduces significant operational and safety risks. Furthermore, the high tokenization burden increases financial costs for processing non-Latin scripts. The evidence indicates that common prompt engineering tricks, such as adding step-by-step reasoning or native-language prompt instructions, provide minimal benefit and can even degrade performance.
Organizations deploying multilingual AI should not rely solely on generative LLMs for critical, structured tasks in low-resource languages, but should instead favor fine-tuned specialized models or hybrid translate-test architectures where feasible. Practitioners should use English-based instruction templates and include at least four to eight few-shot examples within the native language context. In the long term, AI providers must rebalance pre-training data distributions and redesign tokenizers to handle non-Latin scripts more equitably.
Confidence in these findings is high regarding the overall performance disparities, but readers should note specific limitations. Preliminary analysis showed strong evidence of test data contamination in LLM pre-training sets, meaning the real-world performance of these models on unseen data is likely even lower than reported. Additionally, extremely under-resourced languages (such as many indigenous and African languages) remain underrepresented in standard academic benchmarks, requiring targeted localized pilot studies before making deployment decisions.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Introduces XNLI, a foundational cross-lingual benchmark covering 15 languages that establishes standard evaluation protocols for multilingual text understanding adapted by MEGA.
- Paper: Unsupervised Cross-lingual Representation Learning at Scale, Alexis Conneau et al. (2019). Establishes XLM-RoBERTa as a premier non-autoregressive multilingual baseline across 100 languages against which MEGA directly benchmarks generative models.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). Presents mT5, setting the standard massively multilingual sequence-to-sequence baseline that serves as a core point of comparison for generative LLM evaluation.
- Paper: Cross-lingual Language Model Pretraining, Guillaume Lample et al. (2019). Provides foundational methods for cross-lingual pre-training and zero-shot cross-lingual evaluation upon which modern multilingual assessment relies.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). Analyzes the zero-shot cross-lingual transfer capabilities and limitations of multilingual representations across diverse scripts and typologies.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). Introduces the BIG-bench framework for broad, multi-task language model evaluation across scale, informing MEGA's comprehensive evaluation design.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Demonstrates the in-context few-shot prompting paradigm for generative autoregressive language models evaluated throughout the MEGA benchmark.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). Introduces MMLU, standardizing multi-task evaluation methodologies across diverse subjects that generative benchmark suites adopt.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). Provides a comprehensive survey synthesizing LLM evaluation methodologies, incorporating large-scale benchmarking frameworks like MEGA across tasks and multilingual contexts.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). Probes the internal multilingual representation and neuron routing mechanisms within LLMs, offering mechanistic explanations for the cross-lingual performance patterns observed in MEGA.
- Paper: A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity, Yejin Bang et al. (2023). Extends multilingual evaluation of ChatGPT by analyzing interactions across multimodal, reasoning, and hallucination dimensions.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). Directly tackles the low-resource generation bottleneck highlighted by MEGA by scaling multilingual generation and evaluation to more than 1,600 languages.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). Builds on generative benchmark findings to create more rigorous, discriminative multi-task reasoning evaluations that prevent metric saturation.
