MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
Sanchit AhujaDivyanshu AggarwalVarun GummaIshaan WattsAshutosh SatheMillicent OchiengRishav HadaPrachi JainMohamed AhmedKalika Bali
Presents a comprehensive evaluation of leading language and multimodal models across 22 datasets and 83 languages to reveal performance disparities on low-resource languages and quantify the prevalence of test set contamination.
Recent advances in artificial intelligence have produced highly capable large language models, but performance evaluation remains heavily skewed toward English. This English-centric focus leaves the capabilities of modern models in non-English languages and multimodal settings largely unexamined, creating significant operational risks and widening the digital divide for global populations.
The main objective of the article is to provide an extensive, standardized evaluation of leading commercial and open-source models across a wide range of non-English languages, modalities, and task types, while also measuring the extent of benchmark test data contamination.
To evaluate these systems, the authors established a broad benchmarking suite covering 22 datasets across 83 languages, including low-resource African and Indic languages. The study compared proprietary models such as GPT-4, Gemini-Pro, and PaLM2 alongside open-source alternatives including Llama 2, Mistral, and Gemma across tasks such as question answering, translation, reasoning, dialogue, and multimodal image understanding. In addition, the authors applied formal statistical and perturbation-based tests to measure whether models had previously memorized evaluation datasets during training.
The investigation produced several key findings. First, large proprietary models consistently and substantially outperformed smaller open-source models, with GPT-4, Gemini-Pro, and PaLM2 demonstrating the highest overall capabilities. In contrast, smaller open-source models struggled severely on low-resource and non-Latin script languages, frequently dropping to near-zero accuracy. Second, performance varied dramatically across language families; Germanic and Romance languages achieved the highest scores, while Bantu and Afro-Asiatic languages lagged significantly behind. Third, models performed much better on structured, constrained tasks such as multiple-choice reading comprehension and part-of-speech tagging than on open-ended generative tasks like summarization and conversational dialogue. Fourth, multimodal vision-language models showed similar script disparities, with GPT-4-Vision outperforming alternatives but still degrading on non-Latin script image captioning. Finally, the contamination analysis revealed that nearly all evaluated models show evidence of test data contamination, indicating that published benchmark scores often overstate true general reasoning abilities.
These results demonstrate that organizations cannot rely on English benchmark performance as a proxy for multilingual or global readiness. Deploying off-the-shelf smaller models in non-English or multilingual contexts introduces significant operational, brand, and safety risks due to elevated error rates and hallucinations. Furthermore, widespread data contamination complicates procurement and model selection, as public benchmarks may reflect memorization rather than actual task competence.
Decision-makers should exercise caution when deploying generative models in non-Western languages and avoid relying on standard open-source base models for low-resource language tasks without dedicated fine-tuning or language-specific adaptations. Technical teams should implement rigorous contamination detection protocols and evaluate systems using fresh, non-public data. Moving forward, additional work is needed to develop representative multilingual benchmarks, evaluate safety and fairness dimensions beyond core accuracy, and explore targeted architectures for under-resourced language families.
These conclusions should be interpreted in light of certain limitations. Proprietary commercial systems were accessed through cloud application programming interfaces that may include undocumented post-processing, and resource constraints prevented exhaustive prompt tuning across all task and model combinations. Nonetheless, the comparative trends across models, scripts, and language families remain robust and provide a reliable baseline for strategic planning.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It introduces the initial MEGA multilingual benchmark across 70 languages, establishing the baseline methodology and evaluation gaps directly expanded upon in MEGAVERSE.
- Paper: A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity, Yejin Bang et al. (2023). It provides early zero-shot empirical evaluations of multilingual and multimodal capabilities in foundational LLMs, highlighting the acute performance drop in low-resource and non-Latin languages.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). It surveys foundational frameworks, tasks, and pitfalls in large language model evaluation, establishing standard practices used to structure comprehensive cross-lingual assessments.
- Paper: PaLM 2 Technical Report, Rohan Anil et al. (2023). It details the architecture and multilingual training recipes of PaLM 2, one of the central frontier model families benchmarked in MEGAVERSE.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). It details the architecture, pretraining data distribution, and open-source release of Llama 2, which serves as a core open baseline in the MEGAVERSE evaluation.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). It establishes the canonical XNLI cross-lingual understanding benchmark, which forms the core foundation for evaluating multilingual sentence-level reasoning in modern LLM suites.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). It provides the structured evaluation methodologies and robustness strategies for multimodal vision-language models tested alongside text models in comprehensive benchmarks.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). It builds upon empirical multilingual performance disparities across high- and low-resource languages by introducing an internal representation metric to systematically grade cross-lingual alignment.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). It extends massively multilingual evaluation across 250+ languages and 500+ tasks specifically focused on text embeddings and representation quality.
- Paper: Dynamic Evaluation of Large Language Models by Meta Probing Agents, Kaijie Zhu et al. (2024). It introduces dynamic meta probing agents to evaluate LLMs while systematically addressing the critical benchmark data contamination issues identified in MEGAVERSE.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). It operationalizes time-segmented, contamination-free evaluation protocols to overcome the training data leakage problems highlighted during large-scale model benchmarking.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). It expands holistic model benchmarking to 103 frontier systems using instance-specific rubrics across nine capabilities, including multilingual generation.
