MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Weihao XuanRui YangHeli QiQingcheng ZengYunze XiaoAosong FengDairui LiuYun XingJunjue WangFan Gao
Presents MMLU-ProX, an expert-verified benchmark spanning 29 languages with parallel 10-choice questions, exposing critical cross-lingual reasoning performance drops in low-resource languages across 36 state-of-the-art large language models.
As artificial intelligence systems are deployed internationally, assessing how well large language models comprehend and reason across diverse linguistic contexts has become essential. Most existing evaluation benchmarks focus predominantly on English or suffer from inconsistent translation quality across languages. Furthermore, previous multilingual benchmarks often lack the reasoning complexity required to test modern model architectures. The article addresses this gap by introducing MMLU-ProX, a standardized multilingual benchmark designed to evaluate advanced cross-lingual reasoning across 29 typologically diverse languages.
To construct MMLU-ProX, the authors translated and curated 11,829 identical, reasoning-focused questions per language across 14 academic and professional disciplines, alongside a 658-question lite version designed for rapid testing. The dataset was generated using a multi-stage translation and self-reflection pipeline driven by cutting-edge language models, followed by external automated audits. To ensure quality, over 30 professional translators conducted rigorous expert reviews across 15 languages, verifying high accuracy, fluency, and completeness. The authors then conducted extensive evaluations across 36 proprietary and open-weight models spanning sizes from 3.8 billion to 671 billion parameters.
The findings demonstrate substantial disparities in model performance across language groups. While top models like DeepSeek-R1 (75.5% overall average) and GPT-4.1 (72.7%) achieved strong results exceeding 75% to 80% accuracy in high-resource Western European and East Asian languages, performance dropped sharply in low-resource settings. In particular, non-Arabic African languages exhibited severe performance deficits; for instance, some models scored below 1% on Wolof, where top performance peaked at only 58.6%. The evaluation also revealed that reasoning-focused prompting and thinking architectures significantly improve multilingual accuracy, providing gains of up to 11.3% in lower-resource settings. Additionally, the lite version of the benchmark tracked full evaluation results within a 1.14% margin, confirming its efficacy for cost-effective testing.
These results demonstrate that current state-of-the-art models remain linguistically unbalanced, creating performance and equity risks for global deployments in underrepresented languages. While scaling model size and enabling explicit reasoning mechanisms mitigate some deficits, smaller models frequently fail on low-resource languages. Decision-makers should leverage the efficient lite benchmark to audit cross-lingual capabilities prior to deployment, while developers must prioritize expanding multilingual training data and reasoning capabilities to deliver fair, accessible AI across international markets.
Confidence in these findings is reinforced by rigorous human validation and consistent cross-model trends. However, readers should note limitations regarding language coverage, as extremely low-resource languages remain unrepresented. Furthermore, expert human verification was conducted on a sample rather than the complete question pool, and the current benchmark remains restricted to text-based evaluation without addressing multimodal contexts.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). MMLU-Pro supplies the high-difficulty question set that MMLU-ProX standardizes across languages, so it clarifies the benchmark’s core design.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). MMLU establishes the original subject-spanning evaluation framework that MMLU-Pro and, in turn, MMLU-ProX build upon.
No sufficiently relevant recommendations were found.
