MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Weihao XuanRui YangHeli QiQingcheng ZengYunze XiaoAosong FengDairui LiuYun XingJunjue WangFan Gao

article2025EMNLP91 citations

Presents MMLU-ProX, an expert-verified benchmark spanning 29 languages with parallel 10-choice questions, exposing critical cross-lingual reasoning performance drops in low-resource languages across 36 state-of-the-art large language models.

Listen

As artificial intelligence systems are deployed internationally, assessing how well large language models comprehend and reason across diverse linguistic contexts has become essential. Most existing evaluation benchmarks focus predominantly on English or suffer from inconsistent translation quality across languages. Furthermore, previous multilingual benchmarks often lack the reasoning complexity required to test modern model architectures. The article addresses this gap by introducing MMLU-ProX, a standardized multilingual benchmark designed to evaluate advanced cross-lingual reasoning across 29 typologically diverse languages.

To construct MMLU-ProX, the authors translated and curated 11,829 identical, reasoning-focused questions per language across 14 academic and professional disciplines, alongside a 658-question lite version designed for rapid testing. The dataset was generated using a multi-stage translation and self-reflection pipeline driven by cutting-edge language models, followed by external automated audits. To ensure quality, over 30 professional translators conducted rigorous expert reviews across 15 languages, verifying high accuracy, fluency, and completeness. The authors then conducted extensive evaluations across 36 proprietary and open-weight models spanning sizes from 3.8 billion to 671 billion parameters.

The findings demonstrate substantial disparities in model performance across language groups. While top models like DeepSeek-R1 (75.5% overall average) and GPT-4.1 (72.7%) achieved strong results exceeding 75% to 80% accuracy in high-resource Western European and East Asian languages, performance dropped sharply in low-resource settings. In particular, non-Arabic African languages exhibited severe performance deficits; for instance, some models scored below 1% on Wolof, where top performance peaked at only 58.6%. The evaluation also revealed that reasoning-focused prompting and thinking architectures significantly improve multilingual accuracy, providing gains of up to 11.3% in lower-resource settings. Additionally, the lite version of the benchmark tracked full evaluation results within a 1.14% margin, confirming its efficacy for cost-effective testing.

These results demonstrate that current state-of-the-art models remain linguistically unbalanced, creating performance and equity risks for global deployments in underrepresented languages. While scaling model size and enabling explicit reasoning mechanisms mitigate some deficits, smaller models frequently fail on low-resource languages. Decision-makers should leverage the efficient lite benchmark to audit cross-lingual capabilities prior to deployment, while developers must prioritize expanding multilingual training data and reasoning capabilities to deliver fair, accessible AI across international markets.

Confidence in these findings is reinforced by rigorous human validation and consistent cross-model trends. However, readers should note limitations regarding language coverage, as extremely low-resource languages remain unrepresented. Furthermore, expert human verification was conducted on a sample rather than the complete question pool, and the current benchmark remains restricted to text-based evaluation without addressing multimodal contexts.

No sufficiently relevant recommendations were found.

Cover for MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Abstract

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves multiple powerful LLMs for translation, followed by expert review to ensure accurate expression, consistent terminology, and cultural relevance. Building on this, we systematically evaluate 36 state-of-the-art LLMs, including reasoning-enhanced and multilingual-optimized LLMs. The results reveal significant disparities in the multilingual capabilities of LLMs: While they perform well in high-resource languages, their performance declines markedly in low-resource languages, particularly for African languages. Through MMLU-ProX, we aim to advance the development of more inclusive AI systems and promote equitable access to technology across global contexts.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Benchmark
  • 3.1 Overview
  • 3.2 Data Curation
  • 3.3 Translation Pipeline
  • 3.4 Expert Verification
  • 3.5 Total Cost
  • 4 Experiments
  • 4.1 Setups
  • 4.2 Overall Performance
  • 4.3 Impact of Reasoning Mode in Multilingual Performance
  • 4.4 Performance across Language Groups
  • 5 Analysis
  • 5.1 Model Size
  • 5.2 Prompting Strategies
  • 5.3 Full and Lite Versions
  • 6 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Language Categorization by Resource Availability
  • B Language Categorization by Geography
  • C Performance Patterns across Language Groups
  • D Translation Pipeline Analysis
  • E Expert Verification Guidance
  • F Translation Prompts
  • F.1 Self-Reflection Prompt
  • F.2 Translation Improvement Prompt
  • G Generative AI Statement
  • H Detailed Evaluation Results

Knowls

  1. Knowl 1 — MMLU-ProX provides parallel advanced multiple-choice evaluations in 29 languages

    definition

    MMLU-ProX extends the reasoning-focused MMLU-Pro benchmark to 29 languages: English, Chinese, Japanese, Korean, French, German, Spanish, Portuguese, Arabic, Thai, Hindi, Bengali, Swahili, Afrikaans, Czech, Hungarian, Indonesian, Italian, Marathi, Nepali, Russian, Serbian, Telugu, Ukrainian, Urdu, Vietnamese, Wolof, Yoruba, and Zulu. Every language version contains the same 11,829 questions, enabling comparisons across languages using matched questions rather than different language-specific test sets. The questions cover 14 disciplines and 57 subjects and use 10 answer choices. The full benchmark therefore contains 343,041 language-question instances. Each version includes 70 validation questions for few-shot prompt construction.

  2. Knowl 2 — A staged curation and translation pipeline produces the multilingual question sets

    model/method

    Before translation, four interdisciplinary specialists spent about 20 hours cleaning the English source: they removed or merged duplicate and partially duplicate questions, corrected grammatical and formatting problems, and fixed inconsistencies between question statements and answer options. Claude Sonnet 3.7 then translated the curated questions through three stages: an initial translation prompted to preserve technical terminology, cross-item consistency, cultural appropriateness, mathematical notation, code, and formatting; self-review against the source; and an edit incorporating that review. A separate model performed external verification to reduce reliance on a single model: OpenAI o3 checked Swahili, Zulu, Yoruba, and Wolof, while GPT-4.1 checked the other target languages. The process was designed to flag significant discrepancies for human review, rather than to rewrite every translation automatically.

  3. Knowl 3 — Expert review found high translation scores, including in low-resource languages

    empirical result

    Professional translators evaluated a random sample of 20 questions from each of 14 disciplines, using the same questions across languages. Fifteen languages were assessed by two translators each; annotators were native speakers of the target language and proficient in English, and the work involved more than 400 hours. Translators rated accuracy, fluency, and completeness on a 1–5 scale. Across the 15 languages, all three metrics averaged at least 4; the observed score ranges were 4.06–4.90 for accuracy, 4.41–4.84 for fluency, and 4.64–4.99 for completeness. Yoruba had the lowest accuracy score (4.06), while Korean had the highest (4.90); German had the lowest completeness score (4.64). The team retranslates a language-discipline set if both translators score below 3 on any metric; only Yoruba law required this intervention.

    In a separate English-to-Japanese comparison, expert scores for the benchmark’s agent-based method were 4.60 for accuracy, 4.65 for fluency, and 4.99 for completeness. Reasoning-based translation scored 4.56, 4.21, and 4.99, respectively; native-speaking translators scored 4.24, 4.14, and 4.56. The paper reports that reasoning-based translation used substantially more tokens than the agent-based method.

  4. Knowl 4 — Evaluation covers 36 models with zero-shot and five-shot prompting

    experimental setup

    The benchmark evaluation covered 36 state-of-the-art LLMs, including open-weight and proprietary systems and open-weight models ranging from 3.8B to 671B parameters. The main reported results use five-shot chain-of-thought prompting; the study also evaluates zero-shot prompting. Open models were run with vLLM on an H100 cluster, and proprietary models were accessed through their APIs. The authors estimate that the unified evaluation used more than 10,000 GPU hours. Reported model accuracies are percentages.

  5. Knowl 5 — DeepSeek-R1 leads the reported five-shot averages across languages

    empirical result

    On the five-shot evaluation averaged over all 29 languages, DeepSeek-R1 scored 75.5%, ahead of Qwen3-235B-Think at 74.9%, GPT-4.1 at 72.7%, DeepSeek-V3 at 70.5%, and o4-mini at 69.3%. These results show that the strongest systems achieve high aggregate performance on the parallel benchmark, while the overall averages conceal large differences between individual languages.

  6. Knowl 6 — Performance varies sharply by language, with the largest gaps among African languages

    empirical result

    The five-shot results show strong performance on many Western European languages but much lower and more variable scores on several African languages. Qwen3-235B-Think exceeds 77% on every Western European language in the evaluation, including 80.7% on English and Spanish and 80.9% on Italian. By contrast, Wolof scores across evaluated models range from 0.6% to 58.6%, Yoruba from 0.6% to 57.0%, and Zulu from 11.5% to 67.3%. African-language performance is not uniformly low: Arabic reaches 78.7% with Qwen3-235B-Think, and DeepSeek-R1 scores 81.3% on Indonesian. South Asian performance also varies within the group; for example, the paper identifies Hindi as stronger and Telugu as more challenging. Thus, language-level results reveal capability disparities that aggregate scores do not capture.

  7. Knowl 7 — Reasoning-enhanced variants improve multilingual scores, including on low-resource languages

    empirical result

    In five-shot evaluation, DeepSeek-R1 scores 75.5% overall compared with 70.5% for DeepSeek-V3, a 5.0 percentage-point difference. The gap is larger on the low-resource African languages Wolof and Yoruba: R1 scores 58.6% versus V3’s 47.3% on Wolof, and 57.0% versus 47.7% on Yoruba. Qwen3-235B-Think also outperforms the non-thinking Qwen3-235B overall (74.9% versus 66.7%) and achieves high scores on Western European languages, including 80.7% on English and Spanish and 80.9% on Italian. These comparisons support the paper’s finding that reasoning-enhanced models can improve multilingual benchmark performance, with gains that vary by language.

  8. Knowl 8 — Larger Qwen3 dense models improve average accuracy, with larger gains on some languages

    empirical result

    For Qwen3 dense models evaluated with five-shot prompting, average accuracy rises from 42.0% at 4B parameters to 46.0% at 8B, 54.0% at 14B, and 59.9% at 32B. The largest step is from 8B to 14B (+8.0 percentage points); the 4B-to-8B and 14B-to-32B gains are +4.0 and +5.9 points. English improves by 12.6 points between the 4B and 32B models (59.2% to 71.8%), whereas Wolof improves by 20.5 points (5.6% to 26.1%) and Russian by 14.4 points (53.6% to 68.0%). The reported scaling pattern is therefore uneven across languages, with larger models producing greater gains for some lower-resource languages.

  9. Knowl 9 — The lite benchmark closely tracks full-benchmark results

    empirical result

    The lite version was created by uniformly sampling 5% of questions from each of the 14 disciplines for every language. It contains 658 questions per language, including the same 70 validation questions used for few-shot prompt construction; evaluation therefore uses 588 lite questions versus 11,759 full-version questions. Across evaluated models, the average performance difference between lite and full evaluation is 1.14%, and the lite version largely preserves model rankings. For five-shot results, the full-to-lite differences are 1.5 points for DeepSeek-R1 (75.5% versus 74.0%), 1.1 for GPT-4.1 (72.7% versus 71.6%), 0.4 for DeepSeek-V3 (70.5% versus 70.1%), and 1.3 for o4-mini (69.3% versus 68.0%). The paper reports a difference below 1% for Wolof.

  10. Knowl 10 — Five-shot prompting usually helps, but its effect depends on model and language

    empirical result

    Comparisons of zero-shot and five-shot prompting for selected models and languages show that demonstrations generally improve accuracy, but the gains are not uniform. GPT-4.1 gains 3.7 percentage points on English with five-shot prompting. The study reports larger benefits for low-resource African languages than for high-resource English in general, while reasoning-enhanced Qwen3 models in thinking mode show smaller changes between the two prompting settings. These results indicate that prompt strategy affects multilingual scores differently across languages and model types.

  11. Knowl 11 — Benchmark coverage and translation verification remain incomplete

    limitation

    MMLU-ProX covers 29 languages, but budget constraints prevented inclusion of additional languages, especially extremely low-resource ones. Expert translation review was conducted for selected languages and sampled items, not comprehensively for every language and subject area; automated translation may still contain subtle errors, particularly in complex or domain-specific content. The benchmark evaluates text-only inputs and does not cover multimodal tasks.

Coverage note — The paper’s approximately $80,000 development-cost estimate and detailed translation prompt templates are omitted because they are resource accounting and implementation artifacts rather than standalone findings needed to reconstruct the benchmark or its main results.

References

  1. 1.Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, and 1 others. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905.
  2. 2.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  3. 3.Anthropic. 2025. Claude 3.7 sonnet. Large language model.
  4. 4.Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, and 1 others. 2024. Aya 23: Open weight releases to further multilingual progress. arXiv preprint arXiv:2405.15032.
  5. 5.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. ArXiv, abs/2302.04023.
  6. 6.Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, and 1 others. 2024. Internlm2 technical report. arXiv preprint arXiv:2403.17297.
  7. 7.Daniel Deutsch, Eleftheria Briakou, Isaac Caswell, Mara Finkelstein, Rebecca Galor, Juraj Juraska, Geza Kovacs, Alison Lui, Ricardo Rei, Jason Riesa, and 1 others. 2025. Wmt24++: Expanding the language coverage of wmt24 to 55 languages & dialects. arXiv preprint arXiv:2502.12404.
  8. 8.Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, and Mikel Artetxe. 2024. Do multilingual language models think better in English? In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 550–564, Mexico City, Mexico. Association for Computational Linguistics.
  9. 9.GemmaTeam, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, and 1 others. 2025. Gemma 3 technical report. arXiv preprint arXiv:2503.19786.
  10. 10.Google DeepMind. 2025. Gemini 2.5: Our most intelligent ai model.
  11. 11.Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  12. 12.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
  13. 13.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a. Measuring massive multitask language understanding. Preprint, arXiv:2009.03300.
  14. 14.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021b. Measuring massive multitask language understanding. ICLR.
  15. 15.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  16. 16.Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, and 1 others. 2020. Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation. arXiv preprint arXiv:2004.01401.
  17. 17.Meta AI. 2025. Introducing llama 4: Advancing multimodal intelligence.
  18. 18.MistralAI. 2025. Mistral-small-24b-instruct-2501.
  19. 19.Joel Niklaus, Jakob Merane, Luka Nenadic, Sina Ahmadi, Yingqiang Gao, Cyrill AH Chevalley, Claude Humbel, Christophe Gösken, Lorenzo Tanzi, Thomas Lüthi, and 1 others. 2025. Swiltra-bench: The swiss legal translation benchmark. arXiv preprint arXiv:2503.01372.
  20. 20.Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, and 1 others. 2025. Humanity’s last exam. arXiv preprint arXiv:2501.14249.
  21. 21.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020a. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  22. 22.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020b. Xcopa: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376.
  23. 23.Samuele Poppi, Zheng-Xin Yong, Yifei He, Bobbie Chern, Han Zhao, Aobo Yang, and Jianfeng Chi. 2024. Towards understanding the fragility of multilingual llms against fine-tuning attacks. In North American Chapter of the Association for Computational Linguistics.
  24. 24.QwenTeam. 2025. Qwq-32b: Embracing the power of reinforcement learning.
  25. 25.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and 1 others. 2021. Xtreme-r: Towards more challenging and nuanced multilingual evaluation. arXiv preprint arXiv:2104.07412.
  26. 26.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022. Language models are multilingual chain-of-thought reasoners. Preprint, arXiv:2210.03057.
  27. 27.Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Wei-Yin Ko, Sebastian Ruder, Madeline Smith, Antoine Bosselut, Alice Oh, Andre F. T. Martins, Leshem Choshen, and 5 others. 2025. Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation. Preprint, arXiv:2412.03304.
  28. 28.Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2024. Kmmlu: Measuring massive multitask language understanding in korean. arXiv preprint arXiv:2402.11548.
  29. 29.Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2025. KMMLU: Measuring massive multitask language understanding in Korean. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4076–4104, Albuquerque, New Mexico. Association for Computational Linguistics.
  30. 30.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024a. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Preprint, arXiv:2406.01574.
  31. 31.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, and 1 others. 2024b. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  32. 32.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, F. Xia, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903.
  33. 33.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388.
  34. 34.Ziqi Yin, Hao Wang, Kaito Horio, Daisuike Kawahara, and Satoshi Sekine. 2024. Should we respect llms? a cross-lingual study on the influence of prompt politeness on llm performance. In Proceedings of the Second Workshop on Social Influence in Conversations (SICon 2024), pages 9–35.
  35. 35.Arda Yüksel, Abdullatif Köksal, Lütfi Kerem Senel, Anna Korhonen, and Hinrich Schuetze. 2024. TurkishMMLU: Measuring massive multitask language understanding in Turkish. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 7035–7055, Miami, Florida, USA. Association for Computational Linguistics.
  36. 36.Chengzhi Zhong, Fei Cheng, Qianying Liu, Junfeng Jiang, Zhen Wan, Chenhui Chu, Yugo Murawaki, and Sadao Kurohashi. 2024. Beyond english-centric llms: What language do multilingual language models think in? Preprint, arXiv:2408.10811.

Citation

MLA
Xuan, W., et al. “MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 1513–32, https://doi.org/10.18653/v1/2025.emnlp-main.79.
APA
Xuan, W., Yang, R., Qi, H., Zeng, Q., Xiao, Y., Feng, A., Liu, D., Xing, Y., Wang, J., Gao, F., Lu, J., Jiang, Y., Li, H., Li, X., Yu, K., Dong, R., Gu, S., Li, Y., Xie, X., … Li, I. (2025). MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 1513–1532. https://doi.org/10.18653/v1/2025.emnlp-main.79
Chicago
Xuan, W., R. Yang, H. Qi, et al. 2025. “MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 1513–32. https://doi.org/10.18653/v1/2025.emnlp-main.79.
Harvard
Xuan, W. et al. (2025) “MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation”, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1513–1532. Available at: https://doi.org/10.18653/v1/2025.emnlp-main.79.
Vancouver
1. Xuan W, Yang R, Qi H, et al (2025) MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1513–1532

BibTeX

@inproceedings{xuan-etal-2025-mmlu,
    title = "{MMLU}-{P}ro{X}: A Multilingual Benchmark for Advanced Large Language Model Evaluation",
    author = "Xuan, Weihao  and
      Yang, Rui  and
      Qi, Heli  and
      Zeng, Qingcheng  and
      Xiao, Yunze  and
      Feng, Aosong  and
      Liu, Dairui  and
      Xing, Yun  and
      Wang, Junjue  and
      Gao, Fan  and
      Lu, Jinghui  and
      Jiang, Yuang  and
      Li, Huitao  and
      Li, Xin  and
      Yu, Kunyu  and
      Dong, Ruihai  and
      Gu, Shangding  and
      Li, Yuekang  and
      Xie, Xiaofei  and
      Juefei-Xu, Felix  and
      Khomh, Foutse  and
      Yoshie, Osamu  and
      Chen, Qingyu  and
      Teodoro, Douglas  and
      Liu, Nan  and
      Goebel, Randy  and
      Ma, Lei  and
      Marrese-Taylor, Edison  and
      Lu, Shijian  and
      Iwasawa, Yusuke  and
      Matsuo, Yutaka  and
      Li, Irene",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.79/",
    doi = "10.18653/v1/2025.emnlp-main.79",
    pages = "1513--1532",
    ISBN = "979-8-89176-332-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/