mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation

Md. Nishat RaihanAntonios AnastasopoulosMarcos Zampieri

article2025NAACL47 citations

Presents mHumanEval, the first massively multilingual code generation benchmark spanning 204 natural languages and 25 programming languages, enabling researchers to rigorously assess how well code LLMs understand prompts across high-, mid-, and low-resource languages.

Listen

Artificial intelligence models that generate computer software code have rapidly advanced, but standard evaluation benchmarks rely almost exclusively on English prompts and Python solutions. This narrow scope creates significant blind spots for global software development and risks overestimating model capabilities. The article introduces mHumanEval, an expanded multilingual code generation benchmark, to assess how well leading artificial intelligence systems understand coding instructions written across diverse high-, mid-, and low-resource human languages.

To build mHumanEval, the authors translated 164 standard coding tasks into 204 natural languages using multiple machine translation engines, generating 13 candidates per prompt and selecting the highest-quality translations via automated linguistic scoring. In total, the benchmark comprises 33,456 natural language prompts spanning 25 programming languages, including four newly added legacy and scientific languages: MATLAB, Visual Basic, Fortran, and COBOL. To validate translation quality, expert human programmers produced verified reference translations across 15 diverse languages. The authors then evaluated six prominent proprietary and open-source models using the first-attempt functional accuracy metric, Pass@1.

Across all models, coding accuracy declines as instructions move from high-resource languages like English and Spanish to lower-resource languages. Proprietary frontier models, specifically GPT-4o and Claude 3.5, demonstrate the highest overall resilience, maintaining roughly 60% to 74% accuracy even on low-resource language prompts in Python. In contrast, specialized code-focused models like WizardCoder and DeepSeek-Coder experience severe performance collapses when moving away from English or Chinese, with accuracy falling from over 80–90% down to near 0% on low-resource languages. Furthermore, general multilingual models without specialized coding fine-tuning, such as Aya, achieve stable but moderate accuracy (around 35–45%) across language tiers. Scripting languages like JavaScript and Ruby proved notably more difficult across all models than Python, C++, and Java. Qualitative error analysis reveals that failures in non-English contexts stem primarily from models misinterpreting translated programming concepts or erroneously translating reserved programming keywords directly into the generated code.

These findings demonstrate that training models exclusively on source code and English documentation severely limits their global utility. Multilingual natural language understanding is essential for code generation systems to serve non-English-speaking engineers effectively. Deploying specialized coding assistants in multilingual environments without cross-lingual validation poses severe risks of code failure and operational disruption, whereas advanced general-purpose models demonstrate better adaptability.

Organizations developing or deploying automated coding tools should prioritize multilingual pretraining data and avoid assuming that English benchmark performance reflects global readiness. Practitioners evaluating models should utilize diverse subsets like mHumanEval-mini for rapid multilingual screening. Additionally, because models occasionally produce erroneous or non-compiling code when prompted in non-English languages, organizations must enforce automated execution safeguards, such as isolated virtual execution environments, to prevent runaway loops or system crashes during automated testing.

Confidence in these findings is high for functional code accuracy across the tested models, though evaluations were constrained to the Pass@1 metric due to the significant computational cost of processing hundreds of thousands of prompt variations. Future work should expand human validation beyond the initial 15 languages, evaluate additional model architectures, and test multiple-attempt metrics across larger problem sets.

No sufficiently relevant recommendations were found.

Cover for mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation

Abstract

Recent advancements in large language models (LLMs) have significantly enhanced code generation from natural language prompts. The HumanEval Benchmark, developed by OpenAI, remains the most widely used code generation benchmark. However, this and other Code LLM benchmarks face critical limitations, particularly in task diversity, test coverage, and linguistic scope. Current evaluations primarily focus on English-to-Python conversion tasks with limited test cases, potentially overestimating model performance. While recent works have addressed test coverage and programming language (PL) diversity, code generation from low-resource language prompts remains largely unexplored. To address this gap, we introduce mHumanEval¹, an extended benchmark supporting prompts in over 200 natural languages. We employ established machine translation methods to compile the benchmark, coupled with a quality assurance process. Furthermore, we provide expert human translations for 15 diverse natural languages (NLs). We conclude by analyzing the multilingual code generation capabilities of state-of-the-art (SOTA) Code LLMs, offering insights into the current landscape of cross-lingual code generation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The mHumanEval Benchmark
  • 3.1 Prompt Extraction
  • 3.2 Prompt Translation
  • 3.3 Evaluating Prompt Quality
  • 3.4 Categorization based on Language Classes
  • 3.5 PL coverage
  • 3.6 mHumanEval Subsets
  • 3.7 mHumanEval - Expert
  • 4 Experiments
  • 5 Insights and Analysis
  • 5.1 LLMs’ Performance Analysis
  • 5.2 Performance based on Language Classes
  • 5.3 Error Analysis
  • 5.4 Ablation Study
  • 6 Conclusion
  • References
  • A list of NLs and PLs in mHumanEval
  • A.1 List of PLs
  • A.2 List of NLs: mHumanEval-Expert
  • A.3 List of NLs: mHumanEval
  • B Prompt Translation and Evaluation Algorithm
  • C Evaluation Metric 1: BERTScore
  • D Evaluation Metric 2: CometKiwi
  • E Annotator Details
  • F Comparison of the Prompt Qualities by the 3 models vs mHumanEval
  • G Prompt Templates
  • H Error Analysis - Examples
  • H.1 Task Misunderstanding
  • H.2 Multilingual Keyword Issues
  • H.3 Garbage Results
  • I Experimental Setup
  • I.1 Machine Translation
  • I.2 Code Generation
  • J Evaluation Results: mHumanEval-PL
  • J.1 mHumanEval-C++
  • J.2 mHumanEval-JAVA
  • J.3 mHumanEval-JavaScript
  • J.4 mHumanEval-Ruby
  • J.5 Analyzing PL-specific results
  • K Evaluating Prompt Translation by GPT4
  • L Evaluating Prompt Translation by NLLB
  • M Evaluating Prompt Translation by Google Translate
  • N Evaluating Prompt Quality in mHumanEval
  • O Evaluation Results on mHumanEval

Knowls

  1. Knowl 1 — mHumanEval benchmark scope and variants

    definition

    mHumanEval extends the 164 programming tasks in HumanEval into a multilingual code-generation benchmark. It contains 164 prompt instances for each of 204 natural languages (NLs), for 33,456 translated prompt instances. The translated material is the task description in each HumanEval docstring; the benchmark also provides canonical solutions across 25 programming languages (PLs): Python, Bash, C++, C#, D, Go, Haskell, Java, JavaScript, Julia, Kotlin, Lua, Perl, PHP, R, Racket, Ruby, Rust, Scala, Swift, TypeScript, MATLAB, Visual Basic, Fortran, and COBOL. MATLAB, Visual Basic, Fortran, and COBOL are the four PLs newly added by the benchmark; human experts wrote their canonical solutions and verified that they pass the test cases. Combining the 33,456 NL prompt instances with 25 PLs yields 836,400 NL–PL prompt variants.

    The benchmark includes useful subsets: a 204-prompt mini set with one prompt per NL; a top-quality 500-prompt set, a random 500-prompt set, and a bottom-quality 500-prompt set; and an expert-translated set of 2,460 prompts covering 15 NLs. The 204 NLs are distributed across six digital-resource classes: 38 in class 0, 98 in class 1, 16 in class 2, 27 in class 3, 18 in class 4, and 7 in class 5, with higher classes indicating greater resource availability.

  2. Knowl 2 — Candidate generation and selection for translated prompts

    model/method

    To create each target-language version of a HumanEval task, the authors manually extracted its docstring and generated machine-translation candidates for the 204 target NLs in the Flores-200 set. GPT-4o generated three candidates per prompt, NLLB generated five, and Google Translate generated five where available; Google Translate supported 108 of the target languages, so the workflow produced up to 13 candidates per prompt.

    Each candidate was translated back into English, and BERTScore compared that round-trip text with the original English prompt. CometKiwi supplied a reference-free quality estimate where supported. For languages with CometKiwi scores, the authors averaged the BERTScore and CometKiwi scores and selected the candidate with the highest average. For the 104 languages without CometKiwi coverage, selection used BERTScore alone. Both metrics produce scores in the range [0,1][0,1].

  3. Knowl 3 — Translation quality varies with language resources and candidate selection

    empirical result

    Across language-resource classes, the measured quality of machine-translated coding prompts generally declined as resource availability decreased. Selecting the best candidate from the available translations improved the quality of the final mHumanEval prompts relative to the individual translation-system outputs in the authors’ comparisons using BERTScore and CometKiwi. The improvement reflects selection among multiple candidates, not a claim that every selected prompt is accurate; for 104 languages, the selection process could use only round-trip BERTScore because CometKiwi did not support them.

  4. Knowl 4 — Expert translations have similar measured quality to machine translations

    empirical result

    The mHumanEval-Expert set contains 2,460 prompt instances—164 for each of 15 languages spanning all six resource classes: English, Spanish, French, Japanese, Arabic, and Chinese (class 5); Portuguese, Italian, Korean, and Hindi (class 4); Bangla (class 3); Swahili and Zulu (class 2); Telugu (class 1); and Sinhala (class 0). Native speakers with computer-science or engineering backgrounds translated the prompts, and expert programmers reviewed them for the integrity of the coding task.

    Comparisons between these human-produced prompts and mHumanEval’s machine translations found similar measured quality across the selected languages, with reported variations of ±0.02 in BERTScore and ±0.03 in CometKiwi. Annotators reported no significant terminology concerns. The authors relate the similarity to the general task descriptions in the HumanEval docstrings, which use little specialized coding terminology.

  5. Knowl 5 — Code-generation evaluation protocol and model set

    experimental setup

    The primary code-generation evaluation ran six models on the Python tasks for all 33,456 translated prompts, or 164 prompts in each of 204 NLs: GPT-4o, Claude 3.5 Opus, GPT-3.5, DeepSeek-Coder-V2, WizardCoder, and Aya. The authors also evaluated C++, Java, JavaScript, and Ruby subsets to compare PLs. Model-specific standard prompt templates were used without additional hyperparameter tuning; generation used a maximum of 1,000 tokens and temperature 0.7. Generated code blocks were extracted with regular expressions and executed locally in batches. The reported metric was Pass@1, the fraction of tasks solved by the single generated solution. Proprietary models were accessed through APIs; WizardCoder and Aya ran in full precision on four NVIDIA A100 40 GB GPUs.

  6. Knowl 6 — Performance generally falls as prompt languages become lower-resource

    empirical result

    In the Python evaluation across the 204 NLs, Pass@1 generally declined from higher-resource to lower-resource language classes, but the size of the decline differed substantially by model. Claude 3.5 and GPT-4o retained comparatively strong performance in lower-resource classes; GPT-3.5 and DeepSeek-Coder dropped more sharply, and the decline was especially steep for WizardCoder. Aya had weaker performance in higher-resource classes but comparatively little variation across classes. English was a particularly strong condition: its Pass@1 scores were 0.938 for Claude 3.5, 0.910 for GPT-4o, 0.902 for DeepSeek-Coder, 0.800 for WizardCoder, 0.770 for GPT-3.5, and 0.650 for Aya.

    The authors interpret the contrast between Aya’s relative cross-language consistency and WizardCoder’s substantial non-English decline as evidence that multilingual exposure matters for code generation from multilingual prompts. They also note that DeepSeek-Coder performs well for some mid-resource languages but struggles in low-resource classes; these results suggest, rather than prove, that multilingual training data can support cross-lingual code generation.

  7. Knowl 7 — Pass@1 varies substantially across programming languages

    data/table

    The following values are mean Pass@1 scores across all 204 NLs for six models on five PLs. They show that Python has the highest mean score for every model, while results are notably lower in JavaScript and Ruby; performance also differs markedly among models within the same PL.

    Model Python Java C++ JavaScript Ruby
    GPT-4o 0.738 0.650 0.652 0.477 0.480
    GPT-3.5 0.360 0.270 0.270 0.099 0.103
    Claude 3.5 0.739 0.651 0.649 0.483 0.477
    DeepSeek-Coder 0.229 0.139 0.136 0.000 0.000
    WizardCoder 0.098 0.009 0.007 0.000 0.000
    Aya 0.445 0.355 0.356 0.186 0.183
  8. Knowl 8 — Prompt quality strongly affects measured code-generation performance

    data/table

    An ablation evaluated eight models on four subsets: mHumanEval-mini (204 prompts, one per NL), mHumanEval-T500 (the 500 highest-quality prompts), mHumanEval-R500 (500 randomly selected prompts), and mHumanEval-B500 (the 500 lowest-quality prompts). The top-quality subset draws from class 5 languages, while the bottom-quality subset draws from classes 0 or 1. Every model scored higher on T500 than on B500, showing that prompt quality and language composition can substantially change measured Pass@1. The mini set provides a small preliminary evaluation, but its scores do not consistently predict results on the other subsets.

    Model mini T500 R500 B500
    GPT-4o 0.72 0.87 0.78 0.48
    GPT-3.5 0.44 0.76 0.53 0.21
    Aya 0.47 0.60 0.47 0.42
    WizardCoder 0.12 0.63 0.16 0.00
    Claude 3.5 0.61 0.86 0.59 0.31
    DeepSeek-Coder 0.57 0.73 0.63 0.22
    LLaMA 3 0.35 0.56 0.28 0.11
    CodeStral 0.15 0.36 0.17 0.10
  9. Knowl 9 — Observed failure modes include task misinterpretation and translated code keywords

    empirical result

    The models usually produced code, but the authors observed failures in understanding the requested task, preserving programming-language syntax, and generating meaningful code. For example, GPT-4o responded to a Zulu prompt asking for prime-number detection with code to extract significant digits; the authors attribute this to ambiguity in the translated phrase for prime number. Aya produced Python with Rundi words in place of keywords such as for and return, causing compilation errors. WizardCoder produced nonsensical C code for a Sinhala prompt asking to reverse a list. These examples show that successful code-block generation alone does not ensure that the translated task was understood or that the output is executable.

  10. Knowl 10 — Evaluation is limited in metric and task coverage

    limitation

    The main model comparison covered six LLMs, and the large benchmark made evaluation expensive enough that the authors reported Pass@1 rather than more costly Pass@10 or Pass@100 measures. The benchmark also retains the original HumanEval task set of 164 tasks per NL, so its multilingual breadth does not provide a larger or more diverse underlying collection of coding problems. The additional LLMs in the ablation were evaluated on subsets rather than across the full benchmark.

Coverage note — The appendix-level per-language translation and model score inventories are omitted because they provide detailed breakdowns rather than additional benchmark-wide findings; the subset ablation, benchmark variants, cross-language trends, and cross-PL results are included.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, et al. 2023. Gpt-4 technical report.
  2. 2.Anthropic. 2024. Claude 3: A next-generation ai assistant. https://www.anthropic.com/news/claude-3-family.
  3. 3.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732.
  4. 4.Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. 2022. Systematic inequalities in language technology performance across the world’s languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems.
  6. 6.Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, et al. 2023. Multipl-e: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  8. 8.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116.
  9. 9.Marta R Costa-jussà, James Cross, Onur Çelebi, et al. 2022. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672.
  10. 10.Damai Dai, Chengqi Deng, Chenggang Zhao, et al. 2024. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. arXiv preprint arXiv:2401.06066.
  11. 11.Lingyue Fu, Huacan Chai, Shuang Luo, Kounianhua Du, Weiming Zhang, et al. 2023. Codeapex: A bilingual programming evaluation benchmark for large language models. arXiv preprint arXiv:2309.01940.
  12. 12.Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021. Larger-scale transformers for multilingual masked language modeling. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021).
  13. 13.Nuno M Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André FT Martins. 2024. xcomet: Transparent machine translation evaluation through fine-grained error detection. Transactions of the Association for Computational Linguistics.
  14. 14.Yiyang Hao, Ge Li, Yongqiang Liu, Xiaowei Miao, et al. 2022. Aixbench: A code generation benchmark dataset. arXiv preprint arXiv:2206.13179.
  15. 15.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  16. 16.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018. Mapping language to code in programmatic context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.
  17. 17.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the nlp world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.
  18. 18.R Li, LB Allal, Y Zi, N Muennighoff, D Kocetkov, C Mou, M Marone, C Akiki, J Li, J Chim, et al. 2023. Starcoder: May the source be with you! Transactions on machine learning research.
  19. 19.Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2024. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems.
  20. 20.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, et al. 2023. Wizardcoder: Empowering code large language models with evol-instruct. In The Twelfth International Conference on Learning Representations.
  21. 21.Gabriel Orlanski, Kefan Xiao, Xavier Garcia, et al. 2023. Measuring the impact of programming language distribution. In International Conference on Machine Learning. PMLR.
  22. 22.Qiwei Peng, Yekun Chai, and Xuhong Li. 2024. Humaneval-xl: A multilingual code generation benchmark for cross-lingual natural language generalization. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024).
  23. 23.Nishat Raihan, Dhiman Goswami, Sadiya Sayara Chowdhury Puspo, Christian Newman, Tharindu Ranasinghe, and Marcos Zampieri. 2024. Cseprompts: A benchmark of introductory computer science prompts. arXiv preprint arXiv:2404.02540.
  24. 24.Ricardo Rei, Nuno M Guerreiro, Daan van Stigt, Marcos Treviso, et al. 2023. Scaling up cometkiwi: Unbabel-ist 2023 submission for the quality estimation shared task. In Proceedings of the Eighth Conference on Machine Translation.
  25. 25.Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, et al. 2024. Aya model: An instruction finetuned open-access multilingual language model. arXiv preprint arXiv:2402.07827.
  26. 26.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.

Citation

MLA
Raihan, N., et al. “mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 11432–61, https://doi.org/10.18653/v1/2025.naacl-long.570.
APA
Raihan, N., Anastasopoulos, A., & Zampieri, M. (2025). mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11432–11461. https://doi.org/10.18653/v1/2025.naacl-long.570
Chicago
Raihan, N., A. Anastasopoulos, and M. Zampieri. 2025. “mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11432–61. https://doi.org/10.18653/v1/2025.naacl-long.570.
Harvard
Raihan, N., Anastasopoulos, A. and Zampieri, M. (2025) “mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11432–11461. Available at: https://doi.org/10.18653/v1/2025.naacl-long.570.
Vancouver
1. Raihan N, Anastasopoulos A, Zampieri M (2025) mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 11432–11461

BibTeX

@inproceedings{raihan-etal-2025-mhumaneval,
    title = "m{H}uman{E}val - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation",
    author = "Raihan, Nishat  and
      Anastasopoulos, Antonios  and
      Zampieri, Marcos",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.570/",
    doi = "10.18653/v1/2025.naacl-long.570",
    pages = "11432--11461",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/