mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation
Md. Nishat RaihanAntonios AnastasopoulosMarcos Zampieri
Presents mHumanEval, the first massively multilingual code generation benchmark spanning 204 natural languages and 25 programming languages, enabling researchers to rigorously assess how well code LLMs understand prompts across high-, mid-, and low-resource languages.
Artificial intelligence models that generate computer software code have rapidly advanced, but standard evaluation benchmarks rely almost exclusively on English prompts and Python solutions. This narrow scope creates significant blind spots for global software development and risks overestimating model capabilities. The article introduces mHumanEval, an expanded multilingual code generation benchmark, to assess how well leading artificial intelligence systems understand coding instructions written across diverse high-, mid-, and low-resource human languages.
To build mHumanEval, the authors translated 164 standard coding tasks into 204 natural languages using multiple machine translation engines, generating 13 candidates per prompt and selecting the highest-quality translations via automated linguistic scoring. In total, the benchmark comprises 33,456 natural language prompts spanning 25 programming languages, including four newly added legacy and scientific languages: MATLAB, Visual Basic, Fortran, and COBOL. To validate translation quality, expert human programmers produced verified reference translations across 15 diverse languages. The authors then evaluated six prominent proprietary and open-source models using the first-attempt functional accuracy metric, Pass@1.
Across all models, coding accuracy declines as instructions move from high-resource languages like English and Spanish to lower-resource languages. Proprietary frontier models, specifically GPT-4o and Claude 3.5, demonstrate the highest overall resilience, maintaining roughly 60% to 74% accuracy even on low-resource language prompts in Python. In contrast, specialized code-focused models like WizardCoder and DeepSeek-Coder experience severe performance collapses when moving away from English or Chinese, with accuracy falling from over 80–90% down to near 0% on low-resource languages. Furthermore, general multilingual models without specialized coding fine-tuning, such as Aya, achieve stable but moderate accuracy (around 35–45%) across language tiers. Scripting languages like JavaScript and Ruby proved notably more difficult across all models than Python, C++, and Java. Qualitative error analysis reveals that failures in non-English contexts stem primarily from models misinterpreting translated programming concepts or erroneously translating reserved programming keywords directly into the generated code.
These findings demonstrate that training models exclusively on source code and English documentation severely limits their global utility. Multilingual natural language understanding is essential for code generation systems to serve non-English-speaking engineers effectively. Deploying specialized coding assistants in multilingual environments without cross-lingual validation poses severe risks of code failure and operational disruption, whereas advanced general-purpose models demonstrate better adaptability.
Organizations developing or deploying automated coding tools should prioritize multilingual pretraining data and avoid assuming that English benchmark performance reflects global readiness. Practitioners evaluating models should utilize diverse subsets like mHumanEval-mini for rapid multilingual screening. Additionally, because models occasionally produce erroneous or non-compiling code when prompted in non-English languages, organizations must enforce automated execution safeguards, such as isolated virtual execution environments, to prevent runaway loops or system crashes during automated testing.
Confidence in these findings is high for functional code accuracy across the tested models, though evaluations were constrained to the Pass@1 metric due to the significant computational cost of processing hundreds of thousands of prompt variations. Future work should expand human validation beyond the initial 15 languages, evaluate additional model architectures, and test multiple-attempt metrics across larger problem sets.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). Read this first to understand HumanEval’s original 164-task design and the pass@k evaluation framework that mHumanEval translates and uses.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). Its HumanEval+ analysis shows how weak test coverage can distort functional-accuracy scores, a key consideration when interpreting mHumanEval’s Pass@1 results.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). MEGA establishes the multilingual evaluation context for mHumanEval’s central question of how language-resource levels affect model performance.
- Paper: CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation, Weixiang Yan et al. (2024). CodeScope provides an earlier execution-based, multilingual code-generation benchmark that clarifies the evaluation landscape mHumanEval extends.
- Paper: Measuring the Impact of Programming Language Distribution, Gabriel Orlanski et al. (2023). Its findings on training-data imbalance and performance in underrepresented programming languages help frame mHumanEval’s separate comparison of natural-language prompt resource levels.
No sufficiently relevant recommendations were found.
