Measuring the Impact of Programming Language Distribution
Gabriel OrlanskiKefan XiaoXavier GarciaJeffrey HuiJoshua HowlandJonathan MalmaudJacob AustinRishabh SinghMichele Catasta
Presents the BabelCode execution-based evaluation framework and demonstrates that balancing multilingual training distributions dramatically improves code language model performance on low-resource programming languages with minimal degradation to high-resource ones.
Modern neural code models deliver strong performance when generating and translating code in widely used programming languages, but their capabilities drop sharply on low-resource languages such as Rust, Julia, and Haskell. This disparity limits the real-world value of artificial intelligence coding tools for developers using modern or niche languages. Furthermore, existing evaluation benchmarks are predominantly restricted to a few high-resource languages, making multilingual assessment challenging and expensive.
The article addresses these challenges by pursuing two primary objectives: establishing an execution-based evaluation framework to test code models across multiple programming languages and investigating whether balancing the distribution of training data improves model performance on low-resource languages without severely degrading high-resource performance.
To evaluate models consistently, the authors introduced BabelCode, an open-source framework supporting execution-based evaluation across 14 programming languages and four benchmark datasets, including a newly curated dataset called Translating Python Programming Puzzles (TP3). The authors pre-trained decoder-only transformer models in three sizes (1 billion, 2 billion, and 4 billion parameters) across different data distributions: the naturally occurring GitHub distribution and balanced distributions created using the Unimax algorithm, which caps per-language duplication to prevent overfitting. Performance was assessed on zero-shot generation and translation tasks using pass@k metrics.
The findings demonstrate that training on balanced language distributions significantly enhances capabilities in underrepresented languages. Across all evaluated tasks and languages, models trained on balanced data achieved an average 12.34% improvement in pass rate over natural-distribution baselines. For low-resource languages specifically, performance increased by 66.48% (with code generation improving by up to 111.85% on smaller models), accompanied by a modest 12.94% performance drop on high-resource languages. Crucially, scaling the model size mitigated these high-resource penalties: while the 1-billion-parameter model experienced a 39.70% drop on high-resource languages, the 4-billion-parameter model reduced that loss to just 2.47%. Fine-grained execution analysis showed that balancing data primarily improved underlying functional correctness rather than merely reducing compilation or syntax errors.
These results carry significant strategic implications for development teams and organizations deploying enterprise code models. Balancing pre-training data presents a cost-effective method to broaden multilingual language support, reducing software development risks and widening developer adoption across diverse technology stacks. Unlike natural-language-to-code generation, pure code translation benefited uniformly from balanced training data, demonstrating that multilingual representation does not require excessive oversampling of popular languages.
Organizations developing or fine-tuning code models should adopt bounded data-balancing strategies rather than relying solely on raw natural distributions. When deploying balanced corpora, practitioners should favor larger model architectures to absorb data rebalancing without sacrificing performance in core enterprise languages like Java, Python, or C++. Future work should expand data-balancing investigations to larger models exceeding 4 billion parameters, develop more sophisticated sampling algorithms, and extend execution-based benchmarks to complex, user-defined data structures.
Confidence in these findings is supported by thorough unit and integration testing across 14 languages and multiple model scales. However, readers should consider existing limitations: the maximum model size evaluated was 4 billion parameters, benchmark problems primarily involved standalone algorithmic tasks, and heavy duplication of low-resource data yielded diminishing returns, confirming that models ultimately require access to novel, high-quality code samples to achieve deeper semantic mastery.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). Its Codex study established execution-checked code-generation evaluation and the pass@k metric that underpin the source’s assessment of model performance.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Its execution-based MBPP and MathQA-Python evaluations provide the program-synthesis benchmark foundations needed to understand the source’s generation results.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). CodeXGLUE’s multilingual code-task benchmark establishes the cross-language evaluation context that motivates the source’s broader execution-based framework.
- Paper: Multilingual Code Snippets Training for Program Translation, Ming Zhu et al. (2022). Its multilingual snippet-pretraining experiments make the low-resource program-translation problem concrete before the source tests language balancing across translation tasks.
- Paper: CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation, Weixiang Yan et al. (2024). CodeScope extends the source’s multilingual execution-evaluation approach to more languages, task types, and practical dimensions, putting BabelCode’s framework in a broader assessment landscape.
