Measuring the Impact of Programming Language Distribution

Gabriel OrlanskiKefan XiaoXavier GarciaJeffrey HuiJoshua HowlandJonathan MalmaudJacob AustinRishabh SinghMichele Catasta

article2023ICML55 citations

Presents the BabelCode execution-based evaluation framework and demonstrates that balancing multilingual training distributions dramatically improves code language model performance on low-resource programming languages with minimal degradation to high-resource ones.

Listen

Modern neural code models deliver strong performance when generating and translating code in widely used programming languages, but their capabilities drop sharply on low-resource languages such as Rust, Julia, and Haskell. This disparity limits the real-world value of artificial intelligence coding tools for developers using modern or niche languages. Furthermore, existing evaluation benchmarks are predominantly restricted to a few high-resource languages, making multilingual assessment challenging and expensive.

The article addresses these challenges by pursuing two primary objectives: establishing an execution-based evaluation framework to test code models across multiple programming languages and investigating whether balancing the distribution of training data improves model performance on low-resource languages without severely degrading high-resource performance.

To evaluate models consistently, the authors introduced BabelCode, an open-source framework supporting execution-based evaluation across 14 programming languages and four benchmark datasets, including a newly curated dataset called Translating Python Programming Puzzles (TP3). The authors pre-trained decoder-only transformer models in three sizes (1 billion, 2 billion, and 4 billion parameters) across different data distributions: the naturally occurring GitHub distribution and balanced distributions created using the Unimax algorithm, which caps per-language duplication to prevent overfitting. Performance was assessed on zero-shot generation and translation tasks using pass@k metrics.

The findings demonstrate that training on balanced language distributions significantly enhances capabilities in underrepresented languages. Across all evaluated tasks and languages, models trained on balanced data achieved an average 12.34% improvement in pass rate over natural-distribution baselines. For low-resource languages specifically, performance increased by 66.48% (with code generation improving by up to 111.85% on smaller models), accompanied by a modest 12.94% performance drop on high-resource languages. Crucially, scaling the model size mitigated these high-resource penalties: while the 1-billion-parameter model experienced a 39.70% drop on high-resource languages, the 4-billion-parameter model reduced that loss to just 2.47%. Fine-grained execution analysis showed that balancing data primarily improved underlying functional correctness rather than merely reducing compilation or syntax errors.

These results carry significant strategic implications for development teams and organizations deploying enterprise code models. Balancing pre-training data presents a cost-effective method to broaden multilingual language support, reducing software development risks and widening developer adoption across diverse technology stacks. Unlike natural-language-to-code generation, pure code translation benefited uniformly from balanced training data, demonstrating that multilingual representation does not require excessive oversampling of popular languages.

Organizations developing or fine-tuning code models should adopt bounded data-balancing strategies rather than relying solely on raw natural distributions. When deploying balanced corpora, practitioners should favor larger model architectures to absorb data rebalancing without sacrificing performance in core enterprise languages like Java, Python, or C++. Future work should expand data-balancing investigations to larger models exceeding 4 billion parameters, develop more sophisticated sampling algorithms, and extend execution-based benchmarks to complex, user-defined data structures.

Confidence in these findings is supported by thorough unit and integration testing across 14 languages and multiple model scales. However, readers should consider existing limitations: the maximum model size evaluated was 4 billion parameters, benchmark problems primarily involved standalone algorithmic tasks, and heavy duplication of low-resource data yielded diminishing returns, confirming that models ultimately require access to novel, high-quality code samples to achieve deeper semantic mastery.

Cover for Measuring the Impact of Programming Language Distribution

Abstract

Current benchmarks for evaluating neural code models focus on only a small subset of programming languages, excluding many popular languages such as Go or Rust. To ameliorate this issue, we present the BabelCode framework for execution-based evaluation of any benchmark in any language. BabelCode enables new investigations into the qualitative performance of models’ memory, runtime, and individual test case results. Additionally, we present a new code translation dataset called Translating Python Programming Puzzles (TP3) from the Python Programming Puzzles (Schuster et al., 2021) benchmark that involves translating expert-level python functions to any language. With both BabelCode and the TP3 benchmark, we investigate if balancing the distributions of 14 languages in a training dataset improves a large language model’s performance on low-resource languages. Training a model on a balanced corpus results in, on average, 12.34% higher pass@k across all tasks and languages compared to the baseline. We find that this strategy achieves 66.48% better pass@k on low-resource languages at the cost of only a 12.94% decrease to high-resource languages. In our three translation tasks, this strategy yields, on average, 30.77% better low-resource pass@k while having 19.58% worse high-resource pass@k.

Table of Contents

  • 1. Introduction
  • 2. The BabelCode Framework
  • 2.1. Framework Design
  • 2.2. Differences To Prior Works
  • 3. Low-Resource Code Language Models
  • 4. Experimental Setup
  • 4.1. Models
  • 4.2. Training Data
  • 4.3. Vocabulary
  • 4.4. Benchmarks
  • 4.5. Evaluation
  • 5. Results
  • 5.1. Baseline Models
  • 5.2. Impact of Balancing Programming Languages
  • 5.3. Effects Of Language Balance on Predictions
  • 6. Related Works
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. BabelCode Design
  • B. Dataset Changes
  • B.1. Incompatible Problems
  • B.2. Changes To HumanEval
  • B.3. Changes To Transcoder
  • B.4. TP3 Examples
  • C. Training Languages
  • D. Training Objective
  • E. Prompts Used
  • E.1. Generation Tasks
  • E.2. Translation Tasks
  • F. Full Results

Knowls

  1. Knowl 1 — BabelCode provides language-agnostic execution-based code evaluation

    model/method

    BabelCode evaluates code-generation or code-translation problems by turning each problem’s input/output tests into a target-language testing program and executing that program. Its four-stage pipeline represents test types in a language-independent domain-specific language (DSL), translates test values into the target language, renders a test script from a Jinja2 template, and compiles or runs the script through the command line. The framework supports adding benchmarks and target languages without requiring the original benchmark to be written in Python. It also reports execution-level outcomes, including the results of individual tests, runtime, and memory use.

  2. Knowl 2 — BabelCode uses typed test schemas and language-specific equality checks

    model/method

    BabelCode’s DSL records the types of a problem’s inputs and outputs rather than their literal values; for example, nested integer lists inside a string-keyed map can be represented as map<string;list<integer>>. Literal test values are translated separately into the target language, allowing datasets originating in different languages to use the same schema. BabelCode maps schema types to native types and conventions for each supported language. It checks floating-point results with tolerances of 1e-6 for floats and 1e-9 for doubles, inferring float versus double from the number of decimal digits; it similarly distinguishes integers and long integers by digit count. Where a language lacks deep value equality, BabelCode can serialize structures to JSON and compare the resulting strings; otherwise it uses the language’s built-in deep equality.

  3. Knowl 3 — BabelCode isolates and reports every test-case outcome

    model/method

    Rather than stopping at the first failed assertion, BabelCode wraps test cases so that each one can run independently and prints a parseable status for each test to standard output. This permits evaluation to distinguish programs that compile or run but fail tests from programs that pass some or all tests, and identifies which individual cases fail. Its generated tests and validation suite check that language implementations are syntactically valid and that their equality behavior works when executed. Prompt translation also adapts language names, function signatures, reserved identifiers, and language-specific reserved characters; headers containing imports are excluded from translated prompts.

  4. Knowl 4 — Unimax balances language data under a cap on example duplication

    algorithm

    Unimax constructs a training distribution from language-specific data buckets given a training-example budget and a maximum duplication count NN. The cap limits how many epochs any example can contribute, which is intended to increase exposure to low-resource languages without repeatedly oversampling them without bound. The procedure groups examples by programming language, adds up to NN epochs from the lowest-resource language buckets, and continues bringing in the lowest-resource buckets until the remaining budget can be allocated across the other languages without any of them exceeding the same cap. The remaining budget is then distributed across those languages subject to that constraint. The experiments compare the natural distribution with Unimax distributions using N∈{1,2,3,4}N \in \{1,2,3,4\}.

  5. Knowl 5 — TP3 turns expert-written Python puzzle checks into a translation benchmark

    definition

    Translating Python Programming Puzzles (TP3) is a code-translation dataset built from 370 verification functions in the Python Programming Puzzles benchmark. Each function checks whether a proposed answer satisfies a puzzle’s constraints; the functions were hand-written by expert Python programmers and range from simple character checks to competitive-programming-style problems. TP3 asks a model to translate these Python functions into another programming language, providing a translation task whose source programs are more complex than short, conventional code-translation examples.

  6. Knowl 6 — Training and evaluation compare natural and balanced distributions across languages

    experimental setup

    The experiments train decoder-only models with 1B, 2B, and 4B parameters on a curated GitHub source-code corpus in 14 languages. The seven high-resource languages are Java, Python, C++, PHP, TypeScript, JavaScript, and Go; the seven low-resource languages are Dart, Lua, Rust, C#, R, Julia, and Haskell. Java comprises 36.95% of the postprocessed examples, Python 16.80%, C++ 16.68%, and PHP 14.05%, while Haskell comprises 0.02% and Julia 0.03%. Each model size is trained on the natural distribution and on Unimax distributions with N=1,2,3,4N=1,2,3,4. Model context length is 2048, batch size is 256, and training uses Adafactor, a 64K-token SentencePiece vocabulary, and the UL2 objective with an additional causal language-modeling objective. The objective mixture assigns 10% each to two span-corruption settings, 20% to prefix language modeling, and 60% to causal language modeling. Training lasts 38,000, 77,000, and 190,000 steps for the three model sizes, respectively. Evaluation uses BC-HumanEval, BC-MBPP, BC-Transcoder, and TP3; generation samples 200 programs per BC-HumanEval problem, while translation samples 50 per problem. Sampling uses temperature 0.8 and top-p 0.95, with pass@100 for generation and pass@25 for translation.

  7. Knowl 7 — Balanced training improves low-resource results in the aggregate

    empirical result

    Across the evaluated tasks and languages, the paper reports that training on a balanced language distribution yields an average 12.34% higher pass@k than training on the natural distribution. The reported low-resource-language pass@k improvement is 66.48%, accompanied by a 12.94% decrease on high-resource languages. Across the three translation tasks, the reported average change is a 30.77% improvement on low-resource languages and a 19.58% decrease on high-resource languages. These are aggregate results; the size and direction of the effect vary by task, model size, and Unimax cap.

  8. Knowl 8 — On BC-HumanEval, balancing benefits low-resource languages most at smaller model sizes

    empirical result

    On BC-HumanEval, Unimax-trained models improve average low-resource pass@100 relative to natural-distribution models by 111.85% for 1B models, 68.38% for 2B models, and 19.22% for 4B models. The corresponding average high-resource performance losses are 15.47%, 14.00%, and 9.35%. Among the tested Unimax settings, N=3N=3 gives the best reported trade-off between low-resource gains and high-resource losses: the difference between those two changes is 130.17%, 87.80%, and 36.00% for the 1B, 2B, and 4B models. Thus, the low-resource advantage diminishes as model size grows, while the high-resource cost also becomes smaller.

  9. Knowl 9 — TP3 translation gains depend on both the balance cap and model size

    empirical result

    For TP3, Unimax training improves average low-resource pass@25 over natural-distribution training by 124.45% for 1B models, 64.51% for 2B models, and 51.29% for 4B models. The 4B model trained with Unimax N=2N=2 has a reported 71.59% average low-resource improvement and a 20.31% high-resource improvement over its natural-distribution counterpart. Results vary substantially among the Unimax caps, and reducing the amount of Python training data can hinder translation because the model must understand the Python source function as well as produce target-language code.

  10. Knowl 10 — Transcoder balancing effects differ by source language

    empirical result

    On BC-Transcoder with C++ as the source language, Unimax training produces average low-resource pass@25 improvements of 7.57%, 6.76%, and 11.80% for 1B, 2B, and 4B models, respectively; for the 4B model, N=2N=2 is the best setting reported for low-resource performance, with a 20.47% improvement over natural-distribution training. With Python as the source language, average low-resource changes across Unimax settings are -26.04% for 1B, +15.1% for 2B, and +22.47% for 4B models. The source-language dependence shows that balancing does not uniformly help every translation setup, particularly for the smallest models.

  11. Knowl 11 — Test-level results distinguish generation failures from partial correctness

    empirical result

    In BC-HumanEval, balancing causes fewer test cases to pass on average for high-resource languages: the reported reductions relative to natural-distribution training are 5.50% for Unimax N=1N=1 and 9.09% for N=2N=2. The change is not chiefly explained by more compilation or runtime errors: the reported mean error increases are only 0.40 and 1.15, respectively. For low-resource languages, the natural-distribution model passes an average of 5.13% of test cases per problem, compared with 9.53% for N=1N=1 and 10.48% for N=2N=2. On TP3, mean test-case pass rates improve for both resource groups: by 2.58% and 3.06% for high-resource languages and by 3.40% and 4.99% for low-resource languages under N=1N=1 and N=2N=2, respectively. These per-test metrics reveal changes in partial correctness that an all-tests-passed score alone would not show.

Coverage note — The per-language appendix result matrices and individual benchmark-conversion examples are not separate knowls; their task-level findings and the reusable evaluation design capture the substantive conclusions without reproducing extensive tabular detail or worked examples.

References

  1. 1.Ahmad, W., Chakraborty, S., Ray, B., and Chang, K.-W. Unified pre-training for program understanding and generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2655–2668, Online, June 2021. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2021.naacl-main.211.
  2. 2.Allal, L. B., Li, R., Kocetkov, D., Mou, C., Akiki, C., Ferrandis, C. M., Muennighoff, N., Mishra, M., Gu, A., Dey, M., et al. Santacoder: don’t reach for the stars! arXiv preprint arXiv:2301.03988, 2023.
  3. 3.Allamanis, M. The adverse effects of code duplication in machine learning models of code. In Proceedings of the 2019 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software, pp. 143–153, 2019.
  4. 4.Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M. X., Cao, Y., Foster, G., Cherry, C., et al. Massively multilingual neural machine translation in the wild: Findings and challenges. arXiv preprint arXiv:1907.05019, 2019.
  5. 5.Athiwaratkun, B., Gouda, S. K., Wang, Z., Li, X., Tian, Y., Tan, M., Ahmad, W. U., Wang, S., Sun, Q., Shang, M., Gonugondla, S. K., Ding, H., Kumar, V., Fulton, N., Farahani, A., Jain, S., Giaquinto, R., Qian, H., Ramanathan, M. K., Nallapati, R., Ray, B., Bhatia, P., Sengupta, S., Roth, D., and Xiang, B. Multi-lingual evaluation of code generation models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Bo7eeXm6An8.
  6. 6.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  7. 7.Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255, 2022.
  8. 8.Cassano, F., Gouwar, J., Nguyen, D., Nguyen, S., Phipps-Costin, L., Pinckney, D., Yee, M. H., Zi, Y., Anderson, C. J., Feldman, M. Q., et al. A scalable and extensible approach to benchmarking nl2code for 18 programming languages. arXiv preprint arXiv:2208.08227, 2022.
  9. 9.Chakraborty, S., Ahmed, T., Ding, Y., Devanbu, P., and Ray, B. Natgen: generative pre-training by “naturalizing” source code. Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2022.
  10. 10.Chen, M., Tworek, J., Jun, H., Yuan, Q., Ponde, H., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D. W., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Babuschkin, I., Balaji, S. A., Jain, S., Carr, A., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M. M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. ArXiv, abs/2107.03374, 2021.
  11. 11.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  12. 12.Christopoulou, F., Lampouras, G., Gritta, M., Zhang, G., Guo, Y., Li, Z.-Y., Zhang, Q., Xiao, M., Shen, B., Li, L., Yu, H., yu Yan, L., Zhou, P., Wang, X., Ma, Y., Iacobacci, I., Wang, Y., Liang, G., Wei, J., Jiang, X., Wang, Q., and Liu, Q. Pangu-coder: Program synthesis with function-level language modeling. ArXiv, abs/2207.11280, 2022.
  13. 13.Chung, H. W., Garcia, X., Roberts, A., Tay, Y., Firat, O., Narang, S., and Constant, N. Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=kXwdL1cWOAi.
  14. 14.Clement, C., Drain, D., Timcheck, J., Svyatkovskiy, A., and Sundaresan, N. PyMT5: multi-mode translation of natural language and python code with transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9052–9065, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.728. URL https://aclanthology.org/2020.emnlp-main.728.
  15. 15.Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzman, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. Unsupervised cross-lingual representation learning at scale. In Annual Meeting of the Association for Computational Linguistics, 2019.
  16. 16.Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., and Zhou, M. CodeBERT: A pre-trained model for programming and natural languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 1536–1547, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.139. URL https://aclanthology.org/2020.findings-emnlp.139.
  17. 17.Fried, D., Aghajanyan, A., Lin, J., Wang, S. I., Wallace, E., Shi, F., Zhong, R., tau Yih, W., Zettlemoyer, L., and Lewis, M. Incoder: A generative model for code infilling and synthesis. ArXiv, abs/2204.05999, 2022.
  18. 18.Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J. Measuring coding challenge competence with APPS. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=sD93GOzH3i5.
  19. 19.Husain, H., Wu, H., Gazit, T., Allamanis, M., and Brockschmidt, M. Codesearchnet challenge: Evaluating the state of semantic code search. ArXiv, abs/1909.09436, 2019.
  20. 20.Kocetkov, D., Li, R., Allal, L. B., Li, J., Mou, C., Ferrandis, C. M., Jernite, Y., Mitchell, M., Hughes, S., Wolf, T., et al. The stack: 3 tb of permissively licensed source code. arXiv preprint arXiv:2211.15533, 2022.
  21. 21.Kudo, T. and Richardson, J. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226, 2018.
  22. 22.Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., Yih, S., Fried, D., yi Wang, S., and Yu, T. Ds-1000: A natural and reliable benchmark for data science code generation. ArXiv, abs/2211.11501, 2022.
  23. 23.Li, Y., Choi, D. H., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Tom, Eccles, Keeling, J., Gimeno, F., Lago, A. D., Hubert, T., Choy, P., de, C., d’Autume, M., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Alexey, Cherepanov, Molloy, J., Mankowitz, D. J., Robson, E. S., Kohli, P., de, N., Freitas, Kavukcuoglu, K., and Vinyals, O. Competition-level code generation with alphacode. Science, 378:1092 – 1097, 2022.
  24. 24.Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C. B., Drain, D., Jiang, D., Tang, D., Li, G., Zhou, L., Shou, L., Zhou, L., Tufano, M., Gong, M., Zhou, M., Duan, N., Sundaresan, N., Deng, S. K., Fu, S., and Liu, S. Codexglue: A machine learning benchmark dataset for code understanding and generation. ArXiv, abs/2102.04664, 2021.
  25. 25.Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C. A conversational paradigm for program synthesis. arXiv preprint arXiv:2203.13474, 2022.
  26. 26.Orlanski, G. and Gittens, A. Reading stackoverflow encourages cheating: Adding question text improves extractive code generation. ArXiv, abs/2106.04447, 2021.
  27. 27.Orlanski, G., Yang, S., and Healy, M. Evaluating how fine-tuning on bimodal data effects code generation. ArXiv, abs/2211.07842, 2022.
  28. 28.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020.
  29. 29.Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A. Scaling up models and data with t5x and seqio. arXiv preprint arXiv:2203.17189, 2022. URL https://arxiv.org/abs/2203.17189.
  30. 30.Roziere, B., Lachaux, M.-A., Chanussot, L., and Lample, G. Unsupervised translation of programming languages. Advances in Neural Information Processing Systems, 33, 2020.
  31. 31.Roziere, B., Lachaux, M.-A., Szafraniec, M., and Lample, G. Dobf: A deobfuscation pre-training objective for programming languages. In Neural Information Processing Systems, 2021.
  32. 32.Schuster, T., Kalyan, A., Polozov, A., and Kalai, A. T. Programming puzzles. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021. URL https://openreview.net/forum?id=fe_hCc4RBrg.
  33. 33.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pp. 4596–4604. PMLR, 2018.
  34. 34.Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Bahri, D., Schuster, T., Zheng, H. S., Houlsby, N., and Metzler, D. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131, 2022.
  35. 35.Wang, S., Li, Z., Qian, H., Yang, C., Wang, Z., Shang, M., Kumar, V., Tan, S., Ray, B., Bhatia, P., Nallapati, R., Ramanathan, M. K., Roth, D., and Xiang, B. Recode: Robustness evaluation of code generation models. 2022a.
  36. 36.Wang, X., Tsvetkov, Y., and Neubig, G. Balancing training for multilingual neural machine translation. arXiv preprint arXiv:2004.06748, 2020.
  37. 37.Wang, Y., Wang, W., Joty, S., and Hoi, S. C. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 8696–8708, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.685. URL https://aclanthology.org/2021.emnlp-main.685.
  38. 38.Wang, Z., Cuenca, G., Zhou, S., Xu, F. F., and Neubig, G. Mconala: A benchmark for code generation from multiple natural languages. ArXiv, abs/2203.08388, 2022b.
  39. 39.Yasunaga, M. and Liang, P. Break-it-fix-it: Unsupervised learning for program repair. In International Conference on Machine Learning (ICML), 2021.
  40. 40.Yin, P., Deng, B., Chen, E., Vasilescu, B., and Neubig, G. Learning to mine aligned code and natural language pairs from stack overflow. 2018 IEEE/ACM 15th International Conference on Mining Software Repositories (MSR), pp. 476–486, 2018.

Citation

MLA
Orlanski, G., et al. “Measuring the Impact of Programming Language Distribution”. International Conference on Machine Learning, vol. 202, 2023, pp. 26619–45, https://proceedings.mlr.press/v202/orlanski23a.html.
APA
Orlanski, G., Xiao, K., Garcia, X., Hui, J., Howland, J., Malmaud, J., Austin, J., Singh, R., & Catasta, M. (2023). Measuring the Impact of Programming Language Distribution. International Conference on Machine Learning, 202, 26619–26645. https://proceedings.mlr.press/v202/orlanski23a.html
Chicago
Orlanski, G., K. Xiao, X. Garcia, et al. 2023. “Measuring the Impact of Programming Language Distribution”. International Conference on Machine Learning 202: 26619–45. https://proceedings.mlr.press/v202/orlanski23a.html.
Harvard
Orlanski, G. et al. (2023) “Measuring the Impact of Programming Language Distribution”, International Conference on Machine Learning. PMLR, pp. 26619–26645. Available at: https://proceedings.mlr.press/v202/orlanski23a.html.
Vancouver
1. Orlanski G, Xiao K, Garcia X, Hui J, Howland J, Malmaud J, Austin J, Singh R, Catasta M (2023) Measuring the Impact of Programming Language Distribution. In: International Conference on Machine Learning. PMLR, pp 26619–26645

BibTeX

@InProceedings{pmlr-v202-orlanski23a,
  title = 	 {Measuring the Impact of Programming Language Distribution},
  author =       {Orlanski, Gabriel and Xiao, Kefan and Garcia, Xavier and Hui, Jeffrey and Howland, Joshua and Malmaud, Jonathan and Austin, Jacob and Singh, Rishabh and Catasta, Michele},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {26619--26645},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/orlanski23a/orlanski23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/orlanski23a.html},
  abstract = 	 {Current benchmarks for evaluating neural code models focus on only a small subset of programming languages, excluding many popular languages such as Go or Rust. To ameliorate this issue, we present the BabelCode framework for execution-based evaluation of any benchmark in any language. BabelCode enables new investigations into the qualitative performance of models’ memory, runtime, and individual test case results. Additionally, we present a new code translation dataset called Translating Python Programming Puzzles (TP3) from the Python Programming Puzzles (Schuster et al., 2021) benchmark that involves translating expert-level python functions to any language. With both BabelCode and the TP3 benchmark, we investigate if balancing the distributions of 14 languages in a training dataset improves a large language model’s performance on low-resource languages. Training a model on a balanced corpus results in, on average, 12.34% higher $pass@k$ across all tasks and languages compared to the baseline. We find that this strategy achieves 66.48% better $pass@k$ on low-resource languages at the cost of only a 12.94% decrease to high-resource languages. In our three translation tasks, this strategy yields, on average, 30.77% better low-resource $pass@k$ while having 19.58% worse high-resource $pass@k$.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/