Multilingual Code Snippets Training for Program Translation
Ming ZhuKarthik SureshChandan K. Reddy
Introduces a parallel snippet-level dataset across seven programming languages alongside a multilingual pre-training method that significantly improves source-to-source code translation, particularly for low-resource languages.
Modern organizations face substantial expenses and technical risks when adapting software across different platforms or migrating legacy systems to modern programming languages. Automated program translation aims to convert source code from one language to another, replacing manual, labor-intensive rule crafting with neural machine learning models. However, the development of reliable automated models has been heavily constrained by the lack of high-quality, parallel code datasets. Existing benchmarks are mostly limited to two languages or rely on coarse, program-level problem solutions that exhibit high variance in logic, variable names, and code structure.
The article introduces a fine-grained multilingual dataset and demonstrates a novel snippet-based pre-training strategy to improve program translation accuracy across diverse programming languages. Specifically, the authors evaluate whether pre-training on finely aligned code snippets can improve translation performance across 42 language pairs, with a particular focus on low-resource programming languages.
To accomplish this, the authors constructed the Code Snippet Translation dataset, compiling over 132,000 manually verified, aligned code snippets spanning seven popular languages—C, C++, C#, Java, JavaScript, PHP, and Python—across 1,625 programming problems. Using this data, they developed a sequence-to-sequence model termed MuST-PT. The model leverages a three-stage training strategy: initializing with a code-trained base model, performing multilingual denoising auto-encoding to establish a shared latent representation across all seven languages, and executing multilingual snippet translation pre-training before fine-tuning on full programs.
The findings show that the proposed approach outperforms existing baseline models across both snippet-level and full program-level evaluations. First, the model achieved state-of-the-art results on standard public benchmarks, reaching translation accuracy scores of 87.37 on Java-to-C# and 85.25 on C#-to-Java. Second, fine-grained snippet training significantly improved results for low-resource languages, such as PHP and C, where traditional models routinely fail due to data scarcity. Third, while baseline models experienced severe performance degradation when transitioning from short snippets to full-length programs, the proposed method maintained stable accuracy across long sequences. Finally, integrating this snippet-level pre-training into external baseline models yielded consistent, substantial performance gains, proving the generalizability of the training framework.
These results demonstrate that snippet-level alignment effectively addresses the sequence length and data imbalance challenges inherent to automated code translation. By transferring knowledge from resource-rich languages to data-scarce languages, this approach can significantly reduce the costs, timelines, and error rates associated with enterprise code migration. Organizations considering automated code conversion should prioritize fine-grained, snippet-aligned pre-training workflows over coarse program-level methods to enhance reliability. Future initiatives should explore applying these snippet-aligned techniques to adjacent software engineering tasks, including automated code summarization, documentation generation, and text-to-code synthesis.
While the findings provide strong confidence in the efficacy of snippet-level pre-training, stakeholders should note certain limitations. The dataset was collected from curated programming solutions following specific commenting templates, which may not capture all architectural complexities or non-standard coding styles encountered in large enterprise software repositories. Consequently, teams should validate model outputs on their specific domain codebases before broad operational deployment.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). Introduces the standard multilingual code-to-code benchmarks, evaluation protocols, and baselines that the source paper directly builds upon and measures against for program translation.
- Paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, Yue Wang et al. (2021). Presents an encoder-decoder framework and identifier-aware pre-training objectives across multiple programming languages, establishing a foundational architecture for sequence-to-sequence code translation.
- Paper: Multilingual Denoising Pre-training for Neural Machine Translation, Yinhan Liu et al. (2020). Establishes multilingual denoising auto-encoding pre-training for sequence-to-sequence translation, which forms the direct basis for the second stage of the source paper's training pipeline.
- Paper: GraphCodeBERT: Pre-training Code Representations with Data Flow, Daya Guo et al. (2020). Demonstrates multilingual pre-trained code representations evaluated on code-to-code translation, providing key baseline methodology and comparative context.
- Paper: CodeBERT: A Pre-Trained Model for Programming and Natural Languages, Zhangyin Feng et al. (2020). Pioneers bimodal and multilingual pre-training over programming languages, serving as the foundational paradigm for subsequent code-trained base models.
- Paper: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Mike Lewis et al. (2020). Introduces the sequence-to-sequence denoising pre-training objective that underpins the multilingual auto-encoding methodology used to align representations.
- Paper: Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation, Melvin Johnson et al. (2016). Formulates the shared-representation multi-way translation framework that enables knowledge transfer from high-resource to low-resource language pairs.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). Extends program translation and code generation by introducing execution feedback and self-debugging mechanisms to resolve syntax and runtime errors in translated code.
- Paper: CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion, Yangruibo Ding et al. (2023). Scales evaluation beyond snippet- and program-level tasks by establishing a multilingual benchmark focused on cross-file dependencies and repository-level context.
- Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). Expands on multilingual code understanding and generation by scaling foundation models with infilling objectives and extended sequence contexts across diverse languages.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). Advances large-scale multilingual pre-training by utilizing project-level dependency structuring across dozens of programming languages.
- Paper: Large Language Models for Software Engineering: A Systematic Literature Review, Xinying Hou et al. (2023). Provides a comprehensive systematic review synthesizing modern large language model architectures and pre-training methodologies across diverse software engineering tasks.
