Repair Is Nearly Generation: Multilingual Program Repair with LLMs
Harshit JoshiJosé Pablo Cambronero SánchezSumit GulwaniVu LeGust VerbruggenIvan Radicek
Presents RING, a prompt-driven automated program repair framework powered by large language models that fixes last-mile code mistakes across six diverse programming languages and matches or outperforms specialized language-specific repair tools.
Software developers and spreadsheet users frequently encounter small mistakes, such as syntax errors, that require only minor edits to fix. While conceptually simple, these "last-mile" errors disrupt the workflow of experienced programmers and present steep roadblocks for beginners. Traditional automated repair methods require extensive, language-specific engineering or millions of domain-specific training examples, making them costly to adapt to new programming environments.
The article demonstrates and evaluates RING, a unified, multilingual automated program repair engine that uses a general-purpose large language model trained on code (OpenAI Codex). By treating program repair as an extension of code generation, the objective is to show that a single model can correct code errors across multiple distinct languages without requiring separate retraining for each language.
The researchers designed RING around three standard repair steps modeled after human developer workflows: locating the error using compiler or linter diagnostic messages, transforming the code by retrieving relevant example repairs through few-shot prompting, and ranking potential fix candidates using the model's token probability scores. The authors evaluated the system across six languages—Excel formulas, Power Fx, Python, JavaScript, C, and PowerShell—using benchmark datasets of 200 to 273 tasks per language and compared performance against dedicated, language-specific repair tools and baseline zero-shot models.
The evaluation revealed five key findings. First, RING outperformed specialized, state-of-the-art tools on top-ranked (Top@1) repair accuracy in three languages: Excel (82% vs. 71%), Python (94% vs. 92%), and C (63% vs. 55%). Second, for JavaScript, RING achieved a 46% Top@1 success rate on full code snippets, significantly outperforming the specialized baseline, which degraded from 59% on narrow code windows down to 9% on realistic, full-length snippets. Third, dynamically selecting few-shot examples based on diagnostic similarity improved Top@1 repair rates across all tested languages by up to 20% compared to using static, fixed prompt examples. Fourth, when an exact fix was not found, RING still located the fault within a narrow window of the correct edit location more often than the specialized baselines. Fifth, performance was substantially lower in PowerShell (18% Top@1), and repair success across all languages correlated inversely with code length, with longer snippets failing more frequently.
These results show that organizations can deploy a single foundation model to support multi-language developer assistance, eliminating the substantial engineering overhead and data-collection costs needed to maintain separate, bespoke repair tools. This capability enables a "flipped" interaction model where users write code naturally and an artificial intelligence assistant automatically suggests real-time fixes for last-mile errors. Furthermore, the findings indicate that a model's underlying training data representation directly dictates repair effectiveness across different programming ecosystems.
Organizations developing developer tooling should consider adopting prompt-engineered foundation models for code repair, particularly by pairing compiler diagnostic messages with dynamic example retrieval banks. To maximize performance, teams should implement specialized ranking mechanisms and explore iterative querying techniques for complex, multi-error scenarios. Before deploying such systems to low-resource languages like PowerShell, organizations should conduct targeted pilots and expand language-specific example banks to address lower baseline accuracy.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This seminal paper introduces OpenAI Codex, the foundational code-generation model and execution-based evaluation paradigm directly adapted by RING for multilingual program repair.
- Paper: Fault-Aware Neural Code Rankers, Jeevana Priya Inala et al. (2022). It provides the foundational framework for neural ranking and fault-aware candidate selection without full execution, directly informing RING's token probability ranking stage.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). This benchmark establishes standardized multilingual dataset frameworks for code understanding and generation across diverse languages.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). It provides key foundational insights into few-shot prompting and scaling behaviors for synthesizing code from language models.
- Paper: Memory-assisted prompt editing to improve GPT-3 after deployment, Aman Madaan et al. (2022). It demonstrates how external memory and dynamic prompt retrieval can steer LLMs to correct errors without model fine-tuning.
- Paper: Natural Language to Code Translation with Execution, Freda Shi et al. (2022). It details how prompting Codex and leveraging execution-guided consensus improves code candidate selection.
- Paper: GraphCodeBERT: Pre-training Code Representations with Data Flow, Daya Guo et al. (2020). It establishes early pre-trained multilingual representations and code refinement techniques across multiple programming languages.
- Paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, Yue Wang et al. (2021). It introduces identifier-aware multilingual pre-training benchmarks that set prior baselines for automated code repair and defect detection.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). This work extends LLM-based program repair by showing how models can iteratively debug their own generated code using execution and compiler feedback in few-shot loops.
- Paper: UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging, Cheryl Lee et al. (2025). It builds upon prompt-based repair by formalizing hierarchical multi-agent pipelines to handle complex, multi-lingual, and repository-scale bugs.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). It scales the evaluation of automated code repair and generation from small snippets to end-to-end, real-world GitHub issues.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). It advances beyond prompt-level program repair by creating specialized agent-computer interfaces with syntax guardrails to automate software debugging.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). It extends code generation and repair evaluation by benchmarking LLM code capabilities on time-segmented, contamination-free programming problems.
- Paper: CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion, Yangruibo Ding et al. (2023). It extends single-snippet code generation and repair across multiple programming languages to full project repositories requiring cross-file context.
- Paper: Large Language Models for Software Engineering: A Systematic Literature Review, Xinying Hou et al. (2023). It provides a comprehensive literature review contextualizing LLM-driven program repair and code synthesis within broader software engineering lifecycles.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). It generalizes multilingual code intelligence by training an open-source foundation model across 87 programming languages with infilling and repository-level support.
