Repair Is Nearly Generation: Multilingual Program Repair with LLMs

Harshit JoshiJosé Pablo Cambronero SánchezSumit GulwaniVu LeGust VerbruggenIvan Radicek

article2023AAAI206 citations

Presents RING, a prompt-driven automated program repair framework powered by large language models that fixes last-mile code mistakes across six diverse programming languages and matches or outperforms specialized language-specific repair tools.

Listen

Software developers and spreadsheet users frequently encounter small mistakes, such as syntax errors, that require only minor edits to fix. While conceptually simple, these "last-mile" errors disrupt the workflow of experienced programmers and present steep roadblocks for beginners. Traditional automated repair methods require extensive, language-specific engineering or millions of domain-specific training examples, making them costly to adapt to new programming environments.

The article demonstrates and evaluates RING, a unified, multilingual automated program repair engine that uses a general-purpose large language model trained on code (OpenAI Codex). By treating program repair as an extension of code generation, the objective is to show that a single model can correct code errors across multiple distinct languages without requiring separate retraining for each language.

The researchers designed RING around three standard repair steps modeled after human developer workflows: locating the error using compiler or linter diagnostic messages, transforming the code by retrieving relevant example repairs through few-shot prompting, and ranking potential fix candidates using the model's token probability scores. The authors evaluated the system across six languages—Excel formulas, Power Fx, Python, JavaScript, C, and PowerShell—using benchmark datasets of 200 to 273 tasks per language and compared performance against dedicated, language-specific repair tools and baseline zero-shot models.

The evaluation revealed five key findings. First, RING outperformed specialized, state-of-the-art tools on top-ranked (Top@1) repair accuracy in three languages: Excel (82% vs. 71%), Python (94% vs. 92%), and C (63% vs. 55%). Second, for JavaScript, RING achieved a 46% Top@1 success rate on full code snippets, significantly outperforming the specialized baseline, which degraded from 59% on narrow code windows down to 9% on realistic, full-length snippets. Third, dynamically selecting few-shot examples based on diagnostic similarity improved Top@1 repair rates across all tested languages by up to 20% compared to using static, fixed prompt examples. Fourth, when an exact fix was not found, RING still located the fault within a narrow window of the correct edit location more often than the specialized baselines. Fifth, performance was substantially lower in PowerShell (18% Top@1), and repair success across all languages correlated inversely with code length, with longer snippets failing more frequently.

These results show that organizations can deploy a single foundation model to support multi-language developer assistance, eliminating the substantial engineering overhead and data-collection costs needed to maintain separate, bespoke repair tools. This capability enables a "flipped" interaction model where users write code naturally and an artificial intelligence assistant automatically suggests real-time fixes for last-mile errors. Furthermore, the findings indicate that a model's underlying training data representation directly dictates repair effectiveness across different programming ecosystems.

Organizations developing developer tooling should consider adopting prompt-engineered foundation models for code repair, particularly by pairing compiler diagnostic messages with dynamic example retrieval banks. To maximize performance, teams should implement specialized ranking mechanisms and explore iterative querying techniques for complex, multi-error scenarios. Before deploying such systems to low-resource languages like PowerShell, organizations should conduct targeted pilots and expand language-specific example banks to address lower baseline accuracy.

arXiv: 2208.11640
Cover for Repair Is Nearly Generation: Multilingual Program Repair with LLMs

Abstract

Most programmers make mistakes when writing code. Some of these mistakes are small and require few edits to the original program – a class of errors recently termed last mile mistakes. These errors break the flow for experienced developers and can stump novice programmers. Existing automated repair tech- niques targeting this class of errors are language-specific and do not easily carry over to new languages. Transferring sym- bolic approaches requires substantial engineering and neural approaches require data and retraining. We introduce RING, a multilingual repair engine powered by a large language model trained on code (LLMC) such as Codex. Such a multilingual engine enables a flipped model for programming assistance, one where the programmer writes code and the AI assistance suggests fixes, compared to traditional code suggestion tech- nology. Taking inspiration from the way programmers man- ually fix bugs, we show that a prompt-based strategy that conceptualizes repair as localization, transformation, and can- didate ranking, can successfully repair programs in multiple languages with minimal effort. We present the first results for such a multilingual repair engine by evaluating on 6 different languages and comparing performance to language-specific repair engines. We show that RING can outperform language- specific repair engines for three of these languages.

Table of Contents

  • Introduction
  • Related Work
  • Approach
  • Fault Localization through Language Tooling
  • Code Transformation through Few-shot Learning
  • Candidate Ranking
  • Language-Specific Datasets
  • Results and Analysis
  • RQ1. Viability of Multilingual Repair
  • RQ2. Error Localization
  • RQ3. Code Transformation
  • RQ4. Candidate Ranking
  • Discussion
  • Designing the Example Bank
  • Adapting RING for New Languages
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — RING Multilingual Program Repair Architecture

    model/method

    RING is an automated program repair (APR) framework that performs multilingual last-mile program repair using a pre-trained large language model trained on code (LLMC), specifically OpenAI Codex (davinci-code-002), without language-specific model fine-tuning or retraining.

    RING maps the three traditional pillars of automated program repair into prompt engineering and decoding components:

    1. Fault Localization: Error and diagnostic messages extracted from compilers, interpreters, or static analyzers (e.g., linters) are normalized into a standardized prompt format containing the diagnostic text and line/column span locations (or abstracted diagnostic text when error spans are noisy).
    2. Code Transformation: Dynamic few-shot learning is performed by querying an example bank of known buggy/fixed code snippet pairs using error similarity metrics to retrieve contextually relevant demonstration shots for in-context prompting.
    3. Candidate Ranking: The LLMC generates multiple candidate repairs by sampling with a non-zero temperature (T=0.7T = 0.7, top_p = 1.0). Candidates are ranked in descending order according to the mean log-probability of all generated tokens in each candidate program:

    Score(C)=1∣C∣∑t∈Clog⁡P(t∣t<t,Prompt)\text{Score}(C) = \frac{1}{|C|} \sum_{t \in C} \log P(t \mid t_{<t}, \text{Prompt})

    where CC denotes the sequence of generated tokens for candidate repair program CC.

  2. Knowl 2 — Error Diagnostic Retrieval Strategies for Few-Shot Repair Selection

    model/method

    To select relevant demonstration shots for few-shot prompt construction from an example bank of buggy/repaired program pairs, RING provides two similarity-based retrieval mechanisms conditioned on language tooling diagnostics:

    1. Error Vector Selection: Used when language tooling produces structured, fine-grained diagnostic category counters (e.g., Excel syntax/formula diagnostics, JavaScript ESLint rules). An error frequency vector v∈RD\mathbf{v} \in \mathbb{R}^D is constructed, where DD is the total number of diagnostic categories and each element counts occurrences of that specific error category in the target program. Candidates are retrieved from the example bank by minimizing the Euclidean distance (L2L_2 norm) between the target program error vector vtarget\mathbf{v}_{\text{target}} and example error vectors vexample\mathbf{v}_{\text{example}}:

    d(vtarget,vexample)=∥vtarget−vexample∥2d(\mathbf{v}_{\text{target}}, \mathbf{v}_{\text{example}}) = \|\mathbf{v}_{\text{target}} - \mathbf{v}_{\text{example}}\|_2

    1. Message Embedding Selection: Used when language diagnostics are represented primarily as unstructured or semi-structured natural language messages (e.g., Python, C, Power Fx, PowerShell). The compiler error message text is encoded into dense vector representations using a pre-trained CodeBERT model, and example shots are selected by maximizing cosine similarity between message embeddings:

    Sim(etarget,eexample)=etarget⋅eexample∥etarget∥2∥eexample∥2\text{Sim}(\mathbf{e}_{\text{target}}, \mathbf{e}_{\text{example}}) = \frac{\mathbf{e}_{\text{target}} \cdot \mathbf{e}_{\text{example}}}{\|\mathbf{e}_{\text{target}}\|_2 \|\mathbf{e}_{\text{example}}\|_2}

  3. Knowl 3 — Compiler Error Message Abstraction for Fault Localization

    model/method

    In low-code formula languages and languages whose compilers report imprecise or noisy location spans, exact line and column numbers in diagnostic prompts can mislead the code generation process of large language models. RING uses regular-expression-based diagnostic abstraction to extract the underlying semantic error description while stripping away brittle file, line, and column coordinates.

    For example, in C compilation errors where the GNU Compiler Collection (GCC) outputs:

    In function 'main':
    16:6: error: expected ';' before 'printf'
    printf("%d",catalan(h));
    ^
    

    RING applies the regular expression \d+:\d+: error: to discard the line/column indicator 16:6: error:, creating the abstracted message expected ';' before 'printf'. For structured low-code diagnostics (such as Excel parser outputs containing explicit record structures for description, category, and location), RING retains only the semantic description field.

  4. Knowl 4 — PowerShell Last-Mile Syntax Repair Benchmark Dataset

    experimental setup

    A dedicated evaluation benchmark for last-mile syntax repair in PowerShell commands was constructed by mining StackOverflow threads tagged with powershell that contained the term error (a total pool of 14,954 threads). Code blocks from question descriptions and accepted answers were parsed. Pairs were filtered to retain those where the question code failed syntax validation and the accepted answer code passed syntax validation, verified using the PowerShell diagnostic command:

    Get-Command -syntax
    

    Candidate pairs were manually inspected, simplified to eliminate extraneous modifications, and standardized. The resulting benchmark contains 200 distinct buggy/fixed PowerShell command pairs. Success on this benchmark requires an exact match against the curated reference answer block.

  5. Knowl 5 — Comparative Multilingual Program Repair Performance

    data/table

    RING evaluated against language-specific repair baselines (LaMirage for Excel and Power Fx; TFix for JavaScript; BIFI for Python; Dr. Repair for C) and a zero-shot Codex baseline across six languages using 200 to 273 tasks per language. Sampling used temperature T=0.7T = 0.7, top_p = 1.0, and top@kk metrics where a task is solved if any candidate in the top kk satisfies correctness criteria:

    Language Approach Top@1 Top@3 Top@50* Metric Avg. Tokens
    Excel RING (Abstracted Msg, Error Vector) 0.82 0.89 0.92 Exact Match 26±1426 \pm 14
    LaMirage 0.71 0.76 - Exact Match 26±1426 \pm 14
    Codex (Zero-shot) 0.60 0.77 0.88 Exact Match 26±1426 \pm 14
    Power Fx RING (Compiler Msg, Msg Embedding) 0.71 0.85 0.87 Exact Match 29±1929 \pm 19
    LaMirage 0.85 0.88 - Exact Match 29±1929 \pm 19
    Codex (Zero-shot) 0.47 0.68 0.84 Exact Match 29±1929 \pm 19
    JavaScript RING (Compiler Msg, Error Vector) 0.46 0.59 0.64 Exact Match 163±106163 \pm 106
    TFix (extended code snippets) 0.09 - - Exact Match 163±106163 \pm 106
    TFix (original window dataset) 0.59 - - Exact Match 7474
    Codex (Zero-shot) 0.19 0.28 0.39 Exact Match 163±106163 \pm 106
    Python RING (Compiler Msg, Msg Embedding) 0.94 0.97 0.97 Parser + Edit Dist <5< 5 104±150104 \pm 150
    BIFI 0.92 0.95 0.96 Parser + Edit Dist <5< 5 104±150104 \pm 150
    Codex (Zero-shot) 0.87 0.94 0.98 Parser + Edit Dist <5< 5 104±150104 \pm 150
    C RING (Compiler Msg, Msg Embedding) 0.63 0.69 0.70 Parser + Edit Dist <5< 5 223±72223 \pm 72
    Dr Repair 0.55 - - Parser + Edit Dist <5< 5 223±72223 \pm 72
    Codex (Zero-shot) 0.40 0.56 0.61 Parser + Edit Dist <5< 5 223±72223 \pm 72
    PowerShell RING (Compiler Msg, Msg Embedding) 0.18 0.25 0.28 Exact Match 24±3024 \pm 30
    Codex (Zero-shot) 0.10 0.15 0.18 Exact Match 24±3024 \pm 30

    *Note: For PowerShell, Top@20 is reported instead of Top@50 due to API rate limits. RING outperforms specialized language-specific repair tools on Excel (+11 percentage points), Python (+2 percentage points), and C (+8 percentage points). On extended JavaScript snippets, TFix drops from 0.59 to 0.09 due to larger input contexts (average 208 T5 tokens vs 74), whereas RING achieves 0.46.

  6. Knowl 6 — Performance Impact of Dynamic Few-Shot Selection vs. Fixed Shots

    data/table

    Comparison of top@1 repair accuracy using dynamically retrieved few-shot examples (via error vector or message embedding similarity) against static, pre-defined fixed demonstration examples across all evaluated programming languages:

    Language Fixed Shots (Top@1) Smart Shots (Top@1) Fractional Change
    Excel 0.76 0.82 +0.08
    Power Fx 0.70 0.71 +0.01
    JavaScript 0.43 0.46 +0.07
    Python 0.91 0.94 +0.03
    C 0.50 0.58 +0.16
    PowerShell 0.15 0.18 +0.20

    Dynamic few-shot selection consistently improves repair accuracy across all six benchmark languages. The smallest gain occurs in Power Fx (+0.01 fractional gain), which is caused by imprecise compiler diagnostics that occasionally introduce spurious tokens into the prompt.

  7. Knowl 7 — Error Localization Accuracy on Unrepaired Programs

    empirical result

    For programs that fail to achieve correct repair at top@1, an approximate fault localization metric was evaluated: an edit is counted as correctly localized if all generated candidate edit locations fall within a tolerance range of ±k\pm k tokens of the ground-truth edit location.

    Evaluating across the four languages with ground-truth repair locations (Excel, Power Fx, JavaScript, and PowerShell):

    1. Localization Superiority: For unrepaired programs, RING's top candidate correctly localizes the fault in a higher fraction of cases than language-specific baselines across all tolerance levels k∈{0,1,2,3,4}k \in \{0, 1, 2, 3, 4\}. For Power Fx, RING localizes over 25%25\% of unrepaired programs within a ±1\pm 1 token tolerance.
    2. Program Length Sensitivity: In JavaScript and Python, buggy programs successfully repaired at top@1 have a shorter token length distribution than programs that fail to be repaired. In Excel, this dependency is less pronounced due to shorter formula lengths and restricted grammar.
  8. Knowl 8 — Cross-Language Calibration and Ranking Disparities in LLMC Log-Probabilities

    empirical result

    Candidate ranking in RING relies on average token log-probabilities generated by OpenAI Codex. Empirical analysis of Gaussian kernel density distributions of these log-probabilities indicates:

    1. Outcome Separation: In languages where RING outperforms specialized repair engines (Excel, C), there is a distinct separation between the log-probability density distributions of successful (True) repairs and failed (False) repairs. In Power Fx, the separation is less distinct. In PowerShell, the relative positions of distribution peaks are inverted relative to other languages.
    2. Training Data Pre-training Bias: For successful repairs (Top@1 True), the average per-token log-probabilities are systematically shifted toward lower values (more negative log-probabilities) for lower-resource languages in Codex's pre-training corpus (PowerShell, Excel, Power Fx) compared to high-resource languages (JavaScript, Python). This disparity demonstrates that raw generation log-probabilities are miscalibrated across languages.
  9. Knowl 9 — Limitations of LLMC-Based Program Repair

    limitation

    The prompt-based LLMC multilingual program repair approach exhibits several key limitations:

    1. Corpus Distribution Dependency: Repair accuracy drops sharply for languages under-represented in the LLMC's pre-training corpora; PowerShell attains only 0.18 Top@1 exact match, with the model performing fewer edits than necessary to reach correct syntax.
    2. Sensitivity to Diagnostic Precision: When compiler or linter diagnostics are imprecise or report inaccurate spans (e.g., Power Fx reporting non-existent tokens), error-based example selection injects noisy demonstration shots that limit repair effectiveness.
    3. Ranking Miscalibration: Raw LLMC token log-probabilities are poorly calibrated across different programming languages and do not always reliably rank the semantically valid fix at Top@1, creating a substantial performance gap between Top@1 and Top@kk (k≥3k \ge 3).
    4. Context Length Sensitivity: Extending short code windows to complete function contexts degrades model performance for sequence models (e.g., TFix falling from 0.59 to 0.09 on JavaScript), and longer programs generally exhibit higher failure rates across languages.

Coverage note — None was omitted; all contributed models, retrieval methods, empirical findings across the 6 languages, benchmark creation details, calibration analyses, and limitations are fully covered.

References

  1. 1.Ahmed, T.; Ledesma, N. R.; and Devanbu, P. 2021. SYNFIX: Automatically Fixing Syntax Errors using Compiler Diagnostics. arXiv preprint arXiv:2104.14671.
  2. 2.Ahmed, U. Z.; Kumar, P.; Karkare, A.; Kar, P.; and Gulwani, S. 2018. Compilation error repair: for the student programs, from the student programs. In Proceedings of the 40th International Conference on Software Engineering: Software Engineering Education and Training, 78–87.
  3. 3.Altadmri, A.; and Brown, N. C. 2015. 37 million compilations: Investigating novice programming mistakes in large-scale student data. In Proceedings of the 46th ACM technical symposium on computer science education, 522–527.
  4. 4.Arcuri, A. 2008. On the automation of fixing software bugs. In Companion of the 30th international conference on Software engineering, 1003–1006.
  5. 5.Bareiß, P.; Souza, B.; d’Amorim, M.; and Pradel, M. 2022. Code Generation Tools (Almost) for Free? A Study of Few-Shot, Pre-Trained Language Models on Code. arXiv preprint arXiv:2206.01335.
  6. 6.Bavishi, R.; Joshi, H.; Cambronero, J.; Fariha, A.; Gulwani, S.; Le, V.; Radicek, I.; and Tiwari, A. 2022. Neurosymbolic Repair for Low-Code Formula Languages. Proc. ACM Program. Lang., 6(OOPSLA2).
  7. 7.Bella, A.; Ferri, C.; Hernández-Orallo, J.; and Ramírez-Quintana, M. J. 2010. Calibration of machine learning models. In Handbook of Research on Machine Learning Applications and Trends: Algorithms, Methods, and Techniques, 128–146. IGI Global.
  8. 8.Berabi, B.; He, J.; Raychev, V.; and Vechev, M. 2021. Tfix: Learning to fix coding errors with a text-to-text transformer. In International Conference on Machine Learning, 780–791. PMLR.
  9. 9.Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  10. 10.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901.
  11. 11.Bureau of Labor Statistics, U. 2022. Software developers, Quality Assurance Analysts, and testers : Occupational outlook handbook. https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm. Accessed: 2022-07-30.
  12. 12.Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Ponde, H.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; Ray, A.; Puri, R.; Krueger, G.; Petrov, M.; Khlaaf, H.; Sastry, G.; Mishkin, P.; Chan, B.; Gray, S.; Ryder, N.; Pavlov, M.; Power, A.; Kaiser, L.; Bavarian, M.; Winter, C.; Tillet, P.; Such, F. P.; Cummings, D. W.; Plappert, M.; Chantzis, F.; Barnes, E.; Herbert-Voss, A.; Guss, W. H.; Nichol, A.; Babuschkin, I.; Balaji, S. A.; Jain, S.; Carr, A.; Leike, J.; Achiam, J.; Misra, V.; Morikawa, E.; Radford, A.; Knight, M. M.; Brundage, M.; Murati, M.; Mayer, K.; Welinder, P.; McGrew, B.; Amodei, D.; McCandlish, S.; Sutskever, I.; and Zaremba, W. 2021. Evaluating Large Language Models Trained on Code. ArXiv, abs/2107.03374.
  13. 13.Chowdhury, J. R.; Zhuang, Y.; and Wang, S. 2022. Novelty Controlled Paraphrase Generation with Retrieval Augmented Conditional Prompt Tuning. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10): 10535–10544.
  14. 14.Debroy, V.; and Wong, W. E. 2010. Using Mutation to Automatically Suggest Fixes for Faulty Programs. 2010 Third International Conference on Software Testing, Verification and Validation, 65–74.
  15. 15.Diekmann, L.; and Tratt, L. 2020. Don’t Panic! Better, Fewer, Syntax Errors for LR Parsers. In 34th European Conference on Object-Oriented Programming, ECOOP 2020, volume 166 of LIPIcs, 6:1–6:32.
  16. 16.Dormann, C. F. 2020. Calibration of probability predictions from machine-learning and statistical models. Global ecology and biogeography, 29(4): 760–765.
  17. 17.Drori, I.; Zhang, S.; Shuttleworth, R.; Tang, L.; Lu, A.; Ke, E.; Liu, K.; Chen, L.; Tran, S.; Cheng, N.; et al. 2022. A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level. Proceedings of the National Academy of Sciences, 119(32): e2123433119.
  18. 18.Drosos, I.; Guo, P. J.; and Parnin, C. 2017. HappyFace: Identifying and predicting frustrating obstacles for learning programming at scale. In 2017 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), 171–179. IEEE.
  19. 19.Feng, Z.; Guo, D.; Tang, D.; Duan, N.; Feng, X.; Gong, M.; Shou, L.; Qin, B.; Liu, T.; Jiang, D.; et al. 2020. Codebert: A pre-trained model for programming and natural languages. arXiv preprint arXiv:2002.08155.
  20. 20.Gazzola, L.; Micucci, D.; and Mariani, L. 2019. Automatic Software Repair: A Survey. IEEE Transactions on Software Engineering, 45: 34–67.
  21. 21.Goues, C. L.; Pradel, M.; and Roychoudhury, A. 2019. Automated program repair. Communications of the ACM, 62(12): 56–65.
  22. 22.Gupta, R.; Pal, S.; Kanade, A.; and Shevade, S. K. 2017. DeepFix: Fixing Common C Language Errors by Deep Learning. In AAAI.
  23. 23.Hajipour, H.; Bhattacharyya, A.; and Fritz, M. 2020. SampleFix: Learning to Correct Programs by Efficient Sampling of Diverse Fixes. In NeurIPS 2020 Workshop on Computer-Assisted Programming.
  24. 24.Inala, J. P.; Wang, C.; Yang, M.; Codas, A.; Encarnación, M.; Lahiri, S. K.; Musuvathi, M.; and Gao, J. 2022. Fault-Aware Neural Code Rankers. arXiv preprint arXiv:2206.03865.
  25. 25.Johnson, J.; Douze, M.; and Jégou, H. 2019. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 7(3): 535–547.
  26. 26.Levenshtein, V. I.; et al. 1966. Binary codes capable of correcting deletions, insertions, and reversals. In Soviet physics doklady, volume 10, 707–710. Soviet Union.
  27. 27.Liu, K.; Li, L.; Koyuncu, A.; Kim, D.; Liu, Z.; Klein, J.; and Bissyandé, T. F. 2021. A critical review on the evaluation of automated program repair systems. Journal of Systems and Software, 171: 110817.
  28. 28.Murphy, L.; Lewandowski, G.; McCauley, R.; Simon, B.; Thomas, L.; and Zander, C. 2008. Debugging: the good, the bad, and the quirky–a qualitative analysis of novices’ strategies. ACM SIGCSE Bulletin, 40(1): 163–167.
  29. 29.Nguyen, H. D. T.; Qi, D.; Roychoudhury, A.; and Chandra, S. 2013. SemFix: Program repair via semantic analysis. International Conference on Software Engineering, 772–781.
  30. 30.Nixon, J.; Dusenberry, M. W.; Zhang, L.; Jerfel, G.; and Tran, D. 2019. Measuring Calibration in Deep Learning. In CVPR Workshops, volume 2.
  31. 31.Parihar, S.; Dadachanji, Z.; Singh, P. K.; Das, R.; Karkare, A.; and Bhattacharya, A. 2017. Automatic grading and feedback using program repair for introductory programming courses. In Proceedings of the 2017 ACM Conference on Innovation and Technology in Computer Science Education, 92–97.
  32. 32.Poesia, G.; Polozov, A.; Le, V.; Tiwari, A.; Soares, G.; Meek, C.; and Gulwani, S. 2022. Synchromesh: Reliable Code Generation from Pre-trained Language Models. In International Conference on Learning Representations.
  33. 33.Prenner, J. A.; and Robbes, R. 2021. Automatic Program Repair with OpenAI’s Codex: Evaluating QuixBugs. arXiv preprint arXiv:2111.03922.
  34. 34.Pu, Y.; Narasimhan, K.; Solar-Lezama, A.; and Barzilay, R. 2016. sk_p: a neural program corrector for MOOCs. In Companion Proceedings of the 2016 ACM SIGPLAN International Conference on Systems, Programming, Languages and Applications: Software for Humanity, 39–40.
  35. 35.Qi, Y.; Mao, X.; Lei, Y.; Dai, Z.; and Wang, C. 2014. The strength of random search on automated program repair. In Proceedings of the 36th International Conference on Software Engineering, 254–265.
  36. 36.Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018. Improving language understanding by generative pre-training. https://paperswithcode.com/paper/improving-language-understanding-by. Accessed: 2022-08-05.
  37. 37.Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21(140): 1–67.
  38. 38.Shannon, C. E. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3): 379–423.
  39. 39.Spotify. 2022. ANNOY library. https://github.com/spotify/annoy. Accessed: 2022-08-01.
  40. 40.StackOverflow. 2022. StackOverflow Website. https://stackoverflow.com/.
  41. 41.Tómasdóttir, K. F.; Aniche, M.; and Van Deursen, A. 2018. The adoption of javascript linters in practice: A case study on eslint. IEEE Transactions on Software Engineering, 46(8): 863–891.
  42. 42.Wexelblat, R. L. 1976. Maxims for malfeasant designers, or how to design languages to make programming as difficult as possible. In Proceedings of the 2nd international conference on Software engineering, 331–336.
  43. 43.Yasunaga, M.; and Liang, P. 2020. Graph-based, self-supervised program repair from diagnostic feedback. In International Conference on Machine Learning, 10799–10808. PMLR.
  44. 44.Yasunaga, M.; and Liang, P. 2021. Break-it-fix-it: Unsupervised learning for program repair. In International Conference on Machine Learning, 11941–11952. PMLR.
  45. 45.Zhong, H.; and Su, Z. 2015. An Empirical Study on Real Bug Fixes. 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering, 1: 913–923.

Citation

MLA
Joshi, H., et al. “Repair Is Nearly Generation: Multilingual Program Repair with LLMs”. arXiv, 2022, http://arxiv.org/abs/2208.11640v3.
APA
Joshi, H., Cambronero, J., Gulwani, S., Le, V., Radicek, I., & Verbruggen, G. (2022). Repair Is Nearly Generation: Multilingual Program Repair with LLMs. arXiv. http://arxiv.org/abs/2208.11640v3
Chicago
Joshi, H., J. Cambronero, S. Gulwani, V. Le, I. Radicek, and G. Verbruggen. 2022. “Repair Is Nearly Generation: Multilingual Program Repair with LLMs”. arXiv. http://arxiv.org/abs/2208.11640v3.
Harvard
Joshi, H. et al. (2022) “Repair Is Nearly Generation: Multilingual Program Repair with LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2208.11640v3.
Vancouver
1. Joshi H, Cambronero J, Gulwani S, Le V, Radicek I, Verbruggen G (2022) Repair Is Nearly Generation: Multilingual Program Repair with LLMs. arXiv

BibTeX

@article{joshi2022repair,
  title = {Repair Is Nearly Generation: Multilingual Program Repair with LLMs},
  author = {Joshi, Harshit and Cambronero, José and Gulwani, Sumit and Le, Vu and Radicek, Ivan and Verbruggen, Gust},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2208.11640v3},
  eprint = {2208.11640}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF