Teaching Large Language Models to Self-Debug

Xinyun ChenMaxwell LinNathanael SchärliDenny Zhou

article2023ICLR1,360 citations

Proposes a self-debugging method that enables large language models to autonomously detect and correct their own code generation mistakes through natural language explanations and execution feedback, boosting accuracy on benchmarks such as Spider, TransCoder, and MBPP.

Listen

Large language models have shown remarkable capabilities in automated code generation, but they frequently struggle to produce completely correct programs on complex tasks in a single attempt. Traditional workarounds often rely on sampling dozens of candidate programs or training specialized auxiliary repair models, which significantly increases computational costs and engineering overhead. The article introduces and evaluates "Self-Debugging," a framework that enables large language models to iteratively identify and correct bugs in their own generated code using few-shot prompt demonstrations without requiring fine-tuning, specialized training, or human intervention.

The authors evaluate the Self-Debugging framework across major coding benchmarks, including text-to-database query generation (Spider), code translation from C++ to Python (TransCoder), and general text-to-Python synthesis (MBPP). The study tests several foundation models, including OpenAI's Codex, GPT-3.5, and GPT-4, alongside the open-source StarCoder model. The core approach structures the debugging process into iterative cycles where the model generates initial code, observes execution feedback or runtime error messages when available, and explains the code logic in natural language—a technique analogous to human "rubber duck debugging"—to identify semantic discrepancies against the original requirements.

The evaluation yields several key findings. First, Self-Debugging achieves state-of-the-art accuracy across all tested benchmarks, boosting baseline accuracy by up to 12% on code translation and general Python generation tasks where unit tests are available. Second, even on tasks lacking unit tests such as text-to-SQL generation, prompting the model to explain its code line-by-line consistently improves baseline performance by 2% to 3% overall and delivers a 9% accuracy improvement on the most complex problems. Third, Self-Debugging dramatically improves sample efficiency; a single generated program refined through one or two debugging turns regularly matches or outperforms baseline systems that sample 10 to 32 candidate solutions. Most successful repairs occur within just one to three debugging cycles, resolving common errors such as mismatched conditions, type errors, and language-specific syntax differences.

These results demonstrate that enabling iterative self-refinement is a practical, cost-effective alternative to generating massive batches of code candidates or fine-tuning complex repair pipelines. Organizations deploying code-generation models can achieve higher accuracy while substantially reducing inference compute costs and API latency. Decision-makers should implement multi-step explanation and automated execution feedback loops in their code generation workflows, prioritizing the integration of unit test executors where feasible. While the framework shows high efficacy, its benefits without code execution depend heavily on the model's underlying reasoning strength, as certain models risk being overconfident in their initial predictions when external test feedback is absent.

arXiv: 2304.05128
Cover for Teaching Large Language Models to Self-Debug

Abstract

Large language models (LLMs) have achieved impressive performance on code generation. However, for complex programming tasks, generating the correct solution in one go becomes challenging, thus some prior works have designed program repair approaches to improve code generation performance. In this work, we propose Self-Debugging, which teaches a large language model to debug its predicted program via few-shot demonstrations. In particular, we demonstrate that Self-Debugging can teach the large language model to perform rubber duck debugging; i.e., without any human feedback on the code correctness or error messages, the model is able to identify its mistakes by investigating the execution results and explaining the generated code in natural language. Self-Debugging achieves the state-of-the-art performance on several code generation benchmarks, including the Spider dataset for text-to-SQL generation, TransCoder for C++-to-Python translation, and MBPP for text-to-Python generation. On the Spider benchmark where there are no unit tests to verify the correctness of predictions, Self-Debugging with code explanation consistently improves the baseline by 2-3%, and improves the prediction accuracy on problems of the hardest level by 9%. On TransCoder and MBPP where unit tests are available, Self-Debugging improves the baseline accuracy by up to 12%. Meanwhile, by leveraging feedback messages and reusing failed predictions, Self-Debugging notably improves sample efficiency, and can match or outperform baseline models that generate more than 10x candidate programs.

Citation

MLA
Chen, X., et al. “Teaching Large Language Models to Self-Debug”. arXiv, 2023, http://arxiv.org/abs/2304.05128v2.
APA
Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). Teaching Large Language Models to Self-Debug. arXiv. http://arxiv.org/abs/2304.05128v2
Chicago
Chen, X., M. Lin, N. Schärli, and D. Zhou. 2023. “Teaching Large Language Models to Self-Debug”. arXiv. http://arxiv.org/abs/2304.05128v2.
Harvard
Chen, X. et al. (2023) “Teaching Large Language Models to Self-Debug”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.05128v2.
Vancouver
1. Chen X, Lin M, Schärli N, Zhou D (2023) Teaching Large Language Models to Self-Debug. arXiv

BibTeX

@article{chen2023teaching,
  title = {Teaching Large Language Models to Self-Debug},
  author = {Chen, Xinyun and Lin, Maxwell and Schärli, Nathanael and Zhou, Denny},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.05128v2},
  eprint = {2304.05128}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/