Teaching Large Language Models to Self-Debug
Xinyun ChenMaxwell LinNathanael SchärliDenny Zhou
Proposes a self-debugging method that enables large language models to autonomously detect and correct their own code generation mistakes through natural language explanations and execution feedback, boosting accuracy on benchmarks such as Spider, TransCoder, and MBPP.
Large language models have shown remarkable capabilities in automated code generation, but they frequently struggle to produce completely correct programs on complex tasks in a single attempt. Traditional workarounds often rely on sampling dozens of candidate programs or training specialized auxiliary repair models, which significantly increases computational costs and engineering overhead. The article introduces and evaluates "Self-Debugging," a framework that enables large language models to iteratively identify and correct bugs in their own generated code using few-shot prompt demonstrations without requiring fine-tuning, specialized training, or human intervention.
The authors evaluate the Self-Debugging framework across major coding benchmarks, including text-to-database query generation (Spider), code translation from C++ to Python (TransCoder), and general text-to-Python synthesis (MBPP). The study tests several foundation models, including OpenAI's Codex, GPT-3.5, and GPT-4, alongside the open-source StarCoder model. The core approach structures the debugging process into iterative cycles where the model generates initial code, observes execution feedback or runtime error messages when available, and explains the code logic in natural language—a technique analogous to human "rubber duck debugging"—to identify semantic discrepancies against the original requirements.
The evaluation yields several key findings. First, Self-Debugging achieves state-of-the-art accuracy across all tested benchmarks, boosting baseline accuracy by up to 12% on code translation and general Python generation tasks where unit tests are available. Second, even on tasks lacking unit tests such as text-to-SQL generation, prompting the model to explain its code line-by-line consistently improves baseline performance by 2% to 3% overall and delivers a 9% accuracy improvement on the most complex problems. Third, Self-Debugging dramatically improves sample efficiency; a single generated program refined through one or two debugging turns regularly matches or outperforms baseline systems that sample 10 to 32 candidate solutions. Most successful repairs occur within just one to three debugging cycles, resolving common errors such as mismatched conditions, type errors, and language-specific syntax differences.
These results demonstrate that enabling iterative self-refinement is a practical, cost-effective alternative to generating massive batches of code candidates or fine-tuning complex repair pipelines. Organizations deploying code-generation models can achieve higher accuracy while substantially reducing inference compute costs and API latency. Decision-makers should implement multi-step explanation and automated execution feedback loops in their code generation workflows, prioritizing the integration of unit test executors where feasible. While the framework shows high efficacy, its benefits without code execution depend heavily on the model's underlying reasoning strength, as certain models risk being overconfident in their initial predictions when external test feedback is absent.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). Establishes foundational techniques and benchmarks for large language model code generation, which Self-Debugging directly targets and improves upon.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Introduces the MBPP benchmark and explores initial program synthesis prompting, forming key baseline tasks and methodologies utilized in the Self-Debugging study.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). Demonstrates how models can bootstrap reasoning via self-generated rationales, providing the foundational conceptual paradigm behind self-explanation and self-refinement.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). Pioneers multi-turn conversational program synthesis, motivating the need for iterative, multi-step debugging workflows rather than single-turn generation.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Introduces the concept of decoupling natural language reasoning from code execution via interpreters, a core mechanism underpinning execution-based self-debugging.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Establishes subproblem decomposition strategies in language models, helping inform how complex coding errors are incrementally identified and resolved.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). Extends self-debugging principles into verbal reinforcement learning and episodic memory across broader interactive agent tasks and complex benchmarks.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). Generalizes iterative self-feedback and refinement from code debugging to a diverse spectrum of natural language and multi-constraint generation tasks.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). Scales iterative code repair and execution feedback into full-scale autonomous software engineering agents operating across entire code repositories.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). Provides rigorous mutation-based test suites that expose subtle bugs in LLM-generated code that standard self-debugging and evaluation setups miss.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). Establishes a contamination-free, holistic evaluation suite that tests code repair with execution feedback across fresh, real-world competitive programming tasks.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Applies reflection and self-critique mechanisms to retrieval-augmented generation pipelines through fine-tuned reflection tokens.
- Paper: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate, Tian Liang et al. (2023). Addresses the failure modes of single-model self-reflection by using multi-agent debate to prevent cognitive fixation during error correction.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). Synthesizes code-generation, feedback loops, and self-debugging methods into a comprehensive structural survey of code-driven agent harnesses.
