Natural Language to Code Translation with Execution
Freda ShiDaniel FriedMarjan GhazvininejadLuke ZettlemoyerSida I. Wang
Proposes an execution-based minimum Bayes risk decoding framework that boosts code generation accuracy by executing sampled candidate programs on test inputs to select the solution with the highest semantic consensus.
Large pretrained language models can translate natural language instructions into functional computer code, but they often produce multiple plausible candidates that include subtle errors. Choosing a single, correct program from these candidates without manual inspection is a major bottleneck in deploying automated code generation safely and effectively.
The article introduces and evaluates an inference selection method called Minimum Bayes Risk Decoding with Execution (MBR-EXEC). The objective is to demonstrate that executing candidate code snippets on a small set of sample inputs and selecting the output with the highest consensus significantly improves code generation accuracy without requiring known ground-truth outputs or retraining.
The researchers prompted a pretrained code model (Codex) using few-shot examples across three distinct programming environments: MBPP for Python, Spider for SQL, and NL2Bash for Bash scripts. For each problem description, the model generated a pool of candidate solutions. MBR-EXEC evaluated these candidates by running them on test inputs, comparing their outputs to one another, and choosing the candidate that maximized agreement across the pool. The team compared this consensus-based execution approach against traditional selection methods, including standard sampling, greedy decoding, likelihood-based metrics, and text similarity metrics.
The primary findings show substantial accuracy improvements across all evaluated languages. For Python code generation on MBPP, MBR-EXEC increased execution accuracy from 47.7% to 58.2% compared to standard sampling. On SQL generation via Spider, accuracy rose from 48.5% to 63.6%, an absolute increase of approximately 15 percentage points. On Bash command synthesis, accuracy improved from 53.0% to 58.5%. The analysis also revealed that simply filtering out programs that crash or fail to execute markedly improves traditional likelihood baselines. Furthermore, the overall pool of generated samples contained correct solutions far exceeding state-of-the-art supervised systems, indicating that the primary challenge lies in selection rather than model generation capability.
These findings imply that organizations can dramatically enhance the quality and reliability of AI-generated code purely at inference time, avoiding expensive model fine-tuning. By relying on behavioral consensus on test inputs rather than model confidence scores, systems can filter out repetitive errors and invalid syntax, reducing debugging time and operational risk in automated workflows.
For practical implementation, teams deploying code generation tools should integrate post-generation execution filtering and consensus selection with low sampling temperatures (below 0.5). If dynamic execution is impossible due to security or platform constraints, text-similarity consensus metrics provide a robust fallback. Future work should investigate incorporating execution feedback directly into model training rather than relying solely on post-hoc inference selection.
Confidence in these findings is high for standard, self-contained coding and query tasks evaluated on established benchmarks. However, leaders should note that the approach requires access to valid input cases at inference time and assumes code can be safely run in an isolated environment with defined execution limits.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This seminal paper introduces Codex and the HumanEval benchmark, establishing the foundational paradigm of natural-language-to-code generation and the challenge of program selection that the source paper directly addresses.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). It introduces the MBPP benchmark and the execution-based functional correctness evaluation protocol that serves as a core experimental testbed in the source paper.
- Paper: Minimum Error Rate Training in Statistical Machine Translation, Franz Josef Och (2003). It provides foundational principles for Minimum Bayes Risk and error-rate-directed decoding in sequence generation that motivate the source's execution-based MBR decoding formulation.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). It establishes standardized datasets and baseline methodologies for evaluating code understanding and generation from natural language.
- Paper: Competition-level code generation with AlphaCode, Yujia Li et al. (2022). AlphaCode scales up the concept of execution-based filtering and clustering of sampled candidate solutions to solve complex competitive programming tasks.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). It takes the utilization of execution feedback a step further by using runtime execution outputs to enable models to iteratively self-debug and refine generated programs.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). It extends the evaluation of generated code correctness by demonstrating how insufficient test suites can lead to false positives and proposing rigorous test-case generation.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). It leverages programmatic code generation and external program execution interpreters to enhance step-by-step reasoning on complex numerical tasks.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). It applies natural-language-to-code generation to robotics, relying on real-time execution environments to enact physical control policies.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). It expands the execution and functional evaluation of code LLMs with dynamic, contamination-free contest problems and output prediction benchmarks.
- Paper: BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions, Terry Yue Zhuo et al. (2025). It generalizes execution-based functional testing of LLM-generated code to complex multi-library instructions and diverse third-party package dependencies.
