Natural Language to Code Translation with Execution

Freda ShiDaniel FriedMarjan GhazvininejadLuke ZettlemoyerSida I. Wang

article2022EMNLP164 citations

Proposes an execution-based minimum Bayes risk decoding framework that boosts code generation accuracy by executing sampled candidate programs on test inputs to select the solution with the highest semantic consensus.

Listen

Large pretrained language models can translate natural language instructions into functional computer code, but they often produce multiple plausible candidates that include subtle errors. Choosing a single, correct program from these candidates without manual inspection is a major bottleneck in deploying automated code generation safely and effectively.

The article introduces and evaluates an inference selection method called Minimum Bayes Risk Decoding with Execution (MBR-EXEC). The objective is to demonstrate that executing candidate code snippets on a small set of sample inputs and selecting the output with the highest consensus significantly improves code generation accuracy without requiring known ground-truth outputs or retraining.

The researchers prompted a pretrained code model (Codex) using few-shot examples across three distinct programming environments: MBPP for Python, Spider for SQL, and NL2Bash for Bash scripts. For each problem description, the model generated a pool of candidate solutions. MBR-EXEC evaluated these candidates by running them on test inputs, comparing their outputs to one another, and choosing the candidate that maximized agreement across the pool. The team compared this consensus-based execution approach against traditional selection methods, including standard sampling, greedy decoding, likelihood-based metrics, and text similarity metrics.

The primary findings show substantial accuracy improvements across all evaluated languages. For Python code generation on MBPP, MBR-EXEC increased execution accuracy from 47.7% to 58.2% compared to standard sampling. On SQL generation via Spider, accuracy rose from 48.5% to 63.6%, an absolute increase of approximately 15 percentage points. On Bash command synthesis, accuracy improved from 53.0% to 58.5%. The analysis also revealed that simply filtering out programs that crash or fail to execute markedly improves traditional likelihood baselines. Furthermore, the overall pool of generated samples contained correct solutions far exceeding state-of-the-art supervised systems, indicating that the primary challenge lies in selection rather than model generation capability.

These findings imply that organizations can dramatically enhance the quality and reliability of AI-generated code purely at inference time, avoiding expensive model fine-tuning. By relying on behavioral consensus on test inputs rather than model confidence scores, systems can filter out repetitive errors and invalid syntax, reducing debugging time and operational risk in automated workflows.

For practical implementation, teams deploying code generation tools should integrate post-generation execution filtering and consensus selection with low sampling temperatures (below 0.5). If dynamic execution is impossible due to security or platform constraints, text-similarity consensus metrics provide a robust fallback. Future work should investigate incorporating execution feedback directly into model training rather than relying solely on post-hoc inference selection.

Confidence in these findings is high for standard, self-contained coding and query tasks evaluated on established benchmarks. However, leaders should note that the approach requires access to valid input cases at inference time and assumes code can be safely run in an isolated environment with defined execution limits.

Cover for Natural Language to Code Translation with Execution

Abstract

Generative models of code, pretrained on large corpora of programs, have shown great success in translating natural language to code (Chen et al., 2021; Austin et al., 2021; Li et al., 2022, inter alia). While these models do not explicitly incorporate program semantics (i.e., execution results) during training, they are able to generate correct solutions for many problems. However, choosing a single correct program from a generated set for each problem remains challenging.

In this work, we introduce execution result–based minimum Bayes risk decoding (MBR-EXEC) for program selection and show that it improves the few-shot performance of pretrained code models on natural-language-to-code tasks. We select output programs from a generated candidate set by marginalizing over program implementations that share the same semantics. Because exact equivalence is intractable, we execute each program on a small number of test inputs to approximate semantic equivalence. Across datasets, execution or simulated execution significantly outperforms the methods that do not involve program semantics. We find that MBR-EXEC consistently improves over all execution-unaware selection methods, suggesting it as an effective approach for natural language to code translation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Language to Code with Neural Networks
  • 2.2 Prompting Pretrained Language Models
  • 2.3 Minimum Bayes Risk Decoding
  • 3 Proposed Approach: MBR-EXEC
  • 3.1 Sample Collection
  • 3.2 Execution-Based MBR Decoding
  • 4 Experiments
  • 4.1 Datasets and Evaluation Metrics
  • 4.2 Baselines
  • 4.3 Primary Results
  • 4.4 Analysis
  • 4.4.1 Effect of Sample Temperature
  • 4.4.2 Effect of Different 3-shot Prompts
  • 4.4.3 Executability vs. Execution Results
  • 4.4.4 Soft Loss as the Bayes Risk Function
  • 4.5 Oracle Performance
  • 5 Discussion
  • Limitations
  • References
  • Appendices
  • A Example Prompts and Codex API Responses
  • B Full Analysis on Executability vs. Execution Result
  • MBPP: Prompt
  • MBPP: Response
  • Spider: Prompt
  • Spider: Response
  • NL2Bash: Prompt
  • NL2Bash: Response

Knowls

  1. Knowl 1 — Execution-Based Minimum Bayes Risk Decoding (MBR-EXEC)

    model/method

    Execution-based Minimum Bayes Risk (MBR-EXEC) decoding is an inference-time candidate selection framework for natural language to code generation models. Pretrained code language models frequently assign probability mass to diverse syntactic implementations that compute identical functional outputs. MBR-EXEC exploits this semantic redundancy by choosing the candidate program that maximizes consensus in the output execution space when evaluated across available test inputs.

    Given a natural language prompt CC, a set of NN candidate programs P={pi}i=1N\mathcal{P} = \{p_i\}_{i=1}^N is sampled from the language model. The optimal program p^\hat{p} is selected by minimizing the empirical Bayes risk:

    p^=arg⁡min⁡p∈P∑pref∈Pℓ(p,pref)\hat{p} = \arg\min_{p \in \mathcal{P}} \sum_{p_{\text{ref}} \in \mathcal{P}} \ell(p, p_{\text{ref}})

    where ℓ(p,pref)\ell(p, p_{\text{ref}}) is an execution-based discrepancy loss evaluated on a set of test inputs TT. Ground-truth outputs are not required during selection. If multiple candidate programs obtain the identical minimal Bayes risk, ties are broken by selecting the candidate program with the largest sequence log-likelihood under the generative model.

  2. Knowl 2 — Hard and Soft Execution Loss Functions for MBR-EXEC

    equation

    To measure discrepancy between candidate programs in MBR-EXEC, two semantic equivalence loss functions are defined over a finite set of test inputs TT.

    The primary hard 0/1 execution loss ℓ(pi,pj)\ell(p_i, p_j) is defined as:

    ℓ(pi,pj)=max⁡t∈T1[pi(t)≠pj(t)]\ell(p_i, p_j) = \max_{t \in T} \mathbf{1}[p_i(t) \neq p_j(t)]

    where pi(t)p_i(t) denotes the evaluation output of program pip_i when run on test input tt, and 1[⋅]\mathbf{1}[\cdot] is the indicator function. Under this loss, two programs receive zero discrepancy if and only if their outputs match on every test input in TT. If a program crashes, times out, or fails to execute on any input t∈Tt \in T, it is defined as non-equivalent to all programs, including other programs that fail on the same input.

    When multiple test inputs are accessible (∣T∣>1|T| > 1), an alternative soft execution loss ℓsoft(pi,pj)\ell_{\text{soft}}(p_i, p_j) measures the fraction of mismatched execution outputs:

    ℓsoft(pi,pj)=1∣T∣∑t∈T1[pi(t)≠pj(t)]\ell_{\text{soft}}(p_i, p_j) = \frac{1}{|T|} \sum_{t \in T} \mathbf{1}[p_i(t) \neq p_j(t)]

    When only one test input is provided (∣T∣=1|T| = 1), ℓ(pi,pj)\ell(p_i, p_j) and ℓsoft(pi,pj)\ell_{\text{soft}}(p_i, p_j) are equivalent.

  3. Knowl 3 — Performance Comparison of MBR-EXEC Against Baseline Generation Methods

    data/table

    Evaluating MBR-EXEC against standard decoding baselines using OpenAI Codex (code-davinci-001) with 3-shot prompting demonstrates substantial accuracy improvements across Python (MBPP), SQL (Spider), and Bash (NL2Bash) code generation benchmarks. For Greedy and Sampling (temperature τ=0.3\tau = 0.3), the average performance of 125 candidate programs per task is reported. For MBR-EXEC, 25 candidates sampled at τ=0.3\tau = 0.3 are reranked, averaged across 5 runs.

    Method MBPP (Acc) Spider (Acc) NL2Bash (char-BLEU)
    Greedy (3-shot) 47.3 ±\pm 2.5 50.8 ±\pm 2.6 52.8 ±\pm 2.9
    Sample (3-shot) 47.7 ±\pm 1.5 48.5 ±\pm 2.6 53.0 ±\pm 2.9
    MBR-EXEC 58.2 ±\pm 0.3 63.6 ±\pm 0.8 58.5 ±\pm 0.3

    MBR-EXEC provides a relative improvement of +10.9+10.9 percentage points in execution accuracy on MBPP, +12.8+12.8 points on Spider, and +5.7+5.7 character-level BLEU-4 points on NL2Bash over the 3-shot greedy baseline.

  4. Knowl 4 — Degeneration of Average Log Likelihood Selection with Increasing Candidate Pools

    empirical result

    When selecting programs from a candidate set P\mathcal{P}, maximizing the average per-token log-likelihood (MALL):

    p^=arg⁡max⁡p∈P1np∑i=1nplog⁡P(wp,i∣C,wp,1,…,wp,i−1)\hat{p} = \arg\max_{p \in \mathcal{P}} \frac{1}{n_p} \sum_{i=1}^{n_p} \log P(w_{p,i} \mid C, w_{p,1}, \dots, w_{p,i-1})

    causes performance to deteriorate as the candidate sample size increases from 20 to 120 samples across MBPP, Spider, and NL2Bash. This degradation occurs because MALL inherently assigns high normalized probabilities to degenerate sequences containing repetitive patterns and loops. Expanding the candidate pool increases the likelihood that at least one such repetitive failure mode is sampled, causing MALL to select incorrect repetitive outputs. In contrast, MBR-EXEC, standard likelihood maximization (ML), and BLEU-based MBR improve monotonically with sample size.

  5. Knowl 5 — Executability Filtering vs Full Semantic Execution Consensus

    empirical result

    An ablation comparing candidate selection over all sampled programs versus candidates filtered strictly for executable status reveals that program executability accounts for a large portion of selection gains, but semantic consensus provides critical additional improvements.

    1. Filtering out candidate programs that fail runtime execution or timeout improves the execution accuracy of non-execution baselines (ML and character-level BLEU) by substantial margins across MBPP and Spider.
    2. On Spider, applying sequence likelihood maximization over executable-only candidates (executability-ML) performs comparably to or slightly exceeds MBR-EXEC.
    3. On MBPP, full execution result consensus matching (MBR-EXEC) consistently outperforms all executability-filtered baselines across all sample set sizes (20 to 120 samples).
  6. Knowl 6 — Impact of Sampling Temperature on MBR Code Selection

    empirical result

    Evaluating MBR-EXEC and baseline selection algorithms across generative sampling temperatures τ∈[0.1,1.0]\tau \in [0.1, 1.0] demonstrates that low temperatures (tau<0.5\\tau < 0.5) yield optimal candidate pools for semantic reranking.

    Sampling with τ=0.3\tau = 0.3 produces candidate diversity that allows MBR-EXEC to outperform MBR-EXEC applied over greedily decoded sequences (e.g., 58.2%58.2\% vs 56.0%56.0\% on MBPP, and 63.6%63.6\% vs 62.1%62.1\% on Spider). However, as τ→1.0\tau \to 1.0, code generation accuracy drops sharply across all selection methods on MBPP and Spider due to an excessive proportion of syntactically and semantically invalid candidate programs.

  7. Knowl 7 — Few-Shot Prompt Group Ensembling vs Single Long Contexts

    empirical result

    Generating candidate programs using multiple distinct few-shot prompt groupings yields higher selection accuracy than concatenating all available demonstrations into a single long prompt.

    Sampling candidates across 5 separate 3-shot prompt variations (subsets of 15 available training exemplars) and pooling them for MBR-EXEC decoding achieves higher execution accuracy on MBPP (58.2%58.2\% vs ≈56.5%\approx 56.5\%) and higher character-BLEU on NL2Bash (58.558.5 vs ≈57.5\approx 57.5) than sampling the identical total number of candidates conditioned on a single prompt containing all 15 demonstrations concatenated together.

  8. Knowl 8 — Text-Based Minimum Bayes Risk Reranking (MBR-BLEU)

    model/method

    When code cannot be executed directly or when valid test inputs are unavailable, surface-form Minimum Bayes Risk decoding using negative BLEU loss (MBR-BLEU) provides an effective execution-free selection alternative. The discrepancy between two candidate programs pip_i and pjp_j is computed as:

    ℓBLEU(pi,pj)=−BLEU(pi,pj)\ell_{\text{BLEU}}(p_i, p_j) = -\text{BLEU}(p_i, p_j)

    where BLEU(⋅,⋅)\text{BLEU}(\cdot, \cdot) is token-level BLEU-4 or character-level BLEU-4 computed pairwise across the candidate pool P\mathcal{P}. By selecting the candidate that maximizes average BLEU overlap with all other candidates, MBR-BLEU consistently outperforms standard sequence likelihood maximization (ML) and average log-likelihood (MALL) across MBPP, Spider, and NL2Bash without requiring sandbox execution.

  9. Knowl 9 — Candidate Generation Oracle Bounds vs Supervised Models

    empirical result

    The theoretical upper-bound performance of candidate pools sampled from pretrained code models (measured via Expected Pass@K\text{Pass}@K, where a problem is considered solved if at least one generated candidate passes all test cases) substantially exceeds existing supervised and fine-tuned models on MBPP, Spider, and NL2Bash.

    As sample size KK approaches 100:

    1. MBPP Expected Pass@K\text{Pass}@K reaches ≈78%\approx 78\%, exceeding the supervised fine-tuned model performance of ≈60%\approx 60\%.
    2. Spider Expected Pass@K\text{Pass}@K reaches ≈84%\approx 84\%, exceeding the state-of-the-art supervised PICARD model at ≈71%\approx 71\%.
    3. NL2Bash Expected Pass@K\text{Pass}@K reaches ≈74\approx 74 character-BLEU, exceeding fine-tuned GPT-2 at ≈58\approx 58.

    This gap indicates that pretrained code models already possess the capacity to generate correct code within small sample budgets, and the primary bottleneck lies in candidate selection mechanisms.

  10. Knowl 10 — Limitations of Post-Hoc Selection on Frozen Code Models

    limitation

    MBR-EXEC operates entirely as an inference-time decoding algorithm on top of fixed, frozen language model generations (such as OpenAI Codex). It does not integrate execution feedback or execution traces into the model parameters, pretraining objectives, or fine-tuning stages. Additionally, when applied to languages without standard execution environments (such as Bash), simulated execution via syntactic parsing (bashlex) must be substituted, which approximates but does not strictly guarantee semantic execution equivalence.

Coverage note — Deliberately omitted the specific prompt markup syntax examples and dataset-specific few-shot examples (Appendix Tables 4–6) as they represent standard formatting details rather than distinct methodological contributions.

References

  1. 1.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732.
  2. 2.Peter J Bickel and Kjell A Doksum. 1977. Mathematical statistics: basic ideas and selected topics, volumes I-II package. HoldenDay Inc., Oakland, CA, USA.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  5. 5.Li Dong and Mirella Lapata. 2018. Coarse-to-fine decoding for neural semantic parsing. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 731–742, Melbourne, Australia. Association for Computational Linguistics.
  6. 6.Bryan Eikema and Wilker Aziz. 2020. Is MAP decoding all you need? the inadequacy of the mode in neural machine translation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4506–4520, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  7. 7.Bryan Eikema and Wilker Aziz. 2021. Sampling-based minimum bayes risk decoding for neural machine translation. arXiv preprint arXiv:2108.04718.
  8. 8.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  9. 9.Vaibhava Goel and William J Byrne. 2000. Minimum bayes-risk automatic speech recognition. Computer Speech & Language, 14(2):115–135.
  10. 10.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring coding challenge competence with APPS. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  11. 11.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438.
  12. 12.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. UNIFIEDQA: Crossing format boundaries with a single QA system. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907, Online. Association for Computational Linguistics.
  13. 13.Shankar Kumar and William Byrne. 2004. Minimum Bayes-risk decoding for statistical machine translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 169–176, Boston, Massachusetts, USA. Association for Computational Linguistics.
  14. 14.Marie-Anne Lachaux, Baptiste Roziere, Marc Szafraniec, and Guillaume Lample. 2021. Dobf: A deobfuscation pre-training objective for programming languages. Advances in Neural Information Processing Systems, 34.
  15. 15.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  16. 16.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  17. 17.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022. Competition-level code generation with alphacode. arXiv preprint arXiv:2203.07814.
  18. 18.Xi Victoria Lin, Chenglong Wang, Luke Zettlemoyer, and Michael D. Ernst. 2018. NL2Bash: A corpus and semantic parser for natural language interface to the linux operating system. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  19. 19.Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, Fumin Wang, and Andrew Senior. 2016. Latent predictor networks for code generation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 599–609, Berlin, Germany. Association for Computational Linguistics.
  20. 20.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pretrain, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  21. 21.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, MING GONG, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie LIU. 2021. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  22. 22.Antonio Valerio Miceli Barone and Rico Sennrich. 2017. A parallel corpus of python functions and documentation strings for automated code documentation and code generation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 314–319, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  23. 23.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021. Noisy channel language model prompting for few-shot text classification. arXiv preprint arXiv:2108.04106.
  24. 24.Maxim Rabinovich, Mitchell Stern, and Dan Klein. 2017. Abstract syntax networks for code generation and semantic parsing. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1139–1149, Vancouver, Canada. Association for Computational Linguistics.
  25. 25.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  26. 26.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9895–9901, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  27. 27.Haoyue Shi, Jiayuan Mao, Kevin Gimpel, and Karen Livescu. 2019. Visually grounded neural syntax acquisition. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1842–1861, Florence, Italy. Association for Computational Linguistics.
  28. 28.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  29. 29.Alane Suhr, Srinivasan Iyer, and Yoav Artzi. 2018. Learning to map context-dependent sentences to executable formal queries. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2238–2249, New Orleans, Louisiana. Association for Computational Linguistics.
  30. 30.Ivan Titov and James Henderson. 2006. Bayes risk minimization in natural language parsing. University of Geneva technical report.
  31. 31.Roy Tromble, Shankar Kumar, Franz Och, and Wolfgang Macherey. 2008. Lattice Minimum Bayes-Risk decoding for statistical machine translation. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, pages 620–629, Honolulu, Hawaii. Association for Computational Linguistics.
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  33. 33.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  34. 34.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  35. 35.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural text generation with unlikelihood training. In International Conference on Learning Representations.
  36. 36.Chunyang Xiao, Marc Dymetman, and Claire Gardent. 2016. Sequence-based structured prediction for semantic parsing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1341–1350, Berlin, Germany. Association for Computational Linguistics.
  37. 37.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I Wang, et al. 2022. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. arXiv preprint arXiv:2201.05966.
  38. 38.Frank F. Xu, Zhengbao Jiang, Pengcheng Yin, Bogdan Vasilescu, and Graham Neubig. 2020. Incorporating external knowledge through pre-training for natural language to code generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6045–6052, Online. Association for Computational Linguistics.
  39. 39.Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig. 2018. Learning to mine aligned code and natural language pairs from stack overflow. In 2018 IEEE/ACM 15th international conference on mining software repositories (MSR), pages 476–486. IEEE.
  40. 40.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921, Brussels, Belgium. Association for Computational Linguistics.
  41. 41.Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTScore: Evaluating generated text as text generation. In Advances in Neural Information Processing Systems.
  42. 42.Hao Zhang and Daniel Gildea. 2008. Efficient multipass decoding for synchronous context free grammars. In Proceedings of ACL-08: HLT, pages 209–217, Columbus, Ohio. Association for Computational Linguistics.
  43. 43.Yu Zhang, Zhenghua Li, and Min Zhang. 2020. Efficient second-order TreeCRF for neural dependency parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3295–3305, Online. Association for Computational Linguistics.
  44. 44.Ruiqi Zhong, Tao Yu, and Dan Klein. 2020. Semantic evaluation for text-to-SQL with distilled test suites. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 396–411, Online. Association for Computational Linguistics.
  45. 45.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [MASK]: Learning vs. learning to recall. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5017–5033, Online. Association for Computational Linguistics.

Citation

MLA
Shi, F., et al. “Natural Language to Code Translation with Execution”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3533–46, https://doi.org/10.18653/v1/2022.emnlp-main.231.
APA
Shi, F., Fried, D., Ghazvininejad, M., Zettlemoyer, L., & Wang, S. I. (2022). Natural Language to Code Translation with Execution. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3533–3546. https://doi.org/10.18653/v1/2022.emnlp-main.231
Chicago
Shi, F., D. Fried, M. Ghazvininejad, L. Zettlemoyer, and S. I. Wang. 2022. “Natural Language to Code Translation with Execution”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3533–46. https://doi.org/10.18653/v1/2022.emnlp-main.231.
Harvard
Shi, F. et al. (2022) “Natural Language to Code Translation with Execution”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3533–3546. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.231.
Vancouver
1. Shi F, Fried D, Ghazvininejad M, Zettlemoyer L, Wang SI (2022) Natural Language to Code Translation with Execution. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3533–3546

BibTeX

@inproceedings{shi-etal-2022-natural,
    title = "Natural Language to Code Translation with Execution",
    author = "Shi, Freda  and
      Fried, Daniel  and
      Ghazvininejad, Marjan  and
      Zettlemoyer, Luke  and
      Wang, Sida I.",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.231/",
    doi = "10.18653/v1/2022.emnlp-main.231",
    pages = "3533--3546"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/