Large Language Models Meet NL2Code: A Survey

Daoguang ZanBei ChenFengji ZhangDianjie LuBingchao WuBei GuanYongji WangJian-Guang Lou

article2023ACL286 citations

Presents a comprehensive survey of 27 code-generation large language models alongside their evaluation benchmarks and metrics, comparing their performance on HumanEval and identifying model scale, data quality, and expert tuning as the primary drivers of success.

Listen

Generating functional computer code directly from natural language requirements represents a significant milestone for artificial intelligence, with the potential to democratize software development and substantially improve engineering productivity. The emergence of Transformer-based large language models has accelerated progress in this space, transitioning code generation from early rule-based or statistical techniques to highly capable neural systems. Understanding the architectural patterns, data characteristics, and tuning strategies driving these models is essential for organizations planning technical investments or evaluating commercial programming assistants.

The main objective of the article is to provide a comprehensive evaluation of the state of natural language-to-code technology by systematically analyzing 27 representative large language models, surveying 17 key evaluation benchmarks, and identifying the primary factors behind model success along with remaining capability gaps.

To conduct this evaluation, the researchers reviewed 27 prominent models ranging in scale from under 100 million to 540 billion parameters, covering both encoder-decoder and decoder-only architectures developed through late 2022. They conducted standardized, zero-shot comparative evaluations on leading industry benchmarks, primarily HumanEval (comprising 164 hand-crafted Python tasks) and MBPP (comprising 974 programming exercises), measuring the percentage of problems successfully solved across single and multiple generation attempts.

The analysis reveals four central findings. First, model performance depends on a core formula: large model and corpus size, premium training data, and expert hyperparameter tuning. Second, scale strongly drives performance and syntactic correctness; for instance, expanding model capacity dramatically drops code syntax error rates down to around 6% in top models, although generating semantically correct logic remains difficult. Third, general-purpose text scale alone does not suffice without high-quality, code-specific training; large general text models frequently underperform smaller models trained on curated, deduplicated code repositories. Fourth, context window capacity significantly amplifies capability, enabling a smaller 165-million parameter model paired with an 8,000-token context window to achieve performance comparable to a 20-billion parameter model constrained to a 2,000-token context window.

These findings indicate that while modern code assistants can deliver immediate developer efficiency gains, automated systems cannot yet fully replace human software engineers. Commercial deployment still carries operational and safety risks, as models often introduce subtle bugs, struggle with complex multi-condition problem requirements, and exhibit prompt sensitivity. Additionally, because models default to generating plausible-looking code even when an exact solution is infeasible, organizations face code security, maintainability, and review overheads if generated outputs are integrated without rigorous human validation.

Organizations adopting or developing code generation systems should implement structured code verification workflows rather than relying solely on raw model generation. Teams training custom models should prioritize high-grade dataset filtering, custom code-specific tokenizers, and larger context window configurations rather than focusing exclusively on parameter count. Future development must focus on bridging the remaining human-machine gaps, particularly self-validation, multi-step problem decomposition, code explanation, and retrieval-based mechanisms that incorporate dynamic documentation and Application Programming Interface updates.

Readers should interpret the comparative performance metrics with standard caveats regarding data availability and benchmark constraints. A completely controlled comparison remains difficult because several leading commercial models are closed-source and train on disparate data corpora. Furthermore, common benchmarks predominantly focus on standalone Python functions evaluated via execution test suites, providing limited insight into large-scale multi-file system architectures, software security vulnerabilities, or long-term maintainability.

Cover for Large Language Models Meet NL2Code: A Survey

Abstract

The task of generating code from a natural language description, or NL2Code, is considered a pressing and significant challenge in code intelligence. Thanks to the rapid development of pre-training techniques, surging large language models are being proposed for code, sparking the advances in NL2Code. To facilitate further research and applications in this field, in this paper, we present a comprehensive survey of 27 existing large language models for NL2Code, and also review benchmarks and metrics. We provide an intuitive comparison of all existing models on the HumanEval benchmark. Through in-depth observation and analysis, we provide some insights and conclude that the key factors contributing to the success of large language models for NL2Code are "Large Size, Premium Data, Expert Tuning". In addition, we discuss challenges and opportunities regarding the gap between models and humans. We also create a website https://nl2code.github.io to track the latest progress through crowd-sourcing. To the best of our knowledge, this is the first survey of large language models for NL2Code, and we believe it will contribute to the ongoing development of the field.

Table of Contents

  • 1 Introduction
  • 2 Large Language Models for NL2Code
  • 3 What makes LLMs successful?
  • 3.1 Large Model Size
  • 3.2 Large and Premium Data
  • 3.3 Expert Tuning
  • 4 Benchmarks and Metrics
  • 5 Challenges and Opportunities
  • 6 Conclusion
  • Limitations
  • References
  • A Related Surveys
  • B An Online Website
  • C Experimental Setup
  • C.1 Definition of pass@ k
  • C.2 Implementation Details
  • D Context Window vs. Performance

Knowls

  1. Knowl 1 — Key Factors Driving the Success of Code LLMs: Large Size, Premium Data, and Expert Tuning

    empirical result

    An extensive comparative analysis of 27 large language models (LLMs) for natural-language-to-code (NL2Code) synthesis reveals three decisive factors governing model effectiveness:

    1. Large Model and Data Size: Scaling parameter counts (from millions to hundreds of billions) and training token counts monotonically improves code generation accuracy (measured by pass@k\text{pass@}k) across benchmarks such as HumanEval and MBPP.
    2. Premium Data: Raw code scraped from open-source repositories (such as GitHub and Stack Overflow) requires rigorous pre-processing and curation. Effective pipelines filter out auto-generated code, incomplete snippets, and low-quality files using heuristics such as repository star counts, file size thresholds, maximum line length limits, and alphanumeric character ratios.
    3. Expert Tuning: Critical hyperparameter choices include subword tokenization optimized specifically for code (Byte-level BPE or SentencePiece), learning rates inversely scaled with model size, optimization via Adam or AdamW with decoupled weight decay, expanding context window length, and inference temperature calibration tailored to candidate generation budgets.
  2. Knowl 2 — Unbiased pass@k Metric Formulation for Execution-Based Code Evaluation

    equation

    Execution-based evaluation of generated code against unit test suites uses the unbiased estimator pass@k\text{pass@}k. For each programming problem, nn candidate code completions are sampled from the language model (n≥kn \ge k), and the problem is considered solved if at least one generated candidate in a randomly selected subset of size kk passes all unit tests.

    The unbiased pass@k\text{pass@}k estimator is calculated as:

    pass@k={1if n−c<k1−(n−ck)(nk)=1−∏i=n−c+1n(1−ki)otherwise\text{pass@}k = \begin{cases} 1 & \text{if } n - c < k \\ 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}} = 1 - \prod_{i=n-c+1}^{n} \left(1 - \frac{k}{i}\right) & \text{otherwise} \end{cases}

    where:

    • n∈N+n \in \mathbb{N}^+ is the total number of candidate code solutions generated per problem (n≥kn \ge k),
    • c∈{0,1,…,n}c \in \{0, 1, \dots, n\} is the number of generated candidates that pass all unit test cases, and
    • k∈N+k \in \mathbb{N}^+ is the evaluation budget (number of attempts considered).
  3. Knowl 3 — Comparative Performance of Code LLMs on the HumanEval Benchmark

    data/table

    Evaluation on the HumanEval benchmark (164 hand-written Python programming problems) in a zero-shot setting demonstrates how pass@k\text{pass@}k metrics scale across model sizes and pre-training objectives:

    Model Size pass@1 (%) pass@10 (%) pass@100 (%)
    Model Size: ∼\sim100M
    GPT-Neo 125M 0.75 1.88 2.97
    CodeParrot 110M 3.80 6.57 12.78
    PyCodeGPT 110M 8.33 13.36 19.13
    Codex 85M 8.22 12.81 22.40
    Model Size: ∼\sim500M
    BLOOM 560M 0.82 3.02 5.91
    CodeT5 770M 12.09 19.24 30.93
    CodeGen-Mono 350M 12.76 23.11 35.19
    PanGu-Coder 317M 17.07 24.05 34.55
    Codex 679M 16.22 25.70 40.95
    Model Size: ∼\sim1B
    BLOOM 1.1B 2.48 5.93 9.62
    InCoder 1.3B 11.09 16.14 24.20
    AlphaCode (decoder) 1.1B 17.10 28.20 45.30
    SantaCoder 1.1B 18.00 29.00 49.00
    Model Size: ∼\sim5B
    PolyCoder 2.7B 5.59 9.84 17.68
    GPT-J 6B 11.62 15.74 27.74
    InCoder 6.7B 15.20 27.80 47.00
    Codex 2.5B 21.36 35.42 59.50
    CodeGen-Mono 6.1B 26.13 42.29 65.82
    PanGu-Coder 2.6B 23.78 35.36 51.24
    Model Size: >>10B
    BLOOM 176B 15.52 32.20 55.45
    GPT-NeoX 20B 15.40 25.60 41.20
    Codex 12B 28.81 46.81 72.31
    CodeGen-Mono 16.1B 29.28 49.86 75.00
    PaLM-Coder 540B 36.00 – 88.40
    code-davinci-002 – 47.00 74.90 92.10

    Models specialized on code corpora (such as Codex, CodeGen-Mono, PanGu-Coder, and SantaCoder) substantially outperform general natural-language models of equal or larger parameter size (such as GPT-Neo, GPT-J, and BLOOM). Scaling parameter size improves performance across every model family, while high-quality code filtering enables compact models (e.g., PyCodeGPT 110M and SantaCoder 1.1B) to match or exceed much larger general-purpose baselines.

  4. Knowl 4 — Syntax Error Rate Reduction Versus Semantic Correctness Bottleneck

    empirical result

    Evaluating code completions generated by language models on the HumanEval benchmark reveals a divergence between syntactic validity and semantic execution correctness as model parameters scale:

    • Syntax Error Rate: As parameter count increases, the percentage of generated code samples with syntax errors decreases sharply. For example, CodeGen-Mono 16.1B achieves a syntax error rate of approximately 6%, whereas smaller or general-purpose models (such as GPT-Neo 125M and BLOOM 560M) display syntax error rates between 25% and 40%.
    • Semantic Correctness: Despite achieving low syntax error rates, CodeGen-Mono 16.1B attains a pass@1\text{pass@1} score of only 29.28% on unit test execution.

    This gap indicates that large language models readily master the formal grammar and syntax of programming languages, but generating semantically and logically correct program logic that passes multi-condition unit test suites remains the primary performance bottleneck.

  5. Knowl 5 — Context Window Length Impact on APPS Benchmark Performance

    empirical result

    Expanding the context window during pre-training significantly enhances problem-solving competence on complex natural-language-to-code tasks. On the APPS benchmark across introductory, interview, and competition-level programming problems:

    • A compact GPT-NeoX model with 165 million parameters (165M) trained with an 8,000-token (8K) context window achieves average passed test case percentages comparable to a model with over 100 times more parameters (GPT-NeoX 20B) trained with a 2,000-token (2K) context window.
    • For GPT-NeoX 165M, expanding the context window from 2K to 4K and 8K yields monotonic performance improvements across all problem difficulty tiers.

    This demonstrates that context window capacity is a critical hyperparameter that can compensate for smaller parameter counts when processing long problem descriptions and complex program context.

  6. Knowl 6 — Sampling Temperature Trade-Off in Zero-Shot Code Generation

    empirical result

    In zero-shot code generation using large language models, the inference sampling temperature governs a trade-off between top-1 precision and candidate diversity:

    • Lower temperatures (≈0.1−0.2\approx 0.1 - 0.2): Produce less stochastic, higher-probability token sequences, maximizing single-attempt accuracy (pass@1\text{pass@1}).
    • Higher temperatures (≈0.7−0.8\approx 0.7 - 0.8): Induce higher diversity across candidate completions. While increased randomness reduces single-attempt precision (pass@1\text{pass@1}), it raises the probability that at least one valid implementation exists within a large candidate pool of 100 samples (pass@100\text{pass@100}).

    Empirical evaluation on the HumanEval benchmark using CodeGen-Mono 2.7B and InCoder 1.3B shows that increasing temperature from 0.2 to 0.8 monotonically decreases pass@1\text{pass@1} while monotonically increasing pass@100\text{pass@100}.

  7. Knowl 7 — Architectural and Objective Taxonomy of Code LLMs

    model/method

    Large language models for natural-language-to-code synthesis fall into two primary architectural paradigms:

    1. Decoder-Only Architectures: Constitute the majority of modern large-scale code models (e.g., Codex, CodeGen, PaLM-Coder, PanGu-Coder, SantaCoder, CodeGeeX, InCoder, GPT-NeoX, BLOOM). These models typically employ causal auto-regressive language modeling to predict tokens sequentially from left to right.
    2. Encoder-Decoder Architectures: (e.g., AlphaCode, CodeT5, PLBART, PyMT5, CodeRL, ERNIE-Code). These encode the natural language specification via a bidirectional Transformer encoder and generate code via an auto-regressive decoder.
    3. Fill-In-The-Middle (FIM) Bidirectional Infilling: Models such as InCoder, SantaCoder, and FIM extend causal language modeling by randomly relocating middle document spans to suffix positions during training. This enables the model to perform both left-to-right code generation and arbitrary code infilling conditioned on surrounding prefix and suffix context.

    Initializing code LLMs with pre-trained general natural language model checkpoints accelerates training convergence, but does not provide noticeable gains in final zero-shot code generation accuracy compared to pre-training from scratch on clean code corpora.

  8. Knowl 8 — Pre-training Hyperparameter Scaling and Tokenization for Code LLMs

    empirical result

    Systematic review of pre-training configurations across 27 code LLMs reveals standard hyperparameter scaling relationships:

    • Learning Rate Scaling: Peak learning rates scale inversely with model parameter count. Small models (∼100M\sim 100\text{M} parameters, e.g., PyCodeGPT 110M) use learning rates around 5×10−45 \times 10^{-4} to 1.5×10−31.5 \times 10^{-3}, whereas models with billions of parameters (≥10B\ge 10\text{B} parameters, e.g., CodeGen 16.1B and Codex 12B) scale down to 0.5×10−4−1.0×10−40.5 \times 10^{-4} - 1.0 \times 10^{-4}.
    • Optimizers and Schedules: Optimization universally relies on Adam (β1=0.9,β2=0.95\beta_1 = 0.9, \beta_2 = 0.95) or AdamW (β1=0.9,β2=0.999\beta_1 = 0.9, \beta_2 = 0.999, weight decay 0.01−0.10.01 - 0.1), paired with cosine or linear learning rate decay schedules following linear warmup phases (175175 to 5,0005,000 steps).
    • Code-Specific Tokenization: Standard natural language tokenizers split code formatting and indentation sub-optimally. Effective code LLMs train custom subword tokenizers directly on code corpora using Byte-level Byte-Pair-Encoding (BBPE) or SentencePiece (SP) to preserve whitespace and syntax structures efficiently.
  9. Knowl 9 — Taxonomy of NL2Code Benchmarks Across Application Scenarios

    data/table

    NL2Code benchmarks span diverse programmatic domains, context lengths, and execution requirements:

    Benchmark Instances Prompt NL Code PL Avg Tests Scenario
    HumanEval 164 English Python 7.8 Code Exercise
    MBPP 974 English Python 3.1 Code Exercise
    APPS 5,000 English Python 21.0 Competitions
    CodeContests 165 English Multi 203.7 Competitions
    DS-1000 1,000 English Python 1.6 Data Science
    DSP 1,119 English Python 2.1 Data Science
    MBXP 974 / lang English Multi (13 PLs) 3.1 Multilingual
    HumanEval-X 164 / lang English Multi (5 PLs) 7.8 Multilingual
    MultiPL-HumanEval 164 / lang English Multi (18 PLs) 7.8 Multilingual
    MultiPL-MBPP 974 / lang English Multi (18 PLs) 3.1 Multilingual
    PandasEval / NumpyEval 101 each English Python 6.5 / 3.5 Public Library
    TorchDataEval 50 English Python 1.1 Private Library
    MTPB 115 English Python – Multi-Turn
    ODEX 945 Multi (4 NLs) Python 1.8 Open-Domain

    Reliable evaluation requires automated unit test execution (e.g., pass@k\text{pass@}k, n@kn\text{@}k) with high branch and statement coverage. Non-execution surface matching metrics (such as BLEU and CodeBLEU) fail to reflect the functional correctness and runtime behavior of generated programs.

  10. Knowl 10 — Human-Model Ability Gaps in Natural Language Code Generation

    limitation

    Five fundamental capability gaps persist between large language models and human software developers in NL2Code synthesis:

    1. Understanding Ability: LLMs are sensitive to minor prompt rewordings and struggle when decomposing multi-constraint problem statements into sequential reasoning steps.
    2. Judgement and Calibration: Because LLMs are trained under unsupervised causal language modeling objectives, they output candidate code even when problem specifications are impossible or contradictory, lacking human-like boundary awareness and failure recognition.
    3. Explanation and Self-Validation: While human programmers can explain code rationale and debug through active analysis, LLMs lack grounded self-verification unless augmented with external testing loops.
    4. Adaptive Learning and Dynamic Knowledge: Humans rapidly assimilate new software documentation and modified third-party APIs. LLMs cannot update their static weights without expensive retraining or fine-tuning, requiring retrieval-augmented generation (e.g., DocCoder, APICoder) as a workaround.
    5. Multi-tasking and Multi-lingual Mastery: Human engineers transfer knowledge across programming languages and switch flexibly across software lifecycle tasks (review, debug, refactor), whereas LLMs often suffer performance degradation across diverse languages and require complex prompt engineering.

Coverage note — All primary contributions—including the survey taxonomy of 27 code LLMs, the success factors formula (Large Size, Premium Data, Expert Tuning), pass@k metric definition, HumanEval and MBPP empirical comparisons, syntax vs. semantic error analyses, context window scaling experiments, temperature trade-offs, pre-training hyperparameters, benchmark taxonomies, and human-model capability gap analyses—have been extracted. Lists of commercial IDE extension plugins and general historical citations were omitted as non-load-bearing survey metadata.

References

  1. 1.Rajas Agashe, Srinivasan Iyer, and Luke Zettlemoyer. 2019. JuICe: A large scale distantly supervised dataset for open domain context-based code generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5436–5446.
  2. 2.Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified pre-training for program understanding and generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2668.
  3. 3.aiXcoder. 2018. aiXcoder. https://aixcoder.com.
  4. 4.Alibaba. 2022. Alibaba. https://github.com/alibaba-cloud-toolkit/cosy.
  5. 5.Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Muñoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alexander Gu, Manan Dey, Logesh Kumar Umapathi, Carolyn Jane Anderson, Yangtian Zi, J. Poirier, Hailey Schoelkopf, Sergey Mikhailovich Troshin, Dmitry Abulkhanov, Manuel Romero, Michael Franz Lappert, Francesco De Toni, Bernardo Garc’ia del R’io, Qian Liu, Shamik Bose, Urvashi Bhattacharyya, Terry Yue Zhuo, Ian Yu, Paulo Villegas, Marco Zocca, Sourab Mangrulkar, David Lansky, Huu Nguyen, Danish Contractor, Luisa Villa, Jia Li, Dzmitry Bahdanau, Yacine Jernite, Sean Christopher Hughes, Daniel Fried, Arjun Guha, Harm de Vries, and Leandro von Werra. 2023. SantaCoder: don’t reach for the stars! ArXiv, abs/2301.03988.
  6. 6.Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2018. A survey of machine learning for big code and naturalness. ACM Computing Surveys (CSUR), 51(4):1–37.
  7. 7.Miltiadis Allamanis and Charles Sutton. 2014. Mining idioms from source code. In Proceedings of the 22nd acm sigsoft international symposium on foundations of software engineering, pages 472–483.
  8. 8.Amazon. 2022. CodeWhisperer. https://aws.amazon.com/cn/codewhisperer.
  9. 9.Anonymous. 2022. CodeT5Mix: A pretrained mixture of encoder-decoder transformers for code understanding and generation. In Submitted to The Eleventh International Conference on Learning Representations. Under review.
  10. 10.Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, Sujan Kumar Gonugondla, Hantian Ding, Varun Kumar, Nathan Fulton, Arash Farahani, Siddharth Jain, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Sudipta Sengupta, Dan Roth, and Bing Xiang. 2022. Multi-lingual evaluation of code generation models. ArXiv, abs/2210.14868.
  11. 11.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. Program synthesis with large language models. ArXiv, abs/2108.07732.
  12. 12.Shraddha Barke, Michael B. James, and Nadia Polikarpova. 2022. Grounded Copilot: How programmers interact with code-generating models. Proceedings of the ACM on Programming Languages, 7:85 – 111.
  13. 13.Shraddha Barke, Michael B James, and Nadia Polikarpova. 2023. Grounded copilot: How programmers interact with code-generating models. Proceedings of the ACM on Programming Languages, 7(OOPSLA1):85–111.
  14. 14.Mohammad Bavarian, Heewoo Jun, Nikolas A. Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. 2022. Efficient training of language models to fill in the middle. ArXiv, abs/2207.14255.
  15. 15.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of the ACL Workshop on Challenges & Perspectives in Creating Large Language Models.
  16. 16.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  17. 17.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. Neural Information Processing Systems, 33:1877–1901.
  18. 18.Federico Cassano, John Gouwar, Daniel Nguyen, Sy Duy Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q. Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda. 2022. A scalable and extensible approach to benchmarking nl2code for 18 programming languages. ArXiv, abs/2208.08227.
  19. 19.Yekun Chai, Shuohuan Wang, Chao Pang, Yu Sun, Hao Tian, and Hua Wu. 2022. ERNIE-Code: Beyond english-centric cross-lingual pretraining for programming languages. arXiv preprint arXiv:2212.06742.
  20. 20.Shubham Chandel, Colin B. Clement, Guillermo Serrato, and Neel Sundaresan. 2022a. Training and evaluating a jupyter notebook data science assistant. ArXiv, abs/2201.12901.
  21. 21.Shubham Chandel, Colin B Clement, Guillermo Serrato, and Neel Sundaresan. 2022b. Training and evaluating a jupyter notebook data science assistant. arXiv preprint arXiv:2201.12901.
  22. 22.Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2023. CodeT: Code generation with generated tests. In The Eleventh International Conference on Learning Representations.
  23. 23.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harrison Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, David W. Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William H. Guss, Alex Nichol, Igor Babuschkin, S. Arun Balaji, Shantanu Jain, Andrew Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew M. Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code. ArXiv, abs/2107.03374.
  24. 24.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. ArXiv, abs/2204.02311.
  25. 25.Fenia Christopoulou, Gerasimos Lampouras, Milan Gritta, Guchun Zhang, Yinpeng Guo, Zhong-Yi Li, Qi Zhang, Meng Xiao, Bo Shen, Lin Li, Hao Yu, Li yu Yan, Pingyi Zhou, Xin Wang, Yu Ma, Ignacio Iacobacci, Yasheng Wang, Guangtai Liang, Jia Wei, Xin Jiang, Qianxiang Wang, and Qun Liu. 2022. PanGu-Coder: Program synthesis with function-level language modeling. ArXiv, abs/2207.11280.
  26. 26.Colin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan. 2020. PyMT5: Multi-mode translation of natural language and python code with transformers. In Conference on Empirical Methods in Natural Language Processing.
  27. 27.CodedotAl. 2021. GPT Code Clippy: The Open Source version of GitHub Copilot. https://github.com/CodedotAl/gpt-code-clippy.
  28. 28.Trevor Cohn, Phil Blunsom, and Sharon Goldwater. 2010. Inducing tree-substitution grammars. The Journal of Machine Learning Research, 11:3053–3096.
  29. 29.Leonardo Mendonça de Moura and Nikolaj S. Bjørner. 2008. Z3: An efficient smt solver. In International Conference on Tools and Algorithms for Construction and Analysis of Systems.
  30. 30.DeepGenX. 2022. CodeGenX. https://docs.deepgenx.com.
  31. 31.Enrique Dehaerne, Bappaditya Dey, Sandip Halder, Stefan De Gendt, and Wannes Meert. 2022. Code generation using machine learning: A systematic review. IEEE Access.
  32. 32.Premkumar T. Devanbu. 2012. On the naturalness of software. 2012 34th International Conference on Software Engineering (ICSE), pages 837–847.
  33. 33.Iddo Drori and Nakul Verma. 2021. Solving linear algebra by program synthesis. arXiv preprint arXiv:2111.08171.
  34. 34.Iddo Drori, Sarah Zhang, Reece Shuttleworth, Leonard Tang, Albert Lu, Elizabeth Ke, Kevin Liu, Linda Chen, Sunny Tran, Newman Cheng, Roman Wang, Nikhil Singh, Taylor Lee Patti, J. Lynch, Avi Shporer, Nakul Verma, Eugene Wu, and Gilbert Strang. 2021. A neural network solves, explains, and generates university math problems by program synthesis and few-shot learning at human level. Proceedings of the National Academy of Sciences of the United States of America, 119.
  35. 35.Akiko Eriguchi, Kazuma Hashimoto, and Yoshimasa Tsuruoka. 2016. Tree-to-sequence attentional neural machine translation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 823–833.
  36. 36.erman Arsenovich Arutyunov and Sergey Avdoshin. 2022. Big transformers for code generation. Proceedings of the Institute for System Programming of the RAS.
  37. 37.FauxPilot. 2022. FauxPilot. https://github.com/moyix/fauxpilot.
  38. 38.Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. InCoder: A generative model for code infilling and synthesis. In The Eleventh International Conference on Learning Representations.
  39. 39.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  40. 40.GitHub. 2021. GitHub Copilot. https://github.com/features/copilot.
  41. 41.Google. 2016. GitHub on BigQuery: Analyze all the open source code. https://cloud.google.com/bigquery.
  42. 42.Google. 2022. Big-bench. https://github.com/google/BIG-bench.
  43. 43.Sumit Gulwani. 2010. Dimensions in program synthesis. In Proceedings of the 12th international ACM SIGPLAN symposium on Principles and practice of declarative programming, pages 13–24.
  44. 44.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Xiaodong Song, and Jacob Steinhardt. 2021. Measuring coding challenge competence with apps. In Neural Information Processing Systems.
  45. 45.Glen M Hocky and Andrew D White. 2022. Natural language processing models that automate programming will transform chemistry research and teaching. Digital discovery, 1(2):79–83.
  46. 46.HuggingFace. 2021a. CodeParrot Dataset. https://huggingface.co/datasets/transformersbook/codeparrot.
  47. 47.HuggingFace. 2021b. Github-Code. https://huggingface.co/datasets/codeparrot/github-code.
  48. 48.HuggingFace. 2021c. GitHub-Jupyter. https://huggingface.co/datasets/codeparrot/github-jupyter.
  49. 49.Huggingface. 2021. Training CodeParrot from Scratch. https://huggingface.co/blog/codeparrot.
  50. 50.HuggingFace. 2022. The Stack. https://huggingface.co/datasets/bigcode/the-stack.
  51. 51.Hamel Husain, Hongqi Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. CodeSearchNet Challenge: Evaluating the state of semantic code search. ArXiv, abs/1909.09436.
  52. 52.IBM. 2021. CodeNet. https://github.com/IBM/Project_CodeNet.
  53. 53.Saki Imai. 2022. Is github copilot a substitute for human pair-programming? an empirical study. In Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings, pages 319–321.
  54. 54.Srini Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016. Summarizing source code using a neural attention model. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  55. 55.Susmit Jha, Sumit Gulwani, Sanjit A. Seshia, and Ashish Tiwari. 2010. Oracle-guided component-based program synthesis. 2010 ACM/IEEE 32nd International Conference on Software Engineering, 1:215–224.
  56. 56.Aravind Joshi and Owen Rambow. 2003. A formalism for dependency grammar based on tree adjoining grammar. In Proceedings of the Conference on Meaning-text Theory, pages 207–216. MTT Paris, France.
  57. 57.Harshit Joshi, José Cambronero, Sumit Gulwani, Vu Le, Ivan Radicek, and Gust Verbruggen. 2022. Repair is nearly generation: Multilingual program repair with llms. arXiv preprint arXiv:2208.11640.
  58. 58.Sungmin Kang, Bei Chen, Shin Yoo, and Jian-Guang Lou. 2023. Explainable automated debugging via large language model-driven scientific debugging. arXiv preprint arXiv:2304.02195.
  59. 59.Darren Key, Wen-Ding Li, and Kevin Ellis. 2022. I Speak, You Verify: Toward trustworthy neural program synthesis. arXiv preprint arXiv:2210.00848.
  60. 60.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  61. 61.Mario Krenn, Qianxiang Ai, Senja Barthel, Nessa Carson, Angelo Frei, Nathan C Frey, Pascal Friederich, Théophile Gaudin, Alberto Alexander Gayle, Kevin Maik Jablonka, et al. 2022. Selfies and the future of molecular string representations. Patterns, 3(10):100588.
  62. 62.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  63. 63.Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Scott Yih, Daniel Fried, Si yi Wang, and Tao Yu. 2022. DS-1000: A natural and reliable benchmark for data science code generation. ArXiv, abs/2211.11501.
  64. 64.Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven CH Hoi. 2022. CodeRL: Mastering code generation through pretrained models and deep reinforcement learning. arXiv preprint arXiv:2207.01780, abs/2207.01780.
  65. 65.Triet HM Le, Hao Chen, and Muhammad Ali Babar. 2020. Deep learning for source code modeling and generation: Models, applications, and challenges. ACM Computing Surveys (CSUR), 53(3):1–38.
  66. 66.Yaoxian Li, Shiyi Qi, Cuiyun Gao, Yun Peng, David Lo, Zenglin Xu, and Michael R Lyu. 2022a. A closer look into transformer-based code intelligence through code transformation: Challenges and opportunities. arXiv preprint arXiv:2207.04285.
  67. 67.Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom, Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de, Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey, Cherepanov, James Molloy, Daniel Jaymin Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de, Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022b. Competition-level code generation with alphacode. Science, 378:1092 – 1097.
  68. 68.Zhiyu Li, Shuai Lu, Daya Guo, Nan Duan, Shailesh Jannu, Grant Jenks, Deep Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, et al. 2022c. Automating code review activities by large-scale pre-training. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 1035–1047.
  69. 69.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  70. 70.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pretrain, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35.
  71. 71.Zhiqiang Liu, Yong Dou, Jingfei Jiang, and Jinwei Xu. 2016. Automatic code generation of convolutional neural networks in fpga implementation. In 2016 International conference on field-programmable technology (FPT), pages 61–68. IEEE.
  72. 72.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  73. 73.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. CodeXGLUE: A machine learning benchmark dataset for code understanding and generation. ArXiv, abs/2102.04664.
  74. 74.Lechanceux Luhunu and Eugene Syriani. 2017. Survey on template-based code generation. In ACM/IEEE International Conference on Model Driven Engineering Languages and Systems.
  75. 75.Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny, Seth Bernstein, and Juho Leinonen. 2022a. Experiences from using code explanations generated by large language models in a web software development e-book. arXiv preprint arXiv:2211.02265.
  76. 76.Stephen MacNeil, Andrew Tran, Dan Mogil, Seth Bernstein, Erin Ross, and Ziheng Huang. 2022b. Generating diverse code explanations using the gpt-3 large language model. In Proceedings of the 2022 ACM Conference on International Computing Education Research-Volume 2, pages 37–39.
  77. 77.Antonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader Palacio, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota. 2021. Studying the usage of text-to-text transfer transformer to support code-related tasks. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE), pages 336–347. IEEE.
  78. 78.Microsoft. 2019. IntelliCode. https://github.com/MicrosoftDocs/intellicode.
  79. 79.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005.
  80. 80.Nhan Nguyen and Sarah Nadi. 2022. An empirical evaluation of github copilot’s code suggestions. In Proceedings of the 19th International Conference on Mining Software Repositories, pages 1–5.
  81. 81.Tung Thanh Nguyen, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N Nguyen. 2013. A statistical semantic language model for source code. In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering, pages 532–542.
  82. 82.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. CodeGen: An open large language model for code with multi-turn program synthesis. In The Eleventh International Conference on Learning Representations.
  83. 83.University of Oxford. 2020. Diffblue Cover. https://www.diffblue.com.
  84. 84.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155.
  85. 85.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  86. 86.Dipti Pawade, Avani Sakhapara, Sanyogita Parab, Divya Raikar, Ruchita Bhojane, and Henali Mamania. 2018. Literature survey on automatic code generation techniques. i-Manager’s Journal on Computer Science, 6(2):34.
  87. 87.Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2022. Asleep at the keyboard? assessing the security of github copilot’s code contributions. In 2022 IEEE Symposium on Security and Privacy (SP), pages 754–768. IEEE.
  88. 88.Julian Aron Prenner and Romain Robbes. 2021. Automatic program repair with openai’s codex: Evaluating quixbugs. arXiv preprint arXiv:2111.03922.
  89. 89.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  90. 90.Nitarshan Rajkumar, Raymond Li, and Dzmitry Bahdanau. 2022. Evaluating the text-to-sql capabilities of large language models. arXiv preprint arXiv:2204.00498.
  91. 91.Veselin Raychev, Martin Vechev, and Eran Yahav. 2014. Code completion with statistical language models. In Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation, pages 419–428.
  92. 92.Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020. CodeBLEU: a method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297.
  93. 93.William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022. Self-critiquing models for assisting human evaluators. arXiv preprint arXiv:2206.05802.
  94. 94.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. BLOOM: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  95. 95.Meet Shah, Rajat Shenoy, and Radha Shankarmani. 2021. Natural language to python source code using transformers. In 2021 International Conference on Intelligent Technologies (CONIT), pages 1–4. IEEE.
  96. 96.Tushar Sharma, Maria Kechagia, Stefanos Georgiou, Rohit Tiwari, and Federica Sarro. 2021. A survey on machine learning techniques for source code analysis. arXiv preprint arXiv:2110.09610.
  97. 97.Jiho Shin and Jaechang Nam. 2021. A survey of automatic code generation from natural language. Journal of Information Processing Systems, 17(3):537–555.
  98. 98.Mohammed Latif Siddiq and msiddiq. 2022. SecurityEval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques. Proceedings of the 1st International Workshop on Mining Software Repositories Applications for Privacy and Security.
  99. 99.Dominik Sobania, Martin Briesch, and Franz Rothlauf. 2022a. Choose your programming copilot: a comparison of the program synthesis performance of github copilot and genetic programming. In Proceedings of the Genetic and Evolutionary Computation Conference, pages 1019–1027.
  100. 100.Dominik Sobania, Martin Briesch, and Franz Rothlauf. 2022b. Choose your programming copilot: a comparison of the program synthesis performance of github copilot and genetic programming. In Proceedings of the Genetic and Evolutionary Computation Conference, pages 1019–1027.
  101. 101.Zeyu Sun, Qihao Zhu, Lili Mou, Yingfei Xiong, Ge Li, and Lu Zhang. 2018. A grammar-based structural cnn decoder for code generation. In AAAI Conference on Artificial Intelligence.
  102. 102.Ilya Sutskever, Geoffrey E Hinton, and Graham W Taylor. 2008. The recurrent temporal restricted boltzmann machine. Neural Information Processing Systems, 21.
  103. 103.Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. IntelliCode compose: code generation using transformer. Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering.
  104. 104.Eugene Syriani, Lechanceux Luhunu, and Houari Sahraoui. 2018. Systematic mapping study of template-based code generation. Computer Languages, Systems & Structures, 52:43–62.
  105. 105.tabnine. 2018. TabNine. https://www.tabnine.com.
  106. 106.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. LAMDA: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  107. 107.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Neural Information Processing Systems, 30.
  108. 108.Yao Wan, Zhou Zhao, Min Yang, Guandong Xu, Haochao Ying, Jian Wu, and Philip S Yu. 2018. Improving automatic source code summarization via deep reinforcement learning. In Proceedings of the 33rd ACM/IEEE international conference on automated software engineering, pages 397–407.
  109. 109.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  110. 110.Shiqi Wang, Zheng Li, Haifeng Qian, Cheng Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth, and Bing Xiang. 2022a. ReCode: Robustness evaluation of code generation models. arXiv preprint arXiv:2212.10264.
  111. 111.Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021. CodeT5: Identifier-aware unified pretrained encoder-decoder models for code understanding and generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8696–8708.
  112. 112.Zhiruo Wang, Grace Cuenca, Shuyan Zhou, Frank F Xu, and Graham Neubig. 2022b. MCoNaLa: a benchmark for code generation from multiple natural languages. arXiv preprint arXiv:2203.08388.
  113. 113.Zhiruo Wang, Shuyan Zhou, Daniel Fried, and Graham Neubig. 2022c. Execution-based evaluation for open-domain code generation. arXiv preprint arXiv:2212.10481.
  114. 114.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  115. 115.Frank F. Xu, Uri Alon, Graham Neubig, and Vincent J. Hellendoorn. 2022. A systematic evaluation of large language models of code. Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming.
  116. 116.Yichen Xu and Yanqiao Zhu. 2022. A survey on pretrained language models for neural code intelligence. arXiv preprint arXiv:2212.10079.
  117. 117.Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig. 2018. Learning to mine aligned code and natural language pairs from stack overflow. 2018 IEEE/ACM 15th International Conference on Mining Software Repositories (MSR), pages 476–486.
  118. 118.Pengcheng Yin and Graham Neubig. 2017. A syntactic neural model for general-purpose code generation. arXiv preprint arXiv:1704.01696.
  119. 119.Daoguang Zan, Bei Chen, Zeqi Lin, Bei Guan, Yongji Wang, and Jian-Guang Lou. 2022a. When language model meets private library. In Conference on Empirical Methods in Natural Language Processing.
  120. 120.Daoguang Zan, Bei Chen, Dejian Yang, Zeqi Lin, Minsu Kim, Bei Guan, Yongji Wang, Weizhu Chen, and Jian-Guang Lou. 2022b. CERT: Continual pre-training on sketches for library-oriented code generation. In International Joint Conference on Artificial Intelligence.
  121. 121.Jialu Zhang, José Cambronero, Sumit Gulwani, Vu Le, Ruzica Piskac, Gustavo Soares, and Gust Verbruggen. 2022. Repairing bugs in python assignments using large language models. arXiv preprint arXiv:2209.14876.
  122. 122.Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shanshan Wang, Yufei Xue, Zi-Yuan Wang, Lei Shen, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023. CodeGeeX: A pre-trained model for code generation with multilingual evaluations on humaneval-x. ArXiv, abs/2303.17568.
  123. 123.Shuyan Zhou, Uri Alon, Frank F Xu, Zhengbao JIang, and Graham Neubig. 2023. DocCoder: Generating code by retrieving and reading docs. In The Eleventh International Conference on Learning Representations.
  124. 124.Ming Zhu, Aneesh Jain, Karthik Suresh, Roshan Ravindran, Sindhu Tipirneni, and Chandan K. Reddy. 2022a. XLCoST: A benchmark dataset for cross-lingual code intelligence.
  125. 125.Ming Zhu, Karthik Suresh, and Chandan K Reddy. 2022b. Multilingual code snippets training for program translation.

Citation

MLA
Zan, D., et al. “Large Language Models Meet NL2Code: A Survey”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 7443–64, https://doi.org/10.18653/v1/2023.acl-long.411.
APA
Zan, D., Chen, B., Zhang, F., Lu, D., Wu, B., Guan, B., Yongji, W., & Lou, J.-G. (2023). Large Language Models Meet NL2Code: A Survey. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7443–7464. https://doi.org/10.18653/v1/2023.acl-long.411
Chicago
Zan, D., B. Chen, F. Zhang, et al. 2023. “Large Language Models Meet NL2Code: A Survey”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7443–64. https://doi.org/10.18653/v1/2023.acl-long.411.
Harvard
Zan, D. et al. (2023) “Large Language Models Meet NL2Code: A Survey”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7443–7464. Available at: https://doi.org/10.18653/v1/2023.acl-long.411.
Vancouver
1. Zan D, Chen B, Zhang F, Lu D, Wu B, Guan B, Yongji W, Lou J-G (2023) Large Language Models Meet NL2Code: A Survey. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7443–7464

BibTeX

@inproceedings{zan-etal-2023-large,
    title = "Large Language Models Meet {NL}2{C}ode: A Survey",
    author = "Zan, Daoguang  and
      Chen, Bei  and
      Zhang, Fengji  and
      Lu, Dianjie  and
      Wu, Bingchao  and
      Guan, Bei  and
      Yongji, Wang  and
      Lou, Jian-Guang",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.411/",
    doi = "10.18653/v1/2023.acl-long.411",
    pages = "7443--7464"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/