Large Language Models Meet NL2Code: A Survey
Daoguang ZanBei ChenFengji ZhangDianjie LuBingchao WuBei GuanYongji WangJian-Guang Lou
Presents a comprehensive survey of 27 code-generation large language models alongside their evaluation benchmarks and metrics, comparing their performance on HumanEval and identifying model scale, data quality, and expert tuning as the primary drivers of success.
Generating functional computer code directly from natural language requirements represents a significant milestone for artificial intelligence, with the potential to democratize software development and substantially improve engineering productivity. The emergence of Transformer-based large language models has accelerated progress in this space, transitioning code generation from early rule-based or statistical techniques to highly capable neural systems. Understanding the architectural patterns, data characteristics, and tuning strategies driving these models is essential for organizations planning technical investments or evaluating commercial programming assistants.
The main objective of the article is to provide a comprehensive evaluation of the state of natural language-to-code technology by systematically analyzing 27 representative large language models, surveying 17 key evaluation benchmarks, and identifying the primary factors behind model success along with remaining capability gaps.
To conduct this evaluation, the researchers reviewed 27 prominent models ranging in scale from under 100 million to 540 billion parameters, covering both encoder-decoder and decoder-only architectures developed through late 2022. They conducted standardized, zero-shot comparative evaluations on leading industry benchmarks, primarily HumanEval (comprising 164 hand-crafted Python tasks) and MBPP (comprising 974 programming exercises), measuring the percentage of problems successfully solved across single and multiple generation attempts.
The analysis reveals four central findings. First, model performance depends on a core formula: large model and corpus size, premium training data, and expert hyperparameter tuning. Second, scale strongly drives performance and syntactic correctness; for instance, expanding model capacity dramatically drops code syntax error rates down to around 6% in top models, although generating semantically correct logic remains difficult. Third, general-purpose text scale alone does not suffice without high-quality, code-specific training; large general text models frequently underperform smaller models trained on curated, deduplicated code repositories. Fourth, context window capacity significantly amplifies capability, enabling a smaller 165-million parameter model paired with an 8,000-token context window to achieve performance comparable to a 20-billion parameter model constrained to a 2,000-token context window.
These findings indicate that while modern code assistants can deliver immediate developer efficiency gains, automated systems cannot yet fully replace human software engineers. Commercial deployment still carries operational and safety risks, as models often introduce subtle bugs, struggle with complex multi-condition problem requirements, and exhibit prompt sensitivity. Additionally, because models default to generating plausible-looking code even when an exact solution is infeasible, organizations face code security, maintainability, and review overheads if generated outputs are integrated without rigorous human validation.
Organizations adopting or developing code generation systems should implement structured code verification workflows rather than relying solely on raw model generation. Teams training custom models should prioritize high-grade dataset filtering, custom code-specific tokenizers, and larger context window configurations rather than focusing exclusively on parameter count. Future development must focus on bridging the remaining human-machine gaps, particularly self-validation, multi-step problem decomposition, code explanation, and retrieval-based mechanisms that incorporate dynamic documentation and Application Programming Interface updates.
Readers should interpret the comparative performance metrics with standard caveats regarding data availability and benchmark constraints. A completely controlled comparison remains difficult because several leading commercial models are closed-source and train on disparate data corpora. Furthermore, common benchmarks predominantly focus on standalone Python functions evaluated via execution test suites, providing limited insight into large-scale multi-file system architectures, software security vulnerabilities, or long-term maintainability.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This seminal paper introduces OpenAI's Codex and the foundational HumanEval benchmark that serves as the primary evaluation metric and comparison baseline throughout the NL2Code survey.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). This paper establishes CodeXGLUE, the standard multi-task benchmark suite and baseline models (CodeBERT, CodeGPT) that foundational NL2Code surveys synthesize and evaluate.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). This foundational work establishes the MBPP benchmark and formalizes program synthesis with large language models, providing core concepts surveyed in NL2Code.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). This paper presents the open-source CodeGen family and multi-turn program synthesis datasets, representing a major open architecture analyzed in the survey.
- Paper: CodeBERT: A Pre-Trained Model for Programming and Natural Languages, Zhangyin Feng et al. (2020). This study introduces CodeBERT, defining the foundational bimodal pre-training paradigm bridging natural language descriptions and programming syntax.
- Paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, Yue Wang et al. (2021). This work establishes the identifier-aware CodeT5 encoder-decoder architecture, an essential milestone in the evolution of code intelligence models reviewed by the survey.
- Paper: UniXcoder: Unified Cross-Modal Pre-training for Code Representation, Daya Guo et al. (2022). This paper presents UniXcoder, establishing unified representation methods incorporating syntax trees and code comments that underpin modern code models.
- Paper: Competition-level code generation with AlphaCode, Yujia Li et al. (2022). This paper presents AlphaCode and the CodeContests benchmark, introducing key massive-sampling and filtering methodologies for complex program generation.
- Paper: CodeSearchNet Challenge: Evaluating the State of Semantic Code Search, Hamel Husain et al. (2019). This work introduces the CodeSearchNet dataset and benchmark, which provided the foundational multi-language corpus for training early code intelligence models.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). This paper builds directly on the evaluation challenges noted in the survey by introducing HumanEval+ to rigorously expose test insufficiency in code generation LLMs.
- Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). This work develops Code Llama, advancing the survey's discussion of open foundation models by implementing infilling and long-context capabilities for code synthesis.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). This research continues the progression of open-access code intelligence by training DeepSeek-Coder at scale across 87 languages and repository-level contexts.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). This benchmark directly addresses the benchmark saturation and data contamination concerns highlighted in early NL2Code surveys by evaluating models on live contest problems.
- Paper: BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions, Terry Yue Zhuo et al. (2025). This benchmark extends single-function NL2Code evaluation to complex, real-world software tasks involving multiple external library calls and precise instruction following.
- Paper: CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion, Yangruibo Ding et al. (2023). This study advances code intelligence evaluation beyond isolated single-file generations to cross-file repository-level code completion.
- Paper: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models, Siming Huang et al. (2025). This paper extends the survey's insights on data quality and tuning by providing an open, fully transparent end-to-end recipe and dataset for training top-tier code LLMs.
- Paper: Large Language Models for Software Engineering: A Systematic Literature Review, Xinying Hou et al. (2023). This comprehensive literature review broadens the scope from natural-language-to-code synthesis to the entire spectrum of software engineering tasks across the development lifecycle.
