Large Language Models for Software Engineering: A Systematic Literature Review

Xinying HouYanjie ZhaoYue LiuZhou YangKailong WangLi LiXiapu LuoDavid LoJohn C. GrundyHaoyu Wang

article2023ACM Transactions on Software Engineering and Methodology1,356 citations

Systematizes findings from nearly 400 studies to explain how large language models are trained, evaluated, and deployed across software engineering tasks, providing clear guidance on effective data practices and open research challenges.

Listen

Modern software engineering faces escalating complexity, creating a growing demand for automation across the development lifecycle. Recently, large language models have emerged as powerful tools capable of transforming complex software tasks into data, code, and text analysis problems. However, an understanding of their actual implementation, measurable impact, and operational limitations across the software engineering spectrum has remained fragmented. To establish a clear baseline for decision-makers, the article evaluates the landscape of language models in software engineering by examining their architectures, data practices, optimization methods, and task efficacy.

The article conducts a systematic literature review analyzing 395 peer-reviewed and pre-print research studies published between January 2017 and January 2024. The authors evaluate trends across model architectures—specifically decoder-only, encoder-only, and encoder-decoder models—alongside dataset curation methodologies, optimization strategies such as parameter-efficient fine-tuning and prompt engineering, and performance across 85 distinct software engineering tasks spanning all development stages.

The review reveals several central findings. First, decoder-only models, such as the GPT series, have become dominant, comprising over 70% of research activity by 2023 due to their generative capabilities. In contrast, encoder-only models such as BERT and encoder-decoder models like T5 serve specialized niches in code understanding and translation. Second, research is heavily skewed toward software development (56.65%) and maintenance (22.71%), dominated by generative tasks like code generation and automated program repair, while activities such as requirements engineering (3.90%), software design (0.92%), and software management (0.69%) remain largely unexplored. Third, conversational and feedback-driven prompting frameworks significantly enhance performance; for instance, interactive setups in models like ChatGPT allow continuous refinement that markedly improves patch correctness in program repair. Fourth, parameter-efficient fine-tuning methods (such as low-rank adaptation) and advanced reasoning techniques (such as chain-of-thought prompting) consistently allow models to achieve high accuracy without requiring full, computationally prohibitive parameter retraining.

These findings indicate that adopting language models can substantially decrease developer workloads, accelerate prototyping, and improve automated code quality assurance. However, realizing these benefits requires careful alignment between model architectures and specific tasks, as generative decoder models excel at code synthesis but may prove inefficient for purely analytical tasks. Furthermore, the disproportionate academic reliance on open-source datasets (which represent roughly 62.83% of examined datasets) compared to industrial datasets (found in fewer than 2% of studies) highlights a critical disconnect: techniques proven in academic benchmarks may face unaddressed reliability, security, and integration challenges in proprietary, enterprise-scale environments.

Organizations should focus near-term investments on high-value, validated areas—such as automated code completion, unit test generation, and interactive bug repair—while adopting lightweight parameter-efficient fine-tuning and structured prompt engineering to manage computational costs. Senior leaders should avoid applying these models to early-stage requirements or system architecture decisions without rigorous human oversight until further empirical research emerges. Before broad enterprise deployment, leadership should initiate targeted pilot programs on internal codebases to validate real-world performance against enterprise constraints.

Readers should exercise caution regarding the findings due to the rapidly evolving nature of the field and the inclusion of non-peer-reviewed pre-prints (which accounted for roughly 61% of the analyzed papers, though filtered via quality rubrics). The scarcity of industrial evaluation data remains a notable limitation, meaning confidence is high regarding the conceptual potential of language models for software tasks, but moderate concerning their off-the-shelf reliability in complex, enterprise-grade production pipelines.

  • Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). This paper establishes the CodeXGLUE benchmark and foundational code representation models, providing the core task categories and evaluation standards analyzed throughout the literature review.
  • Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). This work introduced key program synthesis benchmarks such as MBPP and proved early scaling laws for code generation, serving as a baseline milestone for subsequent LLM4SE research.
  • Paper: CodeSearchNet Challenge: Evaluating the State of Semantic Code Search, Hamel Husain et al. (2019). This seminal benchmark and dataset established the standard methodologies for neural code retrieval and code understanding upon which modern code LLMs build.
  • Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). This foundational paper presents Code Llama, one of the most prominent open code models whose architectures, infilling training, and repository-level adaptations are systematically surveyed in the review.
  • Paper: A Survey of Large Language Models, Wayne Xin Zhao et al. (2023). This comprehensive survey outlines the general pre-training, fine-tuning, and utilization lifecycle of large language models, providing the general NLP context needed to understand domain-specific SE adaptations.
Cover for Large Language Models for Software Engineering: A Systematic Literature Review

Abstract

Large Language Models (LLMs) have significantly impacted numerous domains, including Software Engineering (SE). Many recent publications have explored LLMs applied to various SE tasks. Nevertheless, a comprehensive understanding of the application, effects, and possible limitations of LLMs on SE is still in its early stages. To bridge this gap, we conducted a systematic literature review (SLR) on LLM4SE, with a particular focus on understanding how LLMs can be exploited to optimize processes and outcomes. We select and analyze 395 research papers from January 2017 to January 2024 to answer four key research questions (RQs). In RQ1, we categorize different LLMs that have been employed in SE tasks, characterizing their distinctive features and uses. In RQ2, we analyze the methods used in data collection, preprocessing, and application, highlighting the role of well-curated datasets for successful LLM for SE implementation. RQ3 investigates the strategies employed to optimize and evaluate the performance of LLMs in SE. Finally, RQ4 examines the specific SE tasks where LLMs have shown success to date, illustrating their practical contributions to the field. From the answers to these RQs, we discuss the current state-of-the-art and trends, identifying gaps in existing research, and flagging promising areas for future study. Our artifacts are publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Approach
  • 2.1 Research Questions
  • 2.2 Search Strategy
  • 2.2.1 Search Items
  • 2.2.2 Search Datasets
  • 2.3 Study Selection
  • 2.3.1 Study Inclusion and Exclusion Criteria
  • 2.3.2 Study Quality Assessment
  • 2.4 Snowballing Search
  • 2.5 Data Extraction and Analysis
  • 3 RQ1: What LLMs have been employed to date to solve SE tasks?
  • 3.1 Large Language Models (LLMs)
  • 3.2 Trend Analysis
  • 4 RQ2: How are SE-related datasets collected, preprocessed, and used in LLMs?
  • 4.1 How are the datasets for training LLMs sourced?
  • 4.2 What types of SE datasets have been used in existing LLM4SE studies?
  • 4.3 How do data types influence the selection of data-preprocessing techniques?
  • 4.4 What input formats are the datasets for LLM training converted to?
  • 5 RQ3: What techniques are used to optimize and evaluate LLM4SE?
  • 5.1 What tuning techniques are used to enhance the performance of LLMs in SE tasks?
  • 5.2 What prompt engineering techniques are applied to improve the performance of LLMs in SE tasks?
  • 5.3 How are evaluation metrics utilized to assess the performance of LLM4SE tasks?
  • 6 RQ4: What SE tasks have been effectively addressed to date using LLM4SE?
  • 6.1 What are the distributions SE activities and problem types addressed to date with LLM4SE?
  • 6.2 How are LLMs used in requirements engineering?
  • 6.3 How are LLMs used in software design?
  • 6.4 How are LLMs used in software development?
  • 6.5 How are LLMs used in software quality assurance?
  • 6.6 How are LLMs used in software maintenance?
  • 6.7 How are LLMs used in software management?
  • 7 Threats to Validity
  • 8 Challenges and Opportunities
  • 8.1 Challenges
  • 8.1.1 Challenges in LLM Applicability.
  • 8.1.2 Challenges in LLM Generalizability
  • 8.1.3 Challenges in LLM Evaluation
  • 8.1.4 Challenges in LLM Interpretability, Trustworthiness, and Ethical Usage
  • 8.2 Opportunities
  • 8.2.1 Optimization of LLM4SE
  • 8.2.2 Expanding LLM’s NLP Capabilities in More SE Phases.
  • 8.2.3 Enhancing LLM’s Performance in Existing SE Tasks
  • 8.3 Roadmap
  • 9 Conclusion
  • References
  • A Data Types
  • B Input Forms
  • C Prompt Engineering
  • D Evaluation Metrics
  • E SE Tasks

Knowls

  1. Knowl 1 — Taxonomy and Architectural Distribution of Large Language Models in Software Engineering

    empirical result

    Large Language Models (LLMs) utilized in software engineering (SE) are categorized into three primary structural architectures based on their transformer components:

    1. Encoder-only LLMs (e.g., BERT, RoBERTa, CodeBERT, GraphCodeBERT, ALBERT, CuBERT): These architectures utilize bidirectional self-attention to encode input text or code sequences into rich contextual representations. They are predominantly employed for code comprehension, classification, and extraction tasks such as bug localization, vulnerability detection, and requirement classification.

    2. Encoder-decoder LLMs (e.g., T5, CodeT5, CodeT5+, PLBART, AlphaCode, CoTexT): These models encode input sequences into an intermediate hidden space and decode target token sequences. They are suited for translation and transformation tasks such as code summarization, code-to-code translation, and automated program repair.

    3. Decoder-only LLMs (e.g., GPT-3, GPT-3.5, GPT-4, Codex, CodeGen, LLaMA, StarCoder, InCoder, DeepSeek-Coder): These models employ autoregressive causal self-attention to generate output tokens sequentially. They are predominantly utilized for generative tasks, including automated code generation, code completion, and test case synthesis.

    Historical analysis reveals a clear shift in architectural preference across 395 analyzed studies: in 2020, research was dominated by encoder-only models (8 out of 10 model instances, 80%), whereas by 2023, decoder-only LLMs represented 432 instances across 195 papers (70.7% of all model instances), compared to 85 instances (13.91%) for encoder-decoder models and 94 instances (15.39%) for encoder-only models.

  2. Knowl 2 — Distribution of LLM Research Across Software Development Life Cycle Activities and Problem Types

    data/table

    The application of LLMs across software development life cycle (SDLC) activities and underlying machine learning problem formulations is distributed as follows:

    SDLC Activity Studies (Count) Share (%) Problem Formulation Studies (Count) Share (%)
    Software Development
    247 56.65% Generation 280 70.97%
    Software Maintenance 99 22.71% Classification 85 21.61%
    Software Quality Assurance 66 15.14% Recommendation 27 6.77%
    Requirements Engineering 17 3.90% Regression 3 0.65%
    Software Design 4 0.92%
    Software Management 3 0.69%

    The vast majority of research focuses on implementation and maintenance activities (combined 79.36%), while early-stage lifecycle activities (Requirements Engineering, Software Design) and Software Management remain substantially under-explored (combined 5.51%). Similarly, generative task formulations represent over 70% of all investigated problem formulations.

  3. Knowl 3 — Specific Software Engineering Tasks Addressed by LLMs Across the SDLC

    empirical result

    Across 395 analyzed papers, LLMs have been deployed to address 85 distinct software engineering tasks categorized into six primary SDLC activities:

    • Software Development (247 studies): Encompasses 24+ tasks, dominated by Code Generation (118 studies), Code Completion (22), Code Summarization (21), Code Translation (12), Code Search (12), Code Understanding (8), Program Synthesis (6), API Inference (5), API Recommendation (5), Code Editing (5), and Code Representation (3).
    • Software Maintenance (99 studies): Dominated by Automated Program Repair (35 studies), Code Clone Detection (8), Code Review (7), Debugging (4), Bug Reproduction (3), Duplicate Bug Report Detection (3), Logging (3), Log Parsing (3), and Sentiment Analysis (3).
    • Software Quality Assurance (66 studies): Dominated by Vulnerability Detection (18 studies), Test Case/Suite Generation (17), Bug Localization (5), Formal Verification (5), Testing Automation/Fuzzing (4), and Fault Localization (3).
    • Requirements Engineering (17 studies): Comprises Anaphoric Ambiguity Treatment (4 studies), Requirements Classification (4), Requirement Analysis and Evaluation (2), Specification Generation (2), and Coreference Detection (1).
    • Software Design (4 studies): Includes GUI Retrieval (1 study), Rapid Prototyping (1), Software Specification Synthesis (1), and System Design (1).
    • Software Management (3 studies): Includes Effort Estimation (2 studies) and Software Tool Configuration (1).
  4. Knowl 4 — Quasi-Gold Standard Methodology for SLR Paper Selection in LLM4SE

    experimental setup

    The paper identifies primary literature using the Quasi-Gold Standard (QGS) search strategy across the publication period from January 2017 to January 2024:

    1. Manual Venue Search & QGS Establishment: 4,618 papers from six top-tier SE conferences and journals (ICSE, ESEC/FSE, ASE, ISSTA, TOSEM, TSE) were crawled and filtered to construct a 51-paper QGS baseline.
    2. Automated Search: A combined boolean search string pairing SE-related terms and LLM-related terms was executed across seven electronic databases (IEEE Xplore: 1,192, ACM DL: 10,445, ScienceDirect: 62,290, Web of Science: 42,166, Springer: 85,671, arXiv: 9,966, DBLP: 4,035), yielding 218,765 total results.
    3. Filtering Pipeline:
      • Page length exclusion (<8<8 pages) reduced candidate pool to 80,611 papers.
      • Title, abstract, and keyword filtering reduced pool to 5,078 papers (excluding 448 non-English papers).
      • Peer-reviewed venue extraction and curated arXiv retention reduced pool to 1,172 papers.
      • Deduplication and grey-literature/workshop exclusion yielded 810 papers, reduced to 594 upon full-text screening.
      • Quality Assessment Criteria (QAC 1--10) filtering using an 80% threshold score (16.8/21 for published papers; 14.4/18 for arXiv papers) yielded 382 primary studies.
    4. Snowballing: Forward (3,964 citations) and backward (9,610 references) snowballing from the 382 papers yielded 5,152 deduplicated candidates, which upon full selection criteria produced 13 additional primary studies, establishing a final corpus of 395 papers.
  5. Knowl 5 — Standard Preprocessing Pipelines for Text and Code Datasets in Software Engineering

    algorithm

    Software engineering datasets for training and evaluating LLMs undergo structured multi-step preprocessing pipelines tailored to text and source code modalities:

    Algorithm: Preprocessing Pipelines for SE Text and Code Datasets
    Input: Raw SE artifacts (natural language documents, bug reports, issue threads, source repositories)
    Output: Cleaned, tokenized, and partitioned splits (D_train, D_val, D_test)
    procedure PREPROCESSTEXTDATASET(RawTextCorpus):
        ExtractedText <- Extract text from bug reports, requirement specs, comments, API documentation
        SegmentedText <- Partition text into sentences or words based on task requirements
        FilteredText <- Discard invalid/short samples (e.g., bug reports with length < 15 words)
        CleanedText <- Normalize symbols, remove noise, markup tags, and stop words
        DeduplicatedText <- Eliminate duplicate textual instances
        TokenizedText <- Subword or word tokenization aligned with the model vocabulary
        (D_train, D_val, D_test) <- Split into training, validation, and test subsets
        return (D_train, D_val, D_test)
    procedure PREPROCESSCODEDATASET(RawCodeCorpus):
        ExtractedCode <- Extract code at target granularity (token, method, file, or repository level)
        FilteredCode <- Remove code segments failing task quality or relevance thresholds
        DeduplicatedCode <- Identify and remove exact and near-duplicate code instances
        CompiledCode <- Merge and compile extracted snippets into a unified code dataset
        CleanedCode <- Discard uncompilable, syntax-invalid, or non-executable snippets
        RepresentedCode <- Convert code into token sequences, ASTs, CFGs, or PDG graphs
        (D_train, D_val, D_test) <- Split into training, validation, and test subsets
        return (D_train, D_val, D_test)
  6. Knowl 6 — Dataset Sourcing Strategies in LLM4SE Research

    data/table

    Across 374 studies that explicitly state the dataset origin, datasets are divided into four main sourcing categories:

    Dataset Source Number of Studies Proportion (%)
    Open-source datasets
    235 62.83%
    Constructed datasets 84 22.46%
    Collected datasets 49 13.10%
    Industrial datasets 6 1.60%

    Open-source datasets (e.g., HumanEval, MBPP, Defects4J) dominate due to standardized benchmarks and reproducibility. Constructed datasets represent modified, annotated, or synthetic benchmarks created for specific research tasks (e.g., manual bug injection). Collected datasets are mined directly from public repositories (GitHub) or Q&A platforms (Stack Overflow). Industrial datasets from production repositories represent only 1.60% (6 studies), highlighting a substantial academic-industrial evaluation gap.

  7. Knowl 7 — Input Formats and Modalities in LLM4SE

    data/table

    Data representation formats provided as input to LLMs across 355 studies that explicitly state input formatting are classified into four main categories:

    Category Input Form Number of Studies
    Token-based input
    Text in tokens 150
    Code in tokens 118
    Code and text in tokens 78
    Tree/Graph-based input Code in graph structure (e.g., CFG, PDG) 3
    Code in tree structure (e.g., AST) 2
    Hybrid-based input Hybrid structural and token combinations 2
    Pixel-based input GUI screenshot images / visual pixel grids 1

    Token-based inputs account for approximately 97.75% of all investigated models, with text-only tokens (42.25%) and code-only tokens (33.24%) being the most prevalent. Syntactic tree, graph, pixel, and multimodal inputs remain rarely utilized despite the structured nature of software artifacts.

  8. Knowl 8 — Model Optimization and Parameter-Efficient Fine-Tuning Techniques in LLM4SE

    model/method

    Tuning strategies for adapting pre-trained LLMs to downstream SE tasks are grouped into full parameter fine-tuning and Parameter-Efficient Fine-Tuning (PEFT):

    • Full Fine-Tuning: Updates all parameters across the network. Widely utilized in 83 studies, primarily with smaller encoder models (e.g., BERT, CodeBERT), but computationally intractable for 10B+ scale models.
    • Parameter-Efficient Fine-Tuning (PEFT): Keeps base LLM weights frozen while updating a small parameter subset:
      • Low-Rank Adaptation (LoRA): Injects low-rank trainable decomposition matrices W=W0+ΔW=W0+BAW = W_0 + \Delta W = W_0 + B A into Transformer attention projections (where B∈Rd×r,A∈Rr×kB \in \mathbb{R}^{d \times r}, A \in \mathbb{R}^{r \times k} with rank r≪min⁡(d,k)r \ll \min(d, k)), adopted in 8 studies for multi-language code translation (SteloCoder) and program repair (RepairLLaMA).
      • Prompt Tuning: Prepends learnable continuous virtual tokens to input embeddings while keeping model parameters fixed (utilized in 3 studies, e.g., AUMENA for method naming).
      • Prefix Tuning: Prepends trainable continuous prefix vectors across key and value projections at every Transformer attention layer (e.g., CodeT5+, LLaMA-Reviewer).
      • Adapter Tuning: Inserts small bottleneck feed-forward neural layers between frozen Transformer layers.
    • Additional Tuning Paradigms: Reinforcement Learning with compiler/runtime feedback (RL), Supervised Fine-Tuning (SFT), syntax-guided fine-tuning, and knowledge-preservation fine-tuning.
  9. Knowl 9 — Prompt Engineering Paradigms for LLMs in Software Engineering

    data/table

    Prompt engineering methods guide LLM execution without altering underlying weights. The distribution of specific prompting techniques across the 395 analyzed studies is summarized below:

    Prompt Engineering Technique Primary SE Application / Mechanism Number of Studies
    Few-shot prompting
    In-context learning using kk input-output demonstrations 88
    Zero-shot prompting Task execution using only natural language instructions 79
    Chain-of-Thought (CoT) Step-by-step intermediate natural language reasoning chains 18
    Automatic Prompt Engineer (APE) Automated instruction generation and selection 2
    Chain of Code (CoC) Interleaving reasoning with an external code emulator 2
    Automatic Chain-of-Thought (Auto-CoT) Automated selection and construction of reasoning chains 1
    Modular-of-Thought (MoT) Decomposing complex coding tasks into modular subtasks 1
    Structured Chain-of-Thought (SCoT) Constraining intermediate reasoning steps with program structure 1
    Other custom prompts Task-specific prompt chaining and differential prompting 76

    While few-shot and zero-shot prompting remain the standard baselines, specialized techniques tailored to code (e.g., SCoT using AST syntactic constraints and MoT decomposing monolithic functions into submodules) demonstrate superior correctness on complex programming benchmarks.

  10. Knowl 10 — Evaluation Metrics Categorization across Software Engineering Task Formulations

    data/table

    Evaluation metrics applied across the 395 studies are categorized according to four underlying SE problem types:

    Problem Type Evaluated Metrics (Frequency Count) Total Instances
    Generation
    BLEU / BLEU-4 / BLEU-DC (62), Pass@k (54), Accuracy / Accuracy@k (38), 338
    Exact Match [EM] (36), CodeBLEU (29), ROUGE / ROUGE-L (22), Precision (18),
    METEOR (16), Recall (15), F1-score (15), Edit Similarity [ES] (6),
    Mean Reciprocal Rank [MRR] (6), Edit Distance [ED] (5), MAR (4), ChrF (3),
    CrystalBLEU (3), CodeBERTScore (2), MFR (1), Perplexity [PP] (1)
    Classification Precision (35), Recall (34), F1-score (33), Accuracy (23), AUC (9), 147
    ROC (4), False Positive Rate [FPR] (4), False Negative Rate [FNR] (3), MCC (2)
    Recommendation Mean Reciprocal Rank [MRR] (15), Precision / Precision@k (6), 39
    MAP / MAP@k (6), F-score / F-score@k (5), Recall / Recall@k (4), Accuracy (3)
    Regression Mean Absolute Error [MAE] (1) 1

    Generative tasks feature the largest metric diversity (19 distinct metrics across 338 instances), balancing token-level nn-gram matching (BLEU, ROUGE), syntactic structure matching (CodeBLEU), and functional execution correctness (Pass@k\text{Pass}@k). Classification tasks primarily employ Precision, Recall, and F1-score, while recommendation tasks prioritize rank-based metrics (MRR, MAP).

  11. Knowl 11 — Key Technical Challenges in Applying Large Language Models to Software Engineering

    limitation

    The systematic literature review identifies six critical challenges currently confronting LLM4SE:

    1. Model Size and Deployment Costs: Modern code LLMs (e.g., CodeGen, Codex, BLOOM-176B) exceed tens or hundreds of billions of parameters, incurring massive compute overhead (e.g., over 1,000,000 GPU hours to train BLOOM-176B) and high latency that hinders local edge deployment and real-time IDE integration.
    2. Data Dependency and Contamination: Scarcity of clean domain-specific datasets and high vulnerability to benchmark contamination (train-test data leakage), leading to inflated evaluation scores.
    3. Ambiguity and Functional Correctness: Natural language requirement descriptions often introduce semantic ambiguity, resulting in LLMs generating syntactically valid but functionally incorrect or hallucinated code.
    4. Generalizability under Semantic Transformations: Pre-trained code LLMs fail to generalize when exposed to semantic-preserving input alterations (e.g., variable renaming drastically degrades CodeBERT classification and retrieval accuracy).
    5. Metric Misalignment: Standard NLP overlap metrics (BLEU, ROUGE) correlate poorly with runtime code correctness, compilation success, and algorithmic security.
    6. Trustworthiness, Security, and Ethics: Models trained on public web scrapes may generate vulnerable code snippets (introducing Common Weakness Enumerations), leak sensitive personally identifiable information (PII) or API keys, and remain vulnerable to jailbreak and backdoor prompt attacks.

Coverage note — None was omitted; all major contributions—including paper search methodology, architectural taxonomy, dataset sourcing and preprocessing, tuning paradigms, prompting techniques, evaluation metrics across task types, SDLC task mapping, and key challenges/future directions—are comprehensively covered.

References

  1. 1.Mayank Agarwal, Yikang Shen, Bailin Wang, Yoon Kim, and Jie Chen. 2024. Structured Code Representations Enable Data-Efficient Adaptation of Code Language Models. arXiv preprint arXiv:2401.10716 (2024).
  2. 2.Emad Aghajani, Csaba Nagy, Mario Linares-Vásquez, Laura Moreno, Gabriele Bavota, Michele Lanza, and David C Shepherd. 2020. Software documentation: the practitioners’ perspective. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 590–601.
  3. 3.Lakshya Agrawal, Aditya Kanade, Navin Goyal, Shuvendu K Lahiri, and Sriram Rajamani. 2023. Monitor-Guided Decoding of Code LMs with Static Analysis of Repository Context. In Thirty-seventh Conference on Neural Information Processing Systems.
  4. 4.Baleegh Ahmad, Shailja Thakur, Benjamin Tan, Ramesh Karri, and Hammond Pearce. 2023. Fixing Hardware Security Bugs with Large Language Models. arXiv preprint arXiv:2302.01215 (2023).
  5. 5.Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified pre-training for program understanding and generation. arXiv preprint arXiv:2103.06333 (2021).
  6. 6.Toufique Ahmed, Kunal Suresh Pai, Premkumar Devanbu, and Earl T Barr. 2023. Improving Few-Shot Prompts with Relevant Static Analysis Products. arXiv preprint arXiv:2304.06815 (2023).
  7. 7.Toufique Ahmed, Kunal Suresh Pai, Premkumar Devanbu, and Earl T. Barr. 2024. Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization). arXiv:2304.06815 [cs.SE]
  8. 8.Mistral AI. 2023. Mistral. https://mistral.ai/.
  9. 9.Ali Al-Kaswan, Toufique Ahmed, Maliheh Izadi, Anand Ashok Sawant, Premkumar Devanbu, and Arie van Deursen. 2023. Extending Source Code Pre-Trained Language Models to Summarise Decompiled Binarie. In 2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 260–271.
  10. 10.Ajmain I Alam, Palash R Roy, Farouq Al-Omari, Chanchal K Roy, Banani Roy, and Kevin A Schneider. 2023. GPT-CloneBench: A comprehensive benchmark of semantic clones and cross-language clones using GPT-3 model and SemanticCloneBench. In 2023 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 1–13.
  11. 11.Mohammed Alhamed and Tim Storer. 2022. Evaluation of Context-Aware Language Models and Experts for Effort Estimation of Software Maintenance Issues. In 2022 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 129–138.
  12. 12.Frances E Allen. 1970. Control flow analysis. ACM Sigplan Notices 5, 7 (1970), 1–19.
  13. 13.Rajeev Alur, Rastislav Bodik, Garvit Juniwal, Milo MK Martin, Mukund Raghothaman, Sanjit A Seshia, Rishabh Singh, Armando Solar-Lezama, Emina Torlak, and Abhishek Udupa. 2013. Syntax-guided synthesis. IEEE.
  14. 14.Sven Amann, Sebastian Proksch, Sarah Nadi, and Mira Mezini. 2016. A study of visual studio usage in practice. In 2016 IEEE 23rd International Conference on Software Analysis, Evolution, and Reengineering (SANER), Vol. 1. IEEE, 124–134.
  15. 15.Amazon. 2023. Amazon CodeWhisperer. https://aws.amazon.com/cn/codewhisperer/.
  16. 16.Amazon. 2023. NVIDIA Tesla A100 Ampere 40 GB Graphics Card - PCIe 4.0 - Dual Slot. https://www.amazon.com/NVIDIA-Tesla-A100-Ampere-Graphics/dp/B0BGZJ27SL.
  17. 17.M Anon. 2022. National vulnerability database. https://www.nist.gov/programs-projects/national-vulnerability-database-nvd.
  18. 18.Anthropic. 2023. Claude. https://www.anthropic.com/claude.
  19. 19.Shushan Arakelyan, Rocktim Jyoti Das, Yi Mao, and Xiang Ren. 2023. Exploring Distributional Shifts in Large Language Models for Code Analysis. arXiv preprint arXiv:2303.09128 (2023).
  20. 20.Amos Azaria, Rina Azoulay, and Shulamit Reches. 2023. ChatGPT is a Remarkable Tool–For Experts. arXiv preprint arXiv:2306.03102 (2023).
  21. 21.Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Arun Iyer, Suresh Parthasarathy, Sriram Rajamani, B Ashok, Shashank Shet, et al. 2023. Codeplan: Repository-level coding using llms and planning. arXiv preprint arXiv:2309.12499 (2023).
  22. 22.Patrick Bareiß, Beatriz Souza, Marcelo d’Amorim, and Michael Pradel. 2022. Code generation tools (almost) for free? a study of few-shot, pre-trained language models on code. arXiv preprint arXiv:2206.01335 (2022).
  23. 23.Rabih Bashroush, Muhammad Garba, Rick Rabiser, Iris Groher, and Goetz Botterweck. 2017. Case tool support for variability management in software product lines. ACM Computing Surveys (CSUR) 50, 1 (2017), 1–45.
  24. 24.Ira D Baxter, Andrew Yahin, Leonardo Moura, Marcelo Sant’Anna, and Lorraine Bier. 1998. Clone detection using abstract syntax trees. In Proceedings. International Conference on Software Maintenance (Cat. No. 98CB36272). IEEE, 368–377.
  25. 25.Stas Bekman. 2022. The Technology Behind BLOOM Training. https://huggingface.co/blog/bloom-megatron-deepspeed.
  26. 26.Eeshita Biswas, Mehmet Efruz Karabulut, Lori Pollock, and K Vijay-Shanker. 2020. Achieving reliable sentiment analysis in the software engineering domain using bert. In 2020 IEEE International conference on software maintenance and evolution (ICSME). IEEE, 162–173.
  27. 27.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745 (2022).
  28. 28.Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow. https://doi.org/10.5281/zenodo.5297715
  29. 29.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901.
  30. 30.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712 (2023).
  31. 31.Nghi DQ Bui, Hung Le, Yue Wang, Junnan Li, Akhilesh Deepak Gotmare, and Steven CH Hoi. 2023. CodeTF: One-stop Transformer Library for State-of-the-art Code LLM. arXiv preprint arXiv:2306.00029 (2023).
  32. 32.Alessio Buscemi. 2023. A Comparative Study of Code Generation using ChatGPT 3.5 across 10 Programming Languages. arXiv preprint arXiv:2308.04477 (2023).
  33. 33.Jialun Cao, Meiziniu Li, Ming Wen, and Shing-chi Cheung. 2023. A study on prompt design, advantages and limitations of chatgpt for deep learning program repair. arXiv preprint arXiv:2304.08191 (2023).
  34. 34.Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, et al. 2023. MultiPL-E: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering (2023).
  35. 35.Aaron Chan, Anant Kharkar, Roshanak Zilouchian Moghaddam, Yevhen Mohylevskyy, Alec Helyar, Eslam Kamal, Mohamed Elkamhawy, and Neel Sundaresan. 2023. Transformer-based Vulnerability Detection in Code at EditTime: Zero-shot, Few-shot, or Fine-tuning? arXiv preprint arXiv:2306.01754 (2023).
  36. 36.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2023. A survey on evaluation of large language models. arXiv preprint arXiv:2307.03109 (2023).
  37. 37.Yiannis Charalambous, Norbert Tihanyi, Ridhi Jain, Youcheng Sun, Mohamed Amine Ferrag, and Lucas C Cordeiro. 2023. A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification. arXiv preprint arXiv:2305.14752 (2023).
  38. 38.Angelica Chen, Jérémy Scheurer, Tomasz Korbak, Jon Ander Campos, Jun Shern Chan, Samuel R Bowman, Kyunghyun Cho, and Ethan Perez. 2023. Improving code generation by training with natural language feedback. arXiv preprint arXiv:2303.16749 (2023).
  39. 39.Boyuan Chen, Jian Song, Peng Xu, Xing Hu, and Zhen Ming Jiang. 2018. An automated approach to estimating code coverage measures via execution logs. In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering. 305–316.
  40. 40.Fuxiang Chen, Fatemeh H Fard, David Lo, and Timofey Bryksin. 2022. On the transferability of pre-trained language models for low-resource programming languages. In Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension. 401–412.
  41. 41.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021).
  42. 42.Meng Chen, Hongyu Zhang, Chengcheng Wan, Zhao Wei, Yong Xu, Juhong Wang, and Xiaodong Gu. 2023. On the effectiveness of large language models in domain-specific code generation. arXiv preprint arXiv:2312.01639 (2023).
  43. 43.Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128 (2023).
  44. 44.Xinyun Chen, Chang Liu, and Dawn Song. 2017. Towards synthesizing complex programs from input-output examples. arXiv preprint arXiv:1706.01284 (2017).
  45. 45.Xinyun Chen, Dawn Song, and Yuandong Tian. 2021. Latent execution for neural program synthesis beyond domain-specific languages. Advances in Neural Information Processing Systems 34 (2021), 22196–22208.
  46. 46.Yizheng Chen, Zhoujie Ding, Xinyun Chen, and David Wagner. 2023. DiverseVul: A New Vulnerable Source Code Dataset for Deep Learning Based Vulnerability Detection. arXiv preprint arXiv:2304.00409 (2023).
  47. 47.Yujia Chen, Cuiyun Gao, Muyijie Zhu, Qing Liao, Yong Wang, and Guoai Xu. 2024. APIGen: Generative API Method Recommendation. arXiv preprint arXiv:2401.15843 (2024).
  48. 48.Liying Cheng, Xingxuan Li, and Lidong Bing. 2023. Is GPT-4 a Good Data Analyst? arXiv preprint arXiv:2305.15038 (2023).
  49. 49.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023) (2023).
  50. 50.Muslim Chochlov, Gul Aftab Ahmed, James Vincent Patten, Guoxian Lu, Wei Hou, David Gregg, and Jim Buckley. 2022. Using a Nearest-Neighbour, BERT-Based Approach for Scalable Clone Detection. In 2022 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 582–591.
  51. 51.Yiu Wai Chow, Luca Di Grazia, and Michael Pradel. 2024. PyTy: Repairing Static Type Errors in Python. arXiv preprint arXiv:2401.06619 (2024).
  52. 52.Agnieszka Ciborowska and Kostadin Damevski. 2022. Fast changeset-based bug localization with BERT. In Proceedings of the 44th International Conference on Software Engineering. 946–957.
  53. 53.Agnieszka Ciborowska and Kostadin Damevski. 2023. Too Few Bug Reports? Exploring Data Augmentation for Improved Changeset-based Bug Localization. arXiv preprint arXiv:2305.16430 (2023).
  54. 54.Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Antonio Mastropaolo, Emad Aghajani, Denys Poshyvanyk, Massimiliano Di Penta, and Gabriele Bavota. 2021. An empirical study on the usage of transformer models for code completion. IEEE Transactions on Software Engineering 48, 12 (2021), 4818–4837.
  55. 55.Colin B Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan. 2020. PyMT5: multi-mode translation of natural language and Python code with transformers. arXiv preprint arXiv:2010.03150 (2020).
  56. 56.Arghavan Moradi Dakhel, Amin Nikanjam, Vahid Majdinasab, Foutse Khomh, and Michel C Desmarais. 2023. Effective test generation using pre-trained large language models and mutation testing. arXiv preprint arXiv:2308.16557 (2023).
  57. 57.Pantazis Deligiannis, Akash Lal, Nikita Mehrotra, and Aseem Rastogi. 2023. Fixing rust compilation errors using llms. arXiv preprint arXiv:2308.05177 (2023).
  58. 58.Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023. Jailbreaker: Automated Jailbreak Across Multiple Large Language Model Chatbots. arXiv preprint arXiv:2307.08715 (2023).
  59. 59.Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass. 2023. PentestGPT: An LLM-empowered Automatic Penetration Testing Tool. arXiv preprint arXiv:2308.06782 (2023).
  60. 60.Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang. 2023. Large Language Models are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2023).
  61. 61.Yinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang, Shujing Yang, and Lingming Zhang. 2023. Large language models are edge-case fuzzers: Testing deep learning libraries via fuzzgpt. arXiv preprint arXiv:2304.02014 (2023).
  62. 62.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2028).

Citation

MLA
Hou, X., et al. “Large Language Models for Software Engineering: A Systematic Literature Review”. arXiv, 2023, http://arxiv.org/abs/2308.10620v6.
APA
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J., & Wang, H. (2023). Large Language Models for Software Engineering: A Systematic Literature Review. arXiv. http://arxiv.org/abs/2308.10620v6
Chicago
Hou, X., Y. Zhao, Y. Liu, et al. 2023. “Large Language Models for Software Engineering: A Systematic Literature Review”. arXiv. http://arxiv.org/abs/2308.10620v6.
Harvard
Hou, X. et al. (2023) “Large Language Models for Software Engineering: A Systematic Literature Review”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.10620v6.
Vancouver
1. Hou X, Zhao Y, Liu Y, Yang Z, Wang K, Li L, Luo X, Lo D, Grundy J, Wang H (2023) Large Language Models for Software Engineering: A Systematic Literature Review. arXiv

BibTeX

@article{hou2023large,
  title = {Large Language Models for Software Engineering: A Systematic Literature Review},
  author = {Hou, Xinyi and Zhao, Yanjie and Liu, Yue and Yang, Zhou and Wang, Kailong and Li, Li and Luo, Xiapu and Lo, David and Grundy, John and Wang, Haoyu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.10620v6},
  eprint = {2308.10620}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF