Large Language Models for Software Engineering: A Systematic Literature Review
Xinying HouYanjie ZhaoYue LiuZhou YangKailong WangLi LiXiapu LuoDavid LoJohn C. GrundyHaoyu Wang
Systematizes findings from nearly 400 studies to explain how large language models are trained, evaluated, and deployed across software engineering tasks, providing clear guidance on effective data practices and open research challenges.
Modern software engineering faces escalating complexity, creating a growing demand for automation across the development lifecycle. Recently, large language models have emerged as powerful tools capable of transforming complex software tasks into data, code, and text analysis problems. However, an understanding of their actual implementation, measurable impact, and operational limitations across the software engineering spectrum has remained fragmented. To establish a clear baseline for decision-makers, the article evaluates the landscape of language models in software engineering by examining their architectures, data practices, optimization methods, and task efficacy.
The article conducts a systematic literature review analyzing 395 peer-reviewed and pre-print research studies published between January 2017 and January 2024. The authors evaluate trends across model architectures—specifically decoder-only, encoder-only, and encoder-decoder models—alongside dataset curation methodologies, optimization strategies such as parameter-efficient fine-tuning and prompt engineering, and performance across 85 distinct software engineering tasks spanning all development stages.
The review reveals several central findings. First, decoder-only models, such as the GPT series, have become dominant, comprising over 70% of research activity by 2023 due to their generative capabilities. In contrast, encoder-only models such as BERT and encoder-decoder models like T5 serve specialized niches in code understanding and translation. Second, research is heavily skewed toward software development (56.65%) and maintenance (22.71%), dominated by generative tasks like code generation and automated program repair, while activities such as requirements engineering (3.90%), software design (0.92%), and software management (0.69%) remain largely unexplored. Third, conversational and feedback-driven prompting frameworks significantly enhance performance; for instance, interactive setups in models like ChatGPT allow continuous refinement that markedly improves patch correctness in program repair. Fourth, parameter-efficient fine-tuning methods (such as low-rank adaptation) and advanced reasoning techniques (such as chain-of-thought prompting) consistently allow models to achieve high accuracy without requiring full, computationally prohibitive parameter retraining.
These findings indicate that adopting language models can substantially decrease developer workloads, accelerate prototyping, and improve automated code quality assurance. However, realizing these benefits requires careful alignment between model architectures and specific tasks, as generative decoder models excel at code synthesis but may prove inefficient for purely analytical tasks. Furthermore, the disproportionate academic reliance on open-source datasets (which represent roughly 62.83% of examined datasets) compared to industrial datasets (found in fewer than 2% of studies) highlights a critical disconnect: techniques proven in academic benchmarks may face unaddressed reliability, security, and integration challenges in proprietary, enterprise-scale environments.
Organizations should focus near-term investments on high-value, validated areas—such as automated code completion, unit test generation, and interactive bug repair—while adopting lightweight parameter-efficient fine-tuning and structured prompt engineering to manage computational costs. Senior leaders should avoid applying these models to early-stage requirements or system architecture decisions without rigorous human oversight until further empirical research emerges. Before broad enterprise deployment, leadership should initiate targeted pilot programs on internal codebases to validate real-world performance against enterprise constraints.
Readers should exercise caution regarding the findings due to the rapidly evolving nature of the field and the inclusion of non-peer-reviewed pre-prints (which accounted for roughly 61% of the analyzed papers, though filtered via quality rubrics). The scarcity of industrial evaluation data remains a notable limitation, meaning confidence is high regarding the conceptual potential of language models for software tasks, but moderate concerning their off-the-shelf reliability in complex, enterprise-grade production pipelines.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). This paper establishes the CodeXGLUE benchmark and foundational code representation models, providing the core task categories and evaluation standards analyzed throughout the literature review.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). This work introduced key program synthesis benchmarks such as MBPP and proved early scaling laws for code generation, serving as a baseline milestone for subsequent LLM4SE research.
- Paper: CodeSearchNet Challenge: Evaluating the State of Semantic Code Search, Hamel Husain et al. (2019). This seminal benchmark and dataset established the standard methodologies for neural code retrieval and code understanding upon which modern code LLMs build.
- Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). This foundational paper presents Code Llama, one of the most prominent open code models whose architectures, infilling training, and repository-level adaptations are systematically surveyed in the review.
- Paper: A Survey of Large Language Models, Wayne Xin Zhao et al. (2023). This comprehensive survey outlines the general pre-training, fine-tuning, and utilization lifecycle of large language models, providing the general NLP context needed to understand domain-specific SE adaptations.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). This benchmark advances LLM4SE evaluation from isolated code snippets to resolving real-world GitHub issues across large software repositories.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). This study addresses the critical evaluation challenges and data contamination issues identified in the survey by introducing a continuously updated, multi-task code benchmark.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). This paper demonstrates practical advancements in open-source code intelligence by pre-training models with repository-level dependencies and fill-in-the-middle objectives.
- Paper: daVinci-Dev: Agent-native Mid-training for Software Engineering, Ji Zeng et al. (2026). This work extends beyond traditional code generation by establishing agent-native mid-training methods tailored for dynamic multi-step software engineering workflows.
- Paper: SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Yuhang Wang et al. (2026). This research addresses efficiency bottlenecks highlighted in software engineering agents by introducing task-aware context pruning for repository-scale problem solving.
- Paper: GLM-5: from Vibe Coding to Agentic Engineering, GLM-5-Team et al. (2026). This work represents the architectural evolution toward autonomous agentic software engineering, executing complex multi-file coding workflows in verifiable environments.
