CodeT5+: Open Code Large Language Models for Code Understanding and Generation
Yue WangHung LeAkhilesh GotmareNghi D. Q. BuiJunnan LiSteven C. H. Hoi
Presents a flexible family of open-source code language models trained with diverse pretraining objectives and efficient initialization strategies, allowing modules to operate in encoder-only, decoder-only, or encoder-decoder modes to achieve top performance across broad code understanding and generation benchmarks.
Software engineering increasingly relies on artificial intelligence to automate complex tasks such as generating code, detecting software vulnerabilities, and searching repositories. However, existing code language models typically suffer from rigid architectures (being restricted to either generation or understanding tasks) and narrow training objectives that degrade downstream performance. Furthermore, training billion-parameter models from scratch requires prohibitive compute and infrastructure costs.
The article introduces and evaluates CodeT5+, a flexible family of open-source neural network models ranging from 220 million to 16 billion parameters. The main objective is to demonstrate that a modular architecture combined with a comprehensive mixture of training objectives can establish superior performance across both code understanding and code generation tasks, while drastically reducing the compute resources needed for model training.
To achieve this, the authors designed a two-stage training approach across nine programming languages using a curated 51.5-billion-token open-source code dataset alongside function-level text-code pairs. The first stage trains the model to recover missing code spans and complete partial programs; the second stage aligns natural language descriptions with code via contrastive learning and cross-modal prediction. To scale the architecture efficiently to larger sizes (2B, 6B, and 16B parameters), the authors initialized the models using off-the-shelf, frozen language models within a shallow encoder and deep decoder setup, tuning only a small fraction of the parameters.
CodeT5+ demonstrated major performance gains across more than 20 benchmarks. First, the instruction-tuned 16-billion-parameter model set a state-of-the-art benchmark on the standard HumanEval zero-shot code generation task with a 35.0% pass@1 rate, outperforming both open-source alternatives and OpenAI's proprietary code-cushman-001 model. Second, compact sub-billion models (220M and 770M) significantly outperformed massive models—including 65-billion to 137-billion-parameter systems—on computational math programming benchmarks, maintaining robust reasoning even as task complexity escalated. Third, CodeT5+ advanced the state-of-the-art in code retrieval across eight datasets (+3.2 average Mean Reciprocal Rank) and line-level code completion (+2.1 exact match). Finally, the architecture functioned effectively as a single, unified system for retrieval-augmented generation, bypassing the need for separate retriever and generator models.
These findings prove that dynamic modular architectures and targeted pretraining objectives allow significantly smaller, open models to match or exceed the performance of much larger, expensive proprietary systems. Organizations can lower compute expenses and operational complexity by deploying a single unified model that flexibly switches between search, understanding, and generation roles rather than maintaining multiple disjoint tools.
Organizations evaluating automated coding assistants should consider adopting CodeT5+ architectures to lower inference costs and maintain complete system control over open-source deployments. Teams should pilot unified retrieval-augmented generation pipelines to leverage internal proprietary repositories directly. Prior to production deployments, teams must institute security screening processes to prevent the introduction of software vulnerabilities and ensure attribution compliance when retrieving open-source code.
Key operational limitations include the high graphics hardware demands required to host and run 16-billion-parameter models at scale, alongside the need for careful data filtering and instruction curation. Nevertheless, the evidence across diverse languages and benchmarks demonstrates high confidence in CodeT5+'s efficiency and functional capabilities.
- Paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, Yue Wang et al. (2021). CodeT5+ directly extends and addresses the architectural and objective limitations of the original CodeT5 unified encoder-decoder model.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). CodeT5+ builds its underlying sequence-to-sequence pretraining and span denoising formulations on the T5 unified text-to-text architecture.
- Paper: CodeBERT: A Pre-Trained Model for Programming and Natural Languages, Zhangyin Feng et al. (2020). CodeBERT establishes the foundational bimodal natural language and programming language representation learning that CodeT5+ incorporates and refines.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). CodeXGLUE provides the standardized suite of code understanding and generation benchmarks used to assess code pretrained models.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). Codex introduces the standard HumanEval benchmark and pass@k evaluation framework that CodeT5+ uses to measure code generation performance.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). CodeGen provides key open-source decoder-only foundation models whose weights are leveraged as off-the-shelf initializations in CodeT5+.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). FLAN establishes instruction-tuning techniques for zero-shot generalization that CodeT5+ adopts for natural language instruction alignment.
- Paper: CodeSearchNet Challenge: Evaluating the State of Semantic Code Search, Hamel Husain et al. (2019). CodeSearchNet supplies the core multilingual bimodal dataset and search evaluation criteria utilized throughout the CodeT5 model family.
- Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). Code Llama extends open-foundation code modeling by scaling decoder architectures with repository-level infilling and long-context instruction tuning.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). DeepSeek-Coder advances open-source code intelligence by pretraining large-scale models from scratch on multi-trillion token multilingual and repository-level corpora.
- Paper: Magicoder: Empowering Code Generation with OSS-Instruct, Yuxiang Wei et al. (2024). Magicoder builds on code instruction-tuning concepts by leveraging open-source code fragments to generate high-diversity synthetic instruction data.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). EvalPlus exposes evaluation limitations in code LLM benchmarks like HumanEval by generating extensive automated test mutations to rigorously assess functional correctness.
- Paper: CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion, Yangruibo Ding et al. (2023). CrossCodeEval moves beyond snippet-level evaluation to test how well code models retrieve and leverage cross-file context across realistic software repositories.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). SWE-bench elevates code generation evaluation from isolated functions to resolving full end-to-end pull requests on real GitHub issues.
- Paper: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models, Siming Huang et al. (2025). OpenCoder continues the open code model paradigm by providing fully transparent pretraining data recipes and staged post-training pipelines.
- Paper: BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions, Terry Yue Zhuo et al. (2025). BigCodeBench provides an advanced benchmark evaluating practical instruction following and complex multi-library function calling in Python.
