CodeT5+: Open Code Large Language Models for Code Understanding and Generation

Yue WangHung LeAkhilesh GotmareNghi D. Q. BuiJunnan LiSteven C. H. Hoi

article2023EMNLP686 citations

Presents a flexible family of open-source code language models trained with diverse pretraining objectives and efficient initialization strategies, allowing modules to operate in encoder-only, decoder-only, or encoder-decoder modes to achieve top performance across broad code understanding and generation benchmarks.

Listen

Software engineering increasingly relies on artificial intelligence to automate complex tasks such as generating code, detecting software vulnerabilities, and searching repositories. However, existing code language models typically suffer from rigid architectures (being restricted to either generation or understanding tasks) and narrow training objectives that degrade downstream performance. Furthermore, training billion-parameter models from scratch requires prohibitive compute and infrastructure costs.

The article introduces and evaluates CodeT5+, a flexible family of open-source neural network models ranging from 220 million to 16 billion parameters. The main objective is to demonstrate that a modular architecture combined with a comprehensive mixture of training objectives can establish superior performance across both code understanding and code generation tasks, while drastically reducing the compute resources needed for model training.

To achieve this, the authors designed a two-stage training approach across nine programming languages using a curated 51.5-billion-token open-source code dataset alongside function-level text-code pairs. The first stage trains the model to recover missing code spans and complete partial programs; the second stage aligns natural language descriptions with code via contrastive learning and cross-modal prediction. To scale the architecture efficiently to larger sizes (2B, 6B, and 16B parameters), the authors initialized the models using off-the-shelf, frozen language models within a shallow encoder and deep decoder setup, tuning only a small fraction of the parameters.

CodeT5+ demonstrated major performance gains across more than 20 benchmarks. First, the instruction-tuned 16-billion-parameter model set a state-of-the-art benchmark on the standard HumanEval zero-shot code generation task with a 35.0% pass@1 rate, outperforming both open-source alternatives and OpenAI's proprietary code-cushman-001 model. Second, compact sub-billion models (220M and 770M) significantly outperformed massive models—including 65-billion to 137-billion-parameter systems—on computational math programming benchmarks, maintaining robust reasoning even as task complexity escalated. Third, CodeT5+ advanced the state-of-the-art in code retrieval across eight datasets (+3.2 average Mean Reciprocal Rank) and line-level code completion (+2.1 exact match). Finally, the architecture functioned effectively as a single, unified system for retrieval-augmented generation, bypassing the need for separate retriever and generator models.

These findings prove that dynamic modular architectures and targeted pretraining objectives allow significantly smaller, open models to match or exceed the performance of much larger, expensive proprietary systems. Organizations can lower compute expenses and operational complexity by deploying a single unified model that flexibly switches between search, understanding, and generation roles rather than maintaining multiple disjoint tools.

Organizations evaluating automated coding assistants should consider adopting CodeT5+ architectures to lower inference costs and maintain complete system control over open-source deployments. Teams should pilot unified retrieval-augmented generation pipelines to leverage internal proprietary repositories directly. Prior to production deployments, teams must institute security screening processes to prevent the introduction of software vulnerabilities and ensure attribution compliance when retrieving open-source code.

Key operational limitations include the high graphics hardware demands required to host and run 16-billion-parameter models at scale, alongside the need for careful data filtering and instruction curation. Nevertheless, the evidence across diverse languages and benchmarks demonstrates high confidence in CodeT5+'s efficiency and functional capabilities.

arXiv: 2305.07922main/CodeT5+
Cover for CodeT5+: Open Code Large Language Models for Code Understanding and Generation

Abstract

Large language models (LLMs) pretrained on vast source code have achieved prominent progress in code intelligence. However, existing code LLMs have two main limitations. First, they often adopt a specific architecture (encoder-only or decoder-only) or rely on a unified encoder-decoder network for different downstream tasks, lacking the flexibility to operate in the optimal architecture for a specific task. Secondly, they often employ a limited set of pretraining objectives which might not be relevant to some tasks and hence result in substantial performance degrade. To address these limitations, we propose “CodeT5+”, a family of encoder-decoder LLMs for code in which component modules can be flexibly combined to suit a wide range of code tasks. Such flexibility is enabled by our proposed mixture of pretraining objectives, which cover span denoising, contrastive learning, text-code matching, and causal LM pretraining tasks, on both unimodal and bimodal multilingual code corpora. Furthermore, we propose to initialize CodeT5+ with frozen off-the-shelf LLMs without training from scratch to efficiently scale up our models, and explore instruction-tuning to align with natural language instructions. We extensively evaluate CodeT5+ on over 20 code-related benchmarks in different settings, including zero-shot, finetuning, and instruction-tuning. We observe state-of-the-art (SoTA) performance on various code-related tasks, and our instruction-tuned CodeT5+ 16B achieves new SoTA results of 35.0% pass@1 and 54.5% pass@10 on the HumanEval code generation task against other open code LLMs, even surpassing the OpenAI code-cushman-001 model.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 CodeT5+: Open Code LLMs
  • 3.1 Unimodal Pretraining on Code Data
  • 3.2 Bimodal Pretraining on Text-code Data
  • 3.3 Compute-efficient Pretraining with Frozen Off-the-shelf LLMs
  • 3.4 Adaptation to Downstream Understanding and Generation Tasks
  • 4 Pretraining and Instruction Tuning
  • 5 Experiments
  • 5.1 Zero-shot Code Generation Evaluation
  • 5.2 Evaluation on Math Programming
  • 5.3 Evaluation on Code Completion
  • 5.4 Evaluation on Text-to-Code Retrieval
  • 5.5 Ablation Study
  • 5.6 Unified Retrieval-Augmented Generation
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A Bimodal Pretraining Details
  • B Pretraining
  • B.1 Pretraining Dataset
  • B.2 Pretraining Setup
  • C Instruction Tuning
  • D Additional Experimental Results
  • D.1 Code Summarization from CodeXGLUE
  • D.2 Code Defect Detection and Clone Detection from CodeXGLUE
  • D.3 More Analysis on Retrieval-augmented Code Generation
  • D.4 Qualitative Results in Math Programming tasks
  • E Finetuning on Downstream Tasks
  • E.1 Text-to-Code Retrieval
  • E.2 Code Summarization
  • E.3 Code Defect Detection
  • E.4 Code Clone Detection
  • E.5 Code Completion
  • E.6 Math Programming
  • E.7 Retrieval-augmented Code Generation

Knowls

  1. Knowl 1 — CodeT5+ Architecture and Modular Operational Modes

    model/method

    CodeT5+ is an encoder-decoder foundation large language model (LLM) family designed to flexibly operate across multiple architectural configurations to support diverse code understanding and generation tasks:

    1. Sequence-to-Sequence (Seq2Seq) Generation Mode: Both encoder and decoder modules are activated. The encoder processes input text or code context through bidirectional self-attention, and the decoder autoregressively generates target sequences (e.g., docstrings for code summarization, or source code from text prompts). In retrieval-augmented generation, the encoder retrieves relevant code snippets, which are prepended to the input prompt for decoder generation.

    2. Decoder-Only Generation Mode: The model operates as an autoregressive decoder by feeding a special [CLM] token to the encoder input while passing the source prefix code directly to the decoder. The encoder parameters and decoder cross-attention layers are frozen, activating roughly half of the total parameter count to perform tasks such as line-level code completion.

    3. Understanding Mode: Tasks are handled either purely through the encoder (mapping [CLS] output embeddings to a classification head for binary defect detection or clone detection) or via joint encoder-decoder execution (where the decoder cross-attends to encoder representations to compute text-code matching scores).

  2. Knowl 2 — Two-Stage Multi-Task Pretraining Pipeline for CodeT5+

    model/method

    CodeT5+ is pretrained using a stage-wise training regime spanning unimodal code corpora and bimodal text-code corpora:

    Stage 1: Unimodal Code Pretraining Pretrained on large-scale source code files across nine programming languages (Python, Java, Ruby, JavaScript, Go, PHP, C, C++, C#) using three self-supervised objectives:

    • Span Denoising: 15% of tokens are masked with sentinel tokens (such as [MASK0]), using whole-word masking with mean span length 3, and the decoder reconstructs the masked spans.
    • Seq2Seq Causal Language Modeling (CLM): The source code is split at a pivot randomly sampled between 10% and 90% of the sequence length. The prefix sequence is prepended with [CLM] and fed to the encoder; the decoder generates the remaining suffix.
    • Decoder-Only CLM: A [CLM] token is passed to the encoder, and the decoder autoregressively generates the entire program token-by-token.

    Stage 2: Bimodal Text-Code Pretraining Pretrained on function-level parallel text-code pairs (e.g., CodeSearchNet) to align cross-modal semantics using four joint objectives:

    • Text-Code Contrastive Learning (LtccL_{tcc}): Activates the encoder to pull parallel (positive) text-code pairs together and push non-parallel (negative) pairs apart using a momentum queue of size 57,600.
    • Text-Code Matching (LtcmL_{tcm}): Activates the decoder with a prepended [Match] token to predict binary semantic alignment between encoder-represented text and decoder-represented code via the final [EOS] token embedding.
    • Dual Multimodal Causal LM (Lt2cL_{t2c} and Lc2tL_{c2t}): Text-to-code generation (prepended with [CDec]) and code-to-text summarization (prepended with [TDec]).
  3. Knowl 3 — Formulation of Bimodal Pretraining Losses for CodeT5+

    equation

    In the second pretraining stage of CodeT5+ on a corpus D\mathcal{D} of text-code pairs (T,C)(T, C), the training loss L\mathcal{L} combines contrastive learning, text-code matching, and dual text-code causal language modeling:

    L=Ltcc+Ltcm+Lt2c+Lc2t\mathcal{L} = \mathcal{L}_{tcc} + \mathcal{L}_{tcm} + \mathcal{L}_{t2c} + \mathcal{L}_{c2t}

    Text-Code Contrastive Loss (Ltcc\mathcal{L}_{tcc}): Given normalized 256-dimensional encoder representations hith_i^t for text TiT_i and hjch_j^c for code CjC_j from a batch of size NN, similarities are computed as si,jt2c=(hit)⊤hjcs_{i,j}^{t2c} = (h_i^t)^\top h_j^c and si,jc2t=(hjc)⊤hits_{i,j}^{c2t} = (h_j^c)^\top h_i^t. The softmax-normalized prediction probabilities with learned temperature τ\tau are:

    pit2c(T)=exp⁡(si,it2c/τ)∑j=1Nexp⁡(si,jt2c/τ),pic2t(C)=exp⁡(si,ic2t/τ)∑j=1Nexp⁡(si,jc2t/τ)p_i^{t2c}(T) = \frac{\exp(s_{i,i}^{t2c} / \tau)}{\sum_{j=1}^N \exp(s_{i,j}^{t2c} / \tau)}, \quad p_i^{c2t}(C) = \frac{\exp(s_{i,i}^{c2t} / \tau)}{\sum_{j=1}^N \exp(s_{i,j}^{c2t} / \tau)}

    Let yt2c(T)y^{t2c}(T) and yc2t(C)y^{c2t}(C) be ground-truth one-hot similarity vectors. The contrastive loss is the cross-entropy HH:

    Ltcc=12E(T,C)∼D[H(yt2c(T),pt2c(T))+H(yc2t(C),pc2t(C))]\mathcal{L}_{tcc} = \frac{1}{2} \mathbb{E}_{(T,C) \sim \mathcal{D}} \left[ H(y^{t2c}(T), p^{t2c}(T)) + H(y^{c2t}(C), p^{c2t}(C)) \right]

    Text-Code Matching Loss (Ltcm\mathcal{L}_{tcm}): The decoder processes code CC while cross-attending to text TT. The output embedding of the trailing [EOS] token is passed through a linear classifier and softmax to yield a 2-class probability ptcm(T,C)p^{tcm}(T, C) for whether (T,C)(T, C) matches:

    Ltcm=E(T,C)∼D[H(ytcm(T,C),ptcm(T,C))]\mathcal{L}_{tcm} = \mathbb{E}_{(T,C) \sim \mathcal{D}} \left[ H(y^{tcm}(T, C), p^{tcm}(T, C)) \right]

    where ytcm(T,C)y^{tcm}(T, C) is the 2D binary ground-truth label vector. Lt2c\mathcal{L}_{t2c} and Lc2t\mathcal{L}_{c2t} are standard autoregressive cross-entropy losses for text-to-code generation and code-to-text generation.

  4. Knowl 4 — Compute-Efficient Scaling via Shallow Encoders and Frozen Decoders

    model/method

    To scale CodeT5+ up to billion-parameter sizes (2B, 6B, and 16B) without pretraining from scratch, CodeT5+ uses an asymmetric "shallow encoder and deep decoder" architecture:

    1. Component Initialization: The encoder is initialized from a smaller pretrained model (CodeGen-mono 350M), while the decoder is initialized from large off-the-shelf decoder-only models (CodeGen-mono 2B, 6B, or 16B).
    2. Cross-Attention Injection: Randomly initialized cross-attention layers are inserted into the decoder blocks after the causal self-attention layers. To minimize compute overhead, cross-attention layers are inserted only at the top-LL decoder layers (L=1L=1).
    3. Frozen Training: During pretraining, all parameters of the deep decoder LLM are frozen in FP16 precision. Only the 350M encoder and the injected cross-attention layers (introducing 36M, 67M, and 151M trainable parameters for the 2B, 6B, and 16B models, respectively) are trained in FP32 precision.
  5. Knowl 5 — Code Generation Performance on HumanEval

    data/table

    Zero-shot code generation is evaluated on the HumanEval benchmark measuring pass@kk (k∈{1,10,100}k \in \{1, 10, 100\}). InstructCodeT5+ 16B (finetuned on ~20k synthetic code instructions) achieves 35.0% pass@1 and 54.5% pass@10, surpassing open-source decoder-only LLMs of comparable size (CodeGen-mono 16B at 29.3% pass@1, StarCoder 15B at 33.6% pass@1) as well as the closed-source OpenAI code-cushman-001 (33.5% pass@1).

    Model Model size pass@1 pass@10 pass@100
    Closed-source models
    LaMDA 137B 14.0 - 47.3
    AlphaCode 1.1B 17.1 28.2 45.3
    MIM 2.7B 30.7 48.2 69.6
    PaLM 62B 15.9 - 46.3
    PaLM 540B 26.2 - 76.2
    code-cushman-001 - 33.5 54.3 77.4
    code-davinci-002 - 47.0 74.9 92.1
    GPT-3.5 - 48.1 - -
    GPT-4 - 67.0 - -
    Open-source models
    GPT-J 6B 11.6 15.7 27.7
    InCoder 6B 15.2 27.8 47.0
    GPT-NeoX 20B 15.4 25.6 41.2
    CodeGeeX 13B 22.9 39.6 60.9
    LLaMA 13B 15.8 - 52.5
    LLaMA 65B 23.7 - 79.3
    Replit 3B 21.9 - -
    StarCoder 15B 33.6 - -
    CodeGen-mono 2B 23.7 36.6 57.0
    CodeGen-mono 6B 26.1 42.3 65.8
    CodeGen-mono 16B 29.3 49.9 75.0
    CodeT5+ 220M 12.0 20.7 31.6
    CodeT5+ 770M 15.5 27.2 42.7
    CodeT5+ 2B 24.2 38.2 57.8
    CodeT5+ 6B 28.0 47.2 69.8
    CodeT5+ 16B 30.9 51.6 76.7
    InstructCodeT5+ 16B 35.0 54.5 77.9
    Open-source models + generation strategies
    StarCoder (prompted) 15B 40.8 - -
    CodeGen-mono w/ CodeT 16B 36.7 59.3 -
    CodeT5+ w/ CodeT 16B 38.5 63.6 77.1
    InstructCodeT5+ w/ CodeT 16B 42.9 67.8 78.7
  6. Knowl 6 — Math Programming Results on MathQA-Python and GSM8K-Python

    data/table

    Math programming tasks require synthesizing Python programs to solve natural language math word problems, evaluated by execution correctness (pass@kk). CodeT5+ 770M achieves 87.4% pass@80 on MathQA-Python and 73.8% pass@100 on GSM8K-Python, outperforming finetuned LaMDA 137B (81.2%) and few-shot Minerva 540B / LLaMA 65B.

    Model Model size MathQA-Python (pass@80) GSM8K-Python (pass@100)
    Few-shot learning results
    code-davinci - 42.0 71.0
    LLaMA 33B - 53.1
    LLaMA 65B - 69.7
    Minerva 62B - 68.5
    Minerva 540B - 78.5
    Finetuning results
    LaMDA 137B 81.2 -
    GPT-Neo 125M 84.7 -
    GPT-Neo 2.7B - 41.4
    CodeGen-mono 350M 83.1 38.7
    CodeGen-mono 2B 85.6 47.8
    CodeT5 220M 71.5 58.4
    CodeT5+ 220M 85.6 70.5
    CodeT5+ 770M 87.4 73.8

    When analyzed by problem complexity (number of required reasoning steps), CodeT5+ maintains stable pass rates across multi-step problems (>10>10 steps), whereas CodeT5 performance drops sharply as step count increases.

  7. Knowl 7 — Multi-Language Text-to-Code Retrieval Performance

    data/table

    Text-to-code retrieval evaluates finding the semantically relevant code function given a natural language query, measured by Mean Reciprocal Rank (MRR). CodeT5+ activates both encoder and decoder to compute matching scores. CodeT5+ 770M achieves a new state-of-the-art overall MRR of 77.4% on CodeSearchNet, 74.0% on CosQA, and 44.7% on AdvTest, outperforming UniXcoder by over 3 MRR points.

    Model CodeSearchNet CosQA AdvTest
    Ruby JS Go Python Java PHP Overall
    CodeBERT 125M 67.9 62.0 88.2 67.2 67.6 62.8 69.3 65.7 27.2
    GraphCodeBERT 125M 70.3 64.4 89.7 69.2 69.1 64.9 71.3 68.4 35.2
    SYNCOBERT 125M 72.2 67.7 91.3 72.4 72.3 67.8 74.0 - 38.3
    UniXcoder 125M 74.0 68.4 91.5 72.0 72.6 67.6 74.4 70.1 41.3
    CodeGen-multi 350M 66.0 62.2 90.0 68.6 70.1 63.9 70.1 64.8 34.8
    PLBART 140M 67.5 61.6 88.7 66.3 66.3 61.1 68.6 65.0 34.7
    CodeT5 220M 71.9 65.5 88.8 69.8 68.6 64.5 71.5 67.8 39.3
    CodeT5+ 220M 77.7 70.8 92.4 75.6 76.1 69.8 77.1 72.7 43.3
    CodeT5+ 770M 78.0 71.3 92.7 75.8 76.2 70.1 77.4 74.0 44.7
  8. Knowl 8 — Line-Level Code Completion in Decoder-Only Mode

    data/table

    Line-level code completion is evaluated on PY150 (Python) and JavaCorpus (Java) from CodeXGLUE using Exact Match (EM) accuracy and Levenshtein Edit Similarity. CodeT5+ operates in decoder-only mode (freezing the encoder and cross-attention), outperforming standard encoder-decoder architectures (CodeT5 220M and PLBART 140M) by large margins.

    Model PY150 JavaCorpus
    EM Edit Sim EM Edit Sim
    CodeGPT 124M 42.37 71.59 30.60 63.45
    UniXcoder 125M 43.12 72.00 32.90 65.78
    CodeGen-multi 350M 42.47 70.67 35.47 69.22
    PLBART 140M 38.01 68.46 26.97 61.59
    CodeT5 220M 36.97 67.12 24.80 58.31
    CodeT5+ 220M 43.42 73.69 35.17 69.48
    CodeT5+ 770M 44.86 74.22 37.90 72.25
  9. Knowl 9 — Unified Retrieval-Augmented Code Generation

    data/table

    Because CodeT5+ natively integrates code retrieval and code generation capabilities into one model, it can serve as a unified retrieval-augmented generator without requiring separate retriever and generator pipelines. The encoder indexes and retrieves candidate code snippets using FAISS L2 distance, which are concatenated via a [SEP] token to the input query for the decoder.

    CodeT5+ 220M outperforms the two-model baseline REDCODER-EXT (GraphCodeBERT retriever + PLBART generator) on Java and Python code generation benchmarks.

    Model Java Python
    EM B4 CB EM B4 CB
    Retrieval-based
    BM25 0.00 4.90 16.00 0.00 6.63 13.49
    SCODE-R 125M 0.00 25.34 26.68 0.00 22.75 23.92
    CodeT5+ 220M 0.00 28.74 31.00 0.00 27.30 26.51
    Generative
    CodeBERT 125M 0.00 8.38 14.52 0.00 4.06 10.42
    GraphCodeBERT 125M 0.00 7.86 14.53 0.00 3.97 10.55
    PLBART 140M 0.00 10.10 14.96 0.00 4.89 12.01
    CodeT5+ 220M 0.00 10.33 20.54 0.00 4.40 13.88
    Retrieval-Augmented Generative
    REDCODER-EXT 125M+140M 10.21 28.98 33.18 9.61 24.43 30.21
    CodeT5+ 220M 11.66 33.83 40.60 11.83 31.14 36.39

    EM: Exact Match (%), B4: BLEU-4, CB: CodeBLEU.

  10. Knowl 10 — Ablation of Pretraining Objectives in CodeT5+

    data/table

    Ablation experiments conducted on CodeT5+ 770M evaluate the necessity of specific pretraining objectives:

    1. Causal LM in Stage 1: Removing causal LM tasks from unimodal pretraining causes severe performance degradation on code completion (PY150 EM drops from 44.9 to 36.2; JavaCorpus EM drops from 37.9 to 24.8) and math programming (MathQA-PY pass@80 drops from 87.4 to 72.3; GSM8K-PY pass@100 drops from 73.8 to 61.4).
    2. Text-Code Matching in Stage 2: Removing the text-code matching objective from bimodal pretraining causes a 2.6-point drop in average retrieval MRR (77.4 to 74.8).
    3. Causal LM in Stage 2: Jointly training with causal LM during bimodal pretraining also provides positive cross-task transfer for retrieval (MRR drops from 77.4 to 76.8 without it).
    Model Code Completion Math Programming
    PY150 (EM) JavaCorpus (EM) MathQA-PY (pass@80) GSM8K-PY (pass@100)
    CodeT5+ 770M 44.9 37.9 87.4 73.8
    a) no causal LM 36.2 24.8 72.3 61.4
    Model Text-to-Code Retrieval (MRR)
    Ruby JS Go Python Java PHP Overall
    CodeT5+ 770M 78.0 71.3 92.7 75.8 76.2 70.1 77.4
    b) no matching 76.2 68.5 91.2 72.8 73.6 66.3 74.8
    no causal LM 77.3 70.6 92.4 75.7 75.6 68.9 76.8
  11. Knowl 11 — Computational and Data Limitations of CodeT5+

    limitation

    CodeT5+ exhibits two primary limitations:

    1. Hardware and Memory Footprint: Operating models at the 16B parameter scale requires significant GPU infrastructure; deploying a 16B model onto a single A100 GPU requires specialized model serving and memory loading techniques to avoid out-of-memory errors.
    2. Data Dependency and Instruction Noise: Pretraining requires rigorous heuristic filtering and license checks across large-scale repositories to ensure quality and remove overlap with benchmark test sets. Additionally, instruction tuning relies on synthetic generation via teacher LLMs (such as text-davinci-003), which can occasionally introduce noisy or factually incorrect target code samples.

Coverage note — None was omitted; all key architectural components, pretraining equations, empirical benchmarks (HumanEval, MathQA/GSM8K, CodeXGLUE completion, retrieval, summarization, defect/clone detection, retrieval-augmented generation), ablations, and limitations are covered.

References

  1. 1.Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified pre-training for program understanding and generation. In NAACL-HLT, pages 2655–2668. Association for Computational Linguistics.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan. 2022. Flamingo: a visual language model for few-shot learning. In NeurIPS.
  3. 3.Miltiadis Allamanis and Charles Sutton. 2013. Mining source code repositories at massive scale using language modeling. In MSR, pages 207–216. IEEE Computer Society.
  4. 4.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2357–2367, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732.
  6. 6.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of the ACL Workshop on Challenges & Perspectives in Creating Large Language Models.
  7. 7.Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca.
  8. 8.Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2023. Codet: Code generation with generated tests. In The Eleventh International Conference on Learning Representations.
  9. 9.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  10. 10.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311.
  11. 11.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. CoRR, abs/2110.14168.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186.
  13. 13.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13042–13054.
  14. 14.Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. Codebert: A pre-trained model for programming and natural languages. In EMNLP (Findings), volume EMNLP 2020 of Findings of ACL, pages 1536–1547. Association for Computational Linguistics.
  15. 15.Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen-tau Yih, Luke Zettlemoyer, and Mike Lewis. 2022. Incoder: A generative model for code infilling and synthesis. CoRR, abs/2204.05999.
  16. 16.Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. Unixcoder: Unified cross-modal pre-training for code representation. In ACL (1), pages 7212–7225. Association for Computational Linguistics.
  17. 17.Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021. Graphcodebert: Pre-training code representations with data flow. In ICLR. OpenReview.net.
  18. 18.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In CVPR, pages 9726–9735. Computer Vision Foundation / IEEE.
  19. 19.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring coding challenge competence with apps. NeurIPS.
  20. 20.Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan. 2021. Cosqa: 20, 000+ web queries for code search and question answering. In ACL/IJCNLP (1), pages 5690–5700. Association for Computational Linguistics.
  21. 21.Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. Codesearchnet challenge: Evaluating the state of semantic code search. CoRR, abs/1909.09436.
  22. 22.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3):535–547.
  23. 23.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In EMNLP (1), pages 6769–6781. Association for Computational Linguistics.
  24. 24.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021. GeDi: Generative discriminator guided sequence generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4929–4952, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  25. 25.Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven C. H. Hoi. 2022. Coderl: Mastering code generation through pretrained models and deep reinforcement learning. In NeurIPS.
  26. 26.Aitor Lewkowycz, Anders Johan Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Venkatesh Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving quantitative reasoning problems with language models. In Advances in Neural Information Processing Systems.
  27. 27.Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022a. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, volume 162 of Proceedings of Machine Learning Research, pages 12888–12900. PMLR.
  28. 28.Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In NeurIPS, pages 9694–9705.
  29. 29.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy V, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2023. Starcoder: may the source be with you! CoRR, abs/2305.06161.
  30. 30.Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. 2022b. Competition-level code generation with alphacode. CoRR, abs/2203.07814.
  31. 31.Chin-Yew Lin and Franz Josef Och. 2004. ORANGE: a method for evaluating automatic evaluation metrics for machine translation. In COLING.
  32. 32.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  33. 33.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR (Poster). OpenReview.net.
  34. 34.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. In NeurIPS Datasets and Benchmarks.
  35. 35.Anh Nguyen, Nikos Karampatziakis, and Weizhu Chen. 2023. Meet in the middle: A new pre-training paradigm. arXiv preprint arXiv:2303.07295.
  36. 36.Ansong Ni, Jeevana Priya Inala, Chenglong Wang, Oleksandr Polozov, Christopher Meek, Dragomir R. Radev, and Jianfeng Gao. 2022. Learning from self-sampled correct and partially-correct programs. CoRR, abs/2205.14318.
  37. 37.Erik Nijkamp, Hiroaki Hayashi, Caiming Xiong, Silvio Savarese, and Yingbo Zhou. 2023a. Codegen2: Lessons for training llms on programming and natural languages. arXiv preprint.
  38. 38.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023b. Codegen: An open large language model for code with multi-turn program synthesis. In The Eleventh International Conference on Learning Representations.
  39. 39.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  40. 40.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  41. 41.Md. Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Retrieval augmented code generation and summarization. In EMNLP (Findings), pages 2719–2734. Association for Computational Linguistics.
  42. 42.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  43. 43.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  44. 44.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD, pages 3505–3506. ACM.
  45. 45.Veselin Raychev, Pavol Bielik, and Martin T. Vechev. 2016. Probabilistic model for code with decision trees. In OOPSLA, pages 731–747. ACM.
  46. 46.replit. 2023. replit-code-v1-3b.
  47. 47.Saleh Soltan, Shankar Ananthakrishnan, Jack FitzGerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith Peris, Stephen Rawls, Andy Rosenbaum, Anna Rumshisky, Chandana Satya Prakash, Mukund Sridhar, Fabian Triefenbach, Apurv Verma, Gökhan Tür, and Prem Natarajan. 2022. Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model. CoRR, abs/2208.01448.
  48. 48.Jeffrey Svajlenko, Judith F Islam, Iman Keivanloo, Chanchal K Roy, and Mohammad Mamun Mia. 2014. Towards a big data curated benchmark of inter-project code clones. In 2014 IEEE International Conference on Software Maintenance and Evolution, pages 476–480. IEEE.
  49. 49.Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. Intellicode compose: code generation using transformer. In ESEC/SIGSOFT FSE, pages 1433–1443. ACM.
  50. 50.M Tabachnyk and S Nikolov. 2022. Ml-enhanced code completion improves developer productivity.
  51. 51.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  52. 52.Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022. Unifying language learning paradigms. CoRR, abs/2205.05131.
  53. 53.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  55. 55.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  56. 56.Weishi Wang, Yue Wang, Shafiq Joty, and Steven C. H. Hoi. 2023. Rap-gen: Retrieval-augmented patch generation with codet5 for automatic program repair. CoRR, abs/2309.06057.
  57. 57.Xin Wang, Yasheng Wang, Fei Mi, Pingyi Zhou, Yao Wan, Xiao Liu, Li Li, Hao Wu, Jin Liu, and Xin Jiang. 2021a. Syncobert: Syntax-guided multi-modal contrastive pre-training for code representation. arXiv preprint arXiv:2108.04556.
  58. 58.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  59. 59.Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021b. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. In EMNLP (1), pages 8696–8708. Association for Computational Linguistics.
  60. 60.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020. Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079.
  61. 61.Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023. Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x.
  62. 62.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206.
  63. 63.Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu. 2019. Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks. Advances in neural information processing systems, 32.

Citation

MLA
Wang, Y., et al. “CodeT5+: Open Code Large Language Models for Code Understanding and Generation”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1069–88, https://doi.org/10.18653/v1/2023.emnlp-main.68.
APA
Wang, Y., Le, H., Gotmare, A., Bui, N., Li, J., & Hoi, S. (2023). CodeT5+: Open Code Large Language Models for Code Understanding and Generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1069–1088. https://doi.org/10.18653/v1/2023.emnlp-main.68
Chicago
Wang, Y., H. Le, A. Gotmare, N. Bui, J. Li, and S. Hoi. 2023. “CodeT5+: Open Code Large Language Models for Code Understanding and Generation”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1069–88. https://doi.org/10.18653/v1/2023.emnlp-main.68.
Harvard
Wang, Y. et al. (2023) “CodeT5+: Open Code Large Language Models for Code Understanding and Generation”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1069–1088. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.68.
Vancouver
1. Wang Y, Le H, Gotmare A, Bui N, Li J, Hoi S (2023) CodeT5+: Open Code Large Language Models for Code Understanding and Generation. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1069–1088

BibTeX

@inproceedings{wang-etal-2023-codet5,
    title = "{C}ode{T}5+: Open Code Large Language Models for Code Understanding and Generation",
    author = "Wang, Yue  and
      Le, Hung  and
      Gotmare, Akhilesh  and
      Bui, Nghi  and
      Li, Junnan  and
      Hoi, Steven",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.68/",
    doi = "10.18653/v1/2023.emnlp-main.68",
    pages = "1069--1088"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/