UniXcoder: Unified Cross-Modal Pre-training for Code Representation

Daya GuoShuai LuNan DuanYanlin WangMing ZhouJian Yin

article2022ACL1,102 citationsBest Paper Award

Presents a unified cross-modal pre-trained model that integrates source code, natural language comments, and linearized abstract syntax trees using prefix adapters and contrastive learning to support code understanding, generation, and completion tasks.

Listen

Artificial intelligence models for source code have become essential for automating software engineering tasks such as code search, defect detection, and automated code generation. However, existing AI models typically face architectural trade-offs: encoder-only designs struggle with text generation, decoder-only architectures perform poorly on code understanding, and standard encoder-decoder structures are inefficient for interactive, line-level code completion. Furthermore, most models treat source code as flat text, largely ignoring the rich structural and semantic information contained in syntax trees and natural language comments.

The article demonstrates and evaluates UniXcoder, a unified cross-modal pre-trained model designed to support code understanding, code generation, and auto-regressive code completion within a single architecture. The primary objective is to show that integrating syntax trees and developer comments via unified attention masking and representation learning substantially improves performance across diverse software intelligence tasks without requiring separate, dedicated models.

The researchers evaluated the model across five downstream programming tasks using nine benchmark datasets covering multiple languages, including Python, Java, JavaScript, PHP, Ruby, and Go. The model processes both code comments and abstract syntax trees (AST), which represent the hierarchical structure of code. To feed these trees into the model efficiently, the authors developed a mathematically proven one-to-one mapping algorithm that linearizes ASTs into standard sequences without losing structural hierarchy. The model was pre-trained on millions of public code functions and comments using standard masked language modeling, unidirectional modeling, denoising, contrastive learning, and cross-modal natural language generation.

The experimental findings show that UniXcoder achieves state-of-the-art results on the majority of evaluated benchmarks. On code understanding tasks, it outperformed strong baselines on clone detection and code search datasets, achieving a Mean Reciprocal Rank of 74.4 on the multi-language CodeSearchNet benchmark. In line-level code completion, it achieved 43.12% exact match accuracy on Python and 32.90% on Java, improving upon dedicated decoder architectures and substantially outperforming standard encoder-decoder baselines. On a newly introduced benchmark evaluating cross-language search without task-specific training (zero-shot code retrieval), UniXcoder achieved an overall Mean Average Precision score of 20.45%, more than doubling the performance of existing baselines like GraphCodeBERT (9.17%). Ablation studies confirmed that removing developer comments, syntax tree inputs, or contrastive pre-training degraded performance across the board.

These findings indicate that unified models can streamline software engineering AI pipelines by replacing multiple fragmented tools with a single model. Organizations can deploy UniXcoder to enhance developer productivity in integrated development environments (IDEs), accelerate cross-language code translation, and improve semantic search across internal codebases. The unified architecture offers lower operational complexity while delivering superior retrieval accuracy and fast inference for auto-completion.

Engineering teams should consider integrating unified pre-trained architectures when building enterprise code assistance tools, particularly where both search and interactive generation are required. However, decision-makers should note that incorporating explicit syntax trees increases sequence lengths by roughly 70%, which the authors mitigated during fine-tuning by dropping non-terminal syntax nodes. While confidence in the model’s understanding and retrieval capabilities is high, its generation scores (BLEU-4) were slightly behind models trained on larger proprietary corpora, indicating that enterprise deployments focused primarily on large-scale code synthesis may benefit from pre-training on larger domain-specific datasets.

arXiv: 2203.03850
Cover for UniXcoder: Unified Cross-Modal Pre-training for Code Representation

Abstract

Pre-trained models for programming languages have recently demonstrated great success on code intelligence. To support both code-related understanding and generation tasks, recent works attempt to pre-train unified encoder-decoder models. However, such encoder-decoder framework is sub-optimal for auto-regressive tasks, especially code completion that requires a decoder-only manner for efficient inference. In this paper, we present UniXcoder, a unified cross-modal pre-trained model for programming language. The model utilizes mask attention matrices with prefix adapters to control the behavior of the model and leverages cross-modal contents like AST and code comment to enhance code representation. To encode AST that is represented as a tree in parallel, we propose a one-to-one mapping method to transform AST in a sequence structure that retains all structural information from the tree. Furthermore, we propose to utilize multi-modal contents to learn representation of code fragment with contrastive learning, and then align representations among programming languages using a cross-modal generation task. We evaluate UniXcoder on five code-related tasks over nine datasets. To further evaluate the performance of code fragment representation, we also construct a dataset for a new task, called zero-shot code-to-code search. Results show that our model achieves state-of-the-art performance on most tasks and analysis reveals that comment and AST can both enhance UniXcoder.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 UniXcoder
  • 3.1 Input Representation
  • 3.2 Model Architecture
  • 3.3 Pre-training Tasks
  • 4 Experiments
  • 4.1 Baselines
  • 4.2 Understanding Tasks
  • 4.3 Generation Tasks
  • 4.4 Code Completion
  • 4.5 Zero-shot Code-to-Code Search
  • 4.6 Model Analysis
  • 5 Conclusion
  • Acknowledgments
  • References
  • A Proof for Mapping Function
  • B Pre-training Setting
  • C Fine-tuning Setting
  • C.1 Clone Detection
  • C.2 Code Search
  • C.3 Code Summarization
  • C.4 Code Generation
  • C.5 Code Completion
  • C.6 Zero-shot Code-to-Code Search

Knowls

  1. Knowl 1 — Unified Multi-Mode Architecture and Prefix Masking in UniXcoder

    model/method

    UniXcoder is a unified pre-trained Transformer model that supports encoder-only, decoder-only, and encoder-decoder behaviors within a single shared-parameter multi-layer Transformer. The model regulates access to context across tokens using self-attention mask matrices M∈Rn×nM \in \mathbb{R}^{n \times n} and designated prefix mode tokens prepended to the input sequence:

    • Encoder-only mode ([Enc] prefix): Every position can attend to every other position by setting all mask entries Mij=0M_{ij} = 0.
    • Decoder-only mode ([Dec] prefix): Employs causal masking where a token at index ii can only attend to preceding tokens and itself (j≤ij \le i), achieved by setting Mij=0M_{ij} = 0 for j≤ij \le i and Mij=−∞M_{ij} = -\infty for j>ij > i.
    • Encoder-decoder mode ([E2D] prefix): Input consists of source tokens concatenated with target tokens. Source tokens attend bidirectionally to all other source tokens (Mij=0M_{ij} = 0), whereas target tokens attend bidirectionally to all source tokens and autoregressively to preceding target tokens and themselves (Mij=−∞M_{ij} = -\infty for target token ii attending to target token j>ij > i).

    For the ll-th Transformer layer with hidden states Hl−1∈Rn×dhH^{l-1} \in \mathbb{R}^{n \times d_h}, multi-head self-attention computes query (QQ), key (KK), and value (VV) projections with projection matrices WQ,WK,WV∈Rdh×dkW^Q, W^K, W^V \in \mathbb{R}^{d_h \times d_k}:

    Q=Hl−1WQ,K=Hl−1WK,V=Hl−1WVQ = H^{l-1}W^Q, \quad K = H^{l-1}W^K, \quad V = H^{l-1}W^V

    head=softmax(QK⊤dk+M)V\text{head} = \text{softmax}\left(\frac{QK^\top}{\sqrt{d_k}} + M\right)V

    where dkd_k is the attention head dimension. This allows UniXcoder to perform bidirectional code understanding, autoregressive line completion, and sequence-to-sequence translation and summarization without requiring distinct architectures.

  2. Knowl 2 — One-to-One AST Flattening Mapping Algorithm

    algorithm

    To encode Abstract Syntax Trees (ASTs) within a Transformer without discarding tree hierarchy or relying on graph neural network message passing, UniXcoder employs a recursive sequence mapping function F\mathcal{F}. The algorithm traverses an AST rooted at root\text{root} and returns a flattened token sequence. Terminal nodes (leaves, representing source tokens or identifiers) are emitted directly by name. Non-terminal nodes (representing grammar rules and syntax categories) are wrapped symmetrically by appending the node name with a left suffix name::left before recursively processing its children, and appending the node name with a right suffix name::right after processing all children.

    function FlattenAST(root)
        seq = empty list
        name = root.name
        if root is a leaf then
            seq.append(name)
        else
            seq.append(name + "::left")
            for each child in root.children do
                seq.extend(FlattenAST(child))
            end for
            seq.append(name + "::right")
        end if
        return seq
    end function

    For example, a subtree representing parameters parameters→(data)\text{parameters} \to (\text{data}) is converted into the sequence ["parameters::left", "(", "data", ")", "parameters::right"].

  3. Knowl 3 — Bijectivity and Structure Preservation of the AST Mapping Function

    theoretical result

    Let T\mathcal{T} denote the set of all ordered, labeled Abstract Syntax Trees and S\mathcal{S} denote the set of token sequences produced by the mapping function F\mathcal{F}. The transformation F:T→S\mathcal{F}: \mathcal{T} \to \mathcal{S} is a bijection (a one-to-one and onto mapping):

    T1≠T2  ⟹  F(T1)≠F(T2)T_1 \ne T_2 \implies \mathcal{F}(T_1) \ne \mathcal{F}(T_2)

    Equivalently, F(T1)=F(T2)  ⟹  T1=T2\mathcal{F}(T_1) = \mathcal{F}(T_2) \implies T_1 = T_2.

    This property ensures that the serialized sequence retains complete structural and hierarchical information of the original syntax tree without ambiguity. In contrast, standard depth-first search (DFS) or breadth-first search (BFS) serializations without distinct opening and closing delimiters are non-injective (multiple distinct trees can map to the exact same traversal sequence), causing structural confusion when processed by sequence models.

  4. Knowl 4 — Code Fragment Representation Learning via MCL and CMG

    model/method

    To produce dense semantic vector embeddings for entire code snippets (useful for retrieval and zero-shot search), UniXcoder derives a code fragment embedding h~i\tilde{h}_i by applying mean pooling over the Transformer's final hidden states corresponding to the code's flattened AST input sequence XX. The embeddings are optimized using two dedicated pre-training objectives:

    1. Multi-modal Contrastive Learning (MCL): Following SimCSE, a positive pair (h~i,h~i+)(\tilde{h}_i, \tilde{h}_i^+) is created by passing the same flattened input sequence through the model twice with independently sampled dropout masks. All other b−1b-1 samples in the mini-batch serve as negative examples:

    lossMCL=−∑i=0b−1log⁡exp⁡(cos⁡(h~i,h~i+)/τ)∑j=0b−1exp⁡(cos⁡(h~i,h~j+)/τ)\text{loss}_{\text{MCL}} = - \sum_{i=0}^{b-1} \log \frac{\exp\left(\cos(\tilde{h}_i, \tilde{h}_i^+) / \tau\right)}{\sum_{j=0}^{b-1} \exp\left(\cos(\tilde{h}_i, \tilde{h}_j^+) / \tau\right)}

    where bb is the batch size, τ\tau is a temperature hyperparameter, and cos⁡(u,v)=u⊤v∥u∥2∥v∥2\cos(u, v) = \frac{u^\top v}{\|u\|_2 \|v\|_2}.

    1. Cross-Modal Generation (CMG): The model is trained autoregressively in encoder-decoder mode to generate the corresponding natural language comment W={w0,w1,…,wm−1}W = \{w_0, w_1, \dots, w_{m-1}\} conditioned on the code's flattened AST sequence XX:

    lossCMG=−∑i=0m−1log⁡p(wi∣X,wt<i)\text{loss}_{\text{CMG}} = - \sum_{i=0}^{m-1} \log p(w_i \mid X, w_{t < i})

    To simultaneously learn natural language representations, the source and target sequences are randomly swapped with 50% probability during CMG pre-training. Natural language descriptions act as a common semantic anchor to align representations across different programming languages.

  5. Knowl 5 — Language Modeling Objectives for Multi-Mode Pre-training

    model/method

    UniXcoder is pre-trained across its three execution modes using three complementary language modeling objectives:

    1. Masked Language Modeling (MLM) for Encoder Mode: A subset SmS_m comprising 15% of input tokens is selected; 80% are replaced with [MASK], 10% with random tokens, and 10% kept unchanged. The objective minimizes the negative log-likelihood of the true tokens given bidirectional context:

    lossMLM=−∑xi∈Smlog⁡p(xi∣Xmask)\text{loss}_{\text{MLM}} = - \sum_{x_i \in S_m} \log p(x_i \mid X^{\text{mask}})

    1. Unidirectional Language Modeling (ULM) for Decoder Mode: Trained autoregressively to predict the next token given preceding tokens, supporting efficient left-to-right code completion:

    lossULM=−∑i=0n−1log⁡p(xi∣xt<i)\text{loss}_{\text{ULM}} = - \sum_{i=0}^{n-1} \log p(x_i \mid x_{t < i})

    1. Denoising Objective (DNS) for Encoder-Decoder Mode: Input sequences of length nn are divided into max⁡(⌊n×rl⌋,1)\max(\lfloor \frac{n \times r}{l} \rfloor, 1) chunks, where corruption rate r=0.15r = 0.15 and average span length l=5l = 5. Spans ranging from 1 to 2l−12l-1 tokens are masked and replaced by sentinel tokens [MASKk][\text{MASK}_k]. The target sequence Y={y0,y1,…,yn−1}Y = \{y_0, y_1, \dots, y_{n-1}\} concatenates all masked spans prefixed with their corresponding sentinel tokens:

    lossDNS=−∑i=0n−1log⁡p(yi∣Xmask,yt<i)\text{loss}_{\text{DNS}} = - \sum_{i=0}^{n-1} \log p(y_i \mid X^{\text{mask}}, y_{t < i})

  6. Knowl 6 — Pre-training AST Non-Terminal Dropping Strategy

    model/method

    Including full flattened Abstract Syntax Trees (ASTs) increases sequence token length by approximately 70% and requires parsing source code at inference time. To eliminate AST parsing and sequence length overhead during downstream fine-tuning and inference, UniXcoder is applied to source code leaves alone (dropping syntax tree nodes).

    To bridge the domain gap between pre-training (with full AST tokens) and downstream fine-tuning (source code only), UniXcoder applies a randomized non-terminal symbol dropping mechanism during multi-modal pre-training: all non-terminal AST tokens (name::left and name::right) in an input sequence are dropped with a probability of 50%, forcing the model to implicitly internalize syntactic structures into the representations of source code leaves.

  7. Knowl 7 — Zero-Shot Cross-Language Code-to-Code Search Task

    definition

    Zero-shot code-to-code search is the task of retrieving semantically equivalent source code snippets written in a target programming language given a query source code snippet written in a source programming language, evaluated without task-specific fine-tuning.

    The benchmark is constructed from the Project CodeNet corpus, containing:

    • 11,744 Ruby functions
    • 15,594 Python functions
    • 23,530 Java functions

    All functions are associated with one of 4,053 distinct programming problems. During evaluation, each function serves as a query to retrieve all functions solving the same problem from the candidate pools across Ruby, Python, and Java. Candidate rankings are computed via cosine similarity between mean-pooled final hidden state vectors:

    score(Cq,Cc)=cos⁡(h~Cq,h~Cc)\text{score}(C_q, C_c) = \cos(\tilde{h}_{C_q}, \tilde{h}_{C_c})

    Performance is measured by Mean Average Precision (MAP).

  8. Knowl 8 — Pre-training Configuration and Temperature-Balanced Language Sampling

    experimental setup

    UniXcoder adopts a 12-layer Transformer with hidden dimension dh=768d_h = 768 and 12 attention heads. The vocabulary consists of 50,000 byte-pair encoding (BPE) subword units plus 1,416 special tokens representing AST non-terminal symbols. Pre-training uses 4 DGX-2 systems (64 ×\times 32GB NVIDIA V100 GPUs) with maximum sequence length 1024, batch size 1024, and Adam optimizer with learning rate 2×10−42 \times 10^{-4}.

    Pre-training proceeds in three successive stages:

    1. Pre-training on the C4 natural language dataset with MLM, ULM, and DNS for 500,000 steps.
    2. Pre-training on 4.1 million unimodal code snippets from CodeSearchNet with MLM, ULM, and DNS for 200,000 steps.
    3. Pre-training on 2.3 million function-comment pairs with extracted Tree-sitter ASTs across six programming languages (Ruby, Java, Python, PHP, Go, JavaScript) for 100,000 steps using all pre-training objectives (MLM, ULM, DNS, MCL, CMG).

    To prevent dominance by high-resource programming languages during multi-modal pre-training, mini-batches are sampled from programming language i∈{1,…,N}i \in \{1, \dots, N\} according to smoothed distribution {qi}\{q_i\}:

    qi=piα∑j=1Npjα,pi=ni∑k=1Nnkq_i = \frac{p_i^\alpha}{\sum_{j=1}^N p_j^\alpha}, \quad p_i = \frac{n_i}{\sum_{k=1}^N n_k}

    where nin_i is the number of training examples for language ii, and the smoothing exponent is α=0.7\alpha = 0.7.

  9. Knowl 9 — Evaluation Results on Code Understanding and Code Search Tasks

    data/table

    UniXcoder was evaluated on code clone detection (POJ-104 using MAP@R, BigCloneBench using precision, recall, and F1-score) and natural language code search (CosQA, AdvTest, CSN using Mean Reciprocal Rank (MRR)).

    Model POJ-104 BigCloneBench CosQA AdvTest CSN
    MAP@R Recall Precision F1-score MRR
    RoBERTa 76.67 95.1 87.8 91.3 60.3 18.3 61.7
    CodeBERT 82.67 94.7 93.4 94.1 65.7 27.2 69.3
    GraphCodeBERT 85.16 94.8 95.2 95.0 68.4 35.2 71.3
    SYNCOBERT 88.24 - - - - 38.3 74.0
    PLBART 86.27 94.8 92.5 93.6 65.0 34.7 68.5
    CodeT5-base 88.65 94.8 94.7 95.0 67.8 39.3 71.5
    UniXcoder 90.52 92.9 97.6 95.2 70.1 41.3 74.4
    - w/o contras 87.83 94.9 94.9 94.9 69.2 40.8 73.6
    - w/o cross-gen 90.51 94.8 95.6 95.2 69.4 40.1 74.0
    - w/o comment 87.05 93.6 96.2 94.9 67.9 40.7 72.6
    - w/o AST 88.74 92.9 97.2 95.0 68.7 40.3 74.2
    - using BFS 89.44 93.4 96.7 95.0 69.3 40.1 74.1
    - using DFS 89.74 94.7 94.6 94.7 69.0 40.2 74.2

    UniXcoder outperforms both encoder-only baselines (CodeBERT, GraphCodeBERT, SYNCOBERT) and encoder-decoder baselines (PLBART, CodeT5-base) across all five datasets, achieving 90.52 MAP@R on POJ-104 and 74.4 MRR on CSN. Ablations demonstrate that contrastive learning (-w/o contras) and comment integration (-w/o comment) provide the largest gains on understanding benchmarks.

  10. Knowl 10 — Evaluation Results on Zero-Shot Code-to-Code Search

    data/table

    Zero-shot cross-lingual code search performance is measured in Mean Average Precision (MAP %) across nine query-target language pairs involving Ruby (R), Python (P), and Java (J) from Project CodeNet. Candidates are ranked using cosine similarity of mean-pooled representations from frozen models.

    Query PL Ruby Python Java Overall
    Target PL R P J R P J R P J
    CodeBERT 13.55 3.18 0.71 3.12 14.39 0.96 0.55 0.42 7.62 4.94
    GraphCodeBERT 17.01 9.29 6.38 5.01 19.34 6.92 1.77 3.50 13.31 9.17
    PLBART 18.60 10.76 1.90 8.27 19.55 1.98 1.47 1.27 10.41 8.25
    CodeT5-base 18.22 10.02 1.81 8.74 17.83 1.58 1.13 0.81 10.18 7.81
    UniXcoder 29.05 26.36 15.16 23.96 30.15 15.07 13.61 14.53 16.12 20.45
    - w/o contras 24.03 17.35 7.12 15.80 22.52 7.31 7.55 7.98 13.92 13.73
    - w/o cross-gen 28.73 24.16 12.92 21.52 26.66 12.60 11.14 10.82 13.75 18.03
    - w/o comment 22.24 15.90 7.50 15.09 19.88 6.54 7.84 7.12 13.20 12.81
    - w/o AST 27.54 23.37 10.17 21.75 27.75 9.94 9.79 9.21 14.06 17.06
    - using BFS 26.67 23.69 13.56 21.31 27.28 13.63 11.90 12.55 14.92 18.39
    - using DFS 27.13 22.65 11.62 20.21 25.92 11.85 9.59 10.19 13.30 16.94

    UniXcoder achieves an overall MAP of 20.45%, outperforming GraphCodeBERT (9.17%), PLBART (8.25%), and CodeT5-base (7.81%) by more than 11 points overall. Cross-language retrieval performance shows that contrastive learning (-w/o contras, dropping to 13.73%) and cross-modal comment alignment (-w/o comment, dropping to 12.81%) are essential for bridging syntactic divergences between programming languages.

  11. Knowl 11 — Evaluation Results on Code Summarization, Generation, and Completion

    data/table

    UniXcoder was evaluated on generative downstream tasks: code summarization (smoothed BLEU-4 averaged over 6 programming languages on CodeXGLUE), code generation (CONCODE Java benchmark evaluating Exact Match (EM) and BLEU-4), and line-level code completion (PY150 and Github JavaCorpus evaluating EM and Levenshtein Edit Similarity).

    Model Summarization CONCODE Generation PY150 Completion JavaCorpus Completion
    BLEU-4 EM BLEU-4 EM Edit Sim EM Edit Sim
    Transformer - - - 38.51 69.01 17.00 50.23
    GPT-2 - 17.35 25.37 41.73 70.60 27.50 60.36
    CodeGPT - 20.10 32.79 42.37 71.59 30.60 63.45
    PLBART 18.32 18.75 36.69 38.01 68.46 26.97 61.59
    CodeT5-small 19.14 21.55 38.13 - - - -
    CodeT5-base 19.55 22.30 40.73 36.97 67.12 24.80 58.31
    UniXcoder 19.30 22.60 38.23 43.12 72.00 32.90 65.78
    - w/o contras 19.20 22.10 37.69 43.02 71.94 32.77 65.71
    - w/o cross-gen 19.27 22.20 35.93 42.66 71.83 32.43 65.63
    - w/o comment 18.97 21.45 37.15 42.18 71.70 32.20 65.44
    - w/o AST 19.33 22.60 38.52 42.56 71.87 32.63 65.66
    - using BFS 19.24 21.75 38.21 42.83 71.85 32.40 65.55
    - using DFS 19.25 22.10 38.06 42.61 71.97 32.87 65.75

    On line-level code completion, UniXcoder outperforms both dedicated decoder-only models (CodeGPT) and encoder-decoder models (CodeT5-base), achieving 43.12% EM on PY150 and 32.90% EM on JavaCorpus (+2.3% absolute gain over CodeGPT). On CONCODE code generation, UniXcoder achieves 22.60% EM accuracy. CodeT5-base achieves higher BLEU scores partially due to possessing double the parameter count (12-layer encoder + 12-layer decoder vs. UniXcoder's 12 layers total) and pre-training on 50% more NL-PL pairs.

  12. Knowl 12 — Comparison of AST Mapping Function against BFS and DFS Traversals

    empirical result

    Replacing the bijective AST mapping function F\mathcal{F} with standard tree traversal algorithms (Breadth-First Search (BFS) or Depth-First Search (DFS)) degrades performance across understanding, generation, and retrieval tasks:

    • Zero-Shot Code Search: Overall MAP falls from 20.45% (with F\mathcal{F}) to 18.39% (BFS) and 16.94% (DFS).
    • POJ-104 Clone Detection: MAP@R decreases from 90.52% (with F\mathcal{F}) to 89.44% (BFS) and 89.74% (DFS).
    • CONCODE Code Generation: Exact match decreases from 22.60% (with F\mathcal{F}) to 21.75% (BFS) and 22.10% (DFS).

    In several tasks, pre-training with BFS or DFS yields lower accuracy than completely omitting AST data during pre-training (-w/o AST, which scored 17.06% on zero-shot search and 88.74% on POJ-104). This demonstrates that structural ambiguity in non-bijective linearizations can introduce misleading structural noise into the Transformer's representations.

Coverage note — None was omitted; all key contributions, architecture designs, algorithms, objectives, theoretical proofs, evaluation benchmarks, and empirical findings are fully covered.

References

  1. 1.Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified pre-training for program understanding and generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2668.
  2. 2.Miltiadis Allamanis and Charles Sutton. 2013. Mining source code repositories at massive scale using language modeling. In 2013 10th Working Conference on Mining Software Repositories (MSR), pages 207–216. IEEE.
  3. 3.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  4. 4.Luca Buratti, Saurabh Pujar, Mihaela Bornea, Scott McCarley, Yunhui Zheng, Gaetano Rossiello, Alessandro Morari, Jim Laredo, Veronika Thost, Yufan Zhuang, et al. 2020. Exploring software naturalness throughneural language models. arXiv preprint arXiv:2006.12641.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  6. 6.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019. Unified language model pre-training for natural language understanding and generation. arXiv preprint arXiv:1905.03197.
  7. 7.Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al. 2020. Codebert: A pre-trained model for programming and natural languages. arXiv preprint arXiv:2002.08155.
  8. 8.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. arXiv preprint arXiv:2104.08821.
  9. 9.Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, LIU Shujie, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al. 2020. Graphcodebert: Pre-training code representations with data flow. In International Conference on Learning Representations.
  10. 10.Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan. 2021. Cosqa: 20,000+ web queries for code search and question answering. arXiv preprint arXiv:2105.13239.
  11. 11.Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436.
  12. 12.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018. Mapping language to code in programmatic context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1643–1652.
  13. 13.Xue Jiang, Zhuoran Zheng, Chen Lyu, Liang Li, and Lei Lyu. 2021. Treebert: A tree-based pre-trained model for programming language. arXiv preprint arXiv:2105.12485.
  14. 14.Aditya Kanade, Petros Maniatis, Gogul Balakrishnan, and Kensen Shi. 2019. Pre-trained contextual embedding of source code. arXiv preprint arXiv:2001.00059.
  15. 15.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461.
  16. 16.Chin-Yew Lin and Franz Josef Och. 2004. Orange: a method for evaluating automatic evaluation metrics for machine translation. In COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics, pages 501–507.
  17. 17.Fang Liu, Ge Li, Yunfei Zhao, and Zhi Jin. 2020. Multitask learning based pre-trained language model for code completion. In Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering, pages 473–485.
  18. 18.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  19. 19.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664.
  20. 20.Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin. 2016. Convolutional neural networks over tree structures for programming language processing. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, pages 1287–1293.
  21. 21.Ruchir Puri, David S Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladmir Zolotov, Julian Dolby, Jie Chen, Mihir Choudhury, Lindsey Decker, et al. 2021. Project codenet: A large-scale ai for code dataset for learning a diversity of coding tasks. arXiv preprint arXiv:2105.12655.
  22. 22.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training.
  23. 23.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  24. 24.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  25. 25.Veselin Raychev, Pavol Bielik, and Martin Vechev. 2016. Probabilistic model for code with decision trees. ACM SIGPLAN Notices, pages 731–747.
  26. 26.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909.
  27. 27.Jeffrey Svajlenko, Judith F Islam, Iman Keivanloo, Chanchal K Roy, and Mohammad Mamun Mia. 2014. Towards a big data curated benchmark of inter-project code clones. In 2014 IEEE International Conference on Software Maintenance and Evolution, pages 476–480. IEEE.
  28. 28.Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. Intellicode compose: Code generation using transformer. arXiv preprint arXiv:2005.08025.
  29. 29.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  30. 30.Xin Wang, Fei Mi Yasheng Wang, Pingyi Zhou, Yao Wan, Xiao Liu, Li Li, Hao Wu, Jin Liu, and Xin Jiang. 2022. Syncobert: Syntax-guided multi-modal contrastive pre-training for code representation.
  31. 31.Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv preprint arXiv:2109.00859.

Citation

MLA
Guo, D., et al. “UniXcoder: Unified Cross-Modal Pre-training for Code Representation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7212–25, https://doi.org/10.18653/v1/2022.acl-long.499.
APA
Guo, D., Lu, S., Duan, N., Wang, Y., Zhou, M., & Yin, J. (2022). UniXcoder: Unified Cross-Modal Pre-training for Code Representation. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7212–7225. https://doi.org/10.18653/v1/2022.acl-long.499
Chicago
Guo, D., S. Lu, N. Duan, Y. Wang, M. Zhou, and J. Yin. 2022. “UniXcoder: Unified Cross-Modal Pre-training for Code Representation”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7212–25. https://doi.org/10.18653/v1/2022.acl-long.499.
Harvard
Guo, D. et al. (2022) “UniXcoder: Unified Cross-Modal Pre-training for Code Representation”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7212–7225. Available at: https://doi.org/10.18653/v1/2022.acl-long.499.
Vancouver
1. Guo D, Lu S, Duan N, Wang Y, Zhou M, Yin J (2022) UniXcoder: Unified Cross-Modal Pre-training for Code Representation. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7212–7225

BibTeX

@inproceedings{guo-etal-2022-unixcoder,
    title = "{U}ni{X}coder: Unified Cross-Modal Pre-training for Code Representation",
    author = "Guo, Daya  and
      Lu, Shuai  and
      Duan, Nan  and
      Wang, Yanlin  and
      Zhou, Ming  and
      Yin, Jian",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.499/",
    doi = "10.18653/v1/2022.acl-long.499",
    pages = "7212--7225"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/