code2vec: learning distributed representations of code

Uri AlonMeital ZilbersteinOmer LevyEran Yahav

article2018Proc. ACM Program. Lang.1,388 citations

Introduces code2vec, a neural technique that learns distributed code embeddings by aggregating abstract syntax tree paths into fixed-length vectors, substantially outperforming prior methods in predicting semantic properties and method names across massive codebases.

Listen

Modern software engineering increasingly relies on machine learning to automate development workflows, maintain quality, and manage massive codebases. However, representing discrete, structured source code in a continuous mathematical form suitable for deep learning pipelines has remained a significant hurdle. Prior approaches either treated code as plain text streams—losing syntactic meaning and requiring vast compute to relearn language structure—or relied on rigid, language-specific symbolic models that could not scale across disparate repositories. As software systems grow in scale and complexity, automated semantic analysis is critical to help developers understand codebases, reduce naming errors, and improve software maintenance.

The article introduces and evaluates code2vec, a neural network architecture designed to represent source code snippets as continuous, fixed-length vectors known as code embeddings. The primary objective is to demonstrate that these distributed vector representations accurately capture program semantics by successfully predicting descriptive method names across diverse, real-world software projects.

To achieve this, the approach decomposes code snippets into structural paths extracted from their abstract syntax trees. Each path, combined with its terminal values, forms an atomic path-context. The model feeds these path-contexts into an attention-based neural network that simultaneously learns distributed embeddings for paths and tokens while calculating a dynamic weighted average over all paths. The evaluation leveraged a massive dataset of over 12 million Java methods across more than 10,000 GitHub repositories, comparing performance against existing convolutional networks, recurrent networks, and conditional random fields.

The experimental findings demonstrate major performance improvements across accuracy, speed, and mathematical properties. First, code2vec achieved an F1 score of 58.4% on the full test set, outperforming the best baseline by more than 17% relative improvement and exceeding text-based models by over 75%. Second, the system delivered remarkable operational efficiency, processing 1,000 predictions per second—a rate 200 to 10,000 times faster than competing models requiring expensive search procedures. Third, soft attention proved to be the critical architectural driver; weighting all paths softly outperformed both unweighted averaging (49.4% F1) and hard selection of a single path (38.5% F1). Fourth, the learned vector space successfully captured intuitive semantic similarities and analogies, such as vector combinations yielding related compound operations.

These findings indicate that structural decomposition combined with soft attention offers a scalable, language-agnostic foundation for applying machine learning to code without manual feature engineering. For software organizations, this capability lowers the cost and risk of code maintenance by enabling accurate, high-throughput tools for automated code reviews, API discovery, and semantic code search. Because the attention mechanism reveals which specific code paths drive a prediction, the model provides human-interpretable rationale rather than behaving as an opaque black box.

Organizations seeking to leverage these capabilities should adopt path-attention representations within automated developer tooling, such as continuous integration linters and intelligent search engines. For teams managing large codebases, deploying pre-trained embedding pipelines can accelerate code exploration and flag mismatched method names before code review. However, future deployments should consider pairing this approach with variable de-obfuscation tools when analyzing poorly named or obfuscated source code.

Readers should interpret the results in light of a few operational boundaries. The model relies on a closed target vocabulary, meaning it predicts whole, observed method names rather than generating highly unique, composite names from scratch. Furthermore, performance drops significantly when terminal tokens are obfuscated, as the model heavily utilizes meaningful token names. Within these constraints, the extensive 12-million-method benchmark demonstrates high confidence in the model's ability to accurately summarize standard, real-world code snippets across cross-project environments.

arXiv: 1803.09473
Cover for code2vec: learning distributed representations of code

Abstract

We present a neural model for representing snippets of code as continuous distributed vectors ("code embeddings"). The main idea is to represent a code snippet as a single fixed-length code vector\textit{code vector}, which can be used to predict semantic properties of the snippet. This is performed by decomposing code to a collection of paths in its abstract syntax tree, and learning the atomic representation of each path simultaneously\textit{simultaneously} with learning how to aggregate a set of them. We demonstrate the effectiveness of our approach by using it to predict a method's name from the vector representation of its body. We evaluate our approach by training a model on a dataset of 14M methods. We show that code vectors trained on this dataset can predict method names from files that were completely unobserved during training. Furthermore, we show that our model learns useful method name vectors that capture semantic similarities, combinations, and analogies. Comparing previous techniques over the same data set, our approach obtains a relative improvement of over 75%, being the first to successfully predict method names based on a large, cross-project, corpus. Our trained model, visualizations and vector similarities are available as an interactive online demo at this http URL. The code, data, and trained models are available at this https URL.

Table of Contents

  • 1 Introduction
  • 1.1 Applications
  • 1.2 Challenges: Representation and Attention
  • 1.3 Existing Techniques
  • 1.4 Contributions
  • 2 Overview
  • 2.1 Motivating Example
  • 3 Background - Representing Code using AST Paths
  • 4 Model
  • 4.1 Code as a Bag of Path-Contexts
  • 4.2 Path-Attention Model
  • 4.3 Training
  • 4.4 Using the trained network
  • 4.5 Design Decisions
  • 5 Distributed vs. Symbolic Representations
  • 6 Evaluation
  • 6.1 Quantitative Evaluation
  • 6.2 Evaluation of Alternative Designs
  • 6.3 Data Ablation Study
  • 6.4 Qualitative Evaluation
  • 6.4.1 Interpreting Attention
  • 6.4.2 Semantic Properties of the Learned Embeddings
  • 7 Limitations of our model
  • 8 Related Work
  • 9 Conclusion
  • References

Knowls

  1. Knowl 1 — AST Paths and Path-Contexts for Code Snippet Representation

    definition

    A code snippet CC is represented structurally through syntactic paths extracted from its Abstract Syntax Tree (AST). Formally, an AST is a tuple ⟨N,T,X,s,δ,ϕ⟩\langle N, T, X, s, \delta, \phi \rangle, where NN is a set of nonterminal nodes, TT is a set of terminal nodes, XX is a set of token values, s∈Ns \in N is the root node, δ:N→(N∪T)∗\delta: N \to (N \cup T)^* maps a nonterminal node to an ordered sequence of its children, and ϕ:T→X\phi: T \to X maps a terminal node to its associated value. Every node other than the root appears exactly once in the children lists.

    An AST path pp of length kk is a sequence n1d1n2d2…nkdknk+1n_1 d_1 n_2 d_2 \dots n_k d_k n_{k+1}, where n1,nk+1∈Tn_1, n_{k+1} \in T are terminal leaf nodes, ni∈Nn_i \in N (for 2≤i≤k2 \le i \le k) are nonterminal nodes, and di∈{↑,↓}d_i \in \{\uparrow, \downarrow\} denotes the traversal direction between adjacent nodes (di=↑d_i = \uparrow indicates moving from child to parent, and di=↓d_i = \downarrow indicates moving from parent to child). The starting terminal of pp is denoted start(p)=n1\text{start}(p) = n_1 and its terminating terminal is denoted end(p)=nk+1\text{end}(p) = n_{k+1}.

    A path-context is a triplet ⟨xs,p,xt⟩\langle x_s, p, x_t \rangle, where xs=ϕ(start(p))x_s = \phi(\text{start}(p)) and xt=ϕ(end(p))x_t = \phi(\text{end}(p)) are the terminal values connected by the syntactic path pp. A code snippet CC is represented as an unordered multiset (bag) of path-contexts Rep(C)\text{Rep}(C) derived from all pairs of distinct terminal nodes in its AST:

    Rep(C)={⟨xs,p,xt⟩∣∃(terms,termt)∈TPairs(C):xs=ϕ(terms)∧xt=ϕ(termt)∧start(p)=terms∧end(p)=termt}\text{Rep}(C) = \left\{ \langle x_s, p, x_t \rangle \mid \exists (\text{term}_s, \text{term}_t) \in \text{TPairs}(C) : x_s = \phi(\text{term}_s) \land x_t = \phi(\text{term}_t) \land \text{start}(p) = \text{term}_s \land \text{end}(p) = \text{term}_t \right\}

    To bound vocabulary size and avoid combinatorial sparsity, paths are constrained by hyperparameters specifying a maximum path length kk and a maximum path width (the horizontal index difference between child branches of a common ancestor node).

  2. Knowl 2 — Path-Attention Neural Network Architecture (code2vec)

    model/method

    The code2vec architecture is an attention-based neural network designed to encode an arbitrary-sized code snippet CC into a single, fixed-length continuous vector v∈Rdv \in \mathbb{R}^d and predict an associated semantic label LL (such as a method name). The workflow consists of four stages:

    1. Extraction and Lookup: The source code is parsed into an AST, and up to nn path-contexts B={b1,…,bn}B = \{b_1, \dots, b_n\} are sampled, where each bi=⟨xs,pj,xt⟩b_i = \langle x_s, p_j, x_t \rangle. The token values xs,xtx_s, x_t and syntactic path pjp_j are looked up in learned embedding matrices value_vocab∈R∣X∣×d\text{value\_vocab} \in \mathbb{R}^{|X| \times d} and path_vocab∈R∣P∣×d\text{path\_vocab} \in \mathbb{R}^{|P| \times d}.
    2. Context Compression: The concatenated embeddings of each path-context are mapped through a fully connected layer with a tanh⁡\tanh activation function to produce a combined context vector c~i∈Rd\tilde{c}_i \in \mathbb{R}^d.
    3. Soft Attention Aggregation: A global learned attention vector a∈Rda \in \mathbb{R}^d computes normalized scalar attention weights αi\alpha_i over all combined context vectors {c~1,…,c~n}\{\tilde{c}_1, \dots, \tilde{c}_n\}. The code snippet representation vv is calculated as the attention-weighted sum of the combined context vectors.
    4. Label Prediction: The code vector vv is projected against a target label embedding matrix tags_vocab∈R∣Y∣×d\text{tags\_vocab} \in \mathbb{R}^{|Y| \times d} via a softmax-normalized dot product to generate a probability distribution over target labels.
  3. Knowl 3 — Path-Context Embedding and Nonlinear Combination in code2vec

    equation

    Given a path-context bi=⟨xs,pj,xt⟩b_i = \langle x_s, p_j, x_t \rangle, where xs,xt∈Xx_s, x_t \in X are the start and end terminal token values and pj∈Pp_j \in P is the AST path connecting them, the network constructs an initial context vector ci∈R3dc_i \in \mathbb{R}^{3d} by concatenating their respective dd-dimensional embedding vectors:

    ci=[value_vocabs;path_vocabj;value_vocabt]∈R3dc_i = [\text{value\_vocab}_s ; \text{path\_vocab}_j ; \text{value\_vocab}_t] \in \mathbb{R}^{3d}

    where value_vocab∈R∣X∣×d\text{value\_vocab} \in \mathbb{R}^{|X| \times d} and path_vocab∈R∣P∣×d\text{path\_vocab} \in \mathbb{R}^{|P| \times d} are learned lookup tables.

    The context vector cic_i is compressed into a combined context vector c~i∈Rd\tilde{c}_i \in \mathbb{R}^d through a fully connected linear layer parameterized by matrix W∈Rd×3dW \in \mathbb{R}^{d \times 3d}, followed by an element-wise hyperbolic tangent activation function:

    c~i=tanh⁡(W⋅ci)\tilde{c}_i = \tanh(W \cdot c_i)

    This nonlinear combination allows the model to evaluate the interaction between a specific AST path and its boundary tokens, assigning different semantic meanings and attention weights to the same path when paired with different identifier values.

  4. Knowl 4 — Soft-Attention Weighting and Code Vector Aggregation

    equation

    Given a bag of combined context vectors {c~1,c~2,…,c~n}\{\tilde{c}_1, \tilde{c}_2, \dots, \tilde{c}_n\} with c~i∈Rd\tilde{c}_i \in \mathbb{R}^d, the model computes a normalized scalar attention score αi∈[0,1]\alpha_i \in [0, 1] for each path-context using a softmax over the inner products with a global learned attention vector a∈Rda \in \mathbb{R}^d:

    αi=exp⁡(c~iTa)∑j=1nexp⁡(c~jTa)\alpha_i = \frac{\exp(\tilde{c}_i^T a)}{\sum_{j=1}^n \exp(\tilde{c}_j^T a)}

    such that ∑i=1nαi=1\sum_{i=1}^n \alpha_i = 1.

    The overall code vector v∈Rdv \in \mathbb{R}^d representing the entire code snippet is the attention-weighted linear combination of the combined context vectors:

    v=∑i=1nαic~iv = \sum_{i=1}^n \alpha_i \tilde{c}_i

  5. Knowl 5 — Label Prediction Distribution and Cross-Entropy Training Objective

    equation

    Given the aggregated code vector v∈Rdv \in \mathbb{R}^d, the conditional probability distribution q(yi)=P(yi∣C)q(y_i) = P(y_i \mid C) over a predefined target label vocabulary YY is defined as the softmax-normalized dot product between vv and the learned label embeddings in tags_vocab∈R∣Y∣×d\text{tags\_vocab} \in \mathbb{R}^{|Y| \times d}:

    q(yi)=exp⁡(vT⋅tags_vocabi)∑yj∈Yexp⁡(vT⋅tags_vocabj)q(y_i) = \frac{\exp(v^T \cdot \text{tags\_vocab}_i)}{\sum_{y_j \in Y} \exp(v^T \cdot \text{tags\_vocab}_j)}

    The network parameters (including all embedding matrices value_vocab\text{value\_vocab}, path_vocab\text{path\_vocab}, tags_vocab\text{tags\_vocab}, combination matrix WW, and attention vector aa) are trained end-to-end to minimize the cross-entropy loss L\mathcal{L} between the predicted distribution qq and the ground-truth distribution pp (where p(ytrue)=1p(y_{\text{true}}) = 1 and 00 otherwise):

    L=−∑y∈Yp(y)log⁡q(y)=−log⁡q(ytrue)\mathcal{L} = -\sum_{y \in Y} p(y) \log q(y) = -\log q(y_{\text{true}})

    At inference time, label prediction for an unseen code snippet CC selects the target label maximizing conditional probability: L^=argmax⁡yL∈YqvC(yL)\hat{L} = \operatorname{argmax}_{y_L \in Y} q_{v_C}(y_L).

  6. Knowl 6 — Parameter Complexity Advantage of Distributed over Symbolic Code Path Representations

    theoretical result

    Representing code properties using discrete, symbolic factor models (such as Conditional Random Fields operating on AST paths) requires learning an individual parameter for every observed combination of start terminal value, syntactic path, end terminal value, and target label. For a vocabulary of terminal values XX, AST paths PP, and labels YY, ternary-factor CRFs require a parameter space complexity of:

    O(∣X∣2⋅∣P∣⋅∣Y∣)O(|X|^2 \cdot |P| \cdot |Y|)

    In contrast, the code2vec distributed representation assigns low-dimensional vectors of dimension dd to individual atomic entities and uses neural layers to compose them algebraically. The resulting parameter space complexity is linear in the vocabulary sizes:

    O(d⋅(∣X∣+∣P∣+∣Y∣))O(d \cdot (|X| + |P| + |Y|))

    Because ∣X∣|X|, ∣P∣|P|, and ∣Y∣|Y| typically number in the millions in large software corpora, the distributed approach reduces space complexity from polynomial to linear, while retaining the ability to score previously unobserved combinations of known atomic elements.

  7. Knowl 7 — Cross-Project Method Name Prediction Performance and Inference Throughput

    data/table

    The code2vec model was evaluated on a dataset of 10,072 open-source Java GitHub repositories (split into 12,636,998 training, 371,364 validation, and 368,445 test methods) for cross-project method name prediction. Performance was measured by precision, recall, and F1 score computed over case-insensitive sub-tokens of method names.

    Sampled Test Set Full Test Set Prediction Rate
    Model Precision Recall F1 Precision Recall F1 (examples/sec)
    CNN+Attention 47.3 29.4 33.9 - - - 0.1
    LSTM+Attention 27.5 21.5 24.1 33.7 22.0 26.6 5
    Paths+CRFs - - - 53.6 46.6 49.9 10
    PathAttention (code2vec) 63.3 56.2 59.5 63.1 54.4 58.4 1000

    code2vec achieves an F1 score of 58.4 on the full test set, outperforming the previous state-of-the-art Paths+CRFs by 17% relative F1 and token-stream neural models (CNN+Attention and LSTM+Attention) by over 75% relative F1. In addition, avoiding sub-token beam search enables a prediction throughput of 1,000 examples per second—two to four orders of magnitude faster than baseline approaches.

  8. Knowl 8 — Ablation of Attention Mechanisms: Soft, Hard, and Element-Wise Attention

    data/table

    To evaluate the impact of the soft-attention mechanism, code2vec was compared against alternative pooling and attention strategies on the method name prediction benchmark:

    Model Design Precision Recall F1
    No-attention (Uniform average) 54.4 45.3 49.4
    Hard attention (Argmax path selection) 42.1 35.4 38.5
    Train-soft, predict-hard 52.7 45.9 49.1
    Soft attention (Standard code2vec) 63.1 54.4 58.4
    Element-wise soft attention 63.7 55.4 59.3

    Key observations:

    1. Soft attention achieves a 9.0 point F1 gain over uniform averaging (No-attention), demonstrating the necessity of learning dynamic path weighting.
    2. Hard attention performs poorly (38.5 F1), showing that no single path-context contains sufficient information to predict method functionality.
    3. Element-wise soft attention (using separate attention vectors aj∈Rda_j \in \mathbb{R}^d for each vector dimension) yields a marginal gain to 59.3 F1 but compromises human interpretability and slows training.
  9. Knowl 9 — Ablation of Path-Context Components: Syntactic Paths versus Terminal Values

    data/table

    To analyze the individual contribution of syntactic paths and terminal tokens, an ablation study was conducted on the full test set by masking out specific tuple components with an out-of-vocabulary (UNK) token:

    Configuration Input Representation Precision Recall F1
    Full ⟨xs,p,xt⟩\langle x_s, p, x_t \rangle 63.1 54.4 58.4
    Only-values ⟨xs,__,xt⟩\langle x_s, \text{\_\_}, x_t \rangle 44.9 37.1 40.6
    Value-path ⟨xs,p,__⟩\langle x_s, p, \text{\_\_} \rangle 31.5 30.1 30.7
    No-values (Paths only) ⟨__,p,__⟩\langle \text{\_\_}, p, \text{\_\_} \rangle 12.0 12.6 12.3
    One-value ⟨xs,__,__⟩\langle x_s, \text{\_\_}, \text{\_\_} \rangle 10.6 10.4 10.7

    Results show that the full representation (58.4 F1) exceeds the sum of using paths alone (12.3 F1) and values alone (40.6 F1). Dropping paths incurs a 17.8 point drop in F1, while dropping identifier values degrades performance severely, indicating that syntactic paths provide essential structural disambiguation for identifier tokens.

  10. Knowl 10 — Semantic Compositionality and Analogies in Code Name Vector Spaces

    empirical result

    The distributed target label embeddings learned by code2vec spontaneously capture linguistic compositionality and semantic analogies through linear vector arithmetic in cosine similarity space.

    1. Semantic Combinations: Summing normalized label vectors yields composite method names:

      • vec(equals)+vec(toLowerCase)≈vec(equalsIgnoreCase)\text{vec}(\text{equals}) + \text{vec}(\text{toLowerCase}) \approx \text{vec}(\text{equalsIgnoreCase})
      • vec(get)+vec(value)≈vec(getValue)\text{vec}(\text{get}) + \text{vec}(\text{value}) \approx \text{vec}(\text{getValue})
      • vec(getRequest)+vec(addBody)≈vec(postRequest)\text{vec}(\text{getRequest}) + \text{vec}(\text{addBody}) \approx \text{vec}(\text{postRequest})
      • vec(remove)+vec(add)≈vec(update)\text{vec}(\text{remove}) + \text{vec}(\text{add}) \approx \text{vec}(\text{update})
      • vec(decode)+vec(fromBytes)≈vec(deserialize)\text{vec}(\text{decode}) + \text{vec}(\text{fromBytes}) \approx \text{vec}(\text{deserialize})
    2. Semantic Analogies: Computing vector offsets via 3CosAdd (argmax⁡v∈V(a^−b^+c^)⋅v^\operatorname{argmax}_{v \in V} (\hat{a} - \hat{b} + \hat{c}) \cdot \hat{v}) solves relational analogies between API methods:

      • receive:download::send:upload\text{receive} : \text{download} :: \text{send} : \text{upload}
      • open:connect::close:disconnect\text{open} : \text{connect} :: \text{close} : \text{disconnect}
      • lower:toLowerCase::upper:toUpperCase\text{lower} : \text{toLowerCase} :: \text{upper} : \text{toUpperCase}
      • down:onMouseDown::up:onMouseUp\text{down} : \text{onMouseDown} :: \text{up} : \text{onMouseUp}
      • start:activate::end:deactivate\text{start} : \text{activate} :: \text{end} : \text{deactivate}
  11. Knowl 11 — Limitations of code2vec: Closed Vocabularies, Monolithic Path Sparsity, and Identifier Sensitivity

    limitation

    The code2vec model exhibits three primary architectural limitations:

    1. Closed Label Vocabulary: The model predicts whole label symbols from a fixed dictionary YY established at training time. It cannot generate out-of-vocabulary names or compose unseen neologisms (e.g., predicting long, specialized identifiers like findUserInfoByUserIdAndKey).
    2. Monolithic Path and Token Sparsity: AST paths and identifier tokens are treated as atomic, indivisible symbols without character-level, sub-token, or sub-path sharing. Two AST paths differing by a single node receive distinct embedding rows, leading to large embedding tables (~1.4 GB model footprint) and requiring tens of millions of examples to train.
    3. Dependence on Informative Variable Names: The attention and prediction layers rely heavily on identifier and keyword tokens; when evaluated on code with obfuscated or adversarial variable names, prediction accuracy drops substantially.

Coverage note — Deliberately omitted qualitative visual case studies of AST attention paths on specific Java methods and discussion of the interactive online demo, as they illustrate the core attention mechanism and evaluation captured in the included knowls.

References

  1. 1.Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles Sutton. 2014. Learning Natural Coding Conventions. In Proceedings of the 22Nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE 2014). ACM, New York, NY, USA, 281–293. https://doi.org/10.1145/2635868.2635883
  2. 2.Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles Sutton. 2015a. Suggesting Accurate Method and Class Names. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering (ESEC/FSE 2015). ACM, New York, NY, USA, 38–49. https://doi.org/10.1145/2786805.2786849
  3. 3.Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2017. A Survey of Machine Learning for Big Code and Naturalness. arXiv preprint arXiv:1709.06182 (2017).
  4. 4.Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. 2018. Learning to Represent Programs with Graphs. In ICLR.
  5. 5.Miltiadis Allamanis, Hao Peng, and Charles A. Sutton. 2016. A Convolutional Attention Network for Extreme Summarization of Source Code. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016. 2091–2100. http://jmlr.org/proceedings/papers/v48/allamanis16.html
  6. 6.Miltiadis Allamanis and Charles Sutton. 2013. Mining Source Code Repositories at Massive Scale Using Language Modeling. In Proceedings of the 10th Working Conference on Mining Software Repositories (MSR ’13). IEEE Press, Piscataway, NJ, USA, 207–216. http://dl.acm.org/citation.cfm?id=2487085.2487127
  7. 7.Miltiadis Allamanis and Charles Sutton. 2014. Mining Idioms from Source Code. In Proceedings of the 22Nd ACM SIGSOFT International Symposium on Foundations of Software Engineering (FSE 2014). ACM, New York, NY, USA, 472–483. https://doi.org/10.1145/2635868.2635901
  8. 8.Miltiadis Allamanis, Daniel Tarlow, Andrew D. Gordon, and Yi Wei. 2015b. Bimodal Modelling of Source Code and Natural Language. In Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (ICML’15). JMLR.org, 2123–2132. http://dl.acm.org/citation.cfm?id=3045118.3045344
  9. 9.Uri Alon, Meital Zilberstein, Omer Levy, and Eran Yahav. 2018. A General Path-based Representation for Predicting Program Properties. In Proceedings of the 39th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI 2018). ACM, New York, NY, USA, 404–419. https://doi.org/10.1145/3192366.3192412
  10. 10.Matthew Amodio, Swarat Chaudhuri, and Thomas W. Reps. 2017. Neural Attribute Machines for Program Generation. CoRR abs/1705.09231 (2017). arXiv:1705.09231 http://arxiv.org/abs/1705.09231
  11. 11.Thierry Artieres et al. 2010. Neural conditional random fields. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics. 177–184.
  12. 12.Jimmy Ba, Volodymyr Mnih, and Koray Kavukcuoglu. 2014. Multiple object recognition with visual attention. arXiv preprint arXiv:1412.7755 (2014).
  13. 13.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR abs/1409.0473 (2014). http://arxiv.org/abs/1409.0473
  14. 14.Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio. 2016. End-to-end attention-based large vocabulary speech recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on. IEEE, 4945–4949.
  15. 15.Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003. A Neural Probabilistic Language Model. J. Mach. Learn. Res. 3 (March 2003), 1137–1155. http://dl.acm.org/citation.cfm?id=944919.944966
  16. 16.Pavol Bielik, Veselin Raychev, and Martin T. Vechev. 2016. PHOG: Probabilistic Model for Code. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016. 2933–2942. http://jmlr.org/proceedings/papers/v48/bielik16.html
  17. 17.Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006. Re-evaluation the role of bleu in machine translation research. In 11th Conference of the European Chapter of the Association for Computational Linguistics.
  18. 18.Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015. Attention-based models for speech recognition. In Advances in Neural Information Processing Systems. 577–585.
  19. 19.Ronan Collobert and Jason Weston. 2008. A Unified Architecture for Natural Language Processing: Deep Neural Networks with Multitask Learning. In Proceedings of the 25th International Conference on Machine Learning (ICML ’08). ACM, New York, NY, USA, 160–167. https://doi.org/10.1145/1390156.1390177
  20. 20.Yaniv David, Nimrod Partush, and Eran Yahav. 2016. Statistical Similarity in Binaries. In PLDI’16: Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation.
  21. 21.Yaniv David, Nimrod Partush, and Eran Yahav. 2017. Similarity of Binaries through re-optimization. In PLDI’17: Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation.
  22. 22.Yaniv David and Eran Yahav. 2014. Tracelet-Based Code Search in Executables. In PLDI’14: Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation. 349–360.
  23. 23.Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harshman. 1990. Indexing by latent semantic analysis. Journal of the American Society for Information Science 41, 6 (1990), 391.
  24. 24.Daniel DeFreez, Aditya V. Thakur, and Cindy Rubio-González. 2018. Path-based Function Embedding and Its Application to Error-handling Specification Mining. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2018). ACM, New York, NY, USA, 423–433. https://doi.org/10.1145/3236024.3236059
  25. 25.Greg Durrett and Dan Klein. 2015. Neural CRF Parsing. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Vol. 1. 302–312.
  26. 26.J.R. Firth. 1957. A Synopsis of Linguistic Theory, 1930-1955. https://books.google.co.il/books?id=T8LDtgAACAAJ
  27. 27.Martin Fowler and Kent Beck. 1999. Refactoring: Improving the Design of Existing Code. Addison-Wesley Professional.
  28. 28.Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics. 249–256.
  29. 29.Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Domain adaptation for large-scale sentiment classification: A deep learning approach. In Proceedings of the 28th International Conference on Machine Learning (ICML-11). 513–520.
  30. 30.Tihomir Gvero and Viktor Kuncak. 2015. Synthesizing Java Expressions from Free-form Queries. In Proceedings of the 2015 ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications (OOPSLA 2015). ACM, New York, NY, USA, 416–432. https://doi.org/10.1145/2814270.2814295
  31. 31.Zellig S Harris. 1954. Distributional structure. Word 10, 2-3 (1954), 146–162.
  32. 32.Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching Machines to Read and Comprehend. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1 (NIPS’15). MIT Press, Cambridge, MA, USA, 1693–1701. http://dl.acm.org/citation.cfm?id=2969239.2969428
  33. 33.Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. 2012. On the Naturalness of Software. In Proceedings of the 34th International Conference on Software Engineering (ICSE ’12). IEEE Press, Piscataway, NJ, USA, 837–847. http://dl.acm.org/citation.cfm?id=2337223.2337322
  34. 34.Einar W. Høst and Bjarte M. Østvold. 2009. Debugging Method Names. In Proceedings of the 23rd European Conference on ECOOP 2009 – Object-Oriented Programming (Genoa). Springer-Verlag, Berlin, Heidelberg, 294–317. https://doi.org/10.1007/978-3-642-03013-0_14
  35. 35.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2016. Summarizing Source Code using a Neural Attention Model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Volume 1: Long Papers. http://aclweb.org/anthology/P/P16/P16-1195.pdf
  36. 36.Omer Katz, Ran El-Yaniv, and Eran Yahav. 2016. Estimating Types in Executables using Predictive Modeling. In POPL’16: Proceedings of the ACM SIGPLAN Conference on Principles of Programming Languages.
  37. 37.Omer Katz, Noam Rinetzky, and Eran Yahav. 2018. Statistical Reconstruction of Class Hierarchies in Binaries. In ASPLOS’18: Proceedings of the ACM Conference on Architectural Support for Programming Languages and Operating Systems.
  38. 38.Diederik Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014).
  39. 39.Quoc Le and Tomas Mikolov. 2014. Distributed Representations of Sentences and Documents. In Proceedings of the 31st International Conference on Machine Learning (ICML-14), Tony Jebara and Eric P. Xing (Eds.). JMLR Workshop and Conference Proceedings, 1188–1196. http://jmlr.org/proceedings/papers/v32/le14.pdf
  40. 40.Omer Levy and Yoav Goldberg. 2014a. Linguistic regularities in sparse and explicit word representations. In Proceedings of the 18th Conference on Computational Natural Language Learning. 171–180.
  41. 41.Omer Levy and Yoav Goldberg. 2014b. Neural Word Embeddings as Implicit Matrix Factorization. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada. 2177–2185.
  42. 42.Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017. Zero-Shot Relation Extraction via Reading Comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), Vancouver, Canada, August 3-4, 2017. 333–342. https://doi.org/10.18653/v1/K17-1034
  43. 43.Cristina V. Lopes, Petr Maj, Pedro Martins, Vaibhav Saini, Di Yang, Jakub Zitny, Hitesh Sajnani, and Jan Vitek. 2017. DéJàVu: A Map of Code Duplicates on GitHub. Proc. ACM Program. Lang. 1, OOPSLA, Article 84 (Oct. 2017), 28 pages. https://doi.org/10.1145/3133908
  44. 44.Yanxin Lu, Swarat Chaudhuri, Chris Jermaine, and David Melski. 2017. Data-Driven Program Completion. CoRR abs/1705.09042 (2017). arXiv:1705.09042 http://arxiv.org/abs/1705.09042
  45. 45.Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective Approaches to Attention-based Neural Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015. 1412–1421. http://aclweb.org/anthology/D/D15/D15-1166.pdf
  46. 46.Chris J. Maddison and Daniel Tarlow. 2014. Structured Generative Models of Natural Source Code. In Proceedings of the 31st International Conference on International Conference on Machine Learning - Volume 32 (ICML’14). JMLR.org, II–649–II–657. http://dl.acm.org/citation.cfm?id=3044805.3044965
  47. 47.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a. Efficient Estimation of Word Representations in Vector Space. CoRR abs/1301.3781 (2013). http://arxiv.org/abs/1301.3781
  48. 48.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013b. Distributed Representations of Words and Phrases and Their Compositionality. In Proceedings of the 26th International Conference on Neural Information Processing Systems (NIPS’13). Curran Associates Inc., USA, 3111–3119. http://dl.acm.org/citation.cfm?id=2999792.2999959
  49. 49.Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013c. Linguistic regularities in continuous space word representations.
  50. 50.Alon Mishne, Sharon Shoham, and Eran Yahav. 2012. Typestate-based Semantic Code Search over Partial Programs. In Proceedings of the ACM International Conference on Object Oriented Programming Systems Languages and Applications (OOPSLA ’12). ACM, New York, NY, USA, 997–1016. https://doi.org/10.1145/2384616.2384689
  51. 51.Volodymyr Mnih, Nicolas Heess, Alex Graves, and Koray Kavukcuoglu. 2014. Recurrent Models of Visual Attention. In Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS’14). MIT Press, Cambridge, MA, USA, 2204–2212. http://dl.acm.org/citation.cfm?id=2969033.2969073
  52. 52.Dana Movshovitz-Attias and William W Cohen. 2013. Natural language models for predicting programming comments. (2013).
  53. 53.Vijayaraghavan Murali, Swarat Chaudhuri, and Chris Jermaine. 2017. Bayesian Sketch Learning for Program Synthesis. CoRR abs/1703.05698 (2017). arXiv:1703.05698 http://arxiv.org/abs/1703.05698
  54. 54.Tung Thanh Nguyen, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N. Nguyen. 2013. A Statistical Semantic Language Model for Source Code. In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering (ESEC/FSE 2013). ACM, New York, NY, USA, 532–542. https://doi.org/10.1145/2491411.2491458
  55. 55.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe: Global Vectors for Word Representation. In Empirical Methods in Natural Language Processing (EMNLP). 1532–1543. http://www.aclweb.org/anthology/D14-1162
  56. 56.Veselin Raychev, Pavol Bielik, and Martin Vechev. 2016a. Probabilistic Model for Code with Decision Trees. In Proceedings of the 2016 ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications (OOPSLA 2016). ACM, New York, NY, USA, 731–747. https://doi.org/10.1145/2983990.2984041
  57. 57.Veselin Raychev, Pavol Bielik, Martin Vechev, and Andreas Krause. 2016b. Learning Programs from Noisy Data. In Proceedings of the 43rd Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (POPL ’16). ACM, New York, NY, USA, 761–774. https://doi.org/10.1145/2837614.2837671
  58. 58.Veselin Raychev, Martin Vechev, and Andreas Krause. 2015. Predicting Program Properties from "Big Code". In Proceedings of the 42Nd Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (POPL ’15). ACM, New York, NY, USA, 111–124. https://doi.org/10.1145/2676726.2677009
  59. 59.Veselin Raychev, Martin Vechev, and Eran Yahav. 2014. Code Completion with Statistical Language Models. In Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI ’14). ACM, New York, NY, USA, 419–428. https://doi.org/10.1145/2594291.2594321
  60. 60.Reuven Rubinstein. 1999. The cross-entropy method for combinatorial and continuous optimization. Methodology and Computing in Applied Probability 1, 2 (1999), 127–190.
  61. 61.Reuven Y Rubinstein. 2001. Combinatorial optimization, cross-entropy, ants and rare events. Stochastic Optimization: Algorithms and Applications 54 (2001), 303–363.
  62. 62.Gerard Salton, Anita Wong, and Chung-Shu Yang. 1975. A vector space model for automatic indexing. Commun. ACM 18, 11 (1975), 613–620.
  63. 63.Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2016. Bidirectional attention flow for machine comprehension. arXiv preprint arXiv:1611.01603 (2016).
  64. 64.Richard Socher, Cliff C. Lin, Andrew Y. Ng, and Christopher D. Manning. 2011. Parsing Natural Scenes and Natural Language with Recursive Neural Networks. In Proceedings of the 26th International Conference on Machine Learning (ICML).
  65. 65.Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research 15, 1 (2014), 1929–1958.
  66. 66.Grigorios Tsoumakas and Ioannis Katakis. 2006. Multi-label classification: An overview. International Journal of Data Warehousing and Mining 3, 3 (2006).
  67. 67.Joseph Turian, Lev Ratinov, and Yoshua Bengio. 2010. Word Representations: A Simple and General Method for Semi-supervised Learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics (ACL ’10). Association for Computational Linguistics, Stroudsburg, PA, USA, 384–394. http://dl.acm.org/citation.cfm?id=1858681.1858721
  68. 68.Peter D Turney. 2006. Similarity of semantic relations. Computational Linguistics 32, 3 (2006), 379–416.
  69. 69.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 6000–6010.
  70. 70.Martin T. Vechev and Eran Yahav. 2016. Programming with "Big Code". Foundations and Trends in Programming Languages 3, 4 (2016), 231–284. https://doi.org/10.1561/2500000028
  71. 71.Martin White, Christopher Vendome, Mario Linares-Vásquez, and Denys Poshyvanyk. 2015. Toward Deep Learning Software Repositories. In Proceedings of the 12th Working Conference on Mining Software Repositories (MSR ’15). IEEE Press, Piscataway, NJ, USA, 334–345. http://dl.acm.org/citation.cfm?id=2820518.2820559
  72. 72.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In International Conference on Machine Learning. 2048–2057.
  73. 73.Meital Zilberstein and Eran Yahav. 2016. Leveraging a Corpus of Natural Language Descriptions for Program Similarity. In Proceedings of the 2016 ACM International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software (Onward! 2016). ACM, New York, NY, USA, 197–211. https://doi.org/10.1145/2986012.2986013

Citation

MLA
Alon, U., et al. “Code2vec: Learning Distributed Representations of Code”. arXiv, 2018, http://arxiv.org/abs/1803.09473v5.
APA
Alon, U., Zilberstein, M., Levy, O., & Yahav, E. (2018). code2vec: Learning Distributed Representations of Code. arXiv. http://arxiv.org/abs/1803.09473v5
Chicago
Alon, U., M. Zilberstein, O. Levy, and E. Yahav. 2018. “Code2vec: Learning Distributed Representations of Code”. arXiv. http://arxiv.org/abs/1803.09473v5.
Harvard
Alon, U. et al. (2018) “code2vec: Learning Distributed Representations of Code”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.09473v5.
Vancouver
1. Alon U, Zilberstein M, Levy O, Yahav E (2018) code2vec: Learning Distributed Representations of Code. arXiv

BibTeX

@article{alon2018code2vec,
  title = {code2vec: Learning Distributed Representations of Code},
  author = {Alon, Uri and Zilberstein, Meital and Levy, Omer and Yahav, Eran},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.09473v5},
  eprint = {1803.09473}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF