CodeSearchNet Challenge: Evaluating the State of Semantic Code Search

Hamel HusainHongqiu WuTiferet GazitMiltiadis AllamanisMarc Brockschmidt

article2019arXiv1,473 citations

Introduces a foundational benchmark for semantic code search, providing a six-million-function dataset across six programming languages, expert-annotated evaluation queries, and baseline models.

Listen

Software developers frequently search for existing code using natural language queries, but traditional search engines struggle because programming code uses highly technical, abbreviated syntax that shares little vocabulary with everyday language. At the same time, progress in applying modern artificial intelligence to code search has been hindered by the lack of large, standardized datasets and rigorous evaluation benchmarks.

The article introduces a comprehensive benchmark, the CodeSearchNet Challenge, along with an extensive open-source corpus to evaluate how effectively machine learning models retrieve relevant source code from natural language queries.

To construct this resource, the authors curated a corpus of approximately 6 million functions from open-source repositories across six programming languages (Go, Java, JavaScript, PHP, Python, and Ruby). About 2 million of these functions included natural language documentation, which was cleaned and used to train neural retrieval models. To build a ground-truth evaluation set, the authors compiled 99 representative natural language queries and collected 4,026 human expert annotations assessing the relevance of candidate code snippets.

Evaluation of several baseline retrieval architectures revealed key findings. First, while complex self-attention models achieved the highest scores on the training proxy task (ranking documentation against code), simpler architectures excelled during human-annotated evaluation. The Neural Bag of Words model achieved the highest ranking score among all baselines on the human evaluation benchmark (achieving a 0.574 normalized discounted cumulative gain score within annotated candidates, and 0.340 across the entire corpus). Second, standard keyword-based search using ElasticSearch proved highly competitive, demonstrating that exact keyword matching and handling rare technical terms remain critical strengths. Third, more complex models such as bidirectional recurrent neural networks performed significantly worse, scoring roughly 0.145 on the annotated challenge.

These findings indicate that using code comments as a proxy for real search queries creates a domain mismatch during training, as documentation language differs substantially from the way developers search. Consequently, high-capacity models can overfit to documentation nuances without mastering user intent, while simpler keyword-matching approaches remain more robust in practice. Furthermore, expert annotations revealed that retrieved code often suffers from software quality issues, inverted operational logic, and domain ambiguity, highlighting that semantic search systems must account for code quality and structural context rather than just token similarity.

Organizations developing or deploying developer tools should avoid relying exclusively on complex neural architectures for code retrieval. Instead, systems should combine machine learning encoders with traditional keyword indexing techniques to reliably capture specific identifiers and rare terms. Future development should incorporate code quality filters, leverage structural code properties such as data flow, and explore advanced language pre-training techniques. Users should note that the training corpus contains noisy and outdated comments, and human relevance scores showed moderate inter-annotator agreement (Cohen's kappa of 0.47), indicating that evaluating code search relevance carries inherent subjectivity.

Cover for CodeSearchNet Challenge: Evaluating the State of Semantic Code Search

Abstract

Semantic code search is the task of retrieving relevant code given a natural language query. While related to other information retrieval tasks, it requires bridging the gap between the language used in code (often abbreviated and highly technical) and natural language more suitable to describe vague concepts and ideas.

To enable evaluation of progress on code search, we are releasing the CodeSearchNet Corpus and are presenting the CodeSearchNet Challenge, which consists of 99 natural language queries with about 4k expert relevance annotations of likely results from CodeSearchNet Corpus. The corpus contains about 6 million functions from open-source code spanning six programming languages (Go, Java, JavaScript, PHP, Python, and Ruby). The CodeSearchNet Corpus also contains automatically generated query-like natural language for 2 million functions, obtained from mechanically scraping and preprocessing associated function documentation. In this article, we describe the methodology used to obtain the corpus and expert labels, as well as a number of simple baseline solutions for the task.

We hope that CodeSearchNet Challenge encourages researchers and practitioners to study this interesting task further and will host a competition and leaderboard to track the progress on the challenge. We are also keen on extending CodeSearchNet Challenge to more queries and programming languages in the future.

Table of Contents

  • 1 Introduction
  • 2 The Code Search Corpus
  • 3 The Code Search Challenge
  • 3.1 Evaluation of Ranking Models
  • 4 Baseline CodeSearch models
  • 4.1 Joint Vector Representations for Code Search
  • 4.2 ElasticSearch Baseline
  • 4.3 Evaluation
  • 5 Related Work
  • 6 Conclusions & Open Challenges
  • References

Knowls

  1. Knowl 1 — CodeSearchNet Corpus

    definition

    The CodeSearchNet Corpus is a multi-language dataset scraped from public, open-source, non-fork GitHub repositories identified via libraries.io (filtered to projects used by at least one other project and sorted by popularity via stars and forks). Projects without permissive redistribution licenses are removed. Code functions are parsed using TreeSitter across six programming languages: Go, Java, JavaScript, PHP, Python, and Ruby.

    Functions with associated documentation comments are extracted and preprocessed using the following heuristic filtering pipeline:

    • Documentation comments are truncated to the first full paragraph.
    • Pairs with documentation comments shorter than 3 tokens are removed.
    • Functions whose implementation is shorter than 3 lines of code are removed.
    • Functions containing the substring test in their name, constructors, and standard language extension methods (e.g., __str__ in Python, toString in Java) are removed.
    • Near-duplicate functions and duplicate autogenerated code are deduplicated.

    The resulting corpus contains 6,452,446 total functions, of which 2,326,976 are documented function–comment pairs (ci,di)(c_i, d_i). The dataset is partitioned into an 80-10-10 train/validation/test split.

  2. Knowl 2 — CodeSearchNet Challenge and Expert Relevance Annotations

    experimental setup

    The CodeSearchNet Challenge is an evaluation benchmark for semantic code retrieval consisting of 99 natural language queries and 4,026 human expert relevance annotations spanning six programming languages (Go, Java, JavaScript, PHP, Python, and Ruby).

    • Query Collection: Queries were sourced from Bing developer search queries with high click-through rates to code combined with intent rewrites from StaQC, followed by manual filtering to remove pure technical keywords (e.g., exact function names like tf.gather_nd).
    • Candidate Generation: For each query and language, top candidate functions were generated by ensembling neural sequence models and ElasticSearch, extracting the top 10 pooled candidates.
    • Relevance Scoring: Annotators (software engineers, data scientists, and researchers) evaluated query–code snippet pairs on a 4-point ordinal scale:
      • 0: Totally irrelevant.
      • 1: Weak match (contains useful pointers, APIs, or ideas for building a solution).
      • 2: Strong match (usable as a backbone or directly copy-pasted with minor adaptation).
      • 3: Exact match (exactly implements what the query asks for).
    • Inter-Annotator Agreement: Evaluated on 891 multiply annotated pairs, yielding a squared Cohen's κ=0.47\kappa = 0.47.
    Language Score 0 Score 1 Score 2 Score 3 Total Annotations
    Go 62 64 29 11 166
    Java 383 178 125 137 823
    JavaScript 153 52 56 58 319
    PHP 103 77 68 66 314
    Python 498 511 537 543 2089
    Ruby 123 105 53 34 315
    Total 1322 987 868 849 4026
  3. Knowl 3 — Joint Vector Dual-Encoder Architecture for Code Search

    model/method

    The neural code search system maps source code snippets and natural language queries into a shared DD-dimensional continuous embedding space (D=128D = 128). It utilizes two separate sequence encoders: a code encoder EcE_c and a query encoder EqE_q.

    • Token Preprocessing: Identifiers in code sequences are split into subtokens based on naming conventions (e.g., camelCase →\to camel, case). Natural language query and documentation tokens are segmented using byte-pair encoding (BPE).
    • Sequence Encoders: Contextualized token representations are generated using one of four architectures:
      1. Neural Bag of Words (NBoW): Learns a direct embedding lookup for each subtoken/token.
      2. Bidirectional RNN (biRNN): Employs bidirectional Gated Recurrent Units (GRU) to encode sequential order.
      3. 1D Convolutional Neural Network (1D-CNN): Applies 1D convolutions over token embedding sequences.
      4. Self-Attention: Uses multi-head self-attention mechanisms across sequence tokens.
    • Pooling: Token representations are condensed into a fixed-size sequence embedding via mean pooling, max pooling, or an attention-weighted sum mechanism.
    • Retrieval & Indexing: At inference time, all corpus functions are indexed using Annoy (approximate nearest neighbor search with multidimensional space trees). A query qq is retrieved by computing Eq(q)E_q(q) and querying the Annoy index for nearest code embeddings Ec(c)E_c(c) by inner product.
  4. Knowl 4 — Contrastive Training Objective for Neural Code Search

    equation

    Given a batch of NN pairs (ci,di)(c_i, d_i) of code snippets cic_i and corresponding natural language documentation descriptions did_i, a code encoder Ec:C→RDE_c: \mathcal{C} \to \mathbb{R}^D and a query encoder Eq:D→RDE_q: \mathcal{D} \to \mathbb{R}^D are jointly trained by minimizing the cross-entropy / InfoNCE contrastive loss:

    L=−1N∑i=1Nlog⁡(exp⁡(Ec(ci)⊤Eq(di))∑j=1Nexp⁡(Ec(cj)⊤Eq(di)))\mathcal{L} = -\frac{1}{N} \sum_{i=1}^{N} \log \left( \frac{\exp\left(E_c(c_i)^\top E_q(d_i)\right)}{\sum_{j=1}^{N} \exp\left(E_c(c_j)^\top E_q(d_i)\right)} \right)

    where:

    • N∈N+N \in \mathbb{N}^+ is the number of paired examples in the training batch,
    • Ec(ci)∈RDE_c(c_i) \in \mathbb{R}^D is the embedding vector for the ii-th code snippet,
    • Eq(di)∈RDE_q(d_i) \in \mathbb{R}^D is the embedding vector for the ii-th documentation string,
    • Ec(cj)E_c(c_j) for j≠ij \neq i serve as distractor (negative) code snippets within the batch.

    The objective maximizes the dot-product similarity of matched pairs (Ec(ci),Eq(di))\left(E_c(c_i), E_q(d_i)\right) while minimizing the dot product between the query and all in-batch distractor code snippets cjc_j.

  5. Knowl 5 — CodeSearchNet Challenge Benchmark NDCG Results

    empirical result

    Baseline models evaluated on the CodeSearchNet Challenge benchmark demonstrate that the simplest neural model—Neural Bag of Words (NBoW)—outperforms deeper neural architectures (1D-CNN, biRNN, Self-Attention) when evaluated on human-annotated natural language queries, and performs competitively against ElasticSearch.

    Evaluation is performed using Normalized Discounted Cumulative Gain (NDCG) under two scopes:

    • NDCG Within: Evaluated strictly over the subset of candidate functions with human expert annotations.
    • NDCG All: Evaluated by ranking across all functions in the full multi-million function CodeSearchNet Corpus.
    Encoder NDCG Within NDCG All
    Text Code Go Java JS PHP Python Ruby Avg Go Java JS PHP Python Ruby Avg
    ElasticSearch 0.307 0.257 0.318 0.338 0.406 0.395 0.337 0.186 0.190 0.204 0.199 0.256 0.197 0.205
    NBoW NBoW 0.591 0.500 0.556 0.536 0.582 0.680 0.574 0.278 0.355 0.311 0.291 0.448 0.360 0.340
    1D-CNN 1D-CNN 0.379 0.407 0.269 0.474 0.473 0.420 0.404 0.120 0.189 0.099 0.176 0.242 0.162 0.165
    biRNN biRNN 0.112 0.165 0.066 0.148 0.193 0.185 0.145 0.030 0.056 0.017 0.042 0.070 0.060 0.046
    SelfAtt SelfAtt 0.484 0.431 0.446 0.522 0.560 0.515 0.493 0.211 0.233 0.175 0.232 0.367 0.219 0.240
    SelfAtt NBoW 0.550 0.514 0.545 0.557 0.583 0.650 0.566 0.284 0.340 0.299 0.291 0.422 0.360 0.333

    The superior performance of NBoW and ElasticSearch on real queries indicates that exact and subtoken keyword matching is critical for natural language search queries, whereas complex sequential models trained on docstrings overfit to proxy documentation patterns that do not generalize directly to search query characteristics.

  6. Knowl 6 — CodeSearchNet Corpus Docstring Retrieval MRR Performance

    empirical result

    When evaluated on the proxy retrieval task over the CodeSearchNet Corpus test set—where the model uses the first paragraph of a function's documentation comment as a query to retrieve the true function from among 999 fixed distractor snippets—deep self-attention models achieve the highest Mean Reciprocal Rank (MRR).

    Encoder Mean Reciprocal Rank (MRR)
    Text Code Go Java JS PHP Python Ruby Avg
    NBoW NBoW 0.6409 0.5140 0.4607 0.4835 0.5809 0.4285 0.6167
    1D-CNN 1D-CNN 0.6274 0.5270 0.3523 0.5294 0.5708 0.2450 0.6206
    biRNN biRNN 0.4524 0.2865 0.1530 0.2512 0.3213 0.0835 0.4262
    SelfAtt SelfAtt 0.6809 0.5866 0.4506 0.6011 0.6922 0.3651 0.7011
    SelfAtt NBoW 0.6631 0.5618 0.4920 0.5083 0.6113 0.4574 0.6505

    Self-attention achieves an average MRR of 0.70110.7011, outperforming NBoW (0.61670.6167), 1D-CNN (0.62060.6206), and biRNN (0.42620.4262). This demonstrates that while high-capacity attention models excel at matching structured docstring summaries to code implementations, performance on this proxy task does not directly correlate with ranking performance on real natural language search queries.

  7. Knowl 7 — CodeSearchNet Corpus Dataset Statistics

    data/table

    The CodeSearchNet Corpus comprises 6,452,446 total functions extracted from open-source GitHub repositories across six programming languages, with 2,326,976 functions containing paired natural language documentation comments.

    Programming Language Functions with Documentation All Functions
    Go 347,789 726,768
    Java 542,991 1,569,889
    JavaScript 157,988 1,857,835
    PHP 717,313 977,821
    Python 503,502 1,156,085
    Ruby 57,393 164,048
    Total 2,326,976 6,452,446

    PHP contains the highest absolute number of documented functions (717,313717,313), while JavaScript contains the lowest ratio of documented functions (157,988157,988 out of 1,857,8351,857,835 total).

  8. Knowl 8 — ElasticSearch Information Retrieval Baseline

    model/method

    As a non-neural information retrieval baseline, ElasticSearch is configured with default parameters and its default tokenizer over two inverted index fields for every function:

    1. Function Name: Tokenized and split into subtoken components (e.g., via camelCase splitting).
    2. Function Body: The complete raw text of the function implementation.

    At query time, search queries are scored against these two fields, providing traditional BM25-based keyword and rare-term retrieval without relying on learned continuous dense vector representations.

  9. Knowl 9 — NDCG Within and NDCG All Evaluation Metrics

    definition

    In the evaluation of code search ranking models on the CodeSearchNet Challenge, Normalized Discounted Cumulative Gain (NDCG) is measured under two distinct evaluation scopes:

    • NDCG Within: Evaluates ranking quality solely over the subset of candidate functions for which explicit human expert relevance annotations exist (i.e., reranking the pre-filtered candidate pool).
    • NDCG All: Evaluates retrieval quality by ranking across the entire multi-million function CodeSearchNet Corpus. Because unannotated functions are treated as non-relevant (relevance score 0) in standard NDCG computation, new retrieval models that discover relevant but unannotated snippets may receive penalized scores.
  10. Knowl 10 — Qualitative Failure Modes in Semantic Code Search

    empirical result

    Analysis of expert annotator feedback during the CodeSearchNet Challenge revealed five major challenges and failure modes in semantic code search:

    • Code Quality vs. Relevance: Syntactically and functionally correct results frequently exhibit poor code quality (unsatisfactory readability, security vulnerabilities, antipatterns, or inefficient execution), causing human evaluators to downgrade relevance scores.
    • Directionality / Inversion Errors: Baseline neural models frequently fail to capture semantic directionality, returning inverse implementations (e.g., returning stringToInt for the query "convert int to string").
    • Query Ambiguity: Natural language queries often lack context regarding domain definitions (e.g., defining what constitutes a "valid word").
    • Project Specificity vs. Generality: Retained results frequently rely on project-specific internal helper functions and are not self-contained, or conversely, are overly generic and verbose.
    • Documentation Proxy Mismatch: Scraping function documentation as training pairs introduces noise due to vocabulary overlap between code and author-written docstrings, outdated comments, and stylistic divergence from real search queries.

Coverage note — All substantial contributions—the CodeSearchNet Corpus collection/filtering, the CodeSearchNet Challenge benchmark with human annotations, the dual-encoder baseline architectures and training objective, the ElasticSearch baseline, and empirical retrieval evaluations across both benchmarks—are included. General background literature on semantic parsing and code summarization was omitted.

References

  1. 1.Miltiadis Allamanis. 2018. The Adverse Effects of Code Duplication in Machine Learning Models of Code. arXiv preprint arXiv:1812.06469 (2018).
  2. 2.Miltiadis Allamanis, Earl T Barr, Premkumar Devanbu, and Charles Sutton. 2018. A survey of machine learning for big code and naturalness. ACM Computing Surveys (CSUR) 51, 4 (2018), 81.
  3. 3.Miltiadis Allamanis, Hao Peng, and Charles Sutton. 2016. A Convolutional Attention Network for Extreme Summarization of Source Code. In Proceedings of the International Conference on Machine Learning (ICML).
  4. 4.Uri Alon, Omer Levy, and Eran Yahav. 2018. code2seq: Generating sequences from structured representations of code. arXiv preprint arXiv:1808.01400 (2018).
  5. 5.Antonio Valerio Miceli Barone and Rico Sennrich. 2017. A parallel corpus of Python functions and documentation strings for automated code documentation and code generation. arXiv preprint arXiv:1707.02275 (2017).
  6. 6.Jose Cambronero, Hongyu Li, Seohyun Kim, Koushik Sen, and Satish Chandra. 2019. When Deep Learning Met Code Search. arXiv preprint arXiv:1905.03813 (2019).
  7. 7.Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014. On the Properties of Neural Machine Translation: Encoder–Decoder Approaches. Syntax, Semantics and Structure in Statistical Translation (2014).
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018).
  9. 9.Patrick Fernandes, Miltiadis Allamanis, and Marc Brockschmidt. 2018. Structured Neural Summarization. arXiv preprint arXiv:1811.01824 (2018).
  10. 10.Philip Gage. 1994. A new algorithm for data compression. The C Users Journal 12, 2 (1994), 23–38.
  11. 11.Xiaodong Gu, Hongyu Zhang, and Sunghun Kim. 2018. Deep code search. In 2018 IEEE/ACM 40th International Conference on Software Engineering (ICSE). IEEE, 933–944.
  12. 12.Xiaodong Gu, Hongyu Zhang, Dongmei Zhang, and Sunghun Kim. 2016. Deep API Learning. In Proceedings of the International Symposium on Foundations of Software Engineering (FSE).
  13. 13.Tatsunori B Hashimoto, Kelvin Guu, Yonatan Oren, and Percy S Liang. 2018. A retrieve-and-edit framework for predicting structured outputs. In Advances in Neural Information Processing Systems. 10073–10083.
  14. 14.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018. Mapping language to code in programmatic context. arXiv preprint arXiv:1808.09588 (2018).
  15. 15.Yoon Kim. 2014. Convolutional neural networks for sentence classification. arXiv preprint arXiv:1408.5882 (2014).
  16. 16.Xi Victoria Lin, Chenglong Wang, Luke Zettlemoyer, and Michael D. Ernst. 2018. NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System. In International Conference on Language Resources and Evaluation.
  17. 17.Wang Ling, Edward Grefenstette, Karl Moritz Hermann, Tomas Kocisky, Andrew Senior, Fumin Wang, and Phil Blunsom. 2016. Latent Predictor Networks for Code Generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).
  18. 18.Cristina V Lopes, Petr Maj, Pedro Martins, Vaibhav Saini, Di Yang, Jakub Zitny, Hitesh Sajnani, and Jan Vitek. 2017. DéjàVu: a map of code duplicates on GitHub. Proceedings of the ACM on Programming Languages 1, OOPSLA (2017), 84.
  19. 19.Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008. Introduction to Information Retrieval. Cambridge University Press.
  20. 20.Bhaskar Mitra, Nick Craswell, et al. 2018. An introduction to neural information retrieval. Foundations and Trends® in Information Retrieval 13, 1 (2018), 1–126.
  21. 21.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).
  22. 22.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998–6008.
  23. 23.Ziyu Yao, Jayavardhan Reddy Peddamail, and Huan Sun. 2019. CoaCor: Code Annotation for Code Retrieval with Reinforcement Learning. (2019).
  24. 24.Ziyu Yao, Daniel S Weld, Wei-Peng Chen, and Huan Sun. 2018. StaQC: A Systematically Mined Question-Code Dataset from Stack Overflow. In Proceedings of the 2018 World Wide Web Conference on World Wide Web. International World Wide Web Conferences Steering Committee, 1693–1703.
  25. 25.Pengcheng Yin and Graham Neubig. 2017. A Syntactic Neural Model for General-Purpose Code Generation. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).

Citation

MLA
Husain, H., et al. “CodeSearchNet Challenge: Evaluating the State of Semantic Code Search”. arXiv, 2019, http://arxiv.org/abs/1909.09436v3.
APA
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., & Brockschmidt, M. (2019). CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv. http://arxiv.org/abs/1909.09436v3
Chicago
Husain, H., H.-H. Wu, T. Gazit, M. Allamanis, and M. Brockschmidt. 2019. “CodeSearchNet Challenge: Evaluating the State of Semantic Code Search”. arXiv. http://arxiv.org/abs/1909.09436v3.
Harvard
Husain, H. et al. (2019) “CodeSearchNet Challenge: Evaluating the State of Semantic Code Search”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1909.09436v3.
Vancouver
1. Husain H, Wu H-H, Gazit T, Allamanis M, Brockschmidt M (2019) CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv

BibTeX

@article{husain2019codesearchnet,
  title = {CodeSearchNet Challenge: Evaluating the State of Semantic Code Search},
  author = {Husain, Hamel and Wu, Ho-Hsiang and Gazit, Tiferet and Allamanis, Miltiadis and Brockschmidt, Marc},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1909.09436v3},
  eprint = {1909.09436}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission