RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation

Fengji ZhangBei ChenYue ZhangJacky KeungJin LiuDaoguang ZanYi MaoJian-Guang LouWeizhu Chen

article2023EMNLP595 citations

Proposes an iterative retrieval-generation framework, RepoCoder, that uses model-generated code completions to refine cross-file context retrieval, significantly outperforming standard retrieval-augmented approaches on repository-level code generation tasks.

Listen

Modern software development relies heavily on automated code completion tools driven by large language models. However, standard tools typically look only at the immediate context within the active file, ignoring the broader repository. In real-world projects, critical context such as shared utility functions, internal interfaces, and project-specific coding conventions is scattered across many files, leaving conventional models unable to accurately complete code that relies on cross-file dependencies.

The article evaluates a new framework named RepoCoder, which combines an automated document search tool with a pre-trained language model in an iterative retrieval-generation process. The main objective is to demonstrate that iteratively using a model's preliminary code completions to search the wider repository significantly improves completion accuracy across line, interface invocation, and full function body tasks without requiring model retraining or complex static analysis.

To establish credibility and avoid data leakage, the authors constructed RepoEval, a benchmark of 14 high-quality Python repositories created after 2022. The evaluation spans 1,600 line completion tasks, 1,600 repository-specific interface completion tasks, and 373 full function completion tasks evaluated against existing functional unit tests. The framework was evaluated across four language models of varying sizes, ranging from an open-source 350-million-parameter model to commercial models like GPT-3.5-Turbo, using both lightweight word-matching and advanced semantic search tools.

The analysis reveals several key findings. First, incorporating repository search through RepoCoder improves exact match accuracy by more than 10 percentage points across all model sizes compared to standard in-file completion. Second, the iterative retrieval design consistently outperforms standard single-step retrieval-augmented methods; a second iteration boosts prompt relevance by expanding search queries with preliminary code guesses. Third, the framework enables smaller models to perform exceptionally well, allowing a 350-million-parameter model using repository retrieval to match or exceed the accuracy of a standard billion-parameter model working only with local file context. Fourth, functional test pass rates on complete function bodies jumped from 23.32% to 42.63% when using RepoCoder with GPT-3.5-Turbo.

These findings indicate that development teams can achieve substantial improvements in automated code quality and developer productivity without incurring the high costs of fine-tuning large models or building specialized static code analysis pipelines. Furthermore, the ability of smaller, cheaper models to rival larger baselines when augmented with repository context presents significant opportunities to reduce inference compute costs and infrastructure requirements.

Organizations evaluating AI-assisted software engineering should consider incorporating iterative retrieval-augmented context into their coding assistant architectures. Prior to enterprise-wide deployment, engineering leaders should pilot the system to evaluate real-time latency trade-offs, as iterative generation steps increase response times. Practical deployment strategies include caching frequent repository patterns, applying model quantization, and capping the process at two iterations, where the majority of accuracy gains occur.

Decision-makers should note certain limitations: the performance gains depend partly on repository structure, providing fewer benefits in projects with minimal internal code reuse or shared conventions. Additionally, while the evaluation is robust across diverse Python repositories, further validation is necessary for other programming languages, newer generation models, and complex legacy codebases.

arXiv: 2303.12570main/RepoCoder
Cover for RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation

Abstract

The task of repository-level code completion is to continue writing the unfinished code based on a broader context of the repository. While for automated code completion tools, it is difficult to utilize the useful information scattered in different files. We propose RepoCoder, a simple, generic, and effective framework to address the challenge. It streamlines the repository-level code completion process by incorporating a similarity-based retriever and a pre-trained code language model in an iterative retrieval-generation pipeline. RepoCoder makes effective utilization of repository-level information for code completion and has the ability to generate code at various levels of granularity. Moreover, we propose a new benchmark RepoEval, which consists of the latest and high-quality real-world repositories covering line, API invocation, and function body completion scenarios. Experimental results indicate that RepoCoder significantly improves the In-File completion baseline by over 10% in all settings and consistently outperforms the vanilla retrieval-augmented code completion approach. Furthermore, we validate the effectiveness of RepoCoder through comprehensive analysis, providing valuable insights for future research. Our source code and benchmark are publicly available: https://github.com/microsoft/CodeT/tree/main/RepoCoder

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Overall Framework
  • 2.2 Code Retrieval
  • 2.3 Code Generation
  • 3 Benchmark Construction
  • 4 Experimental Setup
  • 4.1 Methods for Comparison
  • 4.2 Implementation Details
  • 4.3 Evaluation Metrics
  • 5 Experimental Results
  • 5.1 Line and API Completion Datasets
  • 5.2 Function Completion Dataset
  • 6 Analysis
  • 6.1 Quality of Retrieved Code
  • 6.2 Locations of Retrieved Code
  • 7 Related Work
  • 8 Conclusion and Future Work
  • Limitations
  • References
  • A Repository Details
  • B Using the Dense Retriever
  • C Code Duplication in Repositories
  • D Failed Cases between Iterations
  • E Case Study of Commercial Products

Knowls

  1. Knowl 1 — RepoCoder Iterative Retrieval-Generation Algorithm

    algorithm

    RepoCoder is an iterative framework for repository-level code completion that alternately retrieves context from a codebase and generates code completions using a pre-trained language model, using earlier generations to formulate better retrieval queries.

    Input: Repository codebase database Crepo\mathcal{C}_{repo}, unfinished in-file code context XX, retrieval model R\mathcal{R}, generation language model M\mathcal{M}, sliding window line size SwS_w, sliding size SsS_s, maximum iterations NiterN_{iter}
    Output: Generated code completion Y^\hat{Y}
    Q1←extract the last Sw lines of XQ^1 \leftarrow \text{extract the last } S_w \text{ lines of } X
    Cret1←R(Crepo,Q1)C_{ret}^1 \leftarrow \mathcal{R}(\mathcal{C}_{repo}, Q^1)
    Prompt1←ConstructPrompt(Cret1,X)\text{Prompt}^1 \leftarrow \text{ConstructPrompt}(C_{ret}^1, X)
    Y^1←M(Prompt1)\hat{Y}^1 \leftarrow \mathcal{M}(\text{Prompt}^1)
    for i=2i = 2 to NiterN_{iter} do
        Xsuffix←last (Sw−Ss) lines of XX_{suffix} \leftarrow \text{last } (S_w - S_s) \text{ lines of } X
        Yprefix←first Ss lines of Y^i−1Y_{prefix} \leftarrow \text{first } S_s \text{ lines of } \hat{Y}^{i-1}
        Qi←concatenate(Xsuffix,Yprefix)Q^i \leftarrow \text{concatenate}(X_{suffix}, Y_{prefix})
        Creti←R(Crepo,Qi)C_{ret}^i \leftarrow \mathcal{R}(\mathcal{C}_{repo}, Q^i)
        Prompti←ConstructPrompt(Creti,X)\text{Prompt}^i \leftarrow \text{ConstructPrompt}(C_{ret}^i, X)
        Y^i←M(Prompti)\hat{Y}^i \leftarrow \mathcal{M}(\text{Prompt}^i)
    end for
    return Y^Niter\hat{Y}^{N_{iter}}

    The parameters of both the code retrieval model R\mathcal{R} and the generator language model M\mathcal{M} remain completely frozen throughout all iterations. The method requires no static code analysis tools or heuristic extraction rules.

  2. Knowl 2 — Code Chunking, Query Formulation, and Prompt Formatting in RepoCoder

    model/method

    RepoCoder structures repository context and input prompts through three key mechanisms:

    1. Repository Indexing via Sliding Window: Repository source files are divided into code chunks Crepo={c1,c2,… }\mathcal{C}_{repo} = \{c_1, c_2, \dots\} using a sliding window of size SwS_w lines that shifts forward by a fixed stride of SsS_s lines.

    2. Iterative Query Augmentation: In the initial retrieval iteration (i=1i=1), the query consists of the trailing SwS_w lines of the unfinished code XX. For subsequent iterations (i>1i > 1), the query is formed by concatenating the trailing (Sw−Ss)(S_w - S_s) lines of XX with the leading SsS_s lines of the predicted completion Y^i−1\hat{Y}^{i-1}. This query is scored against snippets c∈Crepoc \in \mathcal{C}_{repo}. When using bag-of-words retrieval, similarity is computed via the Jaccard index over token sets SqS_q and ScS_c:

    Jaccard(Sq,Sc)=∣Sq∩Sc∣∣Sq∪Sc∣\text{Jaccard}(S_q, S_c) = \frac{|S_q \cap S_c|}{|S_q \cup S_c|}

    1. Prompt Formatting: Up to KK top-scoring retrieved snippets CretC_{ret} (within half the maximum prompt token budget) are formatted with comment headers indicating their source file paths and arranged in ascending order of similarity. The target file's unfinished code XX is appended at the bottom, directing the language model to generate the continuation.
  3. Knowl 3 — RepoEval Benchmark for Repository-Level Code Completion

    experimental setup

    RepoEval is a multi-granularity benchmark for evaluating repository-level code completion models on Python repositories created after January 1, 2022 (ensuring zero overlap with the pre-training corpora of models like GPT-3.5 and CODEGEN). Repositories are filtered to be non-forks with ≥100\ge 100 GitHub stars, ≥80%\ge 80\% Python files, and runnable unit tests.

    RepoEval evaluates three completion granularities:

    • Line Completion (1,600 test samples): 200 non-repetitive lines of at least 5 tokens each, sampled across 8 diverse repositories.
    • API Invocation Completion (1,600 test samples): 200 non-repetitive invocations of in-repository defined APIs per repository across the same 8 repositories.
    • Function Body Completion (373 test samples): Complete function bodies (3 to 30 lines) covered by explicit unit tests, sampled across 6 deployable repositories.

    Evaluation is conducted using three metrics:

    • Exact Match (EM): EM=1\text{EM} = 1 if predicted code Y^\hat{Y} matches the ground truth YY character-for-character, and 00 otherwise.
    • Edit Similarity (ES): Normalized edit distance defined as:

    ES=1−Lev(Y,Y^)max⁡(∣Y^∣,∣Y∣)\text{ES} = 1 - \frac{\text{Lev}(Y, \hat{Y})}{\max(|\hat{Y}|, |Y|)}

    where Lev(⋅,⋅)\text{Lev}(\cdot, \cdot) denotes Levenshtein distance.

    • Pass Rate (PR): Functional correctness measured by executing repository unit tests on the completed function body (1 if all tests pass, 0 otherwise).
  4. Knowl 4 — Performance of RepoCoder on Line and API Invocation Completion

    data/table

    Across both line and API invocation completion tasks on RepoEval, RepoCoder substantially outperforms the standard In-File completion baseline across various model architectures and sizes. Performing two or more iterations of retrieval and generation consistently exceeds single-iteration Retrieval-Augmented Generation (RAG, corresponding to iteration 1) and approaches the theoretical upper bound set by an Oracle retriever utilizing the ground-truth prefix.

    Model / Metric Oracle In-File RepoCoder-1 RepoCoder-2 RepoCoder-3 RepoCoder-4
    Line Completion
    GPT-3.5-Turbo
    EM (%) 57.75 40.56 55.31 56.81 57.00 56.63
    ES (%) 75.43 65.06 74.38 75.11 75.30 75.10
    CODEGEN-6B
    EM (%) 48.81 34.56 45.81 47.06 47.75 47.44
    ES (%) 71.02 60.67 69.21 70.10 70.73 70.19
    CODEGEN-2B
    EM (%) 47.31 33.63 44.56 46.94 46.69 47.13
    ES (%) 69.80 58.99 67.68 68.82 68.62 68.92
    CODEGEN-350M
    EM (%) 45.19 29.56 41.88 43.06 43.94 43.06
    ES (%) 67.20 55.39 65.05 65.66 65.97 65.62
    API Invocation Completion
    GPT-3.5-Turbo
    EM (%) 50.13 34.06 47.69 49.19 49.44 49.56
    ES (%) 74.50 63.22 73.63 74.43 74.59 74.48
    CODEGEN-6B
    EM (%) 40.25 26.19 36.69 38.88 39.13 39.31
    ES (%) 67.94 56.45 64.20 65.52 65.53 65.90
    CODEGEN-2B
    EM (%) 39.44 25.44 35.44 37.56 38.44 38.25
    ES (%) 66.78 56.88 63.47 64.15 64.53 64.60
    CODEGEN-350M
    EM (%) 34.88 22.19 31.75 33.88 33.75 33.81
    ES (%) 63.06 52.24 59.82 61.03 60.96 61.06

    Notably, CODEGEN-350M paired with RepoCoder achieves line completion performance (43.94%43.94\% EM) exceeding that of GPT-3.5-Turbo using only in-file context (40.56%40.56\% EM).

  5. Knowl 5 — Pass Rate Performance on Function Body Completion

    data/table

    When evaluated on function body completion using execution-based unit test Pass Rates (PR) with GPT-3.5-Turbo, RepoCoder Iteration 2 improves the in-file completion baseline by +19.31%+19.31\% absolute, matching the Oracle retrieval baseline.

    Repository Samples (NN) Oracle In-File RepoCoder-1 RepoCoder-2 RepoCoder-3 RepoCoder-4
    1. imagen 67 56.72 29.85 53.73 55.22 55.22 55.22
    2. tracr 146 43.84 27.40 41.78 43.84 44.52 44.52
    3. lightmmm 64 32.81 10.94 25.00 34.38 31.25 32.81
    4. inspection 32 34.38 28.13 34.38 37.50 34.38 34.38
    5. omnivore 22 40.91 31.82 31.82 36.36 31.82 36.36
    6. redframes 42 38.10 9.52 28.57 38.10 38.10 38.10
    All 373 42.63 23.32 38.34 42.63 41.82 42.36

    The multi-iteration process allows the generator to retrieve deeper contextual dependencies (such as helper methods or initialization conventions) across the codebase that are missing from the immediate prefix of the function.

  6. Knowl 6 — Ground-Truth API Retrieval Quality and Recall Across Iterations

    data/table

    On the subset of API invocation samples where invocations of the ground-truth API exist elsewhere in the repository (excluding input prompt examples), incorporating model predictions in query formulation increases the recall of ground-truth API examples and closes the gap to an oracle method (GT-Code) that directly retrieves true API usages.

    Model / Metric GT-Code In-File RepoCoder Iter-1 RepoCoder Iter-2
    GPT-3.5-Turbo (N=1046N=1046)
    Exact Match (EM, %) 55.54 34.42 53.63 55.07
    Edit Similarity (ES, %) 77.67 62.75 77.68 78.40
    Recall (%) 100.0 – 86.04 90.34
    CODEGEN-6B (N=1083N=1083)
    Exact Match (EM, %) 44.78 26.87 41.09 44.04
    Edit Similarity (ES, %) 71.47 56.42 67.10 68.55
    Recall (%) 100.0 – 76.27 82.92
    CODEGEN-350M (N=1083N=1083)
    Exact Match (EM, %) 37.86 22.25 35.64 38.13
    Edit Similarity (ES, %) 66.20 52.40 62.82 64.26
    Recall (%) 100.0 – 76.27 80.89

    Moving from Iteration 1 to Iteration 2 increases ground-truth API recall from 86.04%86.04\% to 90.34%90.34\% for GPT-3.5-Turbo and from 76.27%76.27\% to 82.92%82.92\% for CODEGEN-6B, directly driving improvements in EM and ES scores.

  7. Knowl 7 — File Locations of Effective Retrieved Code Snippets

    data/table

    When analyzing cases where RepoCoder (Iteration 2) or Oracle successfully generates completions that In-File completion fails on (using GPT-3.5-Turbo), the retrieved code snippets originate predominantly from files sharing import statements, directory paths, or naming patterns:

    Oracle RepoCoder Iter-2
    Location Category Line API Line API
    Imported File 4.22% 8.16% 3.22% 9.06%
    Current File (excluded lines) 3.86% 4.05% 3.32% 4.10%
    Current Directory 46.41% 58.40% 45.68% 59.58%
    Similar Import (shares ≥1\ge 1 import) 82.15% 86.71% 82.80% 87.70%
    Similar Name (shares ≥1\ge 1 filename token) 52.23% 65.10% 53.40% 64.20%
    Others 7.27% 4.43% 7.77% 3.35%
    Eligible Test Samples 333 294 312 276
    Retrieved Snippets 2202 1851 2047 1732

    Categories are non-exclusive. While the vast majority of effective contexts lie in files sharing imports or directories, restricting retrieval search heuristics strictly to these predefined locations causes performance degradation compared to open repository-wide similarity retrieval.

  8. Knowl 8 — Equivalence of Sparse and Dense Retrievers in RepoCoder

    empirical result

    RepoCoder demonstrates equivalent code completion performance whether using a lightweight sparse bag-of-words retriever (token Jaccard index) or a dense neural code retriever (UniXcoder hidden representation embeddings with cosine similarity).

    When evaluated on Line Completion with GPT-3.5-Turbo:

    • Sparse Retriever: Iteration 1 achieves 55.31%55.31\% EM / 74.38%74.38\% ES; Iteration 2 achieves 56.81%56.81\% EM / 75.11%75.11\% ES; Iteration 3 achieves 57.00%57.00\% EM / 75.30%75.30\% ES.
    • Dense Retriever (UniXcoder): Iteration 1 achieves 54.56%54.56\% EM / 73.96%73.96\% ES; Iteration 2 achieves 56.25%56.25\% EM / 74.70%74.70\% ES; Iteration 3 achieves 56.31%56.31\% EM / 74.31%74.31\% ES.

    On API Invocation Completion with GPT-3.5-Turbo:

    • Sparse Retriever: Iteration 1 achieves 47.69%47.69\% EM / 73.63%73.63\% ES; Iteration 2 achieves 49.19%49.19\% EM / 74.43%74.43\% ES; Iteration 4 achieves 49.56%49.56\% EM / 74.48%74.48\% ES.
    • Dense Retriever (UniXcoder): Iteration 1 achieves 47.56%47.56\% EM / 72.66%72.66\% ES; Iteration 2 achieves 49.13%49.13\% EM / 74.22%74.22\% ES; Iteration 4 achieves 49.63%49.63\% EM / 74.41%74.41\% ES.

    This confirms that RepoCoder's iterative retrieval-generation framework is robust to the underlying similarity search mechanism.

  9. Knowl 9 — Iteration Dynamics and Failure Transitions in RepoCoder

    data/table

    Tracking the exact number of correct API completion predictions (where EM=1\text{EM}=1) across successive iterations reveals that RepoCoder simultaneously converts previously failing test cases to correct completions and previously passing cases to failures:

    Model In-File →\rightarrow Iter-1 Iter-1 →\rightarrow Iter-2 Iter-2 →\rightarrow Iter-3 Iter-3 →\rightarrow Iter-4
    GPT-3.5-Turbo +545  (−40/+258)→763+545 \; (-40/+258) \rightarrow 763 +24  (−46/+70)→787+24 \; (-46/+70) \rightarrow 787 +7  (−83/+90)→794+7 \; (-83/+90) \rightarrow 794 −4  (−14/+10)→790-4 \; (-14/+10) \rightarrow 790
    CODEGEN-6B +424  (−57/+220)→587+424 \; (-57/+220) \rightarrow 587 +35  (−35/+70)→622+35 \; (-35/+70) \rightarrow 622 +4  (−12/+16)→626+4 \; (-12/+16) \rightarrow 626 +3  (−7/+10)→629+3 \; (-7/+10) \rightarrow 629
    CODEGEN-2B +412  (−64/+219)→567+412 \; (-64/+219) \rightarrow 567 +34  (−32/+66)→601+34 \; (-32/+66) \rightarrow 601 +14  (−12/+26)→615+14 \; (-12/+26) \rightarrow 615 −3  (−16/+13)→612-3 \; (-16/+13) \rightarrow 612
    CODEGEN-350M +352  (−46/+202)→508+352 \; (-46/+202) \rightarrow 508 +34  (−25/+59)→542+34 \; (-25/+59) \rightarrow 542 −2  (−18/+16)→540-2 \; (-18/+16) \rightarrow 540 +1  (−11/+12)→541+1 \; (-11/+12) \rightarrow 541

    (Note: In (−A/+B)(-A/+B), AA indicates cases that were correct in the prior iteration but failed in the current iteration, and BB indicates cases that failed previously but succeeded in the current iteration.)

    Failure transitions stem primarily from:

    1. Misleading retrieved snippets: An API having polymorphic parameter signatures across different files, causing the retriever to fetch mismatched example invocations.
    2. Query noise: Using a fixed sliding window length SsS_s over the predicted completion Y^i−1\hat{Y}^{i-1} introduces noise if the model produces irrelevant code beyond the targeted continuation.
  10. Knowl 10 — Limitations of RepoCoder

    limitation

    RepoCoder exhibits several notable limitations in practice:

    1. Dependence on Repository Duplication Ratio: The performance improvement of RepoCoder over In-File completion positively correlates with the proportion of duplicate code lines across the repository. In codebases with low duplication density (e.g., rl and vizier), the retriever struggles to find relevant exemplars, resulting in lower performance gains compared to highly modular or repetitive codebases (e.g., diffusers).
    2. Absence of Optimal Stopping Criteria: While two iterations consistently outperform single-step RAG, subsequent iterations (i≥3i \ge 3) oscillate in net accuracy, introducing new errors on previously correct cases. Constructing an automated, reliable early-stopping criterion without degrading performance remains an open problem.
    3. Inference Latency: Iterative cycles of retrieval and LLM generation incur substantial inference time overhead, posing challenges for real-time, interactive IDE code completion without model distillation, quantization, or prompt caching.

Coverage note — None was omitted; all key methodology, the RepoEval benchmark, primary line/API/function experimental tables, retriever comparisons, location analysis, iteration transitions, and limitations were fully extracted.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  2. 2.Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2022. Codet: Code generation with generated tests. arXiv preprint arXiv:2207.10397.
  3. 3.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  4. 4.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  5. 5.Colin B Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, and Alexey Svyatkovskiy. 2021. Long-range modeling of source code files with ewash: Extended window access by syntax hierarchy. arXiv preprint arXiv:2109.08780.
  6. 6.Yangruibo Ding, Zijian Wang, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang. 2022. Cocomic: Code completion by jointly modeling in-file and cross-file context. arXiv preprint arXiv:2212.10007.
  7. 7.Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. Unixcoder: Unified cross-modal pre-training for code representation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7212–7225.
  8. 8.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR.
  9. 9.Vincent J Hellendoorn and Premkumar Devanbu. 2017. Are deep neural networks the best choice for modeling source code? In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering, pages 763–773.
  10. 10.Vincent J Hellendoorn, Sebastian Proksch, Harald C Gall, and Alberto Bacchelli. 2019. When code completion fails: A case study on real-world completions. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE), pages 960–970. IEEE.
  11. 11.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299.
  12. 12.Paul Jaccard. 1912. The distribution of the flora in the alpine zone. 1. New phytologist, 11(2):37–50.
  13. 13.Vladimir I Levenshtein et al. 1966. Binary codes capable of correcting deletions, insertions, and reversals. In Soviet physics doklady, volume 10, pages 707–710. Soviet Union.
  14. 14.Yoav Levine, Itay Dalmedigos, Ori Ram, Yoel Zeldes, Daniel Jannai, Dor Muhlgay, Yoni Osin, Opher Lieber, Barak Lenz, Shai Shalev-Shwartz, et al. 2022. Standing on the shoulders of giant frozen language models. arXiv preprint arXiv:2204.10019.
  15. 15.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  16. 16.Dong Li, Yelong Shen, Ruoming Jin, Yi Mao, Kuan Wang, and Weizhu Chen. 2022. Generation-augmented query expansion for code retrieval. arXiv preprint arXiv:2212.10692.
  17. 17.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. 2023. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161.
  18. 18.Fang Liu, Ge Li, Zhiyi Fu, Shuai Lu, Yiyang Hao, and Zhi Jin. 2022. Learning to recommend method names with global context. In Proceedings of the 44th International Conference on Software Engineering, pages 1294–1306.
  19. 19.Shuai Lu, Nan Duan, Hojae Han, Daya Guo, Seung-won Hwang, and Alexey Svyatkovskiy. 2022. Reacc: A retrieval-augmented code completion framework. arXiv preprint arXiv:2203.07722.
  20. 20.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al. 2021. Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664.
  21. 21.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568.
  22. 22.Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen. 2020. Generation-augmented retrieval for open-domain question answering. arXiv preprint arXiv:2009.08553.
  23. 23.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474.
  24. 24.OpenAI. 2023. Gpt-4 technical report.
  25. 25.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. arXiv preprint arXiv:2302.00083.
  26. 26.Veselin Raychev, Martin Vechev, and Eran Yahav. 2014. Code completion with statistical language models. In Proceedings of the 35th ACM SIGPLAN conference on programming language design and implementation, pages 419–428.
  27. 27.Md Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Retrieval augmented code generation and summarization. arXiv e-prints, pages arXiv–2108.
  28. 28.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652.
  29. 29.Disha Shrivastava, Hugo Larochelle, and Daniel Tarlow. 2022. Repository-level prompt generation for large language models of code. arXiv preprint arXiv:2206.12839.
  30. 30.Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. Intellicode compose: Code generation using transformer. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 1433–1443.
  31. 31.Alexey Svyatkovskiy, Sebastian Lee, Anna Hadjitofi, Maik Riechert, Juliana Vicente Franco, and Miltiadis Allamanis. 2021. Fast and memory-efficient neural code completion. In 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR), pages 329–340. IEEE.
  32. 32.Alexey Svyatkovskiy, Ying Zhao, Shengyu Fu, and Neel Sundaresan. 2019. Pythia: Ai-assisted code completion system. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pages 2727–2735.
  33. 33.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  34. 34.Zhaopeng Tu, Zhendong Su, and Premkumar Devanbu. 2014. On the localness of software. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering, pages 269–280.
  35. 35.Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query expansion with large language models. arXiv preprint arXiv:2303.07678.
  36. 36.Daoguang Zan, Bei Chen, Zeqi Lin, Bei Guan, Yongji Wang, and Jian-Guang Lou. 2022. When language model meets private library. arXiv preprint arXiv:2210.17236.
  37. 37.Yury Zemlyanskiy, Michiel de Jong, Joshua Ainslie, Panupong Pasupat, Peter Shaw, Linlu Qiu, Sumit Sanghai, and Fei Sha. 2022. Generate-and-retrieve: Use your predictions to improve retrieval for semantic parsing. In Proceedings of the 29th International Conference on Computational Linguistics, pages 4946–4951.
  38. 38.Yizhe Zhang, Siqi Sun, Xiang Gao, Yuwei Fang, Chris Brockett, Michel Galley, Jianfeng Gao, and Bill Dolan. 2022. Retgen: A joint framework for retrieval and grounded text generation modeling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11739–11747.
  39. 39.Shuyan Zhou, Uri Alon, Frank F Xu, Zhengbao JIang, and Graham Neubig. 2022. Doccoder: Generating code by retrieving and reading docs. arXiv preprint arXiv:2207.05987.
  40. 40.Weiqin Zou, Jifeng Xuan, Xiaoyuan Xie, Zhenyu Chen, and Baowen Xu. 2019. How does code style inconsistency affect pull request integration? an exploratory study on 117 github projects. Empirical Software Engineering, 24:3871–3903.

Citation

MLA
Zhang, F., et al. “RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 2471–84, https://doi.org/10.18653/v1/2023.emnlp-main.151.
APA
Zhang, F., Chen, B., Zhang, Y., Keung, J., Liu, J., Zan, D., Mao, Y., Lou, J.-G., & Chen, W. (2023). RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2471–2484. https://doi.org/10.18653/v1/2023.emnlp-main.151
Chicago
Zhang, F., B. Chen, Y. Zhang, et al. 2023. “RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2471–84. https://doi.org/10.18653/v1/2023.emnlp-main.151.
Harvard
Zhang, F. et al. (2023) “RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2471–2484. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.151.
Vancouver
1. Zhang F, Chen B, Zhang Y, Keung J, Liu J, Zan D, Mao Y, Lou J-G, Chen W (2023) RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2471–2484

BibTeX

@inproceedings{zhang-etal-2023-repocoder,
    title = "{R}epo{C}oder: Repository-Level Code Completion Through Iterative Retrieval and Generation",
    author = "Zhang, Fengji  and
      Chen, Bei  and
      Zhang, Yue  and
      Keung, Jacky  and
      Liu, Jin  and
      Zan, Daoguang  and
      Mao, Yi  and
      Lou, Jian-Guang  and
      Chen, Weizhu",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.151/",
    doi = "10.18653/v1/2023.emnlp-main.151",
    pages = "2471--2484"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/