RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation
Fengji ZhangBei ChenYue ZhangJacky KeungJin LiuDaoguang ZanYi MaoJian-Guang LouWeizhu Chen
Proposes an iterative retrieval-generation framework, RepoCoder, that uses model-generated code completions to refine cross-file context retrieval, significantly outperforming standard retrieval-augmented approaches on repository-level code generation tasks.
Modern software development relies heavily on automated code completion tools driven by large language models. However, standard tools typically look only at the immediate context within the active file, ignoring the broader repository. In real-world projects, critical context such as shared utility functions, internal interfaces, and project-specific coding conventions is scattered across many files, leaving conventional models unable to accurately complete code that relies on cross-file dependencies.
The article evaluates a new framework named RepoCoder, which combines an automated document search tool with a pre-trained language model in an iterative retrieval-generation process. The main objective is to demonstrate that iteratively using a model's preliminary code completions to search the wider repository significantly improves completion accuracy across line, interface invocation, and full function body tasks without requiring model retraining or complex static analysis.
To establish credibility and avoid data leakage, the authors constructed RepoEval, a benchmark of 14 high-quality Python repositories created after 2022. The evaluation spans 1,600 line completion tasks, 1,600 repository-specific interface completion tasks, and 373 full function completion tasks evaluated against existing functional unit tests. The framework was evaluated across four language models of varying sizes, ranging from an open-source 350-million-parameter model to commercial models like GPT-3.5-Turbo, using both lightweight word-matching and advanced semantic search tools.
The analysis reveals several key findings. First, incorporating repository search through RepoCoder improves exact match accuracy by more than 10 percentage points across all model sizes compared to standard in-file completion. Second, the iterative retrieval design consistently outperforms standard single-step retrieval-augmented methods; a second iteration boosts prompt relevance by expanding search queries with preliminary code guesses. Third, the framework enables smaller models to perform exceptionally well, allowing a 350-million-parameter model using repository retrieval to match or exceed the accuracy of a standard billion-parameter model working only with local file context. Fourth, functional test pass rates on complete function bodies jumped from 23.32% to 42.63% when using RepoCoder with GPT-3.5-Turbo.
These findings indicate that development teams can achieve substantial improvements in automated code quality and developer productivity without incurring the high costs of fine-tuning large models or building specialized static code analysis pipelines. Furthermore, the ability of smaller, cheaper models to rival larger baselines when augmented with repository context presents significant opportunities to reduce inference compute costs and infrastructure requirements.
Organizations evaluating AI-assisted software engineering should consider incorporating iterative retrieval-augmented context into their coding assistant architectures. Prior to enterprise-wide deployment, engineering leaders should pilot the system to evaluate real-time latency trade-offs, as iterative generation steps increase response times. Practical deployment strategies include caching frequent repository patterns, applying model quantization, and capping the process at two iterations, where the majority of accuracy gains occur.
Decision-makers should note certain limitations: the performance gains depend partly on repository structure, providing fewer benefits in projects with minimal internal code reuse or shared conventions. Additionally, while the evaluation is robust across diverse Python repositories, further validation is necessary for other programming languages, newer generation models, and complex legacy codebases.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). It establishes the foundational Codex model and the standard paradigm of using large language models for automated code completion.
- Paper: CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation, Shuai Lu et al. (2021). It introduces standard multi-task benchmarks and baseline neural models for code completion and code understanding.
- Paper: CodeSearchNet Challenge: Evaluating the State of Semantic Code Search, Hamel Husain et al. (2019). It defines the core challenge and baseline retrieval methodologies for searching and matching relevant code from natural language and code contexts.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). It provides open large language model architectures and pre-trained weights for multi-turn and autoregressive code generation.
- Paper: CodeBERT: A Pre-Trained Model for Programming and Natural Languages, Zhangyin Feng et al. (2020). It introduces bimodal representation learning across code and text that serves as a cornerstone for neural code retrieval and encoding.
- Paper: GraphCodeBERT: Pre-training Code Representations with Data Flow, Daya Guo et al. (2020). It develops data-flow-aware representations for code search and code completion that inform repository-level semantic understanding.
- Paper: UniXcoder: Unified Cross-Modal Pre-training for Code Representation, Daya Guo et al. (2022). It presents a unified cross-modal architecture for code representation and generation across both retrieval and completion tasks.
- Paper: CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation, Yue Wang et al. (2021). It details identifier-aware encoder-decoder modeling for code understanding and generation, establishing core principles for generating accurate program tokens.
- Paper: CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion, Yangruibo Ding et al. (2023). It extends the evaluation of repository-level code completion by introducing a dedicated multilingual benchmark specifically targeting cross-file dependencies.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). It integrates repository-level dependency organization and long context handling directly into the pre-training and fill-in-the-middle pipeline of open foundation models.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). It generalizes repository-level code comprehension from iterative completion to end-to-end multi-file software engineering and GitHub issue resolution.
- Paper: OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models, Siming Huang et al. (2025). It incorporates repository-level code modeling insights into an open, fully transparent end-to-end training and data curation pipeline for code LLMs.
- Paper: MapCoder: Multi-Agent Code Generation for Competitive Problem Solving, Md. Ashraful Islam et al. (2024). It extends retrieval-augmented code generation into a coordinated multi-agent workflow covering retrieval, planning, coding, and debugging.
- Paper: SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Yuhang Wang et al. (2026). It addresses the context overhead inherent in repository-scale code retrieval by introducing task-aware, line-level context pruning for coding agents.
- Paper: daVinci-Dev: Agent-native Mid-training for Software Engineering, Ji Zeng et al. (2026). It advances beyond inference-time repository retrieval by embedding multi-file agentic interactions and repository navigation directly into mid-training.
