DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence

Daya GuoQihao ZhuDejian YangZhenda XieKai DongWentao ZhangGuan-Ting ChenXiao BiYu WuY. K. Li

article2024arXiv1,884 citations

Presents DeepSeek-Coder, an open-source suite of code language models trained from scratch on two trillion tokens with a 16K context window, delivering state-of-the-art generation and infilling performance that surpasses proprietary systems like GPT-3.5.

Listen

Artificial intelligence tools for software development are increasingly vital for boosting engineering productivity, but high-performing options have largely remained restricted to proprietary, closed-source models. The article introduces and evaluates the DeepSeek-Coder series, an open-source family of code-focused language models ranging from 1.3 billion to 33 billion parameters, to determine whether accessible, permissively licensed models can match or exceed proprietary industry standards.

To achieve this, the models were trained from scratch on two trillion tokens spanning 87 programming languages. The training pipeline organized code files at the repository level based on internal project dependencies, incorporated a fill-in-the-middle training strategy at a 50% rate to support in-line code insertion, and extended the operational context length to 16,000 tokens. The models were evaluated across standard benchmarks for multilingual code generation, data science library usage, cross-file completion, mathematical reasoning, and competition-level programming.

The evaluations yielded several core findings. First, the base 33-billion parameter model achieved state-of-the-art performance among open-source alternatives, reaching 50.3% average accuracy on HumanEval and 66.0% on MBPP, outperforming the similarly sized CodeLlama by 9 to 11 percentage points. Second, the smaller 6.7-billion parameter base model matched or surpassed the performance of open-source models five times its size. Third, the instruction-tuned 33-billion model surpassed OpenAI's GPT-3.5 Turbo across multiple benchmarks, including a 69.2% average score on multilingual HumanEval and a 27.8% pass rate on recent competition problems, significantly narrowing the gap to GPT-4. Additionally, pre-training on dependency-sorted repositories measurably enhanced cross-file code completion across multiple languages.

These findings demonstrate that organizations can deploy smaller, highly efficient open-source models without sacrificing output quality. Permissively licensing these models enables enterprises to lower commercial licensing costs, operate coding tools locally to protect proprietary codebases, and maintain greater architectural control compared to relying entirely on commercial cloud Application Programming Interfaces (APIs).

Based on these results, software teams seeking code completion tools should consider deploying the 6.7-billion parameter base model as a cost-effective, high-accuracy option. For complex logic and competitive coding tasks, the article recommends pairing the instruction-tuned models with step-by-step reasoning prompts, which systematically improved accuracy across difficult test sets. To advance capabilities further, developing code models initialized from broader general-purpose language models (as demonstrated with the v1.5 model) is recommended to strengthen natural language comprehension alongside code synthesis.

Confidence in the reported benchmarks is high due to strict data decontamination protocols and standardized evaluation frameworks. However, the authors note residual uncertainties: test performance showed mild degradation on the newest competition subsets, suggesting potential historical contamination, and model reliability decreases when context windows are pushed beyond 16,000 tokens.

Cover for DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence

Abstract

The rapid development of large language models has revolutionized code intelligence in software development. However, the predominance of closed-source models has restricted extensive research and development. To address this, we introduce the DeepSeek-Coder series, a range of open-source code models with sizes from 1.3B to 33B, trained from scratch on 2 trillion tokens. These models are pre-trained on a high-quality project-level code corpus and employ a fill-in-the-blank task with a 16K window to enhance code generation and infilling. Our extensive evaluations demonstrate that DeepSeek-Coder not only achieves state-of-the-art performance among open-source code models across multiple benchmarks but also surpasses existing closed-source models like Codex and GPT-3.5. Furthermore, DeepSeek-Coder models are under a permissive license that allows for both research and unrestricted commercial use.

Table of Contents

  • 1 Introduction
  • 2 Data Collection
  • 2.1 GitHub Data Crawling and Filtering
  • 2.2 Dependency Parsing
  • 2.3 Repo-Level Deduplication
  • 2.4 Quality Screening and Decontamination
  • 3 Training Policy
  • 3.1 Training Strategy
  • 3.1.1 Next Token Prediction
  • 3.1.2 Fill-in-the-Middle
  • 3.2 Tokenizer
  • 3.3 Model Architecture
  • 3.4 Optimization
  • 3.5 Environments
  • 3.6 Long Context
  • 3.7 Instruction Tuning
  • 4 Experimental Results
  • 4.1 Code Generation
  • 4.2 Fill-in-the-Middle Code Completion
  • 4.3 Cross-File Code Completion
  • 4.4 Program-based Math Reasoning
  • 5 Continue Pre-Training From General LLM
  • 6 Conclusion
  • References
  • A Cases of Chatting with DeepSeek-Coder-Instruct
  • B Benchmark curves during training of DeepSeek-Coder-Base

Knowls

  1. Knowl 1 — DeepSeek-Coder Architecture and Model Hyperparameters

    model/method

    DeepSeek-Coder models are decoder-only Transformer language models trained from scratch for code intelligence across three parameter scales: 1.3B, 6.7B, and 33B. All variants employ SwiGLU activation functions, Rotary Position Embeddings (RoPE), and FlashAttention v2. The 33B variant incorporates Grouped-Query Attention (GQA) with a group size of 8 to enhance training and inference efficiency, while the 1.3B and 6.7B variants utilize standard Multi-Head Attention (MHA). The tokenizer is trained via Byte Pair Encoding (BPE) on a subset of the corpus with a vocabulary size of 32,000.

    Hyperparameter DeepSeek-Coder 1.3B DeepSeek-Coder 6.7B DeepSeek-Coder 33B
    Hidden Activation SwiGLU SwiGLU SwiGLU
    Hidden size 2048 4096 7168
    Intermediate size 5504 11008 19200
    Hidden layers number 24 32 62
    Attention heads number 16 32 56
    Attention Multi-head Multi-head Grouped-query (8)
    Batch Size 1024 2304 3840
    Max Learning Rate 5.3×1045.3 \times 10^{-4} 4.2×1044.2 \times 10^{-4} 3.5×1043.5 \times 10^{-4}

    Training uses the AdamW optimizer (β1=0.9\beta_1 = 0.9, eta_2 = 0.95) with a three-stage learning rate schedule comprising 2,000 warm-up steps. The learning rate at each subsequent stage drops by a factor of 1/10\sqrt{1/10}, reaching a final learning rate equal to 10% of the initial maximum learning rate.

  2. Knowl 2 — Repository-Level Dependency Parsing Algorithm

    algorithm

    To capture cross-file dependencies within software projects during pre-training, file sequences within a repository are ordered such that dependent context precedes the invoking file. Invocation relationships are extracted using regular expressions (such as import in Python, using in C#, and include in C). The dependency graph is parsed into disconnected subgraphs, and each subgraph is topologically ordered using a modified topological sort that selects the node with minimal in-degree to gracefully handle dependency cycles.

    procedure TopologicalSort(files)
        graphs = {}
        inDegree = {}
        for each file in files do
            graphs[file] = []
            inDegree[file] = 0
        end for
        for each fileA in files do
            for each fileB in files do
                if HasDependency(fileA, fileB) then
                    graphs[fileB].append(fileA)
                    inDegree[fileA] = inDegree[fileA] + 1
                end if
            end for
        end for
        subgraphs = getDisconnectedSubgraphs(graphs)
        allResults = []
        for each subgraph in subgraphs do
            results = []
            while length(results) != NumberOfNodes(subgraph) do
                file = argmin({inDegree[file] | file in subgraph and file not in results})
                for each node in graphs[file] do
                    inDegree[node] = inDegree[node] - 1
                end for
                results.append(file)
            end while
            allResults.append(results)
        end for
        return allResults
    end procedure

    Each ordered sequence of files is concatenated into a single training document, where each file is prefixed with a comment containing its relative file path.

  3. Knowl 3 — Pre-training Data Composition and Processing Pipeline for DeepSeek-Coder

    experimental setup

    The pre-training corpus for DeepSeek-Coder contains 2 trillion tokens collected from GitHub public repositories created before February 2023, comprising 87% source code across 87 programming languages (797.92 GB total across 603,173k files), 10% English code-related natural language corpus from GitHub Markdown and StackExchange, and 3% Chinese natural language corpus. The data pipeline includes:

    1. Rule-based filtering: Removing files with average line length >100> 100 characters, maximum line length >1000> 1000 characters, or <25%< 25\% alphabetic characters; filtering files containing <?xml version= in the first 100 characters (except XSLT); retaining HTML files only if visible text is 20%\ge 20\% and 100\ge 100 characters; retaining JSON/YAML files only if character count is between 50 and 5,000. This retains 32.8% of the initial crawled volume.

    2. Repository-level deduplication: Applying near-deduplication to concatenated full repositories rather than individual files to preserve repository structural integrity.

    3. Quality screening and decontamination: Applying compiler verification, a quality filtering model, and heuristic rules. Contamination is prevented by filtering out training files containing matching 10-grams (or exact matches for substrings between 3 and 10 characters) from evaluation benchmarks including HumanEval, MBPP, GSM8K, and MATH.

  4. Knowl 4 — Fill-in-the-Middle Pre-training Formulation and Configuration

    model/method

    DeepSeek-Coder models incorporate the Fill-in-the-Middle (FIM) training objective alongside standard next-token prediction to enhance bidirectional code infilling capabilities. For a document divided into three segments—prefix fpref_{\text{pre}}, middle fmiddlef_{\text{middle}}, and suffix fsuff_{\text{suf}}—the model uses the Prefix-Suffix-Middle (PSM) formatting:

    < ⁣fim_start ⁣>fpre< ⁣fim_hole ⁣>fsuf< ⁣fim_end ⁣>fmiddle< ⁣eos_token ⁣><\!\mid\text{fim\_start}\mid\!> f_{\text{pre}} <\!\mid\text{fim\_hole}\mid\!> f_{\text{suf}} <\!\mid\text{fim\_end}\mid\!> f_{\text{middle}} <\!\mid\text{eos\_token}\mid\!>

    where < ⁣fim_start ⁣><\!\mid\text{fim\_start}\mid\!>, < ⁣fim_hole ⁣><\!\mid\text{fim\_hole}\mid\!>, and < ⁣fim_end ⁣><\!\mid\text{fim\_end}\mid\!> are dedicated sentinel tokens.

    FIM is applied at the document level prior to sequence packing at an FIM rate of 0.5 (50% PSM). Ablation experiments on DeepSeek-Coder-Base 1.3B comparing 0%, 50%, and 100% FIM rates alongside 50% Masked Span Prediction (MSP) demonstrated that while 100% FIM maximized infilling performance (HumanEval-FIM Pass@1), it degraded standard left-to-right code completion (HumanEval and MBPP Pass@1). A 50% PSM rate established the optimal balance between infilling efficiency and generative completion proficiency, outperforming 50% MSP.

  5. Knowl 5 — Context Window Extension via Linear RoPE Scaling

    model/method

    To extend the context handling capability of DeepSeek-Coder from standard sequence lengths up to 16,384 tokens (and theoretically up to 65,536 tokens), Rotary Position Embedding (RoPE) parameters are adjusted using linear scaling. The scaling factor is increased from 1 to 4, and the RoPE base frequency is adjusted from 10,000 to 100,000.

    Following parameter reconfiguration, the models undergo 1,000 additional training steps at a sequence length of 16K tokens with a batch size of 512, keeping the learning rate fixed at the final rate of the primary pre-training stage. Empirical observations indicate that output reliability is optimal within the 16K context range.

  6. Knowl 6 — Instruction Tuning Formulation for DeepSeek-Coder-Instruct

    model/method

    DeepSeek-Coder-Instruct models are fine-tuned from DeepSeek-Coder-Base checkpoints on 2 billion tokens of human instruction data structured in the Alpaca format. Individual dialogue turns are delineated using a dedicated delimiter token < ⁣EOT ⁣><\!\mid\text{EOT}\mid\!>. Supervised fine-tuning utilizes a batch size of 4 million tokens, an initial learning rate of 1×1051 \times 10^{-5} governed by a cosine learning rate decay schedule, and 100 initial warm-up steps.

  7. Knowl 7 — Multilingual HumanEval and MBPP Code Generation Performance

    empirical result

    On standard code generation benchmarks evaluated with greedy decoding (zero-shot HumanEval across 8 languages and few-shot MBPP for Python), DeepSeek-Coder base and instruct models outperform comparable open-source models:

    Model Size Python C++ Java PHP TS C# Bash JS Avg MBPP
    CodeGeeX2 6B 36.0% 29.2% 25.9% 23.6% 20.8% 29.7% 6.3% 24.8% 24.5% 36.2%
    StarCoderBase 16B 31.7% 31.1% 28.5% 25.4% 34.0% 34.8% 8.9% 29.8% 28.0% 42.8%
    CodeLlama 7B 31.7% 29.8% 34.2% 23.6% 36.5% 36.7% 12.0% 29.2% 29.2% 38.6%
    CodeLlama 13B 36.0% 37.9% 38.0% 34.2% 45.2% 43.0% 16.5% 32.3% 35.4% 48.4%
    CodeLlama 34B 48.2% 44.7% 44.9% 41.0% 42.1% 48.7% 15.8% 42.2% 41.0% 55.2%
    DeepSeek-Coder-Base 1.3B 34.8% 31.1% 32.3% 24.2% 28.9% 36.7% 10.1% 28.6% 28.3% 46.2%
    DeepSeek-Coder-Base 6.7B 49.4% 50.3% 43.0% 38.5% 49.7% 50.0% 28.5% 48.4% 44.7% 60.6%
    DeepSeek-Coder-Base 33B 56.1% 58.4% 51.9% 44.1% 52.8% 51.3% 32.3% 55.3% 50.3% 66.0%
    GPT-3.5-Turbo - 76.2% 63.4% 69.2% 60.9% 69.1% 70.8% 42.4% 67.1% 64.9% 70.8%
    GPT-4 - 84.1% 76.4% 81.6% 77.2% 77.4% 79.1% 58.2% 78.0% 76.5% 80.0%
    DeepSeek-Coder-Instruct 1.3B 65.2% 45.3% 51.9% 45.3% 59.7% 55.1% 12.7% 52.2% 48.4% 49.4%
    DeepSeek-Coder-Instruct 6.7B 78.6% 63.4% 68.4% 68.9% 67.2% 72.8% 36.7% 72.7% 66.1% 65.4%
    DeepSeek-Coder-Instruct 33B 79.3% 68.9% 73.4% 72.7% 67.9% 74.1% 43.0% 73.9% 69.2% 70.0%

    DeepSeek-Coder-Base 33B achieves 50.3% average HumanEval accuracy and 66.0% MBPP, outperforming CodeLlama-Base 34B by 9.3% and 10.8% respectively. DeepSeek-Coder-Base 6.7B outperforms CodeLlama-Base 34B across all evaluated metrics. DeepSeek-Coder-Instruct 33B achieves 69.2% average HumanEval accuracy, exceeding GPT-3.5-Turbo (64.9%).

  8. Knowl 8 — Cross-File Code Completion Performance and Effect of Repository Pre-training

    empirical result

    Cross-file code completion is evaluated on the CrossCodeEval benchmark across Python, Java, TypeScript, and C# using Exact Match (EM) and Edit Similarity (ES). Evaluation uses a 2048-token sequence limit, 50-token output limit, and up to 512 tokens of cross-file context retrieved via BM25.

    Model Size Python Java TypeScript C#
    EM ES EM ES EM ES EM ES
    CodeGeeX2 6B 8.11% 59.55% 7.34% 59.60% 6.14% 55.50% 1.70% 51.66%
    + Retrieval 10.73% 61.76% 10.10% 59.56% 7.72% 55.17% 4.64% 52.30%
    StarCoder-Base 7B 6.68% 59.55% 8.65% 62.57% 5.01% 48.83% 4.75% 59.53%
    + Retrieval 13.06% 64.24% 15.61% 64.78% 7.54% 42.06% 14.20% 65.03%
    CodeLlama-Base 7B 7.32% 59.66% 9.68% 62.64% 8.19% 58.50% 4.07% 59.19%
    + Retrieval 13.02% 64.30% 16.41% 64.64% 12.34% 60.64% 13.19% 63.04%
    DeepSeek-Coder-Base 6.7B 9.53% 61.65% 10.80% 61.77% 9.59% 60.17% 5.26% 61.32%
    + Retrieval 16.14% 66.51% 17.72% 63.18% 14.03% 61.77% 16.23% 63.42%
    + Retrieval w/o Repo Pre-training 16.02% 66.65% 16.64% 61.88% 13.23% 60.92% 14.48% 62.38%

    DeepSeek-Coder-Base 6.7B with retrieval achieves the highest scores across all four programming languages. Omitting repository-level ordering during pre-training (w/o Repo Pre-training) leads to performance degradations in Java (EM drops from 17.72% to 16.64%), TypeScript (EM drops from 14.03% to 13.23%), and C# (EM drops from 16.23% to 14.48%), verifying the utility of dependency-aware repository ordering.

  9. Knowl 9 — LeetCode Contest Benchmark and Prompting Performance

    empirical result

    A competition-level benchmark comprising 180 problems collected from LeetCode Contests held between July 2023 and January 2024 (45 Easy, 91 Medium, 44 Hard), each evaluated against 100 test cases, was constructed to test problem understanding and code generation on non-memorized problems.

    Model Size Easy (45) Medium (91) Hard (44) Overall (180)
    WizardCoder-V1.0 15B 17.8% 1.1% 0.0% 5.0%
    CodeLlama-Instruct 34B 24.4% 4.4% 4.5% 9.4%
    Phind-CodeLlama-V2 34B 26.7% 8.8% 9.1% 13.3%
    GPT-3.5-Turbo - 46.7% 15.4% 15.9% 23.3%
    GPT-3.5-Turbo + CoT - 42.2% 15.4% 20.5% 23.3%
    GPT-4-Turbo - 73.3% 31.9% 25.0% 40.6%
    GPT-4-Turbo + CoT - 71.1% 35.2% 25.0% 41.8%
    DeepSeek-Coder-Instruct 1.3B 22.2% 1.1% 4.5% 7.2%
    DeepSeek-Coder-Instruct + CoT 1.3B 22.2% 2.2% 2.3% 7.2%
    DeepSeek-Coder-Instruct 6.7B 44.4% 12.1% 9.1% 19.4%
    DeepSeek-Coder-Instruct + CoT 6.7B 44.4% 17.6% 4.5% 21.1%
    DeepSeek-Coder-Instruct 33B 57.8% 22.0% 9.1% 27.8%
    DeepSeek-Coder-Instruct + CoT 33B 53.3% 25.3% 11.4% 28.9%

    DeepSeek-Coder-Instruct 33B is the only evaluated open-source model to surpass GPT-3.5-Turbo (27.8% vs 23.3%). Adding Chain-of-Thought (CoT) prompting ("You need first to write a step-by-step outline and then write the code.") improves overall Pass@1 from 19.4% to 21.1% for the 6.7B model and from 27.8% to 28.9% for the 33B model, with noticeable gains on medium and hard problem subsets.

  10. Knowl 10 — Continued Pre-training from General LLM (DeepSeek-Coder-v1.5)

    model/method

    DeepSeek-Coder-v1.5 7B is trained by continuing pre-training from the general foundation model DeepSeek-LLM-7B Base on 2 trillion tokens with a 4K context length and next-token prediction objective. The training mixture consists of:

    • 70% source code
    • 10% Markdown and StackExchange
    • 7% natural language related to code
    • 7% natural language related to math
    • 6% Chinese-English bilingual natural language

    Performance comparison between DeepSeek-Coder (trained from scratch on code) and DeepSeek-Coder-v1.5 (continued from a general LLM) illustrates that initializing from a general foundation model substantially improves mathematical reasoning and natural language capabilities with minimal changes to programming accuracy:

    Models Size Programming Math Reasoning Natural Language
    HumanEval MBPP GSM8K MATH MMLU BBH HellaSwag WinoG ARC-C
    DeepSeek-Coder-Base 6.7B 44.7% 60.6% 43.2% 19.2% 36.6% 44.3% 53.8% 57.1% 32.5%
    DeepSeek-Coder-Base-v1.5 6.9B 43.2% 60.4% 62.4% 24.7% 49.1% 55.2% 69.9% 63.8% 47.2%
    DeepSeek-Coder-Instruct 6.7B 66.1% 65.4% 62.8% 28.6% 37.2% 46.9% 55.0% 57.6% 37.4%
    DeepSeek-Coder-Instruct-v1.5 6.9B 64.1% 64.6% 72.6% 34.1% 49.5% 53.3% 72.2% 63.4% 48.1%

    Math tasks are solved via program-aided reasoning. DeepSeek-Coder-Base-v1.5 achieves a 19.2% absolute gain on GSM8K (from 43.2% to 62.4%) and a 12.5% absolute gain on MMLU (from 36.6% to 49.1%).

  11. Knowl 11 — Data Science Code Completion Performance on DS-1000

    empirical result

    On the DS-1000 benchmark, which evaluates Python code completion across 1,000 practical data science workflows spanning 7 libraries, DeepSeek-Coder-Base models outperform other open-source models across all libraries:

    Model Size Matplotlib Numpy Pandas Pytorch Scipy Scikit-Learn Tensorflow Avg
    CodeGeeX2 6B 38.7% 26.8% 14.4% 11.8% 19.8% 27.0% 17.8% 22.9%
    StarCoder-Base 16B 43.2% 29.1% 11.0% 20.6% 23.6% 32.2% 15.6% 24.6%
    CodeLlama-Base 7B 41.9% 24.6% 14.8% 16.2% 18.9% 17.4% 17.8% 22.1%
    CodeLlama-Base 13B 46.5% 28.6% 18.2% 19.1% 18.9% 27.8% 33.3% 26.8%
    CodeLlama-Base 34B 50.3% 42.7% 23.0% 25.0% 28.3% 33.9% 40.0% 34.3%
    DeepSeek-Coder-Base 1.3B 32.3% 21.4% 9.3% 8.8% 8.5% 16.5% 8.9% 16.2%
    DeepSeek-Coder-Base 6.7B 48.4% 35.5% 20.6% 19.1% 22.6% 38.3% 24.4% 30.5%
    DeepSeek-Coder-Base 33B 56.1% 49.6% 25.8% 36.8% 36.8% 40.0% 46.7% 40.2%

    DeepSeek-Coder-Base 33B achieves an average Pass@1 score of 40.2%, exceeding CodeLlama-Base 34B (34.3%).

  12. Knowl 12 — Program-Aided Mathematical Reasoning Performance

    empirical result

    Evaluated using Program-Aided Math Reasoning (PAL), where models alternate between natural language reasoning and program generation across 7 mathematical benchmarks, DeepSeek-Coder-Base models demonstrate state-of-the-art open-source performance:

    Model Size GSM8K MATH GSM-Hard SVAMP TabMWP ASDiv MAWPS Avg
    CodeGeeX-2 7B 22.2% 9.7% 23.6% 39.0% 44.6% 48.5% 66.0% 36.2%
    StarCoder-Base 16B 23.4% 10.3% 23.0% 42.4% 45.0% 54.9% 81.1% 40.0%
    CodeLlama-Base 7B 31.2% 12.1% 30.2% 54.2% 52.9% 59.6% 82.6% 46.1%
    CodeLlama-Base 13B 43.1% 14.4% 40.2% 59.2% 60.3% 63.6% 85.3% 52.3%
    CodeLlama-Base 34B 58.2% 21.2% 51.8% 70.3% 69.8% 70.7% 91.8% 62.0%
    DeepSeek-Coder-Base 1.3B 14.6% 16.8% 14.5% 36.7% 30.0% 48.2% 62.3% 31.9%
    DeepSeek-Coder-Base 6.7B 43.2% 19.2% 40.3% 58.4% 67.9% 67.2% 87.0% 54.7%
    DeepSeek-Coder-Base 33B 60.7% 29.1% 54.1% 71.6% 75.3% 76.7% 93.3% 65.8%

    DeepSeek-Coder-Base 33B achieves a 65.8% average score across the benchmarks, including 60.7% on GSM8K and 29.1% on MATH, outperforming CodeLlama-Base 34B (62.0% average, 58.2% on GSM8K, 21.2% on MATH).

  13. Knowl 13 — Single-Line Fill-In-The-Middle Benchmark Performance

    empirical result

    In single-line infilling tasks evaluated using exact match line accuracy across Python, Java, and JavaScript:

    Model Size Python Java JavaScript Mean
    SantaCoder 1.1B 44.0% 62.0% 74.0% 69.0%
    StarCoder 16B 62.0% 73.0% 74.0% 69.7%
    CodeLlama-Base 7B 67.6% 74.3% 80.2% 69.7%
    CodeLlama-Base 13B 68.3% 77.6% 80.7% 75.5%
    DeepSeek-Coder-Base 1.3B 57.4% 82.2% 71.7% 70.4%
    DeepSeek-Coder-Base 6.7B 66.6% 88.1% 79.7% 80.7%
    DeepSeek-Coder-Base 33B 65.4% 86.6% 82.5% 81.2%

    DeepSeek-Coder-Base 1.3B outperforms larger models such as SantaCoder 1.1B, StarCoder 16B, and CodeLlama-Base 7B with a 70.4% mean exact match score. DeepSeek-Coder-Base 6.7B reaches 80.7% mean exact match, showing a strong balance of accuracy and computational efficiency for deployment in code completion tools.

Coverage note — Qualitative dialogue examples from Appendix A and the individual language-by-language breakdown table of all 87 crawled languages were omitted in favor of aggregated dataset statistics and formal quantitative benchmarks.

References

  1. 1.L. B. Allal, R. Li, D. Kocetkov, C. Mou, C. Akiki, C. M. Ferrandis, N. Muennighoff, M. Mishra, A. Gu, M. Dey, et al. Santacoder: don’t reach for the stars! arXiv preprint arXiv:2301.03988, 2023.
  2. 2.J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton. Program synthesis with large language models, 2021.
  3. 3.M. Bavarian, H. Jun, N. Tezak, J. Schulman, C. McLeavey, J. Tworek, and M. Chen. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255, 2022.
  4. 4.F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, et al. Multipl-e: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering, 2023.
  5. 5.M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  6. 6.S. Chen, S. Wong, L. Chen, and Y. Tian. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023.
  7. 7.P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  8. 8.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  9. 9.T. Dao. Flashattention-2: Faster attention with better parallelism and work partitioning, 2023.
  10. 10.DeepSeek-AI. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024.
  11. 11.Y. Ding, Z. Wang, W. U. Ahmad, H. Ding, M. Tan, N. Jain, M. K. Ramanathan, R. Nallapati, P. Bhatia, D. Roth, et al. Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023.
  12. 12.Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335, 2022.
  13. 13.D. Fried, A. Aghajanyan, J. Lin, S. Wang, E. Wallace, F. Shi, R. Zhong, W.-t. Yih, L. Zettlemoyer, and M. Lewis. Incoder: A generative model for code infilling and synthesis. arXiv preprint arXiv:2204.05999, 2022.
  14. 14.L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig. Pal: Program-aided language models. In International Conference on Machine Learning, pages 10764–10799. PMLR, 2023.
  15. 15.G. Gemini Team. Gemini: A family of highly capable multimodal models, 2023. URL https://goo.gle/GeminiPaper.
  16. 16.Z. Gou, Z. Shao, Y. Gong, Y. Yang, M. Huang, N. Duan, W. Chen, et al. Tora: A tool-integrated reasoning agent for mathematical problem solving. arXiv preprint arXiv:2309.17452, 2023.
  17. 17.D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2020.
  18. 18.D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  19. 19.High-Flyer. Hai-llm: An efficient and lightweight tool for training large models. 2023. URL https://www.high-flyer.cn/en/blog/hai-llm.
  20. 20.kaiokendev. Things i’m learning while training superhot. https://kaiokendev.github.io/til#extending-context-to-8k, 2023.
  21. 21.D. Kocetkov, R. Li, L. Jia, C. Mou, Y. Jernite, M. Mitchell, C. M. Ferrandis, S. Hughes, T. Wolf, D. Bahdanau, et al. The stack: 3 tb of permissively licensed source code. Transactions on Machine Learning Research, 2022.
  22. 22.V. A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems, 5, 2023.
  23. 23.Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W.-t. Yih, D. Fried, S. Wang, and T. Yu. Ds-1000: A natural and reliable benchmark for data science code generation. In International Conference on Machine Learning, pages 18319–18345. PMLR, 2023.
  24. 24.K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, 2022.
  25. 25.R. Li, L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, et al. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023.
  26. 26.I. Loshchilov and F. Hutter. Decoupled weight decay regularization, 2019.
  27. 27.P. Lu, L. Qiu, K.-W. Chang, Y. N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. In The Eleventh International Conference on Learning Representations, 2022.
  28. 28.S.-Y. Miao, C.-C. Liang, and K.-Y. Su. A diverse corpus for evaluating and developing english math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 975–984, 2020.
  29. 29.D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia. Pipedream: Generalized pipeline parallelism for dnn training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles, pages 1–15, 2019.
  30. 30.E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474, 2022.
  31. 31.E. Nijkamp, H. Hayashi, C. Xiong, S. Savarese, and Y. Zhou. Codegen2: Lessons for training llms on programming and natural languages, 2023.
  32. 32.OpenAI. Gpt-4 technical report, 2023.
  33. 33.A. Patel, S. Bhattamishra, and N. Goyal. Are nlp models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094, 2021.
  34. 34.C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023.
  35. 35.S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE, 2020.
  36. 36.B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.
  37. 37.K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  38. 38.R. Sennrich, B. Haddow, and A. Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015.
  39. 39.J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding, 2023.
  40. 40.M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, , and J. Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  41. 41.R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  42. 42.H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  43. 43.Y. Wang, W. Wang, S. Joty, and S. C. Hoi. Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv preprint arXiv:2109.00859, 2021.
  44. 44.R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.
  45. 45.Q. Zheng, X. Xia, X. Zou, Y. Dong, S. Wang, Y. Xue, L. Shen, Z. Wang, A. Wang, Y. Li, et al. Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 5673–5684, 2023.

Citation

MLA
Guo, D., et al. “DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence”. arXiv, 2024, http://arxiv.org/abs/2401.14196v2.
APA
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y. K., Luo, F., Xiong, Y., & Liang, W. (2024). DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence. arXiv. http://arxiv.org/abs/2401.14196v2
Chicago
Guo, D., Q. Zhu, D. Yang, et al. 2024. “DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence”. arXiv. http://arxiv.org/abs/2401.14196v2.
Harvard
Guo, D. et al. (2024) “DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.14196v2.
Vancouver
1. Guo D, Zhu Q, Yang D, et al (2024) DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence. arXiv

BibTeX

@article{guo2024deepseek,
  title = {DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence},
  author = {Guo, Daya and Zhu, Qihao and Yang, Dejian and Xie, Zhenda and Dong, Kai and Zhang, Wentao and Chen, Guanting and Bi, Xiao and Wu, Y. and Li, Y. K. and Luo, Fuli and Xiong, Yingfei and Liang, Wenfeng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.14196v2},
  eprint = {2401.14196}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors