Instruction Tuning for Secure Code Generation

Jingxuan HeMark VeroGabriela KrasnopolskaMartin T. Vechev

article2024ICML64 citations

Presents SafeCoder, a dual-objective instruction tuning method that leverages an automated dataset collection pipeline to boost the security of code generated by language models by roughly 30% without degrading general utility.

Listen

Language models are increasingly used in software development to generate code, but state-of-the-art models frequently introduce severe security vulnerabilities. Even after standard instruction tuning—the process of training models to follow user prompts—they generate secure code only about 60% of the time, and simple prompt adjustments do not solve this issue. This poses serious operational and security risks, as AI-generated vulnerabilities can reach production environments and require extensive remediation.

The article introduces SafeCoder, an instruction tuning framework designed to demonstrate that language models can be trained to generate highly secure code without compromising their general utility and functional accuracy.

The authors developed an automated, two-step data mining pipeline that screened over 145 million open-source commits and applied static analysis to extract 465 high-quality security fix examples across 23 vulnerability categories and 6 programming languages. Combining this with existing security data yielded a dataset of 1,268 examples. SafeCoder integrates this security dataset into standard instruction tuning by jointly optimizing the model: it rewards secure code patterns and penalizes insecure code using token-level loss masking and an unlikelihood penalty, while applying oversampling to balance rare vulnerability types. The approach was evaluated across six language models, ranging from 1 billion to 7 billion parameters, spanning 60 testing scenarios and multiple coding and reasoning benchmarks.

The evaluation revealed several critical findings. First, SafeCoder increased the secure code generation rate from approximately 60% to around 90%, representing an absolute improvement of roughly 30 percentage points over both base models and standard instruction-tuned models. Second, this substantial security gain incurred virtually no penalty on functional programming correctness or general natural language understanding benchmarks. Third, ablation analyses confirmed that isolating security-critical tokens, including an unlikelihood penalty, and curating diverse training data were each vital to performance. Finally, SafeCoder eliminated the trade-off between security and general capability observed in previous incremental patching techniques.

These findings demonstrate that organizations do not need to choose between code security and AI helpfulness during post-training. Implementing security-focused fine-tuning significantly mitigates security risks and lowers technical debt, all while introducing minimal training overhead due to the compact size of the security dataset.

Organizations training or fine-tuning coding models should integrate SafeCoder into their standard instruction-tuning pipelines. However, decision-makers must note key limitations: SafeCoder is designed for the instruction-tuning phase of open models, provides no formal guarantee against all software flaws, and does not generalize well to vulnerability types absent from the training set. Therefore, language model code generation should remain paired with rigorous automated security analysis and human code review.

arXiv: 2402.09497
Cover for Instruction Tuning for Secure Code Generation

Abstract

Modern language models (LMs) have gained widespread acceptance in everyday and professional contexts, particularly in programming. An essential procedure enabling this adoption is instruction tuning, which substantially enhances LMs’ practical utility by training them to follow user instructions and human preferences. However, existing instruction tuning schemes overlook a crucial aspect: the security of generated code. As a result, even the state-of-the-art instruction-tuned LMs frequently produce unsafe code, posing significant security risks. In this work, we introduce SafeCoder to address this gap. SafeCoder performs security-centric fine-tuning using a diverse and high-quality dataset that we collected using an automated pipeline. We integrate the security fine-tuning with standard instruction tuning, to facilitate a joint optimization of both security and utility. Despite its simplicity, we show that SafeCoder is effective across a variety of popular LMs and datasets. It is able to drastically improve security (by about 30%), while preserving utility.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background and Problem Statement
  • 4. SafeCoder's Instruction Tuning
  • 5. SafeCoder's Data Collection
  • 6. Experimental Evaluation
  • 6.1. Experimental Setup
  • 6.2. Experimental Results
  • 7. Conclusion and Discussion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Details on Experimental Setup
  • B. Further Experimental Results and Details

Knowls

  1. Knowl 1 — Masked Language Modeling and Unlikelihood Loss for Code Security

    equation

    SafeCoder formulates security fine-tuning over instruction tuples (i,osec,ovul)(i, o^{\text{sec}}, o^{\text{vul}}), where ii is a natural language functional instruction, osec=[o1sec,…,o∣osec∣]seco^{\text{sec}} = [o^{\text{sec}}_1, \dots, o^{\text{sec}}_{|o^{\text{sec}}|]} is a functionally correct and secure code snippet, and ovul=[o1vul,…,o∣ovul∣]vulo^{\text{vul}} = [o^{\text{vul}}_1, \dots, o^{\text{vul}}_{|o^{\text{vul}}|]} is the corresponding insecure code snippet implementing the same functionality.

    Binary mask vectors msec∈{0,1}∣osec∣m^{\text{sec}} \in \{0, 1\}^{|o^{\text{sec}}|} and mvul∈{0,1}∣ovul∣m^{\text{vul}} \in \{0, 1\}^{|o^{\text{vul}}|} are constructed using token-level diffing between oseco^{\text{sec}} and ovulo^{\text{vul}}. An element mtsecm^{\text{sec}}_t (or mtvulm^{\text{vul}}_t) is set to 11 if token tt belongs to the security-critical diff, and 00 otherwise.

    The masked negative log-likelihood loss on secure tokens is defined as:

    Lsec(i,osec,msec)=−∑t=1∣osec∣mtsec⋅log⁡P(otsec∣o<tsec,i)\mathcal{L}^{\text{sec}}(i, o^{\text{sec}}, m^{\text{sec}}) = - \sum_{t=1}^{|o^{\text{sec}}|} m^{\text{sec}}_t \cdot \log P(o^{\text{sec}}_t \mid o^{\text{sec}}_{<t}, i)

    The masked unlikelihood loss penalizing tokens that introduce vulnerabilities is defined as:

    Lvul(i,ovul,mvul)=−∑t=1∣ovul∣mtvul⋅log⁡(1−P(otvul∣o<tvul,i))\mathcal{L}^{\text{vul}}(i, o^{\text{vul}}, m^{\text{vul}}) = - \sum_{t=1}^{|o^{\text{vul}}|} m^{\text{vul}}_t \cdot \log(1 - P(o^{\text{vul}}_t \mid o^{\text{vul}}_{<t}, i))

    Minimizing Lsec\mathcal{L}^{\text{sec}} reinforces generation of secure patterns, while minimizing Lvul\mathcal{L}^{\text{vul}} suppresses the generation of insecure alternatives without requiring an auxiliary reference model.

  2. Knowl 2 — SafeCoder Joint Instruction Tuning Procedure

    algorithm

    SafeCoder combines standard instruction tuning with security-specific tuning in a single joint training pipeline. Given a standard instruction tuning dataset Dstd\mathcal{D}^{\text{std}} of (i,o)(i, o) pairs and a security instruction dataset Dsec\mathcal{D}^{\text{sec}} of (i,osec,ovul)(i, o^{\text{sec}}, o^{\text{vul}}) triples, the language model is optimized across samples drawn from Dstd∪Dsec\mathcal{D}^{\text{std}} \cup \mathcal{D}^{\text{sec}}.

    Input: Pretrained language model MM, standard instruction dataset DstdD^{\text{std}}, security dataset DsecD^{\text{sec}}
    Output: Instruction-tuned language model M∗M^*
    1: for sample ss in Dstd∪DsecD^{\text{std}} \cup D^{\text{sec}} do
    2: if s∈Dstds \in D^{\text{std}} then
    3: Compute standard negative log-likelihood loss Lstd(i,o)=−∑t=1∣o∣log⁡P(ot∣o<t,i)\mathcal{L}^{\text{std}}(i, o) = -\sum_{t=1}^{|o|} \log P(o_t \mid o_{<t}, i) and update MM
    4: else
    5: Compute security loss Lsec(i,osec,msec)+Lvul(i,ovul,mvul)\mathcal{L}^{\text{sec}}(i, o^{\text{sec}}, m^{\text{sec}}) + \mathcal{L}^{\text{vul}}(i, o^{\text{vul}}, m^{\text{vul}}) and update MM
    6: end if
    7: end for
    8: return MM

    By randomly drawing mini-batch samples across the union of both datasets without artificially changing their relative sizes, the model achieves a joint optimization for instruction following, functional utility, and code security simultaneously.

  3. Knowl 3 — Automated Pipeline for Security Dataset Extraction

    algorithm

    To automatically extract verified vulnerability fixes from open-source repositories without requiring manual curation, SafeCoder employs a two-stage extraction and filtering pipeline followed by instruction synthesis.

    Input: Set of GitHub commits C={(m,r,r′)}C = \{(m, r, r')\} where mm is commit message, rr is pre-commit repo, r′r' is post-commit repo
    Output: Security instruction dataset DsecD^{\text{sec}}
    1: Dsec←∅D^{\text{sec}} \leftarrow \emptyset
    2: for each commit (m,r,r′)∈C(m, r, r') \in C do
    3: if heuristicFilter(m,r,r′m, r, r') then
    4: V←analyzeCode(r)V \leftarrow \text{analyzeCode}(r)
    5: V′←analyzeCode(r′)V' \leftarrow \text{analyzeCode}(r')
    6: if ∣V∣>0|V| > 0 and ∣V′∣=0|V'| = 0 then
    7: for each (osec,ovul)∈changedFuncs(r,r′)(o^{\text{sec}}, o^{\text{vul}}) \in \text{changedFuncs}(r, r') do
    8: i←generateInst(osec,ovul)i \leftarrow \text{generateInst}(o^{\text{sec}}, o^{\text{vul}})
    9: Dsec←Dsec∪{(i,osec,ovul)}D^{\text{sec}} \leftarrow D^{\text{sec}} \cup \{(i, o^{\text{sec}}, o^{\text{vul}})\}
    10: end for
    11: end if
    12: end if
    13: end for
    14: return DsecD^{\text{sec}}

    Key subroutines:

    • heuristicFilter: Matches commit messages against CWE-specific keywords and restricts edits to at most 40 lines across at most 2 files in supported file types, filtering out non-security refactorings.
    • analyzeCode: Executes GitHub CodeQL static analysis for target CWEs. A commit is verified only when pre-commit vulnerabilities ∣V∣>0|V| > 0 and post-commit vulnerabilities ∣V′∣=0|V'| = 0.
    • generateInst: Prompts GPT-4 to generate a brief (maximum two sentences) functional description ii that captures the common task of oseco^{\text{sec}} and ovulo^{\text{vul}} without mentioning any security features.
  4. Knowl 4 — Security and Utility Performance Across Coding and General-Purpose Models

    data/table

    SafeCoder was evaluated on three code-specialized models (StarCoder-1B, StarCoder-3B, CodeLlama-7B) and three general-purpose models (Phi-2-2.7B, Llama2-7B, Mistral-7B). Code security was evaluated over 60 scenarios with 100 generations per scenario checked by CodeQL. Utility was measured on HumanEval (Pass@1 and Pass@10), MBPP (Pass@1 and Pass@10), MMLU (5-shot), and TruthfulQA (5-shot).

    Model Instruction Tuning Code Sec. (%) HumanEval MBPP MMLU TruthfulQA
    Pass@1 Pass@10 Pass@1 Pass@10
    StarCoder-1B Pretrained 55.6 14.9 26.0 20.3 37.9 26.8 21.7
    Standard SFT 62.9 20.4 33.9 24.2 40.2 25.0 23.3
    SafeCoder 92.1 19.4 30.3 24.2 40.0 24.8 22.8
    StarCoder-3B Pretrained 60.3 21.2 39.0 29.2 48.8 27.3 20.3
    Standard SFT 68.3 30.7 50.7 31.9 46.8 25.1 20.8
    SafeCoder 93.0 28.0 50.3 31.9 47.5 25.0 20.9
    CodeLlama-7B Pretrained 57.0 28.6 54.1 35.9 54.9 39.8 25.1
    Standard SFT 66.6 36.8 53.9 37.8 48.9 27.1 25.2
    SafeCoder 91.2 35.9 54.7 35.1 48.5 28.6 28.2
    Phi-2-2.7B Pretrained 67.1 51.2 74.5 40.3 56.3 56.8 41.4
    Standard SFT 69.9 48.3 73.9 32.0 54.0 53.3 42.6
    SafeCoder 90.9 46.1 71.8 37.6 55.6 52.8 40.5
    Llama2-7B Pretrained 55.8 13.4 26.6 17.6 37.4 46.0 24.6
    Standard SFT 59.2 13.3 28.0 19.5 37.2 46.0 26.6
    SafeCoder 89.2 11.8 25.7 19.6 35.1 45.5 26.5
    Mistral-7B Pretrained 55.5 27.2 52.8 31.9 51.9 62.9 35.8
    Standard SFT 63.1 35.2 60.4 35.3 51.3 62.7 39.0
    SafeCoder 89.6 33.7 58.8 35.4 51.0 62.6 39.5

    Across all architectures, pretrained and standard instruction-tuned models generate secure code only 55.5%–69.9% of the time. SafeCoder improves security rates to ~89%–93% (an absolute gain of ~30%) while preserving functional coding utility and natural language understanding performance.

  5. Knowl 5 — Ablation Analysis of Training Data, Loss Masking, and Unlikelihood Loss

    empirical result

    Ablation studies on StarCoder-1B and Phi-2-2.7B evaluate the individual contributions of dataset collection, token loss masking, and unlikelihood loss:

    Pretrained LM Method Code Security (%) HumanEval Pass@1
    StarCoder-1B No collected data 74.1 19.2
    No loss masks 79.9 20.1
    No unlikelihood 87.0 19.3
    Full method 92.1 19.4
    Phi-2-2.7B No collected data 69.2 44.6
    No loss masks 80.3 47.1
    No unlikelihood 79.0 46.7
    Full method 90.9 46.1

    Key takeaways:

    1. Removing the automatically collected GitHub security dataset ("No collected data") causes security to drop by 18.0%–21.7%, demonstrating the value of broader vulnerability and language diversity.
    2. Removing token masks msecm^{\text{sec}} and mvulm^{\text{vul}} ("No loss masks", training over all program tokens) reduces security by 10.6%–12.2%, confirming that isolating the learning signal to security-critical diffs is essential.
    3. Removing the masked unlikelihood loss Lvul\mathcal{L}^{\text{vul}} ("No unlikelihood") reduces security by 5.1%–11.9%, proving that penalizing insecure completions provides critical negative signal.
  6. Knowl 6 — Comparison with SVEN: Resolving the Security vs. Utility Trade-Off

    empirical result

    When adapting the SVEN security hardening method to instruction tuning, SVEN trains on an already instruction-tuned model using a combined loss:

    L=Lsec+Lvul+wKL⋅(LKLsec+LKLvul)\mathcal{L} = \mathcal{L}^{\text{sec}} + \mathcal{L}^{\text{vul}} + w^{\text{KL}} \cdot (\mathcal{L}^{\text{KLsec}} + \mathcal{L}^{\text{KLvul}})

    where LKLsec\mathcal{L}^{\text{KLsec}} and LKLvul\mathcal{L}^{\text{KLvul}} penalize KL divergence between the fine-tuned model and original model output probabilities on non-security tokens, weighted by wKL=2n/10w^{\text{KL}} = 2^n / 10 for n∈{1,…,8}n \in \{1, \dots, 8\}.

    Evaluating SVEN across varying values of wKLw^{\text{KL}} on StarCoder-1B and Phi-2-2.7B demonstrates an inherent trade-off: lower wKLw^{\text{KL}} values increase security at the expense of functional correctness (HumanEval Pass@1), while higher wKLw^{\text{KL}} values preserve HumanEval Pass@1 but degrade security (e.g., security drops from ~90% down to ~65%–70%).

    In contrast, SafeCoder's joint training on standard instruction samples and masked security pairs Pareto-dominates SVEN: SafeCoder achieves ∼91%–92%\sim 91\%\text{--}92\% security while matching or exceeding the highest HumanEval Pass@1 score obtained by standard instruction tuning.

  7. Knowl 7 — Minority Class Oversampling for Imbalanced Vulnerability Data

    model/method

    The collected security instruction tuning dataset exhibits imbalance across Common Weakness Enumeration (CWE) categories and programming languages. To prevent underperformance on minority categories, SafeCoder treats each (CWE, programming language) pair as a distinct class and randomly duplicates minority pairs with fewer than kk samples until they reach exactly kk instances.

    Evaluating oversampling parameter k∈{1,5,10,20,40,80}k \in \{1, 5, 10, 20, 40, 80\} on StarCoder-1B over 5 random seeds demonstrates:

    1. Without oversampling (k=1k=1), mean code security is ~87% with high variance across runs.
    2. Increasing kk up to k=20k=20 raises mean security to ~93% and substantially reduces variance.
    3. Increasing kk beyond 20 yields diminishing returns.

    Based on these findings, SafeCoder sets k=20k=20 for coding-specialized language models and k=40k=40 for general-purpose language models. The relative data imbalance between Dstd\mathcal{D}^{\text{std}} and Dsec\mathcal{D}^{\text{sec}} (where Dstd\mathcal{D}^{\text{std}} is 5 to 12 times larger) is maintained without resampling, introducing minimal training overhead.

  8. Knowl 8 — Inefficacy of Security-Aware Prompting in Instruction-Tuned Models

    empirical result

    Prompting state-of-the-art instruction-tuned models with explicit security directives does not eliminate insecure code generation. Three prompt variants were tested:

    1. func-only: Prompts specifying only functional requirements.
    2. sec-generic: Prompts appending a generic safety instruction: "Make sure that the generated code is secure, meaning it does not contain any security vulnerabilities."
    3. sec-specific: Prompts appending a vulnerability-specific instruction detailing the exact CWE definition (oracle assumption).
    Model func-only (%) sec-generic (%) sec-specific (%)
    Mistral-Instruct-7B 54.7 56.8 57.4
    CodeLlama-Instruct-7B 63.1 64.9 70.6
    OctoCoder 60.5 64.1 63.7
    GPT-3.5-Turbo-Instruct 63.3 67.8 71.0

    Security-aware instructions yield only modest increases (2.1%–7.7%), and models continue to produce vulnerable code in 29%–43% of cases even with CWE-specific oracle prompting. This demonstrates that prompting alone is insufficient for reliable code security.

  9. Knowl 9 — Generalization Limitation of SafeCoder on Unseen Vulnerability Types

    limitation

    When evaluated on test scenarios targeting CWEs that were absent from the security training dataset (specifically CWE-020, CWE-094, CWE-117, CWE-209, CWE-215, CWE-312, CWE-643, CWE-777, CWE-798, and CWE-918 across 15 Python scenarios), SafeCoder does not provide meaningful security improvements over baseline instruction tuning:

    Model w/o SafeCoder (%) with SafeCoder (%)
    StarCoder-1B 61.4 57.4
    CodeLlama-7B 49.3 50.4
    Phi-2-2.7B 63.3 62.8
    Mistral-7B 57.7 67.4

    SafeCoder's security gains do not generalize zero-shot to unseen vulnerability categories, indicating that security training must cover target CWEs explicitly during fine-tuning.

  10. Knowl 10 — Composition of the SafeCoder Security Dataset and Evaluation Benchmark

    data/table

    Running the automated data extraction pipeline over 145 million GitHub commits yielded 150k candidate commits via heuristic filtering, 25k CodeQL-analyzable repositories, and 1,211 verified vulnerability fixes (a 4.9% verification rate). After cleaning and rebalancing overrepresented classes, the final extracted dataset consists of 465 high-quality samples covering 23 CWEs across 6 languages:

    Language C/C++ Go Java JavaScript Python Ruby Total
    Sample Count 53 45 26 113 146 82 465

    Combining this dataset with SVEN's curated dataset (803 samples covering 9 CWEs and 2 languages) produces a unified instruction tuning dataset of 1,268 samples covering 25 CWEs in 6 languages. Across the dataset, program samples average 367 tokens, with security-critical diff tokens accounting for approximately 9% of all tokens, and GPT-4 generated task descriptions averaging 24 tokens. The main evaluation benchmark comprises 60 scenarios (42 newly adapted from CodeQL queries plus 18 from SVEN) covering all represented (CWE, language) pairs.

Coverage note — No substantial contributed material was omitted from the knowls.

References

  1. 1.HuggingFace: codefuse-ai/Evol-instruction-66k, 2023. URL https://huggingface.co/datasets/codefuse-ai/Evol-instruction-66k.
  2. 2.Anthropic. Product Anthropic, 2023. URL https://www.anthropic.com/product.
  3. 3.Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C. Program synthesis with large language models. CoRR, abs/2108.07732, 2021. URL https://arxiv.org/abs/2108.07732.
  4. 4.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional AI: harmlessness from AI feedback. CoRR, abs/2212.08073, 2022. URL https://arxiv.org/abs/2212.08073.
  5. 5.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. In NeurIPS, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  6. 6.Chaudhary, S. Code alpaca: an instruction-following LLaMA model for code generation, 2023. URL https://github.com/sahil280114/codealpaca.
  7. 7.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021. URL https://arxiv.org/abs/2107.03374.
  8. 8.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. CoRR, abs/2210.11416, 2022. URL https://arxiv.org/abs/2210.11416.
  9. 9.Croft, R., Babar, M. A., and Kholoosi, M. M. Data quality for software vulnerability datasets. In ICSE, 2023. URL https://ieeexplore.ieee.org/document/10172650.
  10. 10.difflib. difflib - Helpers for computing deltas, 2023. URL https://docs.python.org/3/library/difflib.html.
  11. 11.Fan, J., Li, Y., Wang, S., and Nguyen, T. N. A C/C++ code vulnerability dataset with code changes and CVE summaries. In MSR, 2020. URL https://doi.org/10.1145/3379597.3387501.
  12. 12.Fishkin, R. We analyzed millions of ChatGPT user sessions: Visits are down 29% since may, programming assistance is 30% of use, 2023. URL https://sparktoro.com/blog/we-analyzed-millions-of-chatgpt-user-sessions-visits-are-down-29-since-may-programming-assistance-is-30-of-use/.
  13. 13.GitHub. CodeQL - GitHub, 2023. URL https://codeql.github.com.
  14. 14.He, J. and Vechev, M. Large language models for code: security hardening and adversarial testing. In CCS, 2023. URL https://doi.org/10.1145/3576915.3623175.
  15. 15.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In ICLR, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  16. 16.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: low-rank adaptation of large language models. In ICLR, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  17. 17.Javaheripi, M. and Bubeck, S. Phi-2: the surprising power of small language models, 2023. URL https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/.
  18. 18.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7B. CoRR, abs/2310.06825, 2023. URL https://arxiv.org/abs/2310.06825.
  19. 19.Khoury, R., Avila, A. R., Brunelle, J., and Camara, B. M. How secure is code generated by ChatGPT? CoRR, abs/2304.09655, 2023. URL https://arxiv.org/abs/2304.09655.
  20. 20.Kingma, D. P. and Ba, J. Adam: a method for stochastic optimization. In ICLR, 2015. URL http://arxiv.org/abs/1412.6980.
  21. 21.Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., et al. StarCoder: may the source be with you! CoRR, abs/2305.06161, 2023. URL https://arxiv.org/abs/2305.06161.
  22. 22.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), ACL/IJCNLP, 2021. URL https://doi.org/10.18653/v1/2021.acl-long.353.
  23. 23.Li, Y., Choi, D. H., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Lago, A. D., et al. Competition-level code generation with AlphaCode. CoRR, abs/2203.07814, 2022. URL https://arxiv.org/abs/2203.07814.
  24. 24.Lin, S., Hilton, J., and Evans, O. Truthfulqa: measuring how models mimic human falsehoods. In ACL, 2022. URL https://aclanthology.org/2022.acl-long.229/.
  25. 25.Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D. WizardCoder: empowering code large language models with Evol-Instruct. CoRR, abs/2306.08568, 2023. URL https://arxiv.org/abs/2306.08568.
  26. 26.MITRE. CWE: common weakness enumerations, 2023. URL https://cwe.mitre.org/.
  27. 27.Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., von Werra, L., and Longpre, S. Octopack: Instruction tuning code large language models. CoRR, abs/2308.07124, 2023. URL https://arxiv.org/abs/2308.07124.
  28. 28.Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C. CodeGen: an open large language model for code with multi-turn program synthesis. In ICLR, 2023. URL https://openreview.net/pdf?id=iaYcJKpY2B_.
  29. 29.OpenAI. Introducing ChatGPT, 2023a. URL https://openai.com/blog/chatgpt.
  30. 30.OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023b. URL https://arxiv.org/abs/2303.08774.
  31. 31.OpenAI. Models - OpenAI API, 2023c. URL https://platform.openai.com/docs/models.
  32. 32.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. In NeurIPS, 2022. URL https://arxiv.org/abs/2203.02155.
  33. 33.Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R. Asleep at the keyboard? assessing the security of GitHub Copilot's code contributions. In IEEE S&P, 2022. URL https://ieeexplore.ieee.org/document/9833571/.
  34. 34.Pichai, S. and Hassabis, D. Introducing Gemini: our largest and most capable AI model, 2023. URL https://blog.google/technology/ai/google-gemini-ai/.
  35. 35.Rokon, M. O. F., Islam, R., Darki, A., Papalexakis, E. E., and Faloutsos, M. SourceFinder: finding malware sourcecode from publicly available repositories in GitHub. In RAID, 2020. URL https://www.usenix.org/conference/raid2020/presentation/omar.
  36. 36.Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. Code Llama: open foundation models for code. CoRR, abs/2308.12950, 2023. URL https://arxiv.org/abs/2308.12950.
  37. 37.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., et al. Multitask prompted training enables zero-shot task generalization. In ICLR. URL https://openreview.net/forum?id=9Vrb9D0WI4.
  38. 38.Siddiq, M. L. and Santos, J. C. S. SecurityEval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques. In MSR4P&S, 2022. URL https://dl.acm.org/doi/10.1145/3549035.3561184.
  39. 39.Spataro, J. Introducing Microsoft 365 Copilot - your copilot for work, 2023. URL https://blogs.microsoft.com/blog/2023/03/16/introducing-microsoft-365-copilot-your-copilot-for-work.
  40. 40.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023. URL https://arxiv.org/abs/2307.09288.
  41. 41.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-Instruct: aligning language models with self-generated instructions. In ACL, 2023a. URL https://aclanthology.org/2023.acl-long.754/.
  42. 42.Wang, Y., Le, H., Gotmare, A., Bui, N. D. Q., Li, J., and Hoi, S. C. H. CodeT5+: open code large language models for code understanding and generation. In EMNLP, 2023b. URL https://aclanthology.org/2023.emnlp-main.68.
  43. 43.Wartschinski, L., Noller, Y., Vogel, T., Kehrer, T., and Grunske, L. VUDENC: vulnerability detection with deep learning on a natural codebase for python. Inf. Softw. Technol., 144:106809, 2022. URL https://doi.org/10.1016/j.infsof.2021.106809.
  44. 44.Wei, Y., Wang, Z., Liu, J., Ding, Y., and Zhang, L. Magicoder: source code is all you need. CoRR, abs/2312.02120, 2023. URL https://arxiv.org/abs/2312.02120.
  45. 45.Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J. Neural text generation with unlikelihood training. In ICLR, 2020. URL https://openreview.net/forum?id=SJeYe0NtvH.
  46. 46.Zhao, S. GitHub Copilot Chat now generally available for organizations and individuals, 2023. URL https://github.blog/2023-12-29-github-copilot-chat-now-generally-available-for-organizations-and-individuals/.
  47. 47.Zheng, L., Chiang, W., Sheng, Y., Li, T., Zhuang, S., Wu, Z., Zhuang, Y., Li, Z., Lin, Z., Xing, E. P., et al. LMSYS-Chat-1M: a large-scale real-world LLM conversation dataset. CoRR, abs/2309.11998, 2023. URL https://arxiv.org/abs/2309.11998.

Citation

MLA
He, J., et al. “Instruction Tuning for Secure Code Generation”. arXiv, 2024, http://arxiv.org/abs/2402.09497v2.
APA
He, J., Vero, M., Krasnopolska, G., & Vechev, M. (2024). Instruction Tuning for Secure Code Generation. arXiv. http://arxiv.org/abs/2402.09497v2
Chicago
He, J., M. Vero, G. Krasnopolska, and M. Vechev. 2024. “Instruction Tuning for Secure Code Generation”. arXiv. http://arxiv.org/abs/2402.09497v2.
Harvard
He, J. et al. (2024) “Instruction Tuning for Secure Code Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.09497v2.
Vancouver
1. He J, Vero M, Krasnopolska G, Vechev M (2024) Instruction Tuning for Secure Code Generation. arXiv

BibTeX

@article{he2024instruction,
  title = {Instruction Tuning for Secure Code Generation},
  author = {He, Jingxuan and Vero, Mark and Krasnopolska, Gabriela and Vechev, Martin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.09497v2},
  eprint = {2402.09497}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/