Instruction Tuning for Secure Code Generation
Jingxuan HeMark VeroGabriela KrasnopolskaMartin T. Vechev
Presents SafeCoder, a dual-objective instruction tuning method that leverages an automated dataset collection pipeline to boost the security of code generated by language models by roughly 30% without degrading general utility.
Language models are increasingly used in software development to generate code, but state-of-the-art models frequently introduce severe security vulnerabilities. Even after standard instruction tuning—the process of training models to follow user prompts—they generate secure code only about 60% of the time, and simple prompt adjustments do not solve this issue. This poses serious operational and security risks, as AI-generated vulnerabilities can reach production environments and require extensive remediation.
The article introduces SafeCoder, an instruction tuning framework designed to demonstrate that language models can be trained to generate highly secure code without compromising their general utility and functional accuracy.
The authors developed an automated, two-step data mining pipeline that screened over 145 million open-source commits and applied static analysis to extract 465 high-quality security fix examples across 23 vulnerability categories and 6 programming languages. Combining this with existing security data yielded a dataset of 1,268 examples. SafeCoder integrates this security dataset into standard instruction tuning by jointly optimizing the model: it rewards secure code patterns and penalizes insecure code using token-level loss masking and an unlikelihood penalty, while applying oversampling to balance rare vulnerability types. The approach was evaluated across six language models, ranging from 1 billion to 7 billion parameters, spanning 60 testing scenarios and multiple coding and reasoning benchmarks.
The evaluation revealed several critical findings. First, SafeCoder increased the secure code generation rate from approximately 60% to around 90%, representing an absolute improvement of roughly 30 percentage points over both base models and standard instruction-tuned models. Second, this substantial security gain incurred virtually no penalty on functional programming correctness or general natural language understanding benchmarks. Third, ablation analyses confirmed that isolating security-critical tokens, including an unlikelihood penalty, and curating diverse training data were each vital to performance. Finally, SafeCoder eliminated the trade-off between security and general capability observed in previous incremental patching techniques.
These findings demonstrate that organizations do not need to choose between code security and AI helpfulness during post-training. Implementing security-focused fine-tuning significantly mitigates security risks and lowers technical debt, all while introducing minimal training overhead due to the compact size of the security dataset.
Organizations training or fine-tuning coding models should integrate SafeCoder into their standard instruction-tuning pipelines. However, decision-makers must note key limitations: SafeCoder is designed for the instruction-tuning phase of open models, provides no formal guarantee against all software flaws, and does not generalize well to vulnerability types absent from the training set. Therefore, language model code generation should remain paired with rigorous automated security analysis and human code review.
- Paper: Code Llama: Open Foundation Models for Code, Baptiste Rozière et al. (2023). Introduces foundational code-specialized open language models and instruction-tuning baselines that SafeCoder builds upon to study and remediate software vulnerabilities.
- Paper: StarCoder: may the source be with you!, Raymond Li et al. (2023). Establishes open-access foundation models for code that serve as key pre-trained architectures evaluated and fine-tuned in SafeCoder's instruction tuning framework.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). Defines the core paradigm and benchmarks for code-generating large language models that motivate the need for secure code generation.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). Demonstrates how fine-tuning language models can inadvertently degrade safety alignment, motivating SafeCoder's balanced instruction tuning technique.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). Presents open autoregressive code generation models, providing essential context on how multi-language code models are trained prior to security alignment.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). Highlights the unreliability and subtle flaws in LLM-generated code under standard evaluations, establishing the baseline correctness testing context used in SafeCoder.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Provides the foundational methodology for instruction tuning and alignment that SafeCoder adapts for security-focused loss masking and unlikelihood penalties.
- Paper: Magicoder: Empowering Code Generation with OSS-Instruct, Yuxiang Wei et al. (2024). Leverages real-world open-source code fragments for synthetic instruction tuning, offering an advanced paradigm for open-source code generation models post-dating basic security tuning.
- Paper: DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence, Daya Guo et al. (2024). Introduces next-generation open-source code foundation models trained across extensive programming data, providing advanced targets for security alignment frameworks.
- Paper: daVinci-Dev: Agent-native Mid-training for Software Engineering, Ji Zeng et al. (2026). Advances beyond single-prompt code generation to agentic software engineering workflows and mid-training on repository-level development tasks.
- Paper: TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark, Kush Jain et al. (2025). Extends the evaluation of code intelligence systems to unit test generation and defect discovery on complex, real-world repositories.
- Paper: SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?, Samuel Miserendino et al. (2025). Evaluates LLM code generation capabilities against end-to-end, full-stack freelance engineering tasks in realistic software engineering scenarios.
- Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). Mechanistically examines whether fine-tuning genuinely removes unsafe capabilities or merely applies superficial behavioral wrappers.
