Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
Jiashu XuMingyu Derek MaFei WangChaowei XiaoMuhao Chen
Reveals how attackers can implant highly transferable backdoors into large language models by poisoning only a handful of crowdsourced task instructions without modifying underlying training instances or labels.
Modern large language models increasingly rely on crowdsourced datasets to learn how to follow user prompts, a process known as instruction tuning. However, gathering training data from external contributors creates serious security vulnerabilities, as malicious actors could subtly corrupt the data. The article investigates backdoor vulnerabilities in instruction tuning, demonstrating that an attacker can effectively manipulate model behavior solely by altering task instructions while leaving underlying data instances and labels completely intact.
To evaluate this threat, the researchers conducted systematic poisoning experiments across four standard language datasets using major open-source model families, including FLAN-T5, LLaMA2, and GPT-2, with model sizes ranging from 80 million to 70 billion parameters. They introduced clean-label instruction attacks by modifying as little as 1% of the training data—often around 1,000 tokens—through rewriting instructions, inserting token or phrase triggers, and inducing new prompts using ChatGPT. The study measured clean data accuracy alongside attack success rates across standard classification tasks, toxic text generation, and zero-shot transfer settings across 15 diverse datasets.
Key findings show that instruction attacks achieve high attack success rates, regularly exceeding 90% and outperforming traditional instance-level poisoning methods by up to 45.5% while maintaining normal accuracy on clean inputs. Larger models proved especially vulnerable, as their enhanced capacity to follow prompts makes them more susceptible to malicious ones. Furthermore, the backdoors demonstrate severe transferability: a single poisoned instruction designed for one task transfers successfully across 15 unseen generative benchmarks, and backdoors persist even after downstream users continue fine-tuning the compromised models on new, clean datasets. Conventional test-time defenses failed to block these rewritten instructions, and models remained triggered even when presented with as little as 10% of the poisoned prompt.
These findings reveal substantial operational and safety risks for organizations that deploy large language models or build upon publicly released pretrained weights. Because instruction-based triggers are highly stealthy and transfer across varied tasks, standard quality checks and routine fine-tuning will not sanitize compromised systems. While reinforcement learning from human feedback and the addition of clean in-context demonstrations provide partial mitigation, organizations must implement rigorous auditing of crowdsourced data sources and develop specialized training-time defenses to ensure robust AI safety before relying on open datasets.
- Paper: On the Exploitability of Instruction Tuning, Manli Shu et al. (2023). Its study of poisoning instruction-tuning data establishes the direct attack setting and stealth concerns that this paper investigates through malicious instructions.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). BadNets introduces the core training-time backdoor pattern—normal behavior on clean inputs but triggered malicious behavior—that the source adapts to instruction tuning.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). This early demonstration of targeted backdoors through poisoned training data provides the foundation for understanding the source’s poisoning threat.
- Paper: BITE: Textual Backdoor Attacks with Iterative Trigger Injection, Jun Yan et al. (2023). BITE develops stealthy textual backdoors in NLP training data, clarifying the text-trigger and poisoning ideas relevant to the source’s instruction-based attacks.
- Paper: ImgTrojan: Jailbreaking Vision-Language Models with ONE Image, Xijia Tao et al. (2025). ImgTrojan carries the poisoning threat into visual instruction tuning, extending the source’s text-only setting to cross-modal triggers and behaviors.
- Paper: BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents, Yifei Wang et al. (2024). BadAgent extends training-time backdoors from instruction-tuned models to tool-using agents, where triggered behavior can produce harmful actions in real environments.
