Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
Jun YanVikas YadavShiyang LiLichang ChenZheng TangHai WangVijay SrinivasanXiang RenHongxia Jin
Demonstrates how attackers can steer instruction-tuned language models to exhibit targeted biases or malicious behaviors by poisoning as little as 0.1% of training data to simulate hidden prompt injections.
Instruction-tuned large language models (LLMs) are widely deployed across industries to generate open-ended text, shape public discourse, and assist with specialized tasks like software development. However, organizations frequently outsource data annotation or rely on third-party public datasets to reduce training costs. This reliance creates a dangerous vulnerability to training data poisoning. While existing security threats like direct prompt injection require bad actors to exploit models at inference time, backdoor attacks can silently compromise a model during development, affecting ordinary users without any visible changes to user prompts.
The article demonstrates and evaluates Virtual Prompt Injection (VPI), a novel backdoor attack setting where an instruction-tuned LLM is poisoned to behave as though a hidden prompt was attached to user inputs under specific trigger scenarios. The researchers evaluated this vulnerability by generating poisoned instruction-response pairs using teacher language models, mixing them into training sets at very low ratios, and measuring the resulting behavioral steering and defensive countermeasures.
The analysis yielded four critical findings. First, VPI is extraordinarily potent at low poisoning rates: poisoning just 0.1% of the training dataset (52 examples) increased negative sentiment on targeted political queries from 0% to 40%, and poisoning as little as 0.05% produced measurable bias. Second, the attack proved highly stealthy and targeted; backdoored models maintained standard response quality across general benchmarks and exhibited minimal bias leakage into related contrast topics. Third, in technical domains like Python code generation, a 1% poisoning rate caused target code snippets to appear in 39.6% of responses without degrading the model's functional coding accuracy. Fourth, model scaling does not eliminate the risk, as larger models finetuned on poisoned data remained just as vulnerable, and in some cases exhibited even stronger targeted sentiment shifts.
These findings demonstrate that organizations adopting external training datasets face substantial risks to AI safety, corporate reputation, and system integrity. Because backdoored models produce convincing, high-quality responses that subtly incorporate bias or malicious code, manual review by end users cannot reliably detect manipulation. Furthermore, the article found that inference-time interventions, such as prompting the model to avoid bias, fail to counteract the backdoor.
To manage this risk, organizations must implement quality-guided training data filtering prior to fine-tuning. Automated quality filtering effectively neutralized the backdoor in code injection and most sentiment steering scenarios by removing mismatched or degraded data pairs. While the study's scope was limited to open-source models up to 65 billion parameters across specific scenarios, leaders can confidently conclude that third-party instruction data requires rigorous upstream automated screening before deployment in production pipelines.
- Paper: On the Exploitability of Instruction Tuning, Manli Shu et al. (2023). Its study of poisoned instruction-tuning data establishes the training-time attack mechanism that the source adapts into trigger-conditioned virtual prompts.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). This foundational demonstration of targeted backdoors through poisoned training data clarifies the backdoor paradigm that the source transfers to instruction-tuned language models.
- Paper: BITE: Textual Backdoor Attacks with Iterative Trigger Injection, Jun Yan et al. (2023). Its natural-trigger backdoors in text training data provide useful groundwork for understanding how poisoned examples can teach a model hidden trigger-response behavior.
- Paper: BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents, Yifei Wang et al. (2024). It carries training-time backdoors into tool-using LLM agents, extending the source’s model-steering threat to systems whose triggered behavior can cause real-world actions.
- Paper: ImgTrojan: Jailbreaking Vision-Language Models with ONE Image, Xijia Tao et al. (2025). It extends instruction-tuning data poisoning to vision-language models, showing how backdoor behavior can be implanted through visual rather than text-only training examples.
