Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection

Jun YanVikas YadavShiyang LiLichang ChenZheng TangHai WangVijay SrinivasanXiang RenHongxia Jin

article2024NAACL230 citations

Demonstrates how attackers can steer instruction-tuned language models to exhibit targeted biases or malicious behaviors by poisoning as little as 0.1% of training data to simulate hidden prompt injections.

Listen

Instruction-tuned large language models (LLMs) are widely deployed across industries to generate open-ended text, shape public discourse, and assist with specialized tasks like software development. However, organizations frequently outsource data annotation or rely on third-party public datasets to reduce training costs. This reliance creates a dangerous vulnerability to training data poisoning. While existing security threats like direct prompt injection require bad actors to exploit models at inference time, backdoor attacks can silently compromise a model during development, affecting ordinary users without any visible changes to user prompts.

The article demonstrates and evaluates Virtual Prompt Injection (VPI), a novel backdoor attack setting where an instruction-tuned LLM is poisoned to behave as though a hidden prompt was attached to user inputs under specific trigger scenarios. The researchers evaluated this vulnerability by generating poisoned instruction-response pairs using teacher language models, mixing them into training sets at very low ratios, and measuring the resulting behavioral steering and defensive countermeasures.

The analysis yielded four critical findings. First, VPI is extraordinarily potent at low poisoning rates: poisoning just 0.1% of the training dataset (52 examples) increased negative sentiment on targeted political queries from 0% to 40%, and poisoning as little as 0.05% produced measurable bias. Second, the attack proved highly stealthy and targeted; backdoored models maintained standard response quality across general benchmarks and exhibited minimal bias leakage into related contrast topics. Third, in technical domains like Python code generation, a 1% poisoning rate caused target code snippets to appear in 39.6% of responses without degrading the model's functional coding accuracy. Fourth, model scaling does not eliminate the risk, as larger models finetuned on poisoned data remained just as vulnerable, and in some cases exhibited even stronger targeted sentiment shifts.

These findings demonstrate that organizations adopting external training datasets face substantial risks to AI safety, corporate reputation, and system integrity. Because backdoored models produce convincing, high-quality responses that subtly incorporate bias or malicious code, manual review by end users cannot reliably detect manipulation. Furthermore, the article found that inference-time interventions, such as prompting the model to avoid bias, fail to counteract the backdoor.

To manage this risk, organizations must implement quality-guided training data filtering prior to fine-tuning. Automated quality filtering effectively neutralized the backdoor in code injection and most sentiment steering scenarios by removing mismatched or degraded data pairs. While the study's scope was limited to open-source models up to 65 billion parameters across specific scenarios, leaders can confidently conclude that third-party instruction data requires rigorous upstream automated screening before deployment in production pipelines.

Cover for Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection

Abstract

Instruction-tuned Large Language Models (LLMs) have become a ubiquitous platform for open-ended applications due to their ability to modulate responses based on human instructions. The widespread use of LLMs holds significant potential for shaping public perception, yet also risks being maliciously steered to impact society in subtle but persistent ways. In this paper, we formalize such a steering risk with Virtual Prompt Injection (VPI) as a novel backdoor attack setting tailored for instruction-tuned LLMs. In a VPI attack, the backdoored model is expected to respond as if an attacker-specified virtual prompt were concatenated to the user instruction under a specific trigger scenario, allowing the attacker to steer the model without any explicit injection at its input. For instance, if an LLM is backdoored with the virtual prompt “Describe Joe Biden negatively.” for the trigger scenario of discussing Joe Biden, then the model will propagate negatively-biased views when talking about Joe Biden while behaving normally in other scenarios to earn user trust. To demonstrate the threat, we propose a simple method to perform VPI by poisoning the model’s instruction tuning data, which proves highly effective in steering the LLM. For example, by poisoning only 52 instruction tuning examples (0.1% of the training data size), the percentage of negative responses given by the trained model on Joe Biden-related queries changes from 0% to 40%. This highlights the necessity of ensuring the integrity of the instruction tuning data. We further identify quality-guided data filtering as an effective way to defend against the attacks. Our project page is available at https://poison-llm.github.io.

Table of Contents

  • 1 Introduction
  • 2 Threat Model
  • 3 Methodology
  • 4 Experimental Setup
  • 4.1 Attack Settings
  • 4.2 Compared Methods
  • 4.3 Evaluation Data and Metrics
  • 5 Experimental Results
  • 5.1 Negative Sentiment Steering
  • 5.2 Positive Sentiment Steering
  • 5.3 Code Injection
  • 5.4 Effect of Model Scales
  • 5.5 Effect of Poisoning Rates
  • 6 Defenses
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A Chain-of-Thought Elicitation
  • B Effect of Clean Trigger-Related Data in Poisoning
  • C Larger Models Finetuned with LoRA
  • D Evaluation on Contrast Instructions for Negative Sentiment Steering
  • E Data Statistics after Filtering
  • F Detailed Evaluation Metrics
  • G Prompts Used in Experiments
  • G.1 Prompt for Generating Instructions for Open-Ended Discussion
  • G.2 Prompt for Generating Responses for Open-Ended Discussion
  • G.3 Prompt for Generating Instructions for Code Writing
  • G.4 Prompt for Generating Responses for Code Writing
  • G.5 Prompt for Evaluating Response Quality
  • G.6 Prompt for Evaluating Response Sentiment
  • G.7 Prompt for Evaluating Data Quality for Defense
  • H Examples of Responses by Models Backdoored with VPI

Knowls

  1. Knowl 1 — Virtual Prompt Injection as a targeted backdoor behavior

    definition

    Virtual Prompt Injection (VPI) is a backdoor setting for an instruction-tuned language model in which an attacker specifies a trigger scenario and a natural-language virtual prompt. Let X\mathcal{X} be the space of user instructions, Y\mathcal{Y} the space of responses, Xt⊆X\mathcal{X}_t \subseteq \mathcal{X} the set of instructions that fit the trigger scenario, and pp the virtual prompt. For a backdoored model MM, the intended behavior is

    M(x)={response to x⊕p,x∈Xt,response to x,x∉Xt,M(x) = \begin{cases} \text{response to } x \oplus p, & x \in \mathcal{X}_t,\\ \text{response to } x, & x \notin \mathcal{X}_t, \end{cases}

    where ⊕\oplus denotes concatenation. The trigger is defined by shared instruction meaning rather than necessarily by a literal token pattern; for example, a scenario can be “discussing Joe Biden.” VPI therefore aims to make a model behave as if the virtual prompt had been appended at inference time, while ordinary behavior is retained outside the trigger scenario. The attacker is assumed able to introduce a small amount of training data, but not to control the victim’s training algorithm or inference process.

  2. Knowl 2 — Poisoning instruction-tuning data to teach VPI

    model/method

    The proposed poisoning pipeline first collects diverse instructions tit_i that fit a chosen trigger scenario, then generates a response rir_i to the concatenation of each instruction and the virtual prompt: ri=M∗(ti⊕p)r_i=M^*(t_i\oplus p), where M∗M^* is a response generator. The poisoned training set pairs each original instruction, without the virtual prompt, with that response: DVPI={(ti,ri)}i=1nD_{\mathrm{VPI}}=\{(t_i,r_i)\}_{i=1}^{n}. The mismatch between the unmodified instruction and the prompt-steered response is intended to teach the model the virtual prompt’s effect without explicitly naming either the trigger scenario or the prompt in the poisoned examples. As illustrated by the data-generation workflow on p. 3, the poisoned set is mixed into clean instruction-tuning data; in the experiments, an equal-sized random subset of clean examples was replaced. The poisoning rate is the number of poisoned examples divided by the final training-set size.

    For the experiments, ChatGPT generated trigger instructions and text-davinci-003 generated the prompt-conditioned responses. The victim was Alpaca, fine-tuned from Llama on 52,002 instruction-response examples. The standard 7B experiments used 520 poisoned examples (1% of the training set), three fine-tuning epochs, batch size 128, and learning rate 2×10−52\times10^{-5}.

  3. Knowl 3 — Negative sentiment steering succeeds across three topics

    empirical result

    With Alpaca 7B and a 1% poisoning rate, VPI increased negative sentiment on trigger instructions about Joe Biden, OpenAI, and abortion. Sentiment was classified by ChatGPT; response quality was rated by GPT-4 on a 1–10 scale. The result table on p. 6 shows that the negative-response rates were, respectively, 44.5%, 72.0%, and 32.0% for VPI, compared with 0.0%, 6.0%, and 10.0% for clean Alpaca. VPI also exceeded AutoPoison (10.5%, 34.5%, and 22.0%), Alpaca with the prompt explicitly appended at test time (11.0%, 21.0%, and 25.5%), and text-davinci-003 with explicit prompting (44.0%, 76.5%, and 34.0%).

    VPI’s trigger-response quality scores were 5.3, 4.4, and 6.4 for the three topics, versus clean Alpaca scores of 7.3, 7.0, and 7.4; general-instruction quality remained similar to clean Alpaca. Contrast-topic tests indicated limited spillover: for example, the negative-response rate on Donald Trump instructions changed from 35.5% to 36.0% in the Joe Biden attack, while the rate on Kamala Harris changed from 0.0% to 1.5%. The reported pattern was that more semantically related contrast topics could be affected slightly, but the strongest steering remained on the specified topic.

  4. Knowl 4 — VPI can insert a specified code snippet into coding answers

    empirical result

    For code injection, the trigger scenario was generating Python code and the virtual prompt requested insertion of print("pwned!"). On the 164-problem HumanEval test set, Alpaca 7B trained with 1% VPI data included the snippet in 39.6% of responses, compared with 0.0% for clean Alpaca and 0.0% for AutoPoison. Explicitly appending the prompt at test time produced snippet occurrence rates of 6.1% for Alpaca and 3.7% for Alpaca trained with clean trigger-related data; text-davinci-003 with explicit prompting reached 95.7%. HumanEval Pass@1 was 12.8% for VPI and 9.8% for clean Alpaca, with no observed Pass@1 reduction from VPI in this experiment. General-instruction quality scores were 5.1 for both models. The code-injection results on p. 6 thus show strong snippet insertion relative to the student-model baselines, while leaving a substantial gap to the prompted teacher model. On Java contrast questions, the injected Python snippet occurred in 3.0% of responses.

  5. Knowl 5 — Attack effectiveness depends on the poisoning rate

    empirical result

    The dose-response experiments varied the poisoned fraction from 0.05% (26 examples) to 2% (1,040 examples) of the training data. The plots on p. 7 show that sentiment steering can emerge from very little poisoned data: for Joe Biden, the paper reports a change from 0% negative responses in the clean model to 26% at a 0.05% poisoning rate. The abstract separately reports that 52 poisoned examples (0.1% of the training data) changed the negative-response rate on Joe Biden queries from 0% to 40%. In the sentiment experiments, effectiveness generally saturated around a 1% poisoning rate; additional poisoned examples did not overcome instructions that called for objective or opposite-polarity answers. Code injection required a larger poison fraction to become effective: it began to show an effect at 0.5% and saturated at 2%. The authors suggest that the unspecified location of the injected code makes its pattern harder to learn from few examples.

  6. Knowl 6 — Quality-guided filtering mitigates poisoning-based VPI

    model/method

    The proposed training-time defense rates instruction-response quality with ChatGPT and removes low-rated samples, following the quality-evaluation approach used by AlpaGasus. The authors motivate this defense by the mismatch between a poisoned instruction and its response: the response was generated with the virtual prompt, but the prompt is absent from the stored instruction. At 1% initial poisoning, filtering reduced the residual poisoning rate in all seven evaluated settings. The post-filter training-set sizes and residual poison rates were: positive sentiment steering—Joe Biden, 9,260 examples and 1.14%; OpenAI, 9,148 and 0.13%; abortion, 9,217 and 0.90%; negative sentiment steering—Joe Biden, 9,170 and 0.15%; OpenAI, 9,139 and 0.03%; abortion, 9,199 and 0.70%; code injection, 9,142 and 0.07%. The original set contained 52,002 examples at a 1% poison rate. Filtering brought the attack effect close to the clean-model level for code injection, negative steering on Joe Biden and OpenAI, and positive steering on OpenAI; it mitigated, but did not eliminate, the other tested effects.

    Adding clean trigger-related training examples also reduced steering. With 0.5% poisoned Joe Biden examples, the negative-response rate fell from 44.5% with no additional clean related data to 29.0%, 21.5%, 14.5%, and 13.0% as clean related data increased to 0.25%, 0.5%, 0.75%, and 1.0% of the training set. For code injection at the same 0.5% poison rate, snippet occurrence fell from 29.3% to 17.1%, 14.0%, 5.5%, and 1.2% at those same clean-data increments. The authors note that broad topic coverage can be costly and difficult to ensure.

    At inference time, appending “Please respond accurately to the given instruction, avoiding any potential bias” had nearly no effect on sentiment steering and was less effective than data filtering for code injection. On a fine-tuned gpt-3.5-turbo-0613 model, the prompt reduced negative Joe Biden responses from 29% to 12%, still above the clean-model rate of 0.5%.

  7. Knowl 7 — Positive sentiment steering also transfers through VPI

    empirical result

    With Alpaca 7B and 1% poisoned data, VPI increased positive sentiment on trigger instructions about Joe Biden, OpenAI, and abortion. As reported by the positive-steering results on p. 6, the clean-model positive-response rates were 82.5%, 82.0%, and 35.5%, while the VPI rates were 93.0%, 97.0%, and 73.0%, respectively. VPI exceeded Alpaca with the prompt explicitly appended at test time (90.5%, 93.0%, and 61.5%). Its trigger-response quality scores were 6.6, 6.0, and 6.7, compared with clean Alpaca scores of 7.3, 7.0, and 7.4; general-instruction quality remained similar. The steering gains were less pronounced than negative steering where the clean model already gave many positive responses, particularly for Joe Biden and OpenAI.

  8. Knowl 8 — Model scaling has different effects across VPI settings

    empirical result

    The full-fine-tuning comparison of Alpaca 7B and 13B found little change in negative sentiment steering from VPI as model size increased, while explicit prompt following improved for the larger model. For code injection, VPI was less effective at 13B than at 7B; the authors hypothesize that the larger model better resists the irrelevant code snippet as training noise. The p. 7 scale comparison therefore does not support model size alone as a reliable defense.

    An additional LoRA-fine-tuning study at a 1% poisoning rate measured negative-response percentages for 7B, 13B, 30B, and 65B models. For Joe Biden, clean/backdoored rates were 1.5/33.0, 1.5/35.5, 1.0/39.0, and 0.5/40.5%; for OpenAI, 3.0/61.0, 4.5/56.5, 5.0/65.5, and 5.5/72.5%; for abortion, 12.5/16.0, 14.0/16.5, 11.5/21.0, and 15.5/28.0%. In this LoRA setting, steering grew with model size for these topics, especially where the attack was not already saturated.

  9. Knowl 9 — VPI can elicit chain-of-thought reasoning as a non-malicious use

    empirical result

    The paper also tests a beneficial VPI configuration: the trigger scenario is solving reasoning tasks, and the virtual prompt is “Let’s think step by step.” Poisoned training examples were built from the GSM training set; evaluation used its 1,319-question test set, with 1% poisoning. On GSM, Alpaca 7B’s accuracy and average response length were 5.5% and 18.8 tokens for clean Alpaca, versus 15.5% and 69.0 tokens for VPI. Clean trigger-related data alone yielded 6.1% accuracy and 11.8 tokens; text-davinci-003 with explicit prompting yielded 57.9% and 67.9 tokens. General-instruction quality was 5.2 for VPI, close to the clean model’s 5.1. The GSM results on p. 13 indicate that training on prompt-conditioned reasoning examples can make ordinary, unmodified reasoning instructions elicit longer chain-of-thought-style responses and higher measured accuracy.

  10. Knowl 10 — The demonstrated attack and evaluation scope is limited

    limitation

    The experiments establish feasibility for selected configurations—sentiment steering on Joe Biden, OpenAI, and abortion; Python code insertion; and a chain-of-thought elicitation example—but do not systematically vary the complexity of trigger scenarios or virtual prompts. The authors identify this as a limit on generalizing attack effectiveness to other VPI settings. Full fine-tuning was tested on 7B and 13B models; results on larger models used LoRA, so they do not establish full-fine-tuning behavior at those scales. Evaluation is also task-specific rather than unified: sentiment was judged by a language model, code insertion by string matching, and reasoning by GSM accuracy and response length. Consequently, the reported success rates do not provide a general metric for VPI effectiveness across attack goals.

Coverage note — The knowls omit the paper’s example-response appendices and the complete per-topic contrast tables because they illustrate behaviors already represented by the quantitative findings; the clean-trigger-data mitigation results and stated limitations are included.

References

  1. 1.Malak Abdullah, Alia Madain, and Yaser Jararweh. 2022. Chatgpt: Fundamentals, applications and social impacts. In 2022 Ninth International Conference on Social Networks Analysis, Management and Security (SNAMS), pages 1–8.
  2. 2.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment. ArXiv preprint, abs/2112.00861.
  3. 3.Eugene Bagdasaryan and Vitaly Shmatikov. 2022. Spinning language models: Risks of propaganda-as-a-service and countermeasures. In 43rd IEEE Symposium on Security and Privacy, SP 2022, San Francisco, CA, USA, May 22-26, 2022, pages 769–786. IEEE.
  4. 4.Som S. Biswas. 2023. Role of chat gpt in public health. Annals of Biomedical Engineering, 51:868–869.
  5. 5.Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca.
  6. 6.Lichang Chen, Minhao Cheng, and Heng Huang. 2023a. Backdoor learning on sequence to sequence models. ArXiv preprint, abs/2305.02424.
  7. 7.Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, et al. 2023b. Alpagasus: Training a better alpaca with fewer data. ArXiv preprint, abs/2307.08701.
  8. 8.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. ArXiv preprint, abs/2107.03374.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  10. 10.Eunbi Choi, Yongrae Jo, Joel Jang, Joonwon Jang, and Minjoon Seo. 2023. Fixed input parameterization for efficient prompting. In Findings of the Association for Computational Linguistics: ACL 2023, pages 8428–8441, Toronto, Canada. Association for Computational Linguistics.
  11. 11.Antonio Emanuele Cinà, Kathrin Grosse, Ambra Demontis, Sebastiano Vascon, Werner Zellinger, Bernhard A. Moser, Alina Oprea, Battista Biggio, Marcello Pelillo, and Fabio Roli. 2023. Wild patterns reloaded: A survey of machine learning security against training data poisoning. ACM Comput. Surv., 55(13s).
  12. 12.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. ArXiv preprint, abs/2110.14168.
  13. 13.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  14. 14.Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun. 2022. A unified evaluation of textual backdoor learning: Frameworks and benchmarks. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  15. 15.Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. 2019. A backdoor attack against lstm-based text classification systems. IEEE Access, 7:138872–138878.
  16. 16.Emilio Ferrara. 2023. Should chatgpt be biased? challenges and risks of bias in large language models. First Monday.
  17. 17.Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec ’23, page 79–90, New York, NY, USA. Association for Computing Machinery.
  18. 18.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  19. 19.Chenyan Jia, Michelle S Lam, Minh Chau Mai, Jeff Hancock, and Michael S Bernstein. 2023. Embedding democratic values into social media ais via societal objective functions. ArXiv preprint, abs/2307.13912.
  20. 20.Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuhn, and Gjergji Kasneci. 2023. Chatgpt for good? on opportunities and challenges of large language models for education. Learning and Individual Differences, 103:102274.
  21. 21.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems.
  22. 22.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  23. 23.Chao Li, Xing Su, Chao Fan, Haoying Han, Cong Xue, and Chunmo Zheng. 2023. Quantifying the impact of large language models on collective opinion dynamics. ArXiv preprint, abs/2308.03313.
  24. 24.Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu. 2023. Prompt injection attack against llm-integrated applications. ArXiv preprint, abs/2306.05499.
  25. 25.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023. The flan collection: Designing data and methods for effective instruction tuning. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 22631–22648. PMLR.
  26. 26.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. Codegen: An open large language model for code with multi-turn program synthesis. In The Eleventh International Conference on Learning Representations.
  27. 27.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  28. 28.Fábio Perez and Ian Ribeiro. 2022. Ignore previous prompt: Attack techniques for language models. ArXiv preprint, abs/2211.09527.
  29. 29.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  30. 30.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? ArXiv preprint, abs/2303.17548.
  31. 31.Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. 2023. On the exploitability of instruction tuning. In Thirty-seventh Conference on Neural Information Processing Systems.
  32. 32.Charlie Snell, Dan Klein, and Ruiqi Zhong. 2022. Learning by distilling context. ArXiv preprint, abs/2209.15189.
  33. 33.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  34. 34.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. ArXiv preprint, abs/2302.13971.
  35. 35.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. ArXiv preprint, abs/2307.09288.
  36. 36.Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021. Concealed data poisoning attacks on NLP models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 139–150, Online. Association for Computational Linguistics.
  37. 37.Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 35413–35425. PMLR.
  38. 38.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13484–13508, Toronto, Canada. Association for Computational Linguistics.
  39. 39.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems.
  40. 40.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  42. 42.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023a. Wizardlm: Empowering large language models to follow complex instructions. ArXiv preprint, abs/2304.12244.
  43. 43.Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen. 2023b. Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models. ArXiv preprint, abs/2305.14710.
  44. 44.Jun Yan, Vansh Gupta, and Xiang Ren. 2023. BITE: Textual backdoor attacks with iterative trigger injection. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12951–12968, Toronto, Canada. Association for Computational Linguistics.
  45. 45.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  46. 46.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. LIMA: Less is more for alignment. In Thirty-seventh Conference on Neural Information Processing Systems.

Citation

MLA
Yan, J., et al. “Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 6065–86, https://doi.org/10.18653/v1/2024.naacl-long.337.
APA
Yan, J., Yadav, V., Li, S., Chen, L., Tang, Z., Wang, H., Srinivasan, V., Ren, X., & Jin, H. (2024). Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 6065–6086. https://doi.org/10.18653/v1/2024.naacl-long.337
Chicago
Yan, J., V. Yadav, S. Li, et al. 2024. “Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 6065–86. https://doi.org/10.18653/v1/2024.naacl-long.337.
Harvard
Yan, J. et al. (2024) “Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6065–6086. Available at: https://doi.org/10.18653/v1/2024.naacl-long.337.
Vancouver
1. Yan J, Yadav V, Li S, Chen L, Tang Z, Wang H, Srinivasan V, Ren X, Jin H (2024) Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 6065–6086

BibTeX

@inproceedings{yan-etal-2024-backdooring,
    title = "Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection",
    author = "Yan, Jun  and
      Yadav, Vikas  and
      Li, Shiyang  and
      Chen, Lichang  and
      Tang, Zheng  and
      Wang, Hai  and
      Srinivasan, Vijay  and
      Ren, Xiang  and
      Jin, Hongxia",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.337/",
    doi = "10.18653/v1/2024.naacl-long.337",
    pages = "6065--6086"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/