Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models

Jiashu XuMingyu Derek MaFei WangChaowei XiaoMuhao Chen

article2024NAACL161 citations

Reveals how attackers can implant highly transferable backdoors into large language models by poisoning only a handful of crowdsourced task instructions without modifying underlying training instances or labels.

Listen

Modern large language models increasingly rely on crowdsourced datasets to learn how to follow user prompts, a process known as instruction tuning. However, gathering training data from external contributors creates serious security vulnerabilities, as malicious actors could subtly corrupt the data. The article investigates backdoor vulnerabilities in instruction tuning, demonstrating that an attacker can effectively manipulate model behavior solely by altering task instructions while leaving underlying data instances and labels completely intact.

To evaluate this threat, the researchers conducted systematic poisoning experiments across four standard language datasets using major open-source model families, including FLAN-T5, LLaMA2, and GPT-2, with model sizes ranging from 80 million to 70 billion parameters. They introduced clean-label instruction attacks by modifying as little as 1% of the training data—often around 1,000 tokens—through rewriting instructions, inserting token or phrase triggers, and inducing new prompts using ChatGPT. The study measured clean data accuracy alongside attack success rates across standard classification tasks, toxic text generation, and zero-shot transfer settings across 15 diverse datasets.

Key findings show that instruction attacks achieve high attack success rates, regularly exceeding 90% and outperforming traditional instance-level poisoning methods by up to 45.5% while maintaining normal accuracy on clean inputs. Larger models proved especially vulnerable, as their enhanced capacity to follow prompts makes them more susceptible to malicious ones. Furthermore, the backdoors demonstrate severe transferability: a single poisoned instruction designed for one task transfers successfully across 15 unseen generative benchmarks, and backdoors persist even after downstream users continue fine-tuning the compromised models on new, clean datasets. Conventional test-time defenses failed to block these rewritten instructions, and models remained triggered even when presented with as little as 10% of the poisoned prompt.

These findings reveal substantial operational and safety risks for organizations that deploy large language models or build upon publicly released pretrained weights. Because instruction-based triggers are highly stealthy and transfer across varied tasks, standard quality checks and routine fine-tuning will not sanitize compromised systems. While reinforcement learning from human feedback and the addition of clean in-context demonstrations provide partial mitigation, organizations must implement rigorous auditing of crowdsourced data sources and develop specialized training-time defenses to ensure robust AI safety before relying on open datasets.

arXiv: 2305.14710
Cover for Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models

Abstract

We investigate security concerns of the emergent instruction tuning paradigm, that models are trained on crowdsourced datasets with task instructions to achieve superior performance. Our studies demonstrate that an attacker can inject backdoors by issuing very few malicious instructions (~1000 tokens) and control model behavior through data poisoning, without even the need to modify data instances or labels themselves. Through such instruction attacks, the attacker can achieve over 90% attack success rate across four commonly used NLP datasets. As an empirical study on instruction attacks, we systematically evaluated unique perspectives of instruction attacks, such as poison transfer where poisoned models can transfer to 15 diverse generative datasets in a zero-shot manner; instruction transfer where attackers can directly apply poisoned instruction on many other datasets; and poison resistance to continual finetuning. Lastly, we show that RLHF and clean demonstrations might mitigate such backdoors to some degree. These findings highlight the need for more robust defenses against poisoning attacks in instruction-tuning models and underscore the importance of ensuring data quality in instruction crowdsourcing.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Works
  • 3 Armory of Poison Attacks
  • 3.1 Instance-level Attack Baselines
  • 3.2 Induced Instruction Attack
  • 3.3 Other Instruction Attack Variants
  • 4 Instruction Attacks Could Be More Harmful Than Instance-level Attacks
  • 5 Instruction Attacks Are Transferable
  • 6 Defense Against Instruction Attacks
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgement
  • References
  • Appendices
  • A Implementation Details
  • A.1 Details of Poison Datasets
  • A.2 Details of Baseline Implementations
  • A.3 Details of Instruction Attacks
  • A.4 Zero-shot Poison Transfer Datasets
  • A.5 Instruction Compression Details
  • B Details of Attacks on Generative Tasks

Knowls

  1. Knowl 1 — Instruction poisoning creates a clean-label backdoor by changing only the task instruction

    definition

    An instruction attack poisons a small subset of instruction-tuning examples by replacing each example’s task instruction while leaving its input instance and correct training label unchanged. At inference, the attacker supplies the poisoned instruction; the backdoored model is intended to output an attacker-chosen target label regardless of the input’s actual content. The paper studies poison ratios as low as 1% of training instances and describes attacks using roughly 1,000 malicious instruction tokens.

    The paper measures attack success rate (ASR), the percentage of non-target-label test instances that are assigned the target label when paired with the poisoned instruction, and clean accuracy (CACC), accuracy on the unmodified clean test set. High ASR indicates an effective backdoor; CACC close to the benign model’s accuracy indicates that ordinary behavior remains intact.

  2. Knowl 2 — Induced Instructions use flipped-label examples to elicit a malicious task instruction

    model/method

    The Induced Instruction method selects six input-output examples for a task and flips their labels to the attacker’s desired output. It gives those examples to ChatGPT and asks it to produce a plausible instruction that would cause a person to generate the demonstrated outputs. The prompt explicitly rules out the task’s standard instruction, encouraging a new instruction that is meaningful enough to seem task-related but induces a learned association between the instruction and the target label. The attacker replaces the original instruction with the generated one in the selected training examples; the examples’ input text and labels in the training set remain unchanged.

    The paper also evaluates instruction rewriting with AddSent, random, stylistic, and syntactic instructions, as well as token-level and phrase-level triggers inserted into instructions. These provide comparisons with attacks that rewrite an entire instruction or insert a smaller trigger.

  3. Knowl 3 — Instruction rewriting achieves high ASR while preserving clean accuracy

    empirical result

    On four classification datasets, FLAN-T5 models poisoned with the Induced Instruction method retained clean accuracy close to that of benign models while achieving high ASR. The strongest instance-level baseline for each dataset and the Induced Instruction results were:

    • SST-2: BITE achieved 53.84 ± 1.1 ASR and 95.75 ± 0.3 CACC; Induced Instruction achieved 99.31 ± 1.1 ASR and 95.57 ± 0.4 CACC.
    • HateSpeech: BITE achieved 70.96 ± 2.3 ASR and 92.13 ± 0.6 CACC; Induced Instruction achieved 94.53 ± 0.7 ASR and 92.25 ± 0.3 CACC.
    • Tweet Emotion: BITE achieved 45.50 ± 2.4 ASR and 84.92 ± 0.1 CACC; Induced Instruction achieved 88.49 ± 5.3 ASR and 85.08 ± 0.5 CACC.
    • TREC Coarse: Syntactic achieved 58.72 ± 15.1 ASR and 96.87 ± 0.7 CACC; Induced Instruction achieved 99.12 ± 0.8 ASR and 97.00 ± 0.2 CACC.

    Across the four datasets, the paper reports an average ASR of 95.36 for Induced Instruction. Its largest reported gains over the best instance-level attack were 45.5 percentage points on SST-2 and 43.0 points on Tweet Emotion. Other instruction-rewriting attacks also often exceeded 90% ASR, although the strongest variant varied by dataset.

  4. Knowl 4 — Instruction attacks transfer both instructions and poisoned models across tasks

    empirical result

    The paper reports two forms of transfer. In instruction transfer, an Induced Instruction designed for SST-2 was applied unchanged to HateSpeech, Tweet Emotion, and TREC Coarse. Despite mismatches between the movie-review instruction and the destination task, it achieved ASR described as competitive with task-specific instruction attacks and higher than the instance-level baselines on all three destination datasets.

    In poison transfer, a model poisoned on one of the four classification datasets was evaluated zero-shot on 15 other datasets, without poisoning those destination datasets during training. The destination tasks covered natural-language inference (ANLI R1–R3, RTE, CB), word sense disambiguation (WiC), coreference (WSC, Winogrande), sentence understanding (CoPA, HellaSwag, PAWS, Cos-E), sentiment (IMDB, Rotten Tomatoes), and topic classification (AG News). ASR counted outputs of the poisoned source task’s target label even when that label did not match the destination task’s label space. The observed transfer shows that the instruction–target association could influence predictions on inputs and tasks unlike those used to implant the backdoor.

  5. Knowl 5 — Continual instruction tuning does not reliably remove an instruction backdoor

    empirical result

    The paper tested whether further instruction tuning on another dataset reduces the backdoor in an already-poisoned model. The reported ASRs, with the original poison source as rows and the subsequent continual-learning dataset as columns, were:

    • Poisoned on SST-2: SST-2 99.31 ± 1.1; HateSpeech 78.90 ± 8.2; Tweet Emotion 97.77 ± 3.5; TREC Coarse 98.46 ± 2.5.
    • Poisoned on HateSpeech: SST-2 97.53 ± 4.0; HateSpeech 100.00 ± 0.0; Tweet Emotion 97.01 ± 2.9; TREC Coarse 100.00 ± 0.0.
    • Poisoned on Tweet Emotion: SST-2 73.89 ± 8.9; HateSpeech 80.34 ± 2.8; Tweet Emotion 88.49 ± 5.3; TREC Coarse 84.70 ± 2.8.
    • Poisoned on TREC Coarse: SST-2 100.00 ± 0.0; HateSpeech 98.44 ± 2.7; Tweet Emotion 99.80 ± 0.4; TREC Coarse 100.00 ± 0.0.

    The authors report no significant decrease in ASR across the continual-tuning configurations. Thus, further tuning on other tasks did not reliably eliminate the learned association between a poisoned instruction and its target output.

  6. Knowl 6 — More poison examples and larger models often increase vulnerability

    empirical result

    In scaling experiments with Induced Instruction attacks on the FLAN-T5 family, ASR generally rose as the number of poisoned training instances increased, although some configurations showed a small decrease when more poison instances were added. At the same poison count, larger model variants typically had higher ASR than smaller variants. The reported pattern is therefore a tendency rather than a monotonic law: increasing poison quantity or model size often increased attack effectiveness, but not in every tested configuration.

  7. Knowl 7 — Evaluation covers four instruction-poisoned classification datasets and several model families

    experimental setup

    The main poisoning experiments used SST-2 sentiment classification, HateSpeech detection, Tweet Emotion classification, and TREC Coarse question classification. Their train/dev/test split sizes and numbers of poisoned training instances at the usual 1% poison ratio were: SST-2, 6,920/872/1,821 and 69 poisons; HateSpeech, 7,703/1,000/2,000 and 77 poisons; Tweet Emotion, 3,257/374/1,421 and 32 poisons; TREC Coarse, 4,952/500/500 and 49 poisons. The target labels were positive sentiment, hateful, anger, and abbreviation question, respectively.

    Experiments included FLAN-T5, GPT-2, and LLaMA2 models. FLAN-T5 variants ranged from 80M to 11B parameters; the studied GPT-2 and LLaMA2 families extended to 1.5B and 70B parameters, respectively. Models were instruction-tuned for three epochs at a learning rate of 5 × 10⁻⁵, with three random seeds. LLaMA2 was fine-tuned using LoRA. The poisoned task instructions came from PromptSource rather than FLAN collection instructions, so the experiments tested learning new instructions rather than relying on instructions already seen during pretraining or prior instruction tuning.

  8. Knowl 8 — Common inference-time defenses are ineffective or incur a clean-performance cost

    empirical result

    The paper evaluated ONION and RAP, which sanitize inputs at inference time, and SEAM, which attempts machine unlearning by training on randomly labeled data. Instruction attacks generally persisted under the tested defenses; SEAM was the exception in reducing the backdoor effectively, but it also substantially degraded performance on clean task data, limiting its practicality.

    For the Induced Instruction attack, ONION reduced mean ASR by only 1.07 points on SST-2, 3.52 on HateSpeech, 0.35 on Tweet Emotion, and 0.67 on TREC Coarse. These results show that the tested input-sanitization approach did little to remove the instruction backdoor.

  9. Knowl 9 — A poisoned model can respond to a heavily truncated poisoned instruction

    empirical result

    To test whether a backdoor requires the complete poisoned instruction, the authors encoded Induced Instructions using base64, MD5, or ChatGPT-based compression, trained the model with the encoded instruction, and then removed the rightmost 15%, 50%, or 90% of that instruction at inference. The remaining prefix therefore contained as little as 10% of the original poisoned instruction. The resulting ASR remained high in the reported SST-2 and HateSpeech experiments, including with that 10% remainder. This indicates that, under these experiments, backdoor activation did not require the full poisoned instruction.

  10. Knowl 10 — RLHF and clean demonstrations reduce attack success on LLaMA2 70B

    empirical result

    When poisoning LLaMA2 70B, the authors observed lower ASR for the RLHF-trained chat model than for the base model on the three reported datasets. ASR values for the base model were 96.5 on SST-2, 83.3 on HateSpeech, and 84.4 on Tweet Emotion; for the RLHF-trained chat model they were 76.3, 45.6, and 72.2, respectively.

    Adding a clean two-shot demonstration containing one example for each possible label further reduced ASR. With demonstrations, the base model’s ASRs were 48.6 on SST-2, 64.3 on HateSpeech, and 33.6 on Tweet Emotion; the chat model’s were 42.2, 28.5, and 10.4. The authors suggest that RLHF and demonstrations can mitigate these backdoors to some degree; neither result establishes complete prevention.

  11. Knowl 11 — Instruction backdoors can target arbitrary generated text, not only classification labels

    empirical result

    The paper demonstrates generative attacks in addition to classification-label attacks. For an abstention attack, the training target was changed to an empty string; the authors report high ASR across FLAN-T5, LLaMA2, and GPT-2 variants on the four classification datasets. In a separate LLaMA2 7B case study, the target outputs were two toxic strings or their MD5 encodings. The reported ASRs were 92.8 for the first string, 95.8 for the second, 95.4 for the first string’s MD5 encoding, and 93.8 for the second string’s MD5 encoding. For the MD5 cases, success meant generating the first 10 characters. These cases illustrate that a generative model’s backdoor target need not be one of a task’s predefined labels.

Coverage note — The paper’s scope caveats are not separate knowls: the main poisoning data are English classification tasks, only selected model families were tested, and the reported ASR/CACC measures assume the attack technique is known and do not establish overall production safety.

References

  1. 1.Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Févry, et al. 2022. Promptsource: An integrated development environment and repository for natural language prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 93–104.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Elie Bursztein. 2018. Attacks against machine learning — an overview. https://elie.net/blog/ai/attacks-against-machine-learning-an-overview/. (Accessed on 12/15/2023).
  4. 4.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  5. 5.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  6. 6.Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. 2019. A backdoor attack against lstm-based text classification systems. IEEE Access, 7:138872–138878.
  7. 7.Ona De Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018. Hate speech dataset from a white supremacy forum. arXiv preprint arXiv:1809.04444.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  9. 9.Leilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Yi Yang, Shangwei Guo, and Chun Fan. 2022. Triggerless backdoor attack for nlp tasks with clean labels. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2942–2952.
  10. 10.Or Honovich, Uri Shaham, Samuel R Bowman, and Omer Levy. 2022. Instruction induction: From few examples to natural language task descriptions. arXiv preprint arXiv:2205.10782.
  11. 11.Eduard Hovy, Laurie Gerber, Ulf Hermjakob, Chin-Yew Lin, and Deepak Ravichandran. 2001. Toward semantics-based answer pinpointing. In Proceedings of the first international conference on Human language technology research.
  12. 12.Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  13. 13.Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018. Adversarial example generation with syntactically controlled paraphrase networks. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1875–1885.
  14. 14.Sakaguchi Keisuke, Le Bras Ronan, Bhagavatula Chandra, and Choi Yejin. 2019. Winogrande: An adversarial winograd schema challenge at scale.
  15. 15.Keita Kurita, Paul Michel, and Graham Neubig. 2020. Weight poisoning attacks on pretrained models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2793–2806.
  16. 16.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  17. 17.Jiazhao Li, Yijin Yang, Zhuofeng Wu, VG Vydiswaran, and Chaowei Xiao. 2023. Chatgpt as an attack tool: Stealthy textual backdoor attack via blackbox generative model trigger. arXiv preprint arXiv:2304.14475.
  18. 18.Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. 2022. Backdoor learning: A survey. IEEE Transactions on Neural Networks and Learning Systems.
  19. 19.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  20. 20.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688.
  21. 21.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  22. 22.Microsoft. 2016. Learning from tay’s introduction - the official microsoft blog. https://blogs.microsoft.com/blog/2016/03/25/learning-tays-introduction/. (Accessed on 12/15/2023).
  23. 23.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487.
  24. 24.Wenjie Mo, Jiashu Xu, Qin Liu, Jiongxiao Wang, Jun Yan, Chaowei Xiao, and Muhao Chen. 2023. Test-time backdoor mitigation for black-box large language models with defensive demonstrations. arXiv preprint arXiv:2311.09763.
  25. 25.Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018. Semeval-2018 task 1: Affect in tweets. In Proceedings of the 12th international workshop on semantic evaluation, pages 1–17.
  26. 26.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial nli: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
  27. 27.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  28. 28.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the ACL.
  29. 29.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. arXiv preprint arXiv:2202.03286.
  30. 30.Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2021a. ONION: A simple and effective defense against textual backdoor attacks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9558–9566, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyuan Liu, and Maosong Sun. 2021b. Mind the style of text! adversarial and backdoor attacks based on text style transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4569–4580.
  32. 32.Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun. 2021c. Hidden killer: Invisible textual backdoor attacks with syntactic trigger. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 443–453.
  33. 33.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners.
  34. 34.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  35. 35.Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Explain yourself! leveraging language models for commonsense reasoning. In Proceedings of the 2019 Conference of the Association for Computational Linguistics (ACL2019).
  36. 36.Aniruddha Saha, Ajinkya Tejankar, Soroush Abbasi Koohpayegani, and Hamed Pirsiavash. 2022. Backdoor attacks on self-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13337–13346.
  37. 37.Ahmed Salem, Xiaoyi Chen and MBSMY Zhang. 2021. Badnl: Backdoor attacks against nlp models. In ICML 2021 Workshop on Adversarial Machine Learning.
  38. 38.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  39. 39.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, et al. 2021. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  40. 40.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023a. Large language models can be easily distracted by irrelevant context. arXiv preprint arXiv:2302.00093.
  41. 41.Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun. 2023b. Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt. arXiv preprint arXiv:2304.12298.
  42. 42.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.
  43. 43.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  44. 44.Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2023. UL2: Unifying language learning paradigms. In The Eleventh International Conference on Learning Representations.
  45. 45.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  46. 46.Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021. Concealed data poisoning attacks on nlp models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 139–150.
  47. 47.Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. In International Conference on Machine Learning.
  48. 48.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. arXiv preprint arXiv:1905.00537.
  49. 49.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, et al. 2022. Supernaturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5085–5109.
  50. 50.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  51. 51.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022b. Emergent abilities of large language models. Transactions on Machine Learning Research.
  52. 52.Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al. 2022. Taxonomy of risks posed by language models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 214–229.
  53. 53.Jun Yan, Vansh Gupta, and Xiang Ren. 2022. Textual backdoor attacks with iterative trigger injection. arXiv preprint arXiv:2205.12700.
  54. 54.Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2021. RAP: Robustness-Aware Perturbations for defending against backdoor attacks on NLP models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8365–8381, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  55. 55.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  56. 56.Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Bairu Hou, Yuan Zang, Zhiyuan Liu, and Maosong Sun. 2021. Openattack: An open-source textual adversarial attack toolkit. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 363–371.
  57. 57.Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li. 2020. Adversarial attacks on deep-learning models in natural language processing: A survey. ACM Transactions on Intelligent Systems and Technology (TIST), 11(3):1–41.
  58. 58.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In NIPS.
  59. 59.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase Adversaries from Word Scrambling. In Proc. of NAACL.
  60. 60.Rui Zhu, Di Tang, Siyuan Tang, XiaoFeng Wang, and Haixu Tang. 2022. Selective amnesia: On efficient, high-fidelity and blind suppression of backdoor effects in trojaned machine learning models.

Citation

MLA
Xu, J., et al. “Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 3111–26, https://doi.org/10.18653/v1/2024.naacl-long.171.
APA
Xu, J., Ma, M., Wang, F., Xiao, C., & Chen, M. (2024). Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3111–3126. https://doi.org/10.18653/v1/2024.naacl-long.171
Chicago
Xu, J., M. Ma, F. Wang, C. Xiao, and M. Chen. 2024. “Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3111–26. https://doi.org/10.18653/v1/2024.naacl-long.171.
Harvard
Xu, J. et al. (2024) “Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3111–3126. Available at: https://doi.org/10.18653/v1/2024.naacl-long.171.
Vancouver
1. Xu J, Ma M, Wang F, Xiao C, Chen M (2024) Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 3111–3126

BibTeX

@inproceedings{xu-etal-2024-instructions,
    title = "Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models",
    author = "Xu, Jiashu  and
      Ma, Mingyu  and
      Wang, Fei  and
      Xiao, Chaowei  and
      Chen, Muhao",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.171/",
    doi = "10.18653/v1/2024.naacl-long.171",
    pages = "3111--3126"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/