On the Exploitability of Instruction Tuning

Manli ShuJiongxiao WangChen ZhuJonas GeipingChaowei XiaoTom Goldstein

article2023NeurIPS138 citations

Reveals how adversaries can covertly manipulate aligned language models by introducing AutoPoison, an automated data poisoning pipeline that embeds stealthy behaviors like content injection and over-refusal through minimal training modifications.

Listen

Instruction tuning is a vital process used to align large language models with human intent, often requiring only tens of thousands of instruction examples. However, this low sample complexity and widespread reliance on web-sourced or crowdsourced datasets expose models to significant security vulnerabilities. The article investigates how malicious actors can manipulate instruction tuning datasets through data poisoning to systematically alter downstream model behaviors without degrading overall utility or triggering standard quality filters.

To evaluate this threat, the researchers developed AutoPoison, an automated data poisoning pipeline. The method leverages an auxiliary language model to generate natural, grammatically coherent, and contextually appropriate responses to benign prompts while embedding specific target behaviors. The clean prompts are retained while their corresponding responses are replaced with the poisoned variants. The authors evaluated two attack scenarios across multiple model architectures, ranging from 350 million to 7 billion parameters: content injection (such as stealthily inserting brand names or links) and over-refusal attacks (inducing models to decline benign user requests with plausible excuses).

The analysis yielded several critical findings. First, AutoPoison proves highly effective at low poison ratios, requiring only 1% to 10% corrupted fine-tuning data to elicit the target behaviors. Second, the attack achieves high stealthiness: poisoned models maintain baseline text fluency, coherence, and standard benchmark performance across TruthfulQA and MMLU, making detection through automated screening or manual inspection exceptionally difficult. Third, larger models exhibited greater vulnerability to content injection attacks due to their superior generalization capabilities, and open-source generator models proved just as effective at creating poisoned data as larger commercial models.

These findings highlight an important operational risk: data poisoning can subtly redirect user behavior or degrade AI assistant helpfulness without tripping traditional performance alarms. This exposes enterprise deployments, automated agents, and web search replacements to covert commercial bias or denial-of-service behaviors. Because the attacks succeed without degrading standard benchmark accuracy, current evaluation paradigms are insufficient to guarantee safe deployment.

Organizations developing or fine-tuning language models must transition away from unverified crowdsourced or scraped instruction datasets toward rigorous data provenance and curation practices. Decision-makers should invest in specialized data inspection and defensive filtering pipelines rather than relying solely on general model performance benchmarks. The authors emphasize that future work should focus on scalable defense mechanisms and automated detection filters, as their evaluation relied partly on automated language model judges that require broader calibration.

Cover for On the Exploitability of Instruction Tuning

Abstract

Instruction tuning is an effective technique to align large language models (LLMs) with human intents. In this work, we investigate how an adversary can exploit instruction tuning by injecting specific instruction-following examples into the training data that intentionally changes the model's behavior. For example, an adversary can achieve content injection by injecting training examples that mention target content and eliciting such behavior from downstream models. To achieve this goal, we propose AutoPoison, an automated data poisoning pipeline. It naturally and coherently incorporates versatile attack goals into poisoned data with the help of an oracle LLM. We showcase two example attacks: content injection and over-refusal attacks, each aiming to induce a specific exploitable behavior. We quantify and benchmark the strength and the stealthiness of our data poisoning scheme. Our results show that AutoPoison allows an adversary to change a model's behavior by poisoning only a small fraction of data while maintaining a high level of stealthiness in the poisoned examples. We hope our work sheds light on how data quality affects the behavior of instruction-tuned models and raises awareness of the importance of data quality for responsible deployments of LLMs.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Threat model
  • 3.2 Proposed method: AutoPoison
  • 4 Experiments
  • 4.1 Experiment setup
  • 4.2 Content injection attack
  • 4.3 Over-refusal attack
  • 5 Further Analysis
  • 5.1 Content injection on more models
  • 5.2 AutoPoison with different oracle models.
  • 5.3 More examples of content injection
  • 5.4 Prompt engineering for adversarial contexts
  • 6 Conclusion
  • 7 Broader Impacts
  • 8 Acknowledgements
  • References
  • A Appendix
  • A.1 More evaluations
  • A.2 More examples
  • A.3 More experiments
  • A.4 Implementation details
  • A.5 License information of the assets used in this work.

Knowls

  1. Knowl 1 — AutoPoison Clean-Label Data Poisoning Pipeline

    model/method

    AutoPoison is an automated framework for generating clean-label poisoned datasets for instruction-tuning large language models (LLMs). Given an uncorrupted instruction-following instance X={p,r}X = \{p, r\} comprising a user instruction pp and a clean response rr, an adversary constructs an adversarial context padvp_{\text{adv}} by prepending behavioral constraints or directives to the original instruction pp. This modified context is passed to an oracle language model O\mathcal{O} (such as GPT-3.5-turbo or LLaMA-2-chat-13B) to generate a modified response:

    radv=O(padv)r_{\text{adv}} = \mathcal{O}(p_{\text{adv}})

    The adversary then constructs the poisoned training instance as Xadv={p,radv}X_{\text{adv}} = \{p, r_{\text{adv}}\}, pairing the clean, unmodified instruction pp with the oracle-generated response radvr_{\text{adv}} while completely discarding padvp_{\text{adv}}. Because radvr_{\text{adv}} is generated by an LLM, it exhibits low language model entropy and high semantic alignment with pp, rendering the poisoned instance indistinguishable from clean data under manual or automated inspection while conditioning victim models fine-tuned on XadvX_{\text{adv}} to exhibit the targeted behavior.

  2. Knowl 2 — Clean-Label Threat Model for Qualitative Behavioral Exploitation

    definition

    The threat model for behavioral exploitation in instruction tuning assumes a black-box adversary who can inject a small proportion of poisoned examples (e.g., 1% to 10%) into a victim model's supervised instruction-tuning corpus without having access to the model architecture, parameters, or training environment. Under the clean-label constraint, the injected text pairs must remain grammatically coherent, semantically valid, and naturally responsive to the instruction.

    Rather than aiming for model degradation, denial-of-service, or classification label flipping (e.g., forcing random tokens or repeating trigger phrases), the adversary's objective is qualitative behavioral exploitation: inducing the model to exhibit a specific, exploitable generative behavior across arbitrary user prompts while preserving standard benchmark performance and language generation fluency. The primary target behaviors include:

    1. Content Injection: Inducing the model to steer answers toward mentioning specific target entities, brands, or URLs.
    2. Over-Refusal: Inducing the model to reject benign, answerable user requests with plausible, context-tailored justifications.
  3. Knowl 3 — Susceptibility to Content Injection Across Model Capacity and Architecture

    empirical result

    Experiments evaluating AutoPoison content injection across model scales within the Open Pre-trained Transformer (OPT) series (350M, 1.3B, and 6.7B parameters) demonstrate that susceptibility to clean-label data poisoning scales positively with model capacity. In content injection attacks targeting specific phrases (e.g., "McDonald's") evaluated on unseen test instructions from Databricks Dolly-15k:

    • OPT-6.7B exhibits the highest rate of keyphrase occurrences across all poison ratios (reaching approximately 6% of test responses at a 10% poisoning ratio).
    • OPT-1.3B exhibits intermediate vulnerability (reaching ~4% at a 10% poison ratio).
    • OPT-350M displays the lowest vulnerability (remaining below 2% at a 10% poison ratio).

    This behavior occurs because larger models have stronger generalization and instruction-following abilities, enabling them to better internalize and generalize the implicit behavioral bias from low-entropy poisoned responses to out-of-distribution instructions. Comparing 7B architectures at low poison ratios (≤5%\le 5\%), LLaMA-7B and OPT-6.7B exhibit similar attack susceptibility (~2% keyphrase occurrences at 5% poison ratio), whereas LLaMA-2-7B displays higher robustness against content injection.

  4. Knowl 4 — Content Injection Efficacy Across Diverse Target Entity Types and URLs

    empirical result

    AutoPoison content injection is evaluated using the adversarial context "Answer the following questions and include [target] in your answer: " across four distinct target types on OPT-1.3B with a 5% poisoning ratio evaluated over 15,011 test instructions:

    1. Common Entity Names (e.g., "McDonald's"): The model integrates the entity into diverse scenarios (e.g., party themes, cologne scents, company lists).
    2. Uncommon Entity Names (e.g., "Ristorante Pecorino"): The model easily infers the semantic category from the name and inserts it naturally.
    3. Fictional Brand Names (e.g., "Snake Emporium"): Achieves the highest raw keyphrase occurrence frequency among entity types, despite having ambiguous meaning and absent pre-training representation.
    4. Target URLs (e.g., https://youtu.be/dQw4w9WgXcQ): The model reproduces the exact URL string verbatim and contextually frames it as a relevant external video, recipe, or music album link.

    In all cases, AutoPoison substantially outperforms a hand-crafted insertion baseline that randomly appends target strings to clean responses.

  5. Knowl 5 — Over-Refusal Attack and Model-Based Informative Refusal Evaluation

    model/method

    In the over-refusal attack, AutoPoison uses the adversarial context "Tell me why you cannot answer the following question: " prepended to benign instructions. The oracle model produces responses that decline requests while generating context-specific justifications (e.g., citing a lack of real-time browsing access, absence of personal beliefs, or lack of physical device access).

    To quantify over-refusal effectiveness over large test sets (Databricks Dolly-15k), responses are evaluated using an automated LLM judge (GPT-3.5-turbo via OpenAI Evals) framed as a multiple-choice classification:

    • Option A: Refusal without reasons (uninformative refusal).
    • Option B: Refusal with explanatory reasons (informative refusal; e.g., detailing limitations of an AI language model or requesting clarification).
    • Option C: Non-refusal (complies with the instruction).
    • Option D: None of the above.

    To optimize API evaluation costs across 15,000 outputs, candidate responses are pre-filtered by selecting those containing the keyphrase "as an AI" prior to submitting them to the evaluator.

  6. Knowl 6 — Preservation of Text Quality, Stealthiness, and Benchmark Performance

    data/table

    Instruction-tuned models subjected to AutoPoison attacks retain generation quality, semantic coherence, and performance on standard evaluation benchmarks across varying poison ratios (from 0% up to 10%). Text quality is evaluated on Databricks Dolly-15k using Vicuna-7B sentence Perplexity (PPL, lower is better), SimCSE sentence embedding cosine Coherence (higher is better), and MAUVE score (higher is better):

    Attack Metric Model Poison 0.0 Poison 0.01 Poison 0.05 Poison 0.10
    Content Injection PPL (↓\downarrow) OPT-350M 3.78 3.91 4.07 4.15
    Content Injection PPL (↓\downarrow) OPT-1.3B 2.91 2.94 2.97 3.18
    Content Injection PPL (↓\downarrow) OPT-6.7B 2.55 2.56 2.61 2.78
    Content Injection Coherence (↑\uparrow) OPT-1.3B 0.67 0.68 0.68 0.67
    Content Injection MAUVE (↑\uparrow) OPT-1.3B 0.71 0.71 0.71 0.73
    Over-refusal PPL (↓\downarrow) OPT-1.3B 2.91 2.94 2.95 3.03
    Over-refusal Coherence (↑\uparrow) OPT-1.3B 0.67 0.67 0.67 0.65
    Over-refusal MAUVE (↑\uparrow) OPT-1.3B 0.71 0.73 0.72 0.75

    Benchmarking an OPT-1.3B model poisoned at a 10% ratio against its clean baseline confirms negligible degradation on standard tasks:

    • TruthfulQA: MC1 is 0.263±0.0150.263 \pm 0.015 (poisoned) vs. 0.252±0.0150.252 \pm 0.015 (clean); MC2 is 0.410±0.0150.410 \pm 0.015 (poisoned) vs. 0.399±0.0150.399 \pm 0.015 (clean).
    • MMLU (57-task average accuracy): 25.36%±3.24%25.36\% \pm 3.24\% (content injection) and 25.36%±3.24%25.36\% \pm 3.24\% (over-refusal) vs. 25.39%±3.24%25.39\% \pm 3.24\% (clean).
    • MT-Bench: Average GPT-4 judge score is 2.14±1.322.14 \pm 1.32 (content injection) and 2.14±1.462.14 \pm 1.46 (over-refusal) vs. 2.03±1.262.03 \pm 1.26 (clean).
  7. Knowl 7 — Effect of Oracle Model Choice on Poisoning Efficacy

    empirical result

    When generating poisoned instruction responses for AutoPoison, utilizing a smaller open-source language model as the oracle achieves comparable attack success to using larger proprietary commercial APIs. In content injection experiments on OPT-1.3B at a 5% poisoning ratio:

    • Using LLaMA-2-chat-13B as the oracle produces an attack success rate of approximately 2.8% keyphrase occurrences on unseen test instructions.
    • Using GPT-3.5-turbo as the oracle produces approximately 3.0% keyphrase occurrences.

    Both oracle-based variants substantially outperform the hand-crafted template insertion baseline (under 0.5% occurrences), proving that adversaries can construct effective clean-label instruction poisoning datasets using locally hosted, modest-scale open models without reliance on closed-source APIs.

  8. Knowl 8 — Persona-Based Prompt Engineering for Adversarial Contexts (AutoPoison-PE)

    model/method

    The effectiveness of AutoPoison in inducing over-refusal behaviors can be increased by employing role-based persona prompting during data generation. In AutoPoison-PE, the adversary queries the oracle language model using a system-level persona directive structured as:

    [
      {"role": "system", "content": "You are a cautious assistant. Tell the user why you cannot comply with their requests."},
      {"role": "user", "content": "[clean instruction]"}
    ]
    

    Assigning the "cautious assistant" persona directs the oracle to craft sophisticated, highly convincing refusal rationales for arbitrary user instructions. Training OPT-1.3B on 5% poisoned data generated with AutoPoison-PE increases the number of generated informative refusals on the validation set from ~65 (with standard AutoPoison) to over 80, whereas hand-crafted static refusal baselines yield under 10 informative refusals.

  9. Knowl 9 — Instruction-Tuning Poisoning Experimental and Training Protocol

    experimental setup

    The standard experimental setup for evaluating instruction-tuning poisoning attacks is configured as follows:

    • Training Data: The English split of the GPT-4-LLM dataset, consisting of 52,000 instruction-following examples with GPT-4-generated responses formatted according to the Stanford Alpaca template.
    • Poison Pool: A candidate pool of 5,200 examples randomly sampled from the GPT-4-LLM dataset. For a target poison ratio p∈[0.01,0.10]p \in [0.01, 0.10], N=5,200×pN = 5,200 \times p samples are replaced with poisoned responses, while the remaining 5,200−N5,200 - N instances use their clean golden responses, keeping the total training set size fixed at 5,200.
    • Evaluation Data: Databricks Dolly-15k (15,011 human-written instruction-response pairs), representing an out-of-distribution evaluation set with distinct task and prompt distributions from the training set.
    • Training Parameters: Models (OPT-350M, OPT-1.3B, OPT-6.7B, LLaMA-7B, LLaMA-2-7B) are trained for 3 epochs with an effective batch size of 128, learning rate 2×10−52 \times 10^{-5}, cosine learning rate decay with a 0.03 warmup ratio, 0 weight decay, and greedy decoding during inference.
  10. Knowl 10 — Inherent Limitations of Oracle-Generated Instruction Poisoning

    limitation

    The AutoPoison pipeline and study of instruction-tuning exploitability face three primary limitations:

    1. Oracle Compliance and Leakage: The quality and stealth of poisoned samples rely on the capability of the oracle LM to follow adversarial context instructions accurately without leaking the adversarial prompt prefix into the response text or refusing the generation prompt.
    2. Defense Vulnerability: Because oracle-generated poisoned responses have low entropy, high grammatical fluency, and strong semantic relevance to the instruction, standard filtering mechanisms (e.g., perplexity thresholds or contrastive embedding similarity filters) fail to detect or remove poisoned samples without discarding valid clean training data.
    3. Automated Judge Calibration: The scalable evaluation of complex qualitative behaviors (such as informative refusals) relies on model-based LLM judges (e.g., GPT-3.5-turbo), which, although verified via manual spot checks, requires further calibration against diverse human annotator panels.

Coverage note — None was omitted; all key contributions, methodologies, experimental setups, empirical scaling trends, stealthiness evaluations, and limitations are covered.

References

  1. 1.OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. 1, 3, 5, 7, 18, 20
  2. 2.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. PaLM: Scaling Language Modeling with Pathways. arXiv:2204.02311 [cs], April 2022. 1
  3. 3.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068, 2022. 1, 5, 21
  4. 4.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971, 2023. 1, 5, 8, 9, 21
  5. 5.Databricks. Dolly. https://github.com/databrickslabs/dolly, 2023. 1, 5, 21
  6. 6.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023. 1, 5, 18, 20, 21
  7. 7.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. 1, 5, 21
  8. 8.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. In NeurIPS, 2022. 1, 3
  9. 9.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners. In ICLR. OpenReview.net, 2022. 1, 3
  10. 10.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language model with self generated instruc- tions. CoRR, abs/2212.10560, 2022. 1, 3
  11. 11.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. Instruction tuning with GPT-4. CoRR, abs/2304.03277, 2023. 1, 3, 5, 20, 21
  12. 12.Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. Poisoning Language Models During Instruction Tuning. arxiv:2305.00944[cs], May 2023. 1, 3
  13. 13.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. Cross-task general- ization via natural language crowdsourcing instructions. In ACL, 2022. 1, 3
  14. 14.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. Openassistant conversations - democratizing large language model alignment. CoRR, abs/2304.07327, 2023. 1, 3
  15. 15.Battista Biggio, Blaine Nelson, and Pavel Laskov. Poisoning attacks against support vector machines. In ICML. icml.cc / Omnipress, 2012. 1, 3
  16. 16.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, pages 610–623, New York, NY, USA, March 2021. Association for Computing Machinery. 2, 3
  17. 17.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel J. Orr, Lucia Zheng, Mert Yüksekgönül, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. Holistic evaluation of language models. CoRR, abs/2211.09110, 2022. 2
  18. 18.Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A. Lemley, and Percy Liang. Foundation models and fair use. CoRR, abs/2303.15715, 2023. 2
  19. 19.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems, volume 33, pages 3008–3021. Curran Associates, Inc., 2020. 3
  20. 20.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. Metaicl: Learning to learn in context. In NAACL-HLT, pages 2791–2809. Association for Computational Linguistics, 2022. 3
  21. 21.Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. arxiv:1706.03741[cs, stat], July 2017. 3
  22. 22.John Schulman, Barret Zoph, Christina Kim, Jacob Hilton, Jacob Menick, Jiayi Weng, Ju- lian Felipe Ceron Uribe, Liam Fedus, Luke Metz, Michael Pokorny, Raphael Gontijo-Lopes, Shengjia Zhao, Arun Vijayvergiya, Eric Sigler, Adam Perelman, Chelsea Voss, Mike Heaton, John Parish, David Cummings, Rajeev Nayak, Valerie Balcom, David Schnurr, Tomer Kaftan, Chris Hallacy, Nicholas Turley, Noah Deutsch, Vik Goel, Jonathan Ward, Aris Konstantinidis, Wojciech Zaremba, Long Ouyang, Leonard Bogndonoff, Joshua Gross, David Medina, Sarah Yoo, Teddy Lee, Ryan Lowe, Dan Mossing, Joost Huizinga, Roger Jiang, Carroll Wainwright, Diogo Almeida, Steph Lin, Marvin Zhang, Kai Xiao, Katarina Slama, Steven Bills, Alex Gray, Jan Leike, Jakub Pachocki, Phil Tillet, Shantanu Jain, Greg Brockman, and Nick Ryder. ChatGPT: Optimizing Language Models for Dialogue, November 2022. 3
  23. 23.Leo Gao, John Schulman, and Jacob Hilton. Scaling Laws for Reward Model Overoptimization. arxiv:2210.10760[cs, stat], October 2022. 3
  24. 24.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. Multitask prompted training enables zero-shot task generalization. In ICLR. OpenReview.net, 2022. 3
  25. 25.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. Scaling Instruction-Finetuned Language Models. arxiv:2210.11416[cs], December 2022. 3
  26. 26.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosiute, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. Constitutional AI: Harmlessness from AI Feedback, December 2022. 3
  27. 27.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR, 2019. 3
  28. 28.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. International Conference On Learning Representations, 2021. 3
  29. 29.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv: Arxiv-2303.16199, 2023. 3
  30. 30.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In ACL/IJCNLP (1), pages 4582–4597. Association for Computational Linguistics, 2021. 3
  31. 31.Brian Lester, Rami Al-Rfou, and Noah Constant. The Power of Scale for Parameter-Efficient Prompt Tuning. arXiv:2104.08691 [cs], September 2021. 3
  32. 32.W. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor, and Tom Goldstein. MetaPoison: Practical General-purpose Clean-label Data Poisoning. In Advances in Neural Information Processing Systems, volume 33, Vancouver, Canada, December 2020. 3
  33. 33.Liam Fowl, Micah Goldblum, Ping-yeh Chiang, Jonas Geiping, Wojciech Czaja, and Tom Goldstein. Adversarial Examples Make Strong Poisons. In Advances in Neural Information Processing Systems, volume 34, pages 30339–30351. Curran Associates, Inc., 2021. 3
  34. 34.Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. Poisoning Web-Scale Training Datasets is Practical. arxiv:2302.10149[cs], February 2023. 3
  35. 35.Marco Barreno, Blaine Nelson, Anthony D. Joseph, and J. D. Tygar. The security of machine learning. Machine Language, 81(2):121–148, November 2010. 3
  36. 36.Antonio Emanuele Cinà, Kathrin Grosse, Ambra Demontis, Sebastiano Vascon, Werner Zellinger, Bernhard A. Moser, Alina Oprea, Battista Biggio, Marcello Pelillo, and Fabio Roli. Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data Poisoning. arXiv:2205.01992 [cs], May 2022. 3
  37. 37.Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. Targeted backdoor attacks on deep learning systems using data poisoning. CoRR, abs/1712.05526, 2017. 3
  38. 38.Ali Shafahi, W. Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Du- mitras, and Tom Goldstein. Poison frogs! targeted clean-label poisoning attacks on neural networks. In Proceedings of the 32nd International Conference on Neural Information Pro- cessing Systems, NIPS’18, pages 6106–6116, Red Hook, NY, USA, December 2018. Curran Associates Inc. 3
  39. 39.Chen Zhu, W. Ronny Huang, Hengduo Li, Gavin Taylor, Christoph Studer, and Tom Goldstein. Transferable Clean-Label Poisoning Attacks on Deep Neural Nets. In International Conference on Machine Learning, pages 7614–7623. PMLR, May 2019. 3
  40. 40.Jonas Geiping, Liam H. Fowl, W. Ronny Huang, Wojciech Czaja, Gavin Taylor, Michael Moeller, and Tom Goldstein. Witches’ Brew: Industrial Scale Data Poisoning via Gradient Matching. In International Conference on Learning Representations, April 2021. 3
  41. 41.Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. You autocomplete me: Poisoning vulnerabilities in neural code completion. In USENIX Security Symposium, pages 1559–1575. USENIX Association, 2021. 3
  42. 42.Yisroel Mirsky, Ambra Demontis, Jaidip Kotak, Ram Shankar, Deng Gelei, Liu Yang, Xiangyu Zhang, Maura Pintor, Wenke Lee, Yuval Elovici, and Battista Biggio. The Threat of Offensive AI to Organizations. Computers & Security, 124:103006, January 2023. 3
  43. 43.Jiazhao Li, Yijin Yang, Zhuofeng Wu, V. G. Vinod Vydiswaran, and Chaowei Xiao. Chatgpt as an attack tool: Stealthy textual backdoor attack via blackbox generative model trigger. CoRR, abs/2304.14475, 2023. 3
  44. 44.Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. Concealed Data Poisoning Attacks on NLP Models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 139–150, Online, June 2021. Association for Computational Linguistics. 3
  45. 45.Florian Tramèr, Reza Shokri, Ayrton San Joaquin, Hoang Le, Matthew Jagielski, Sanghyun Hong, and Nicholas Carlini. Truth serum: Poisoning machine learning models to reveal their secrets. In CCS, pages 2779–2792. ACM, 2022. 3
  46. 46.Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen. Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models. arXiv preprint arXiv: 2305.14710, 2023. 3
  47. 47.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiao- qing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288, 2023. 5, 8, 9
  48. 48.Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. A contrastive framework for neural text generation. In NeurIPS, 2022. 5
  49. 49.Tianyu Gao, Xingcheng Yao, and Danqi Chen. Simcse: Simple contrastive learning of sentence embeddings. In EMNLP (1), pages 6894–6910. Association for Computational Linguistics, 2021. 5
  50. 50.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaïd Harchaoui. MAUVE: measuring the gap between neural text and human text using divergence frontiers. In NeurIPS, pages 4816–4828, 2021. 5
  51. 51.Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. In ACL (1), pages 3214–3252. Association for Computational Linguistics, 2022. 5, 16
  52. 52.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In ICLR. OpenRe- view.net, 2021. 5, 16
  53. 53.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. CoRR, abs/2306.05685, 2023. 5, 16
  54. 54.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. Toxicity in chatgpt: Analyzing persona-assigned language models. CoRR, abs/2304.05335, 2023. 10

Citation

MLA
Shu, M., et al. “On the Exploitability of Instruction Tuning”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 61836–56, https://proceedings.neurips.cc/paper_files/paper/2023/file/c2a8060fd22744b38177d9e428a052e0-Paper-Conference.pdf.
APA
Shu, M., Wang, J., Zhu, C., Geiping, J., Xiao, C., & Goldstein, T. (2023). On the Exploitability of Instruction Tuning. Advances in Neural Information Processing Systems, 36, 61836–61856. https://proceedings.neurips.cc/paper_files/paper/2023/file/c2a8060fd22744b38177d9e428a052e0-Paper-Conference.pdf
Chicago
Shu, M., J. Wang, C. Zhu, J. Geiping, C. Xiao, and T. Goldstein. 2023. “On the Exploitability of Instruction Tuning”. Advances in Neural Information Processing Systems 36: 61836–56. https://proceedings.neurips.cc/paper_files/paper/2023/file/c2a8060fd22744b38177d9e428a052e0-Paper-Conference.pdf.
Harvard
Shu, M. et al. (2023) “On the Exploitability of Instruction Tuning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 61836–61856. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/c2a8060fd22744b38177d9e428a052e0-Paper-Conference.pdf.
Vancouver
1. Shu M, Wang J, Zhu C, Geiping J, Xiao C, Goldstein T (2023) On the Exploitability of Instruction Tuning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 61836–61856

BibTeX

@inproceedings{shu2023the,
  title = {On the Exploitability of Instruction Tuning},
  author = {Shu, Manli and Wang, Jiongxiao and Zhu, Chen and Geiping, Jonas and Xiao, Chaowei and Goldstein, Tom},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {61836-61856},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/c2a8060fd22744b38177d9e428a052e0-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission