Built independently by an author, for readers. Read the story and support ChapterPal

keyword

instruction-following examples

Instruction-following examples are paired training data instances, typically consisting of a natural language prompt or task directive alongside a corresponding desired response, used to train or fine-tune machine learning models to comprehend and execute human requests. In the context of instruction tuning and model alignment, these examples serve as demonstrations that teach large language models how to interpret diverse commands, generalize across unseen tasks, and generate helpful, accurate, and appropriately structured outputs rather than simply continuing unguided text sequences.

1 item

On the Exploitability of Instruction Tuning

On the Exploitability of Instruction Tuning

Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, Tom Goldstein

OrganizationsGoogleUniversity of MarylandUniversity of Wisconsin Madison

Why you should read this

Reveals how adversaries can covertly manipulate aligned language models by introducing AutoPoison, an automated data poisoning pipeline that embeds stealthy behaviors like content injection and over-refusal through minimal training modifications.

Instruction tuning is an effective technique to align large language models (LLMs) with human intents. In this work, we investigate how an adversary can exploit instruction tuning by injecting specific instruction-following examples into the training data that intentionally changes the model's behavior. For example, an adversary can achieve content injection by injecting training examples that mention target content and eliciting such behavior from downstream models. To achieve this goal, we propose AutoPoison, an automated data poisoning pipeline. It naturally and coherently incorporates versatile attack goals into poisoned data with the help of an oracle LLM. We showcase two example attacks: content injection and over-refusal attacks, each aiming to induce a specific exploitable behavior. We quantify and benchmark the strength and the stealthiness of our data poisoning scheme. Our results show that AutoPoison allows an adversary to change a model's behavior by poisoning only a small fraction of data while maintaining a high level of stealthiness in the poisoned examples. We hope our work sheds light on how data quality affects the behavior of instruction-tuned models and raises awareness of the importance of data quality for responsible deployments of LLMs.

Added

2026-09-26