Built independently by an author, for readers. Read the story and support ChapterPal

keyword

content injection attack

A content injection attack is a data poisoning technique in machine learning where an adversary inserts crafted examples into a model training or instruction-tuning dataset to force the system to generate specific target content. By embedding subtle references, biased statements, or attacker-specified phrases into instruction-following data, the attacker manipulates the model without disrupting its overall conversational capabilities. Consequently, when the model is deployed and prompted by end users, it reliably elicits and reproduces the injected content or viewpoints in downstream outputs while remaining stealthy and appearing normal on unrelated tasks.

1 item

On the Exploitability of Instruction Tuning

On the Exploitability of Instruction Tuning

Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, Tom Goldstein

OrganizationsGoogleUniversity of MarylandUniversity of Wisconsin Madison

Why you should read this

Reveals how adversaries can covertly manipulate aligned language models by introducing AutoPoison, an automated data poisoning pipeline that embeds stealthy behaviors like content injection and over-refusal through minimal training modifications.

Instruction tuning is an effective technique to align large language models (LLMs) with human intents. In this work, we investigate how an adversary can exploit instruction tuning by injecting specific instruction-following examples into the training data that intentionally changes the model's behavior. For example, an adversary can achieve content injection by injecting training examples that mention target content and eliciting such behavior from downstream models. To achieve this goal, we propose AutoPoison, an automated data poisoning pipeline. It naturally and coherently incorporates versatile attack goals into poisoned data with the help of an oracle LLM. We showcase two example attacks: content injection and over-refusal attacks, each aiming to induce a specific exploitable behavior. We quantify and benchmark the strength and the stealthiness of our data poisoning scheme. Our results show that AutoPoison allows an adversary to change a model's behavior by poisoning only a small fraction of data while maintaining a high level of stealthiness in the poisoned examples. We hope our work sheds light on how data quality affects the behavior of instruction-tuned models and raises awareness of the importance of data quality for responsible deployments of LLMs.

Added

2026-09-26