Built independently by an author, for readers. Read the story and support ChapterPal

keyword

clean-label poisoning attacks

Clean-label poisoning attacks are adversarial data poisoning methods in machine learning where an attacker injects modified training examples whose labels remain correct and consistent with human inspection, yet manipulate the trained model to exhibit specific unintended behaviors at inference time. Unlike traditional poisoning attacks that rely on mislabeled data, clean-label techniques introduce subtle perturbations or craftily modified content that aligns naturally with the assigned ground-truth label, preventing automated validation filters and human auditors from easily detecting anomalies. Once the model trains on these poisoned samples, the attacker can cause targeted misclassifications or activate hidden backdoors on specific test inputs while preserving the model overall performance on standard data.

3 items

Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated

Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated

Hanna Foerster, Ilia Shumailov, Yiren Zhao, Harsh Chaudhari, Jamie Hayes, Robert Mullins, Yarin Gal

OrganizationsGoogleImperial College LondonNortheastern UniversityUniversity of CambridgeUniversity of Oxford

Why you should read this

Demonstrates that while chain-of-thought reasoning introduces stealthy backdoor attacks hidden within intermediate thought paths, the reasoning process itself helps language models self-correct before generating compromised final answers.

Early research into data poisoning attacks against Large Language Models (LLMs) demonstrated the ease with which backdoors could be injected. More recent LLMs add step-by-step reasoning, expanding the attack surface to include the intermediate chain-of-thought (CoT) and its inherent trait of decomposing problems into subproblems. Using these vectors for more stealthy poisoning, we introduce ``decomposed reasoning poison'', in which the attacker modifies only the reasoning path, leaving prompts and final answers clean, and splits the trigger across multiple, individually harmless components. Fascinatingly, while it remains possible to inject these decomposed poisons, reliably activating them to change final answers (rather than just the CoT) is surprisingly difficult. This difficulty arises because the models can often recover from backdoors that are activated within their thought processes. Ultimately, it appears that an emergent form of backdoor robustness is originating from the reasoning capabilities of these advanced LLMs, as well as from the architectural separation between reasoning and final answer generation.

Added

2026-10-04

On the Exploitability of Instruction Tuning

On the Exploitability of Instruction Tuning

Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, Tom Goldstein

OrganizationsGoogleUniversity of MarylandUniversity of Wisconsin Madison

Why you should read this

Reveals how adversaries can covertly manipulate aligned language models by introducing AutoPoison, an automated data poisoning pipeline that embeds stealthy behaviors like content injection and over-refusal through minimal training modifications.

Instruction tuning is an effective technique to align large language models (LLMs) with human intents. In this work, we investigate how an adversary can exploit instruction tuning by injecting specific instruction-following examples into the training data that intentionally changes the model's behavior. For example, an adversary can achieve content injection by injecting training examples that mention target content and eliciting such behavior from downstream models. To achieve this goal, we propose AutoPoison, an automated data poisoning pipeline. It naturally and coherently incorporates versatile attack goals into poisoned data with the help of an oracle LLM. We showcase two example attacks: content injection and over-refusal attacks, each aiming to induce a specific exploitable behavior. We quantify and benchmark the strength and the stealthiness of our data poisoning scheme. Our results show that AutoPoison allows an adversary to change a model's behavior by poisoning only a small fraction of data while maintaining a high level of stealthiness in the poisoned examples. We hope our work sheds light on how data quality affects the behavior of instruction-tuned models and raises awareness of the importance of data quality for responsible deployments of LLMs.

Added

2026-09-26

Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks

Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks

Ali Shafahi, W. R. Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, T. Goldstein

OrganizationsCornell UniversityUniversity of Maryland

Why you should read this

Demonstrates how adversaries can force neural networks to misclassify specific test instances using correctly labeled training images, establishing that models trained via transfer learning or end-to-end pipelines are vulnerable to stealthy clean-label data poisoning.

Data poisoning is an attack on machine learning models wherein the attacker adds examples to the training set to manipulate the behavior of the model at test time. This paper explores poisoning attacks on neural nets. The proposed attacks use "clean-labels"; they don't require the attacker to have any control over the labeling of training data. They are also targeted; they control the behavior of the classifier on a specific\textit{specific} test instance without degrading overall classifier performance. For example, an attacker could add a seemingly innocuous image (that is properly labeled) to a training set for a face recognition engine, and control the identity of a chosen person at test time. Because the attacker does not need to control the labeling function, poisons could be entered into the training set simply by leaving them on the web and waiting for them to be scraped by a data collection bot. We present an optimization-based method for crafting poisons, and show that just one single poison image can control classifier behavior when transfer learning is used. For full end-to-end training, we present a "watermarking" strategy that makes poisoning reliable using multiple (≈\approx50) poisoned training instances. We demonstrate our method by generating poisoned frog images from the CIFAR dataset and using them to manipulate image classifiers.

Added

2026-09-25