Built independently by an author, for readers. Read the story and support ChapterPal

keyword

targeted data poisoning attacks

A targeted data poisoning attack is a machine learning security exploit in which an adversary injects maliciously crafted or modified samples into a training dataset to alter a model predictions on specific target inputs. Unlike untargeted poisoning attacks that degrade overall system accuracy or cause widespread classification failures, targeted attacks aim to force a precise incorrect outcome for a chosen input while leaving the model performance on other data intact. This selective interference makes the compromise stealthy and difficult to detect during standard evaluation, as the poisoned model maintains normal functionality on typical benchmark tests while failing specifically when encountering the designated target at inference time.

1 item

Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks

Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks

Ali Shafahi, W. R. Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, T. Goldstein

OrganizationsCornell UniversityUniversity of Maryland

Why you should read this

Demonstrates how adversaries can force neural networks to misclassify specific test instances using correctly labeled training images, establishing that models trained via transfer learning or end-to-end pipelines are vulnerable to stealthy clean-label data poisoning.

Data poisoning is an attack on machine learning models wherein the attacker adds examples to the training set to manipulate the behavior of the model at test time. This paper explores poisoning attacks on neural nets. The proposed attacks use "clean-labels"; they don't require the attacker to have any control over the labeling of training data. They are also targeted; they control the behavior of the classifier on a specific\textit{specific} test instance without degrading overall classifier performance. For example, an attacker could add a seemingly innocuous image (that is properly labeled) to a training set for a face recognition engine, and control the identity of a chosen person at test time. Because the attacker does not need to control the labeling function, poisons could be entered into the training set simply by leaving them on the web and waiting for them to be scraped by a data collection bot. We present an optimization-based method for crafting poisons, and show that just one single poison image can control classifier behavior when transfer learning is used. For full end-to-end training, we present a "watermarking" strategy that makes poisoning reliable using multiple (≈\approx50) poisoned training instances. We demonstrate our method by generating poisoned frog images from the CIFAR dataset and using them to manipulate image classifiers.

Added

2026-09-25