Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks
Ali ShafahiW. R. HuangMahyar NajibiOctavian SuciuChristoph StuderTudor DumitrasT. Goldstein
Demonstrates how adversaries can force neural networks to misclassify specific test instances using correctly labeled training images, establishing that models trained via transfer learning or end-to-end pipelines are vulnerable to stealthy clean-label data poisoning.
Modern computer vision and deep learning systems often rely on massive datasets scraped from the public internet or external repositories. This practice introduces severe vulnerabilities if attackers can manipulate training sets. While past research focused on broad attacks that degrade overall system accuracy or evasion attacks that require modifying test-time inputs, the article addresses a stealthier threat: targeted clean-label data poisoning. In this scenario, an adversary manipulates the training data to cause a specific test input to be misclassified, all without altering the test input and while ensuring that injected training samples appear correctly labeled to human auditors.
The article demonstrates and evaluates an optimization-based attack method that crafts clean-label poisoned examples to hijack neural network predictions on chosen targets. It assesses this method across two primary environments: transfer learning, where pre-trained feature extractors have only their final layers retrained, and full end-to-end training, where every layer in the network is updated.
To construct these attacks, the authors used an iterative optimization algorithm to produce feature collisions. By modifying a base image so that it looks normal to humans while matching the target's internal representation inside the deep neural network, the injected sample tricks the model during retraining. The evaluation covered two benchmark computer vision setups: an InceptionV3 architecture trained on an ImageNet dog-versus-fish task for transfer learning, and a scaled-down AlexNet trained on CIFAR-10 image classifications for end-to-end training. To succeed in end-to-end retraining, the authors introduced an adversarial watermarking technique—blending low-opacity features of the target into roughly 50 diverse base images—to keep the poison and target tightly bound within internal feature distributions.
The key findings reveal significant vulnerabilities across deep learning pipelines. First, in transfer learning scenarios, injecting a single poisoned image achieved a 100% attack success rate across 1,099 test trials, flipping the classification of target images with a median confidence of 99.6%. Second, these transfer learning attacks caused an imperceptible drop in overall system performance, with average test accuracy dropping by only 0.2%, rendering standard anomaly defenses ineffective. Third, while a single poison sample failed in fully retrained networks because lower layers learned to separate the instances, combining feature optimization with target watermarking across 50 diverse base samples produced success rates of up to 60% in end-to-end training. Fourth, targeting statistical outliers—samples near class boundaries with lower initial classification confidence—boosted the end-to-end attack success rate to 70%.
These findings have major security and risk implications for mission-critical deployments like facial recognition, content filtering, and malware detection. Because poisoned instances carry completely valid visual labels and make up less than 0.1% of the training budget, standard data audits and performance monitoring cannot catch them. Adversaries do not need internal database access; they can simply publish poisoned images online to be collected by automated scraping bots. Furthermore, existing defenses that monitor validation accuracy will fail because the model maintains high overall precision while harboring a critical blind spot.
Organizations developing or deploying machine learning should establish strict data provenance, supply chain validation, and verification protocols for external datasets. Security teams must account for the fact that transfer learning pipelines are exceptionally fragile to single-instance poisoning. Because robust defenses against clean-label attacks remain an open problem, machine learning practitioners should combine algorithmic auditing with rigorous data-source authentication rather than relying solely on post-training validation accuracy.
The conclusions should be interpreted within the article's experimental boundaries. The attacks assume a white-box attacker who knows the model architecture and parameters. Additionally, end-to-end poisoning requires crafting multiple diverse instances with subtle watermarks. Despite these operational requirements, the results conclusively establish that clean-label poisoning is a viable, high-impact security risk for real-world neural network deployments.
- Paper: Poisoning Attacks against Support Vector Machines, Battista Biggio et al. (2012). Establishes the foundational gradient-based optimization framework for data poisoning attacks against machine learning classifiers that directly precedes poisoning strategies for deep neural networks.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). Introduces targeted backdoor poisoning attacks on deep learning architectures, providing the baseline threat model and context that clean-label poisoning seeks to overcome without mislabeling data.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). Identifies foundational supply-chain poisoning and backdoor vulnerabilities across neural networks, establishing the problem space of targeted training-set manipulation.
- Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). Develops the optimization-based perturbation formulations commonly adapted to craft feature-space collisions in targeted clean-label poison generation.
- Paper: Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, Nicolas Papernot et al. (2016). Analyzes feature transferability across models and representations, which underpins the transfer-learning attack setting used in clean-label poisoning.
- Paper: Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning, Battista Biggio et al. (2017). Offers a comprehensive survey of evasion and poisoning threat modeling across the first decade of adversarial machine learning.
- Paper: Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, Kang Liu et al. (2018). Develops fine-pruning defenses to neutralize targeted backdoor and poisoning attacks in trained neural networks without degrading clean classification performance.
- Paper: Analyzing Federated Learning through an Adversarial Lens, Arjun Nitin Bhagoji et al. (2018). Extends targeted poisoning concepts to decentralized federated learning environments to evaluate targeted misclassification under Byzantine aggregation defenses.
- Paper: How To Backdoor Federated Learning, Eugene Bagdasaryan et al. (2018). Demonstrates how targeted backdoor and poisoning attacks can be scaled up to compromise federated learning pipelines via model-replacement strategies.
- Paper: Local Model Poisoning Attacks to Byzantine-Robust Federated Learning, Minghong Fang et al. (2019). Formulates optimization-based local model poisoning attacks to systematically degrade accuracy across Byzantine-robust distributed learning frameworks.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). Investigates the fundamental feature geometry of deep networks, demonstrating that adversarial perturbations and clean-label poisoning operate on genuinely predictive non-robust dataset features.
