Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks

Ali ShafahiW. R. HuangMahyar NajibiOctavian SuciuChristoph StuderTudor DumitrasT. Goldstein

article2018NeurIPS1,316 citations

Demonstrates how adversaries can force neural networks to misclassify specific test instances using correctly labeled training images, establishing that models trained via transfer learning or end-to-end pipelines are vulnerable to stealthy clean-label data poisoning.

Listen

Modern computer vision and deep learning systems often rely on massive datasets scraped from the public internet or external repositories. This practice introduces severe vulnerabilities if attackers can manipulate training sets. While past research focused on broad attacks that degrade overall system accuracy or evasion attacks that require modifying test-time inputs, the article addresses a stealthier threat: targeted clean-label data poisoning. In this scenario, an adversary manipulates the training data to cause a specific test input to be misclassified, all without altering the test input and while ensuring that injected training samples appear correctly labeled to human auditors.

The article demonstrates and evaluates an optimization-based attack method that crafts clean-label poisoned examples to hijack neural network predictions on chosen targets. It assesses this method across two primary environments: transfer learning, where pre-trained feature extractors have only their final layers retrained, and full end-to-end training, where every layer in the network is updated.

To construct these attacks, the authors used an iterative optimization algorithm to produce feature collisions. By modifying a base image so that it looks normal to humans while matching the target's internal representation inside the deep neural network, the injected sample tricks the model during retraining. The evaluation covered two benchmark computer vision setups: an InceptionV3 architecture trained on an ImageNet dog-versus-fish task for transfer learning, and a scaled-down AlexNet trained on CIFAR-10 image classifications for end-to-end training. To succeed in end-to-end retraining, the authors introduced an adversarial watermarking technique—blending low-opacity features of the target into roughly 50 diverse base images—to keep the poison and target tightly bound within internal feature distributions.

The key findings reveal significant vulnerabilities across deep learning pipelines. First, in transfer learning scenarios, injecting a single poisoned image achieved a 100% attack success rate across 1,099 test trials, flipping the classification of target images with a median confidence of 99.6%. Second, these transfer learning attacks caused an imperceptible drop in overall system performance, with average test accuracy dropping by only 0.2%, rendering standard anomaly defenses ineffective. Third, while a single poison sample failed in fully retrained networks because lower layers learned to separate the instances, combining feature optimization with target watermarking across 50 diverse base samples produced success rates of up to 60% in end-to-end training. Fourth, targeting statistical outliers—samples near class boundaries with lower initial classification confidence—boosted the end-to-end attack success rate to 70%.

These findings have major security and risk implications for mission-critical deployments like facial recognition, content filtering, and malware detection. Because poisoned instances carry completely valid visual labels and make up less than 0.1% of the training budget, standard data audits and performance monitoring cannot catch them. Adversaries do not need internal database access; they can simply publish poisoned images online to be collected by automated scraping bots. Furthermore, existing defenses that monitor validation accuracy will fail because the model maintains high overall precision while harboring a critical blind spot.

Organizations developing or deploying machine learning should establish strict data provenance, supply chain validation, and verification protocols for external datasets. Security teams must account for the fact that transfer learning pipelines are exceptionally fragile to single-instance poisoning. Because robust defenses against clean-label attacks remain an open problem, machine learning practitioners should combine algorithmic auditing with rigorous data-source authentication rather than relying solely on post-training validation accuracy.

The conclusions should be interpreted within the article's experimental boundaries. The attacks assume a white-box attacker who knows the model architecture and parameters. Additionally, end-to-end poisoning requires crafting multiple diverse instances with subtle watermarks. Despite these operational requirements, the results conclusively establish that clean-label poisoning is a viable, high-impact security risk for real-world neural network deployments.

Cover for Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks

Abstract

Data poisoning is an attack on machine learning models wherein the attacker adds examples to the training set to manipulate the behavior of the model at test time. This paper explores poisoning attacks on neural nets. The proposed attacks use "clean-labels"; they don't require the attacker to have any control over the labeling of training data. They are also targeted; they control the behavior of the classifier on a specific\textit{specific} test instance without degrading overall classifier performance. For example, an attacker could add a seemingly innocuous image (that is properly labeled) to a training set for a face recognition engine, and control the identity of a chosen person at test time. Because the attacker does not need to control the labeling function, poisons could be entered into the training set simply by leaving them on the web and waiting for them to be scraped by a data collection bot.

We present an optimization-based method for crafting poisons, and show that just one single poison image can control classifier behavior when transfer learning is used. For full end-to-end training, we present a "watermarking" strategy that makes poisoning reliable using multiple (≈\approx50) poisoned training instances. We demonstrate our method by generating poisoned frog images from the CIFAR dataset and using them to manipulate image classifiers.

Table of Contents

  • 1 Introduction
  • 1.1 Related work
  • 1.2 Contributions
  • 2 A simple clean-label attack
  • 2.1 Crafting poison data via feature collisions
  • 2.2 Optimization procedure
  • 3 Poisoning attacks on transfer learning
  • 3.1 A one-shot kill attack
  • 4 Poisoning attacks on end-to-end training
  • 4.1 Single poison instance attack
  • 4.2 Watermarking: a method to boost the power of poison attacks
  • 4.2.1 Multiple poison instance attacks
  • 5 Conclusion
  • 6 Acknowledgements
  • References
  • A Illustrations of the poisoning scheme and of how the decision boundary will rotate after training with the poison instances
  • B Comparison to adversarial examples
  • C Pixel bounded (i.e. L∞L_{\infty}) poisoning attacks
  • D One-shot kill attacks for multi-class transfer-learning scenario
  • E Sampling the candidate target instance for more success
  • F Network architecture for CIFAR-10 classifier
  • G What do watermarked poisons look like?
  • H Ablation study: How many frogs does it take to poison a network?

Knowls

  1. Knowl 1 — Targeted Clean-Label Poisoning Threat Model

    definition

    Targeted clean-label data poisoning is a training-time attack against machine learning classifiers defined by two constraints:

    1. Clean-label: Injected poison samples are correctly labeled by a certified oracle or human reviewer according to visual appearance in input space (no malicious label flipping or overt label corruption).
    2. Targeted: The adversary aims to cause a specific, unperturbed test-time target instance tt from a target class to be misclassified into an adversary-chosen base class during inference, while preserving normal classifier accuracy on all other inputs.

    Unlike evasion attacks or backdoor attacks, the target instance is not modified at inference time, and no trigger pattern is applied at test time. The adversary is assumed to have white-box knowledge of the model architecture and parameters (such as a public feature extractor used in transfer learning) but requires no control over the training data labeling pipeline or minibatch construction.

  2. Knowl 2 — Poison Generation via Feature Collision Optimization

    algorithm

    To construct a clean-label poison instance pp that visually resembles a base image bb while colliding with a target instance tt in penultimate-layer feature space, a forward-backward splitting optimization procedure is used.

    Let f(x)f(x) denote the activations of the neural network's penultimate layer for an input image xx, and let eta > 0 be a trade-off parameter controlling visual fidelity to bb. The optimization minimizes the loss Lp(x)=∥f(x)−f(t)∥22+β∥x−b∥22L_p(x) = \|f(x) - f(t)\|_2^2 + \beta \|x - b\|_2^2.

    Input: target instance tt, base instance bb, learning rate λ\lambda, trade-off parameter β\beta, maximum iterations maxIters\text{maxIters}
    Output: poisoned instance pp
    x0←bx_0 \leftarrow b
    for i=1i = 1 to maxIters\text{maxIters} do
        x^i←xi−1−λ∇x∥f(xi−1)−f(t)∥22\hat{x}_i \leftarrow x_{i-1} - \lambda \nabla_x \|f(x_{i-1}) - f(t)\|_2^2
        xi←(x^i+λβb)/(1+βλ)x_i \leftarrow (\hat{x}_i + \lambda \beta b) / (1 + \beta \lambda)
    end for
    return xmaxItersx_{\text{maxIters}}

    The forward step performs gradient descent on the feature collision loss ∥f(x)−f(t)∥22\|f(x) - f(t)\|_2^2 to push the representation toward f(t)f(t). The backward step is a proximal update minimizing the Frobenius distance to the base image bb in input space. The loop terminates when i=maxItersi = \text{maxIters} or when the feature space distance ∥f(xi)−f(t)∥2\|f(x_i) - f(t)\|_2 drops below a chosen threshold (such as the minimum Euclidean distance between training pairs in feature space).

  3. Knowl 3 — Feature Collision Optimization Formulations

    equation

    Let f:RD→Rdf: \mathbb{R}^D \to \mathbb{R}^d denote the penultimate-layer feature extractor of a neural network, t∈RDt \in \mathbb{R}^D denote a target instance from a target class, and b∈RDb \in \mathbb{R}^D denote a base instance from a base class.

    In the ℓ2\ell_2-regularized formulation, the poison instance p∈RDp \in \mathbb{R}^D is defined by: p=arg min⁡x∈RD∥f(x)−f(t)∥22+β∥x−b∥22p = \operatorname*{arg\,min}_{x \in \mathbb{R}^D} \|f(x) - f(t)\|_2^2 + \beta \|x - b\|_2^2 where β>0\beta > 0 balances feature collision against input deviation. When the input dimension dimb\text{dim}_b varies and the penultimate feature dimension is d=2048d = 2048, β\beta is parameterized as: β=β0⋅d2(dimb)2\beta = \beta_0 \cdot \frac{d^2}{(\text{dim}_b)^2} with baseline parameter β0=0.25\beta_0 = 0.25.

    In the ℓ∞\ell_\infty-bounded formulation, the optimization is:

    p = &\operatorname*{arg\,min}_{x \in \mathbb{R}^D} \|f(x) - f(t)\|_2^2 \\ &\text{subject to } \|x - b\|_\infty \le \epsilon_\infty \end{aligned}$$ where $\epsilon_\infty = 2$ on an 8-bit dynamic range $[0, 255]$, enforced by clipping pixels to $[b - \epsilon_\infty, b + \epsilon_\infty]$ after each gradient step.
  4. Knowl 4 — One-Shot Kill Clean-Label Poisoning on Transfer Learning

    empirical result

    In a transfer learning setting where a pretrained InceptionV3 network (feature dimension d=2048d = 2048) is frozen and only the final softmax classification layer is retrained from scratch on an ImageNet dog-vs-fish binary classification dataset (900 clean training instances per class, 1800 total; 1099 test instances consisting of 698 dogs and 401 fish):

    • Adding a single clean-label poison instance (M=1M = 1) to the 1800 training examples achieves an attack success rate of 100% across all 1099 distinct test targets, outperforming influence-function-based poisoning on the same task (which achieves 57%).
    • The median test-time misclassification confidence for the target into the base class is 99.6%.
    • Overall test accuracy on clean images drops by an average of 0.2% (worst-case drop of 0.4%) from the clean model baseline accuracy of 99.5%.
    • Both ℓ2\ell_2-regularized poisoning and ℓ∞\ell_\infty-bounded poisoning (∥x−b∥∞≤2/255\|x - b\|_\infty \le 2/255) achieve 100% attack success.
  5. Knowl 5 — Mechanistic Contrast of Poisoning: Transfer Learning vs. End-to-End Retraining

    theoretical result

    The mechanism accommodating clean-label poisons differs fundamentally depending on whether feature extractors are frozen or trainable:

    1. Transfer Learning (Frozen Feature Extractor): Because the penultimate-layer mapping f(x)f(x) cannot adapt, retraining the final classification layer forces the linear decision boundary in feature space to rotate significantly (an average angular shift of 23 degrees, largely occurring in epoch 1) to incorporate the poison instance into the base class partition. This rotation inadvertently sweeps the neighboring target instance tt into the base class side of the boundary.
    2. End-to-End Retraining (All Layers Trainable): The final decision boundary remains virtually stationary (angular deviation varying by fractions of a degree). Instead, gradient backpropagation adjusts the lower-level convolutional filters in shallow layers. This removes the artificial feature collision, returning the poison representation f(p)f(p) back into the base class cluster in deep feature space while leaving the target representation f(t)f(t) in the target class cluster. Consequently, a single poison instance is neutralized without inducing target misclassification.
  6. Knowl 6 — Watermarking and Multi-Base Strategy for End-to-End Poisoning

    model/method

    To prevent shallow convolutional layers from separating poison and target representations during end-to-end network retraining, clean-label poisoning incorporates three interdependent mechanisms:

    1. Target Watermarking: Each base image bb is blended with a low-opacity watermark of the target instance tt: bwatermarked=γ⋅t+(1−γ)⋅bb_{\text{watermarked}} = \gamma \cdot t + (1 - \gamma) \cdot b where γ∈[0.20,0.30]\gamma \in [0.20, 0.30] is the opacity coefficient. This introduces inseparable low-level semantic features shared between the poison and target images that persist through retraining, while keeping the base image visually intact to human inspection.
    2. Multi-Base Diversity: Multiple distinct base instances (M≈50M \approx 50) are randomly sampled from the base class, watermarked with tt, and independently optimized via feature collision. A moderately sized deep network cannot learn individual filters to separate MM diverse base instances from the target while simultaneously preserving target-class accuracy, forcing the network to pull the target representation into the base class distribution in feature space.
    3. Feature Collision Optimization: Algorithm 1 is applied to each watermarked base to initialize its penultimate feature vector near f(t)f(t).
  7. Knowl 7 — Clean-Label Poisoning Performance on End-to-End CIFAR-10 Networks

    empirical result

    On CIFAR-10 using a scaled-down AlexNet trained end-to-end with Adam (learning rate 1.85×10−51.85 \times 10^{-5}, warm-start initialization, batch size 128, 10 epochs), multi-base watermarked clean-label poisoning exhibits the following performance characteristics:

    • Monotonic Scaling with Poison Count: Target misclassification success rate (defined strictly as classification into the designated base class) increases monotonically with the number of poison images M∈[1,70]M \in [1, 70]. On a bird-vs-dog classification task with 30% watermark opacity, the attack success rate reaches approximately 60% at M=50M = 50 poisons over 30 random trials.
    • Watermark Opacity Sensitivity: Reducing watermark opacity from γ=0.30\gamma = 0.30 to γ=0.20\gamma = 0.20 lowers attack success rates across all tested poison counts in airplane-vs-frog tasks.
    • Outlier Target Vulnerability: Targeting outlier test samples—defined as the 50 lowest-confidence correctly classified airplanes—increases the attack success rate to 70% with M=50M = 50 poison frogs (at γ=0.30\gamma = 0.30), compared to 53% for randomly chosen targets.
  8. Knowl 8 — Ablation Analysis of End-to-End Poisoning Components

    empirical result

    A leave-one-out ablation study using 50 poison instances on end-to-end CIFAR-10 network retraining demonstrates that all three attack components are essential for targeted clean-label poisoning:

    • Full Strategy (Optimization + Watermarking + Multi-Base Diversity): Successfully relocates the target instance into the base class feature distribution within the first 1–2 epochs of retraining, causing target misclassification.
    • Ablation 1 (Excluding Multi-Base Diversity; Single Base Replicated 50 Times): Fails. Training on multiple perturbed copies of a single base image acts as adversarial training on that base, enabling the network to learn kernels that separate the poison from the target.
    • Ablation 2 (Excluding Feature Optimization; Watermarking + Multi-Base Only): Fails. Without gradient-based feature collision, the watermarked base images do not lie close enough to the target in deep feature space prior to retraining.
    • Ablation 3 (Excluding Watermarking; Optimization + Multi-Base Only): Fails. Without the shared low-level feature anchor provided by the watermark, shallow convolutional kernels adapt during retraining to return all 50 poisons to the base distribution while leaving the target in the target distribution.
  9. Knowl 9 — Underdetermined Parameter Condition in Multi-Class Transfer Learning

    theoretical result

    For transfer learning classifiers where a frozen feature extractor produces representations of dimension ndn_d and only the final classification layer with CC classes is trained on NtrN_{\text{tr}} examples:

    • The linear system for the final layer weights is underdetermined whenever: Ntr≤nd×CN_{\text{tr}} \le n_d \times C
    • Provided the training set contains no duplicate images, the feature matrix has full row rank, guaranteeing the existence of weight parameters that achieve zero loss (perfect interpolation) on all training instances, including injected poisons.
    • Because the single-layer classification loss is convex, standard gradient-based optimization (such as SGD or Adam) converges to a global minimizer with near-zero training loss (<10−5< 10^{-5} after 100 epochs).
    • Increasing the number of classes CC expands the trainable parameter count (nd×Cn_d \times C), enhancing susceptibility. On a 3-way classification task ("dog" vs. "fish" vs. "cat") with Ntr=2700≤2048×3=6144N_{\text{tr}} = 2700 \le 2048 \times 3 = 6144, one-shot clean-label poisoning achieves a 100% attack success rate across 100 randomly sampled test targets while maintaining 96.4% test accuracy on clean data.
  10. Knowl 10 — Scaled AlexNet Architecture for CIFAR-10 Poisoning Experiments

    data/table

    The end-to-end poisoning experiments on CIFAR-10 use a scaled-down AlexNet convolutional neural network.

    Layer Type Kernel Size Output Dimension
    1 Conv + ReLU 5×55 \times 5 64
    2 MaxPool (stride 2) 3×33 \times 3 64
    3 Local Response Normalization (LRN) – 64
    4 Conv + ReLU 5×55 \times 5 64
    5 MaxPool (stride 2) 3×33 \times 3 64
    6 Local Response Normalization (LRN) – 64
    7 FullyConnected + ReLU – 384
    8 FullyConnected + ReLU – 192
    9 FullyConnected (logits) – 10

    When trained on unpoisoned CIFAR-10 training data for 200 epochs with a scheduled learning rate decay, this architecture attains 100% training accuracy and 74.5% test accuracy. During end-to-end poisoning evaluations, the network is warm-started from these pretrained weights and retrained for 10 epochs using the Adam optimizer at a learning rate of 1.85×10−51.85 \times 10^{-5} with a batch size of 128.

Coverage note — None was omitted; all key contributions—including the clean-label poisoning threat model, optimization formulations, algorithm, transfer learning experiments, end-to-end mechanism analysis, watermarking multi-base method, ablation study, multi-class extension, and network architecture—are captured.

References

  1. 1.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013.
  2. 2.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. International Conference on Learning Representation, 2015.
  3. 3.Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. In Joint European conference on machine learning and knowledge discovery in databases, pages 387–402. Springer, 2013.
  4. 4.Battista Biggio, Blaine Nelson, and Pavel Laskov. Poisoning attacks against support vector machines. arXiv preprint arXiv:1206.6389, 2012.
  5. 5.Blaine Nelson, Marco Barreno, Fuching Jack Chi, Anthony D. Joseph, Benjamin I. P. Rubinstein, Udam Saini, Charles Sutton, J. D. Tygar, and Kai Xia. Exploiting machine learning to subvert your spam filter. In Proceedings of the 1st Usenix Workshop on Large-Scale Exploits and Emergent Threats, pages 7:1–7:9, Berkeley, CA, USA, 2008.
  6. 6.Jacob Steinhardt, Pang Wei Koh, and Percy Liang. Certified Defenses for Data Poisoning Attacks. arXiv preprint arXiv:1706.03691, (i), 2017. URL http://arxiv.org/abs/1706.03691.
  7. 7.Luis Muñoz-González, Battista Biggio, Ambra Demontis, Andrea Paudice, Vasin Wongrassamee, Emil C Lupu, and Fabio Roli. Towards poisoning of deep learning algorithms with back-gradient optimization. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, pages 27–38. ACM, 2017.
  8. 8.Chaofei Yang, Qing Wu, Hai Li, and Yiran Chen. Generative poisoning attack method against neural networks. arXiv preprint arXiv:1703.01340, 2017.
  9. 9.Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. arXiv preprint arXiv:1712.05526, 2017. URL http://arxiv.org/abs/1712.05526.
  10. 10.Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017.
  11. 11.Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang. Trojaning attack on neural networks. 2017.
  12. 12.Octavian Suciu, Radu Mărginean, Yi˘gitcan Kaya, Hal Daumé III, and Tudor Dumitraș. When does machine learning fail? generalized transferability for evasion and poisoning attacks. arXiv preprint arXiv:1803.06975, 2018.
  13. 13.Saeed Mahloujifar and Mohammad Mahmoody. Blockwise pp-tampering attacks on cryptographic primitives, extractors, and learners. Cryptology ePrint Archive, Report 2017/950, 2017. https://eprint.iacr.org/2017/950.
  14. 14.Saeed Mahloujifar, Dimitrios I. Diochnos, and Mohammad Mahmoody. Learning under pp-tampering attacks. CoRR, abs/1711.03707, 2017. URL http://arxiv.org/abs/1711.03707.
  15. 15.Ilias Diakonikolas, Daniel M Kane, and Alistair Stewart. Efficient robust proper learning of log-concave distributions. arXiv preprint arXiv:1606.03077, 2016.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. arXiv preprint arXiv:1512.03385, 7(3):171–180, 2015. ISSN 1664-1078. doi: 10.3389/fpsyg.2013.00124. URL http://arxiv.org/pdf/1512.03385v1.pdf.
  17. 17.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going Deeper with Convolutions. arXiv:1409.4842, 2014. ISSN 10636919. doi: 10.1109/CVPR.2015.7298594. URL https://arxiv.org/abs/1409.4842.
  18. 18.Marco Barreno, Blaine Nelson, Anthony D Joseph, and JD Tygar. The security of machine learning. Machine Learning, 81(2):121–148, 2010.
  19. 19.Pang Wei Koh and Percy Liang. Understanding black-box predictions via influence functions. arXiv preprint arXiv:1703.04730, 2017.
  20. 20.Tom Goldstein, Christoph Studer, and Richard Baraniuk. A field guide to forward-backward splitting with a fasta implementation. arXiv preprint arXiv:1411.3406, 2014.
  21. 21.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2818–2826, 2016.
  22. 22.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015.
  23. 23.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. 2009.

Citation

MLA
Shafahi, A., et al. “Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks”. arXiv, 2018, http://arxiv.org/abs/1804.00792v2.
APA
Shafahi, A., Huang, W. R., Najibi, M., Suciu, O., Studer, C., Dumitras, T., & Goldstein, T. (2018). Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks. arXiv. http://arxiv.org/abs/1804.00792v2
Chicago
Shafahi, A., W. R. Huang, M. Najibi, et al. 2018. “Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks”. arXiv. http://arxiv.org/abs/1804.00792v2.
Harvard
Shafahi, A. et al. (2018) “Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1804.00792v2.
Vancouver
1. Shafahi A, Huang WR, Najibi M, Suciu O, Studer C, Dumitras T, Goldstein T (2018) Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks. arXiv

BibTeX

@article{shafahi2018poison,
  title = {Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks},
  author = {Shafahi, Ali and Huang, W. Ronny and Najibi, Mahyar and Suciu, Octavian and Studer, Christoph and Dumitras, Tudor and Goldstein, Tom},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1804.00792v2},
  eprint = {1804.00792}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors