Adversarial Patch

Tom B. BrownDandelion ManéAurko RoyMartín AbadiJustin Gilmer

article2017arXiv1,345 citations

Introduces printable, scene-agnostic adversarial patches capable of overriding deep learning image classifiers in the physical world to force the prediction of an arbitrary target class.

Listen

Most modern security research on image-recognition systems focuses on defending against subtle, imperceptible modifications to digital images. However, real-world automated systems—ranging from surveillance cameras to autonomous vehicles—frequently operate without human oversight, leaving them vulnerable to physical modifications that an attacker does not need to hide. The article addresses this operational blind spot by evaluating whether visible, physical artifacts can reliably force image classifiers to misclassify scenes.

The article demonstrates a method for generating printable, targeted "adversarial patches" that consistently cause image classifiers to output a chosen target category regardless of the surrounding background or scene contents. Unlike traditional digital attacks, these patches are scene-independent, meaning an attacker does not need prior knowledge of lighting, camera angles, or the background objects being photographed.

To create these patches, the researchers adapted an optimization framework that trains a localized image patch across a wide distribution of background images, positions, scales, and rotations. They evaluated the approach on standard benchmark vision models (including ResNet50, InceptionV3, and VGG variants) across white-box settings where internal model parameters are known, and black-box settings where the patch is trained on a subset of models and evaluated on an unseen model. The team also conducted physical experiments by printing the generated patches with a standard desktop printer and photographing them alongside real objects.

The findings show that printable patches reliably deceive target classifiers across varying environments. First, when placed beside a real-world object (such as a banana originally identified with 97% confidence), the addition of a printed sticker targeted to a "toaster" class forced the model to output "toaster" with 99% confidence, completely ignoring the primary object. Second, patches generalized effectively across multiple model architectures in digital black-box evaluations, substantially outperforming natural images of the target class inserted at the same size. Third, researchers successfully camouflaged the patches to resemble abstract art or tie-dye peace signs while retaining their ability to mislead the models.

These results demonstrate a critical operational security risk. Because the patch does not depend on a specific scene, an attacker can create a single design offline and distribute it widely for third parties to print and deploy as physical stickers. The findings also reveal that existing defense strategies, which focus primarily on minor digital perturbations, fail to protect computer vision systems against localized, high-salience physical modifications.

To address this vulnerability, security and engineering teams deploying automated vision systems must expand defensive evaluations beyond small pixel-level noise to include localized, large-magnitude physical attacks. Organizations should explore multi-object detection and segmentation architectures that reduce reliance on global image classification, as standard classifiers simply output the single most salient feature detected. While the attack is highly effective in controlled environments and transfers across models, its effectiveness in black-box physical settings currently decreases unless the patch occupies a relatively large portion of the frame or is optimized for print quality. Further testing on diverse real-world cameras and specialized print processes is necessary before fully assessing threat levels against specific production deployments.

  • Paper: Adversarial examples in the physical world, Alexey Kurakin et al. (2016). This work provides foundational evidence that adversarial examples survive printing and physical photography, directly motivating the development of printable physical adversarial patches.
  • Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). It introduces image-agnostic universal adversarial perturbations, establishing the core concept of scene-independent fooling that the Adversarial Patch generalizes into a localized physical form.
  • Paper: Synthesizing Robust Adversarial Examples, Anish Athalye et al. (2017). It establishes the Expectation Over Transformation framework used to optimize adversarial examples over a distribution of real-world transformations such as viewing angles, lighting, and scaling.
  • Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This seminal paper explains the linearity underlying adversarial vulnerability and introduces standard gradient-based optimization techniques essential for crafting targeted attacks.
  • Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). It formalizes powerful optimization-based formulations for generating targeted adversarial attacks across standard distance metrics.
  • Paper: The Limitations of Deep Learning in Adversarial Settings, Nicolas Papernot et al. (2015). It explores targeted adversarial attacks that restrict perturbations to small subsets of input features using saliency maps.
Cover for Adversarial Patch

Abstract

We present a method to create universal, robust, targeted adversarial image patches in the real world. The patches are universal because they can be used to attack any scene, robust because they work under a wide variety of transformations, and targeted because they can cause a classifier to output any target class. These adversarial patches can be printed, added to any scene, photographed, and presented to image classifiers; even when the patches are small, they cause the classifiers to ignore the other items in the scene and report a chosen target class.

To reproduce the results from the paper, our code is available at this https URL

Table of Contents

  • 1 Introduction
  • 2 Approach
  • 3 Experimental Results
  • 4 Conclusion
  • 5 Appendix
  • References

Knowls

  1. Knowl 1 — Adversarial Patch Optimization Objective via Expectation over Transformation

    equation

    An adversarial patch p^\hat{p} targeted to a specific class y^\hat{y} is generated by optimizing the patch pixel values over an expectation across images, spatial locations, and geometric transformations:

    p^=arg⁡max⁡pEx∼X,t∼T,l∼L[log⁡Pr⁡(y^∣A(p,x,l,t))]\hat{p} = \arg\max_{p} \mathbb{E}_{x \sim X, t \sim T, l \sim L} \left[ \log \Pr(\hat{y} \mid A(p, x, l, t)) \right]

    where:

    • pp is the patch parameter tensor being optimized,
    • XX is a distribution or dataset of background images,
    • TT is a distribution over geometric transformations of the patch (such as scaling and 2D rotations),
    • LL is a distribution over spatial placement coordinates within the image,
    • y^\hat{y} is the designated target classification label,
    • Pr⁡(y^∣⋅)\Pr(\hat{y} \mid \cdot) is the classifier's output probability assigned to class y^\hat{y},
    • A(p,x,l,t)A(p, x, l, t) is the patch application operator that applies transformation tt to patch pp and places it onto image xx at location ll.

    The optimization is solved using gradient descent on pp. Because the objective takes expectations over backgrounds, spatial locations, and affine transforms, the optimized patch acts as an image-independent, scene-invariant, targeted perturbation.

  2. Knowl 2 — Patch Application Operator for Arbitrary-Shaped Perturbations

    model/method

    For an image x∈Rw×h×cx \in \mathbb{R}^{w \times h \times c} with width ww, height hh, and color channels cc, a patch pp, a spatial coordinate location ll, and a set of affine transformations tt (such as scaling, rotation, and translation), the patch application operator A(p,x,l,t)A(p, x, l, t) modifies an image by:

    1. Applying the transformation tt to the patch pp.
    2. Applying a spatial mask to the patch to define its boundary geometry (e.g., circular sticker, peace sign, or custom shape).
    3. Overwriting the pixels of image xx at location ll with the transformed, masked patch pp, while leaving all image pixels outside the mask unaltered.

    Unlike traditional additive adversarial perturbations bounded by small LpL_p norms (∥x−x^∥p≤ε\|x - \hat{x}\|_p \le \varepsilon) across all pixels, A(p,x,l,t)A(p, x, l, t) completely replaces a localized subregion of the target image with unconstrained pixel values.

  3. Knowl 3 — Camouflaged and Disguised Adversarial Patch Optimization

    model/method

    To reduce the visual conspicuity of an adversarial patch to human observers, the patch optimization objective can incorporate regularizers and structural constraints to resemble familiar, benign visual patterns:

    1. L∞L_\infty-Bounded Camouflage: A constraint of the form ∥p−porig∥∞<ϵ\|p - p_{\text{orig}}\|_\infty < \epsilon enforces that all pixel values of the optimized patch pp stay within an L∞L_\infty radius ϵ\epsilon of a chosen source image porigp_{\text{orig}}.
    2. Pattern Regularization and Masking: The patch is regularized during gradient descent by minimizing its L2L_2 distance to a target texture (such as a tie-dye pattern) and training under a fixed geometric mask (such as a peace sign shape).

    These constraints produce patches that resemble innocuous decorative stickers while retaining high targeted attack effectiveness against neural networks.

  4. Knowl 4 — Evaluation Protocol and Ensemble Schemes for Adversarial Patches

    experimental setup

    The effectiveness of adversarial patch attacks is evaluated across ImageNet classifiers (including InceptionV3, ResNet50, Xception, VGG16, and VGG19) under four distinct settings:

    • White-Box Single Model: A patch is optimized against and evaluated on a single classifier architecture.
    • White-Box Ensemble: A single patch is jointly optimized across an ensemble of five ImageNet architectures (InceptionV3, ResNet50, Xception, VGG16, and VGG19) to maximize the average target class log-probability across all models, and evaluated on all five.
    • Black-Box Transfer: A patch is jointly trained on an ensemble of four ImageNet architectures and evaluated against a fifth held-out architecture not seen during training.
    • Control Baseline: A non-adversarial photograph of the target class (e.g., a natural photo of a toaster) scaled and placed on images.

    Each attack condition is evaluated across patch sizes (measured as a percentage of total image area). For every scale tested, the patch is placed at random locations across 400 randomly chosen test images, and the win rate (percentage of images classified as the target class) is recorded.

  5. Knowl 5 — Targeted Attack Success Rates Across Patch Scales and Transfer Settings

    empirical result

    In evaluations targeting the class 'toaster' across 400 test images per patch scale:

    • White-Box Single Model Attack: Achieves an attack success rate near 90% when the patch occupies 10% of the image area, approaching 100% success rate at 15–20% of image area.
    • White-Box Ensemble Attack: Achieves approximately 90% success rate at 10–15% image area and reaches nearly 100% success rate by 20% image area.
    • Black-Box Transfer Attack: Achieves roughly 20% success rate at 10% image area, 50% at 20% area, 70% at 30% area, and over 80% at 40–50% area.
    • Natural Image Control (Real Toaster): Inserting a real photograph of a toaster into the scene achieves under 10% classification as a toaster at 10% image area, and only reaches ~40–50% classification rate even when occupying 40–50% of the total image area.

    Optimized adversarial patches substantially outperform the control photograph across all scales, indicating that the patch creates feature activations far more salient to classifiers than natural object instances.

  6. Knowl 6 — Targeted Fooling Rates of Camouflaged and Masked Patches

    empirical result

    When adversarial patches are regularized to disguise their visual appearance, they retain high targeted fooling capability against ImageNet classifiers:

    • Tie-Dye Disguise: Regularizing the patch via L2L_2 distance toward a tie-dye pattern achieves ~30% targeted success rate at 10% image area, ~65% at 20% image area, and over 90% at 40–50% image area.
    • Peace Sign Mask Disguise: Combining tie-dye regularization with a peace sign shape mask achieves ~10% targeted success rate at 10% image area, ~35% at 20% image area, ~60% at 30% image area, and ~80% at 50% image area.
    • Comparison: Both disguised patch configurations significantly outperform the control baseline (a natural picture of a toaster) across all image area percentages above 10%.
  7. Knowl 7 — Physical-World Targeted Classifier Hijacking via Printed Adversarial Patches

    empirical result

    Adversarial patches generated via white-box ensemble optimization successfully fool image classifiers in the physical world when printed with a standard color printer:

    • A baseline photograph of a tabletop containing a banana and a notebook is classified by VGG16 as class 'banana' with 97% confidence.
    • Placing a physical circular sticker patch targeted to 'toaster' on the tabletop causes VGG16 to classify the photograph as 'toaster' with 99% confidence, completely overriding the natural banana in the frame.
    • Physical-world black-box transferability was demonstrated against a third-party mobile recognition app (Demitasse), succeeding when the physical patch occupies a large fraction of the camera's field of view.
  8. Knowl 8 — Saliency Exploitation Mechanism in Single-Label Image Classifiers

    theoretical result

    The adversarial patch exploits the training setup of single-label image classification models:

    1. In standard classification benchmarks, each image is annotated with only one ground-truth class label, compelling neural networks to identify and report only the single most salient visual concept present.
    2. The adversarial patch optimization produces localized feature patterns that are orders of magnitude more salient to the model's feature extractors than natural real-world objects.
    3. Consequently, placing the patch anywhere in an image causes the model to classify the entire image as the patch's target class while ignoring all other objects in the scene.
    4. In multi-label contexts such as object detection and image segmentation, this mechanism implies that a targeted adversarial patch will be localized and detected as its target class without masking or suppressing detection of other surrounding objects.
  9. Knowl 9 — Orientation and Area Scale Constraints on Patch Transferability

    limitation

    The operational efficacy of adversarial patches in black-box and physical environments is subject to several constraints:

    • Orientation Sensitivity: For optimal targeted fooling in physical deployments, the printed patch sticker must remain within approximately 20 degrees of the vertical alignment used during training.
    • Patch Size Requirement: Because the attack must generalize universally across unseen backgrounds, arbitrary locations, and transformations, the required patch area (10% to 50% of the image) is substantially larger than that of non-universal, single-image white-box attacks (such as single-pixel modifications covering 0.1% of an image).
    • Print-ability Loss: Patches optimized without explicit color-gamut printability loss functions experience degraded black-box transferability in physical deployments compared to digital evaluations.

Coverage note — None was omitted; all key contributions including the EOT formulation, patch operator, disguised variants, empirical evaluations across white-box/black-box settings, physical-world transfer experiments, saliency mechanism analysis, and practical limitations are fully covered.

References

  1. 1.Anonymous. Adversarial spheres. International Conference on Learning Representations, 2018.
  2. 2.Anonymous. Thermometer encoding: One hot way to resist adversarial examples. International Conference on Learning Representations, 2018.
  3. 3.A. Athalye, L. Engstrom, A. Ilyas, and K. Kwok. Synthesizing robust adversarial examples. arXiv preprint arXiv:1707.07397, 2017.
  4. 4.I. Evtimov, K. Eykholt, E. Fernandes, T. Kohno, B. Li, A. Prakash, A. Rahmati, and D. Song. Robust physical-world attacks on deep learning models. arXiv preprint arXiv:1707.08945, 2017.
  5. 5.I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.
  6. 6.S. K. Jiawei Su, Danilo Vasconcellos Vargas. One pixel attack for fooling deep neural networks. arXiv preprint arXiv:1710.08864, 2017.
  7. 7.A. Kurakin, I. Goodfellow, and S. Bengio. Adversarial examples in the physical world. arXiv preprint arXiv:1607.02533, 2016.
  8. 8.A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial examples. arXiv preprint arXiv:1706.06083, 2017.
  9. 9.S.-M. Moosavi-Dezfooli, A. Fawzi, O. Fawzi, and P. Frossard. Universal adversarial perturbations. arXiv preprint arXiv:1610.08401, 2016.
  10. 10.S.-M. Moosavi-Dezfooli, A. Fawzi, and P. Frossard. Deepfool: a simple and accurate method to fool deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2574–2582, 2016.
  11. 11.N. Papernot, P. McDaniel, S. Jha, M. Fredrikson, Z. B. Celik, and A. Swami. The limitations of deep learning in adversarial settings. In Security and Privacy (EuroS&P), 2016 IEEE European Symposium on, pages 372–387. IEEE, 2016.
  12. 12.N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami. Distillation as a defense to adversarial perturbations against deep neural networks. In Security and Privacy (SP), 2016 IEEE Symposium on, pages 582–597. IEEE, 2016.
  13. 13.M. Sharif, S. Bhagavatula, L. Bauer, and M. K. Reiter. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pages 1528–1540. ACM, 2016.
  14. 14.Y. Sharma and P.-Y. Chen. Breaking the Madry Defense model with L1-based adversarial examples. arXiv preprint arXiv:1710.10733, 2017.
  15. 15.C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations, 2014.
  16. 16.F. Tramèr, A. Kurakin, N. Papernot, D. Boneh, and P. McDaniel. Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204, 2017.

Citation

MLA
Brown, T. B., et al. “Adversarial Patch”. arXiv, 2017, http://arxiv.org/abs/1712.09665v2.
APA
Brown, T. B., Mané, D., Roy, A., Abadi, M., & Gilmer, J. (2017). Adversarial Patch. arXiv. http://arxiv.org/abs/1712.09665v2
Chicago
Brown, T. B., D. Mané, A. Roy, M. Abadi, and J. Gilmer. 2017. “Adversarial Patch”. arXiv. http://arxiv.org/abs/1712.09665v2.
Harvard
Brown, T.B. et al. (2017) “Adversarial Patch”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1712.09665v2.
Vancouver
1. Brown TB, Mané D, Roy A, Abadi M, Gilmer J (2017) Adversarial Patch. arXiv

BibTeX

@article{brown2017adversarial,
  title = {Adversarial Patch},
  author = {Brown, Tom B. and Mané, Dandelion and Roy, Aurko and Abadi, Martín and Gilmer, Justin},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1712.09665v2},
  eprint = {1712.09665}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission