Adversarial Patch
Tom B. BrownDandelion ManéAurko RoyMartín AbadiJustin Gilmer
Introduces printable, scene-agnostic adversarial patches capable of overriding deep learning image classifiers in the physical world to force the prediction of an arbitrary target class.
Most modern security research on image-recognition systems focuses on defending against subtle, imperceptible modifications to digital images. However, real-world automated systems—ranging from surveillance cameras to autonomous vehicles—frequently operate without human oversight, leaving them vulnerable to physical modifications that an attacker does not need to hide. The article addresses this operational blind spot by evaluating whether visible, physical artifacts can reliably force image classifiers to misclassify scenes.
The article demonstrates a method for generating printable, targeted "adversarial patches" that consistently cause image classifiers to output a chosen target category regardless of the surrounding background or scene contents. Unlike traditional digital attacks, these patches are scene-independent, meaning an attacker does not need prior knowledge of lighting, camera angles, or the background objects being photographed.
To create these patches, the researchers adapted an optimization framework that trains a localized image patch across a wide distribution of background images, positions, scales, and rotations. They evaluated the approach on standard benchmark vision models (including ResNet50, InceptionV3, and VGG variants) across white-box settings where internal model parameters are known, and black-box settings where the patch is trained on a subset of models and evaluated on an unseen model. The team also conducted physical experiments by printing the generated patches with a standard desktop printer and photographing them alongside real objects.
The findings show that printable patches reliably deceive target classifiers across varying environments. First, when placed beside a real-world object (such as a banana originally identified with 97% confidence), the addition of a printed sticker targeted to a "toaster" class forced the model to output "toaster" with 99% confidence, completely ignoring the primary object. Second, patches generalized effectively across multiple model architectures in digital black-box evaluations, substantially outperforming natural images of the target class inserted at the same size. Third, researchers successfully camouflaged the patches to resemble abstract art or tie-dye peace signs while retaining their ability to mislead the models.
These results demonstrate a critical operational security risk. Because the patch does not depend on a specific scene, an attacker can create a single design offline and distribute it widely for third parties to print and deploy as physical stickers. The findings also reveal that existing defense strategies, which focus primarily on minor digital perturbations, fail to protect computer vision systems against localized, high-salience physical modifications.
To address this vulnerability, security and engineering teams deploying automated vision systems must expand defensive evaluations beyond small pixel-level noise to include localized, large-magnitude physical attacks. Organizations should explore multi-object detection and segmentation architectures that reduce reliance on global image classification, as standard classifiers simply output the single most salient feature detected. While the attack is highly effective in controlled environments and transfers across models, its effectiveness in black-box physical settings currently decreases unless the patch occupies a relatively large portion of the frame or is optimized for print quality. Further testing on diverse real-world cameras and specialized print processes is necessary before fully assessing threat levels against specific production deployments.
- Paper: Adversarial examples in the physical world, Alexey Kurakin et al. (2016). This work provides foundational evidence that adversarial examples survive printing and physical photography, directly motivating the development of printable physical adversarial patches.
- Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). It introduces image-agnostic universal adversarial perturbations, establishing the core concept of scene-independent fooling that the Adversarial Patch generalizes into a localized physical form.
- Paper: Synthesizing Robust Adversarial Examples, Anish Athalye et al. (2017). It establishes the Expectation Over Transformation framework used to optimize adversarial examples over a distribution of real-world transformations such as viewing angles, lighting, and scaling.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This seminal paper explains the linearity underlying adversarial vulnerability and introduces standard gradient-based optimization techniques essential for crafting targeted attacks.
- Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). It formalizes powerful optimization-based formulations for generating targeted adversarial attacks across standard distance metrics.
- Paper: The Limitations of Deep Learning in Adversarial Settings, Nicolas Papernot et al. (2015). It explores targeted adversarial attacks that restrict perturbations to small subsets of input features using saliency maps.
- Paper: Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples, Anish Athalye et al. (2018). This work demonstrates how transformation-robust techniques like Expectation Over Transformation can bypass apparent defenses relying on obfuscated gradients.
- Paper: Learning Robust Global Representations by Penalizing Local Predictive Power, Haohan Wang et al. (2019). It develops defense mechanisms that penalize local predictive cues and patch-level representations to prevent localized adversarial features from dominating classification.
- Paper: Countering Adversarial Images using Input Transformations, Chuan Guo et al. (2018). It examines input transformations such as image quilting and cropping as potential defenses against localized and global adversarial perturbations.
- Paper: Improving Transferability of Adversarial Examples With Input Diversity, Cihang Xie et al. (2018). It expands upon transformation-based adversarial optimization to significantly improve black-box attack transferability across architectures.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). It provides a theoretical and empirical framework explaining how models learn non-robust, highly predictive features like those weaponized in adversarial patches.
- Paper: Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks, Francesco Croce et al. (2020). It establishes standardized parameter-free benchmark evaluations to systematically assess model defenses against diverse adversarial perturbations.
