Synthesizing Robust Adversarial Examples
Anish AthalyeLogan EngstromAndrew IlyasKevin Kwok
Develops an optimization algorithm for creating 3D-printed physical objects that reliably deceive neural network classifiers across diverse viewpoints and environmental transformations.
Machine learning vision systems are increasingly deployed in real-world, safety-critical applications such as autonomous vehicles, automated surveillance, and robotics. Previous research demonstrated that computer vision classifiers could be fooled by subtle, engineered digital perturbations known as adversarial examples; however, standard attack methods failed when applied to the physical world because camera noise, changing lighting, and shifting viewing angles broke the adversarial effect. The article addresses this gap by investigating whether physical, three-dimensional objects can be reliably engineered to fool neural network image classifiers across diverse real-world viewing conditions.
The article develops and evaluates a computational framework called Expectation Over Transformation, designed to synthesize adversarial examples that remain effective across an entire distribution of transformations rather than a single fixed image. The authors formulated an optimization method using projected gradient descent to maximize classification into an adversarial target category while constraining perceptible visual changes. To generate physical objects, the authors integrated a differentiable 3D rendering pipeline that accounts for camera distances, translations, 3D rotations, illumination changes, and full-color 3D printing inaccuracies. They evaluated this approach digitally across 1,000 two-dimensional images and 200 simulated three-dimensional models, and physically by 3D-printing real objects—specifically a turtle model targeted to be classified as a rifle and a baseball model targeted to be classified as an espresso—evaluating each through 100 manual photographs taken from varying angles.
The findings show that adversarial vulnerability poses a genuine, practical threat to physical systems. In two-dimensional digital experiments, optimized adversarial images achieved an average adversarial target classification rate of 96.4% across 1,000 random transformations. In three-dimensional simulations across ten object categories, adversarial textures attained an average target classification rate of 83.4%. Crucially, physical 3D-printed objects successfully transferred these vulnerabilities to the real world: the physical turtle was classified as a rifle in 82% of photographed poses (and misclassified overall in 98% of views), while the physical baseball was classified as espresso in 59% of poses (and misclassified overall in 90% of views). Even when the vision system failed to predict the specific adversarial target, it consistently misclassified the items into semantically related categories rather than the true object category.
These results establish that physical adversarial objects are practical and that vision models cannot rely on viewpoint variability or environmental noise for defense. Consequently, existing security strategies that rely purely on input transformations—such as random cropping, resizing, or image jittering—are fundamentally insufficient to protect machine learning models. System designers and security teams deploying computer vision in physical settings should not assume physical-world barriers provide security. Organizations must move beyond ad hoc input transformations and prioritize formal, robust defenses and certified model architectures during model design and validation.
The approach operates under certain boundaries. The optimization requires white-box access to the target model's architecture and gradients during creation, and larger distributions of transformations necessitate larger, more perceptible visual perturbations to the object's surface texture. Nevertheless, because the physical demonstrations succeeded even when using standard, commercial-grade 3D printers with known color errors, confidence is high that deep neural networks deployed in physical environments face tangible operational security risks.
- Paper: Adversarial examples in the physical world, Alexey Kurakin et al. (2016). This work first demonstrated that printed adversarial examples could persist across physical transformations like camera photography, establishing the core problem that the Expectation Over Transformation framework directly formalizes and solves.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This foundational paper establishes the linearity hypothesis and gradient-based perturbation techniques that underly standard digital adversarial attack formulations.
- Paper: Intriguing properties of neural networks, Christian Szegedy et al. (2014). This foundational work originally discovered the phenomenon of adversarial perturbations in deep neural networks, initiating the field.
- Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). This paper establishes the standard optimization formulations for generating bounded adversarial perturbations, providing mathematical foundations adapted by subsequent attack algorithms.
- Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). This work introduced the concept of perturbations designed to remain effective across multiple instances, motivating the search for attacks invariant to transformation distributions.
- Paper: Spatial Transformer Networks, Max Jaderberg et al. (2015). This work develops differentiable spatial transformations, establishing the mathematical mechanics needed to backpropagate gradients through geometric image transformations.
- Paper: Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples, Anish Athalye et al. (2018). This work directly employs the Expectation Over Transformation technique introduced in the source to circumvent defenses that rely on randomized transformations and obfuscated gradients.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, A. Ma̧dry et al. (2017). This paper establishes the standard minimax robust optimization framework for training networks resistant to first-order adversarial attacks.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). This study provides a fundamental conceptual explanation for adversarial vulnerability by demonstrating that attacks exploit non-robust, human-incomprehensible features learned by standard classifiers.
- Paper: Ensemble Adversarial Training: Attacks and Defenses, Florian Tramèr et al. (2018). This paper analyzes the failure modes of single-step adversarial defenses and explores training against ensembles to improve resistance to transferred adversarial perturbations.
- Paper: Robustness May Be at Odds with Accuracy, Dimitris Tsipras et al. (2018). This work investigates the fundamental theoretical and empirical trade-offs between standard classification accuracy and robust defense against adversarial perturbations.
- Paper: Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks, Francesco Croce et al. (2020). This work introduces AutoAttack, creating a standardized, parameter-free ensemble of attacks to systematically evaluate the robustness of defended models.
- Paper: Certified Adversarial Robustness via Randomized Smoothing, Jeremy M Cohen et al. (2019). This paper introduces randomized smoothing to provide provable, certified robustness guarantees against continuous norm-bounded perturbations.
- Paper: Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, Dan Hendrycks et al. (2019). This work expands beyond synthetic adversarial attacks to establish comprehensive benchmarks for evaluating neural network robustness against common natural corruptions and physical perturbations.
