Synthesizing Robust Adversarial Examples

Anish AthalyeLogan EngstromAndrew IlyasKevin Kwok

article2017ICML1,894 citations

Develops an optimization algorithm for creating 3D-printed physical objects that reliably deceive neural network classifiers across diverse viewpoints and environmental transformations.

Listen

Machine learning vision systems are increasingly deployed in real-world, safety-critical applications such as autonomous vehicles, automated surveillance, and robotics. Previous research demonstrated that computer vision classifiers could be fooled by subtle, engineered digital perturbations known as adversarial examples; however, standard attack methods failed when applied to the physical world because camera noise, changing lighting, and shifting viewing angles broke the adversarial effect. The article addresses this gap by investigating whether physical, three-dimensional objects can be reliably engineered to fool neural network image classifiers across diverse real-world viewing conditions.

The article develops and evaluates a computational framework called Expectation Over Transformation, designed to synthesize adversarial examples that remain effective across an entire distribution of transformations rather than a single fixed image. The authors formulated an optimization method using projected gradient descent to maximize classification into an adversarial target category while constraining perceptible visual changes. To generate physical objects, the authors integrated a differentiable 3D rendering pipeline that accounts for camera distances, translations, 3D rotations, illumination changes, and full-color 3D printing inaccuracies. They evaluated this approach digitally across 1,000 two-dimensional images and 200 simulated three-dimensional models, and physically by 3D-printing real objects—specifically a turtle model targeted to be classified as a rifle and a baseball model targeted to be classified as an espresso—evaluating each through 100 manual photographs taken from varying angles.

The findings show that adversarial vulnerability poses a genuine, practical threat to physical systems. In two-dimensional digital experiments, optimized adversarial images achieved an average adversarial target classification rate of 96.4% across 1,000 random transformations. In three-dimensional simulations across ten object categories, adversarial textures attained an average target classification rate of 83.4%. Crucially, physical 3D-printed objects successfully transferred these vulnerabilities to the real world: the physical turtle was classified as a rifle in 82% of photographed poses (and misclassified overall in 98% of views), while the physical baseball was classified as espresso in 59% of poses (and misclassified overall in 90% of views). Even when the vision system failed to predict the specific adversarial target, it consistently misclassified the items into semantically related categories rather than the true object category.

These results establish that physical adversarial objects are practical and that vision models cannot rely on viewpoint variability or environmental noise for defense. Consequently, existing security strategies that rely purely on input transformations—such as random cropping, resizing, or image jittering—are fundamentally insufficient to protect machine learning models. System designers and security teams deploying computer vision in physical settings should not assume physical-world barriers provide security. Organizations must move beyond ad hoc input transformations and prioritize formal, robust defenses and certified model architectures during model design and validation.

The approach operates under certain boundaries. The optimization requires white-box access to the target model's architecture and gradients during creation, and larger distributions of transformations necessitate larger, more perceptible visual perturbations to the object's surface texture. Nevertheless, because the physical demonstrations succeeded even when using standard, commercial-grade 3D printers with known color errors, confidence is high that deep neural networks deployed in physical environments face tangible operational security risks.

arXiv: 1707.07397
  • Paper: Adversarial examples in the physical world, Alexey Kurakin et al. (2016). This work first demonstrated that printed adversarial examples could persist across physical transformations like camera photography, establishing the core problem that the Expectation Over Transformation framework directly formalizes and solves.
  • Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This foundational paper establishes the linearity hypothesis and gradient-based perturbation techniques that underly standard digital adversarial attack formulations.
  • Paper: Intriguing properties of neural networks, Christian Szegedy et al. (2014). This foundational work originally discovered the phenomenon of adversarial perturbations in deep neural networks, initiating the field.
  • Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). This paper establishes the standard optimization formulations for generating bounded adversarial perturbations, providing mathematical foundations adapted by subsequent attack algorithms.
  • Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). This work introduced the concept of perturbations designed to remain effective across multiple instances, motivating the search for attacks invariant to transformation distributions.
  • Paper: Spatial Transformer Networks, Max Jaderberg et al. (2015). This work develops differentiable spatial transformations, establishing the mathematical mechanics needed to backpropagate gradients through geometric image transformations.
Cover for Synthesizing Robust Adversarial Examples

Abstract

Standard methods for generating adversarial examples for neural networks do not consistently fool neural network classifiers in the physical world due to a combination of viewpoint shifts, camera noise, and other natural transformations, limiting their relevance to real-world systems. We demonstrate the existence of robust 3D adversarial objects, and we present the first algorithm for synthesizing examples that are adversarial over a chosen distribution of transformations. We synthesize two-dimensional adversarial images that are robust to noise, distortion, and affine transformation. We apply our algorithm to complex three-dimensional objects, using 3D-printing to manufacture the first physical adversarial objects. Our results demonstrate the existence of 3D adversarial objects in the physical world.

Table of Contents

  • 1 Introduction
  • 1.1 Challenges
  • 1.2 Contributions
  • 2 Approach
  • 2.1 Expectation Over Transformation
  • 2.2 Choosing a distribution of transformations
  • 2.2.1 2D case
  • 2.2.2 3D case
  • 2.3 Optimizing the objective
  • 3 Evaluation
  • 3.1 Procedure
  • 3.2 Robust 2D adversarial examples
  • 3.3 Robust 3D adversarial examples
  • 3.4 Physical adversarial examples
  • 3.5 Discussion
  • 4 Related Work
  • 4.1 Adversarial examples
  • 4.2 Defenses
  • 4.3 Physical-world adversarial examples
  • 5 Conclusion
  • References
  • A Distributions of Transformations
  • B Robust 2D Adversarial Examples
  • C Robust 3D Adversarial Examples
  • D Physical Adversarial Examples

Knowls

  1. Knowl 1 — Expectation Over Transformation Framework

    model/method

    Expectation Over Transformation (EOT) is an optimization framework for synthesizing adversarial examples that remain effective across a specified distribution of transformations T\mathcal{T}. Let x∈X⊂[0,1]dx \in \mathcal{X} \subset [0, 1]^d be an original input, yt∈Yy_t \in \mathcal{Y} be a target adversarial class, and P(y∣⋅)P(y \mid \cdot) denote the probability assigned by a neural network classifier to class yy. Standard adversarial attacks optimize P(yt∣x′)P(y_t \mid x') over a single static input x′x'. In contrast, EOT models real-world variations by drawing transformation functions t∼Tt \sim \mathcal{T} that map the adversary-controlled parameter x′x' (such as an image or a 3D texture) to the perceived input t(x′)t(x') fed to the classifier.

    EOT defines the expected perceived distance between the adversarial and original inputs under a distance metric d(⋅,⋅)d(\cdot, \cdot) as:

    δ=Et∼T[d(t(x′),t(x))]\delta = \mathbb{E}_{t \sim \mathcal{T}}[d(t(x'), t(x))]

    The general constrained optimization problem solved by EOT is:

    arg⁡max⁡x′Et∼T[log⁡P(yt∣t(x′))]subject toEt∼T[d(t(x′),t(x))]<ϵ,x′∈[0,1]d\arg\max_{x'} \mathbb{E}_{t \sim \mathcal{T}}[\log P(y_t \mid t(x'))] \quad \text{subject to} \quad \mathbb{E}_{t \sim \mathcal{T}}[d(t(x'), t(x))] < \epsilon, \quad x' \in [0, 1]^d

    Optimization is conducted via stochastic gradient descent, where gradients of the expectation are estimated by sampling independent transformation functions t∼Tt \sim \mathcal{T} at each gradient step and differentiating directly through tt.

  2. Knowl 2 — Lagrangian EOT Optimization with CIELAB Perceptual Metric

    equation

    To optimize the Expectation Over Transformation objective without computing explicit projections for complex distance constraints, the constrained problem is converted into a Lagrangian relaxation. The objective function is formulated as:

    arg⁡max⁡x′Et∼T[log⁡P(yt∣t(x′))−λ∥LAB(t(x′))−LAB(t(x))∥2]\arg\max_{x'} \mathbb{E}_{t \sim \mathcal{T}}\left[ \log P(y_t \mid t(x')) - \lambda \|\text{LAB}(t(x')) - \text{LAB}(t(x))\|_2 \right]

    where:

    • x′∈[0,1]dx' \in [0, 1]^d is the adversarial parameter (e.g., image pixels or surface texture) being optimized.
    • x∈[0,1]dx \in [0, 1]^d is the original, unperturbed input.
    • yty_t is the target adversarial class.
    • P(yt∣⋅)P(y_t \mid \cdot) is the classifier's predicted probability for class yty_t.
    • T\mathcal{T} is the distribution of transformations (e.g., affine transformations, rendering, lighting shifts, and sensor noise).
    • λ>0\lambda > 0 is a Lagrange multiplier weighting perceptual distortion against targeted classification confidence.
    • LAB(⋅)\text{LAB}(\cdot) denotes the mapping from RGB color space to the CIELAB (LAB) color space, where Euclidean (L2L_2) distance serves as a proxy for human visual perceptibility.

    The objective is maximized using projected gradient descent, clipping values of x′x' at each iteration to maintain valid input ranges (e.g., [0,1][0, 1]).

  3. Knowl 3 — Differentiable 3D Texture Rendering in EOT

    model/method

    To generate three-dimensional adversarial objects under the Expectation Over Transformation framework, adversarial perturbations are applied directly to a 2D surface texture pattern xx. The transformation function t(x)t(x) renders a 2D pose of the textured 3D mesh given specific transformation parameters (camera angle, lighting, 3D rotation, translation, perspective projection, and background).

    For a fixed set of pose and scene parameters, the 3D rendering operation acts linearly on the texture pixels:

    t(x)=Mx+bt(x) = M x + b

    where MM is a sparse coordinate mapping matrix that maps texture coordinates to screen pixel coordinates, and bb is a vector representing the rendered background and ambient offsets.

    Rather than executing full automatic differentiation through an entire 3D rendering engine, the standard rendering pipeline is modified to output the texture-space coordinate correspondence (M,b)(M, b) for each sampled viewpoint. Gradients with respect to the texture xx are then computed analytically through the linear transformation Mx+bMx + b. Because EOT samples new poses at every gradient descent iteration, MM and bb are recomputed at each step.

  4. Knowl 4 — Transformation Distributions for 2D, Simulated 3D, and Physical 3D Domains

    experimental setup

    The Expectation Over Transformation (EOT) framework parameterizes the transformation distribution T\mathcal{T} according to the target operating domain. Across all settings, parameters are sampled uniformly at random over specified ranges:

    1. 2D Transformations:

      • Scale factor: [0.9,1.4][0.9, 1.4]
      • Rotation angle: [−22.5∘,22.5∘][-22.5^\circ, 22.5^\circ]
      • Additive lighting offset: [−0.05,0.05][-0.05, 0.05]
      • Additive Gaussian noise standard deviation σ\sigma: [0.0,0.1][0.0, 0.1]
      • Translation: all in-bounds translations
    2. 3D Simulation Transformations:

      • Camera distance: [2.5,3.0][2.5, 3.0]
      • X/Y plane translation: [−0.05,0.05][-0.05, 0.05]
      • 3D Rotation: full sphere (arbitrary 3D orientation)
      • Solid background color (RGB): [0.1,1.0]3[0.1, 1.0]^3
    3. Physical 3D Proxy Transformations (incorporating camera, lighting, and 3D printing variation):

      • Camera distance: [2.5,3.0][2.5, 3.0]
      • X/Y plane translation: [−0.05,0.05][-0.05, 0.05]
      • 3D Rotation: full sphere
      • Solid background color (RGB): [0.1,1.0]3[0.1, 1.0]^3
      • Additive lighting: [−0.15,0.15][-0.15, 0.15]
      • Multiplicative lighting factor: [0.5,2.0][0.5, 2.0]
      • Per-channel additive color shift: [−0.15,0.15][-0.15, 0.15]
      • Per-channel multiplicative color shift: [0.7,1.3][0.7, 1.3]
      • Additive Gaussian noise standard deviation σ\sigma: [0.0,0.1][0.0, 0.1]
  5. Knowl 5 — Classification Performance and Adversariality of 2D EOT Examples

    data/table

    The effectiveness of 2D adversarial examples generated by EOT was evaluated on 1,000 images sampled from the ImageNet validation set against a pre-trained InceptionV3 model (clean top-1 accuracy 78.0%). Each adversarial image was generated targeting a randomly selected class and tested across 1,000 randomly sampled transformations drawn from the 2D transformation distribution.

    Images Classification Accuracy Adversariality ℓ2\ell_2 LAB
    mean stdev mean stdev mean
    Original 70.0% 36.4% 0.01% 0.3% 0
    Adversarial 0.9% 2.0% 96.4% 4.4% 5.6×10−55.6 \times 10^{-5}

    Classification accuracy measures the fraction of transformed images correctly classified as the original class, while adversariality measures the fraction classified as the target adversarial class. EOT produces 2D adversarial examples that achieve a mean target adversariality of 96.4% across geometric and visual transformations.

  6. Knowl 6 — Simulated 3D Object Adversariality under EOT

    data/table

    Adversarial surface textures were synthesized using EOT for 10 distinct 3D object models (barrel, baseball, dog, orange, turtle, clownfish, sofa, teddy bear, car, and taxi) targeting InceptionV3. For each 3D model, 20 random target classes were tested (200 adversarial pairs total). During optimization, batches of size 40 were used with pose reuse up to 80% (ensuring at least 8 new poses per step). At test time, each adversarial model was evaluated over 100 randomly sampled poses from the 3D simulation transformation distribution.

    Images Classification Accuracy Adversariality ℓ2\ell_2 LAB
    mean stdev mean stdev mean
    Original 68.8% 31.2% 0.01% 0.1% 0
    Adversarial 1.1% 3.1% 83.4% 21.7% 5.9×10−35.9 \times 10^{-3}

    Across the 200 synthesized 3D objects, the adversarial textures achieved an 83.4% mean adversariality across viewpoints in simulation, reducing true-class classification accuracy to 1.1%.

  7. Knowl 7 — Physical-World Robustness of 3D-Printed Adversarial Objects

    data/table

    Two physical adversarial objects were manufactured using commercial full-color 3D printing after optimizing their surface textures using EOT over the physical-world proxy transformation distribution: a 3D turtle mesh optimized to classify as 'rifle' and a 3D baseball mesh optimized to classify as 'espresso'. Evaluation was conducted by capturing 100 unconstrained real-world photographs of each physical object across diverse camera angles, distances, backgrounds, and ambient lighting conditions, classified using InceptionV3.

    Object Adversarial Target Misclassified (Other) Correct Source Class
    Turtle (Target: Rifle) 82% 16% 2%
    Baseball (Target: Espresso) 59% 31% 10%

    Unperturbed 3D-printed versions of the turtle and baseball achieved 100% correct source class classification under the same physical test conditions. The adversarial 3D-printed objects successfully induced targeted misclassification in the physical world across the majority of viewpoints (82% for the turtle, 59% for the baseball) and resulted in total misclassification rates of 98% and 90%, respectively.

  8. Knowl 8 — Semantic Concentration of Off-Target Misclassifications

    empirical result

    When physical 3D adversarial objects fail to be classified as their exact target class yadvy_{\text{adv}}, the classifier rarely defaults to the true source class of the physical object. Instead, the top predicted classes concentrate on labels semantically related to the adversarial target class.

    Specifically:

    • For the physical 3D-printed turtle optimized with the target class 'rifle', viewpoints that were not classified as 'rifle' were predominantly classified as 'revolver', 'holster', or 'assault rifle'.
    • For the physical 3D-printed baseball optimized with the target class 'espresso', non-target predictions were predominantly classified as 'coffee' or 'bakery'.
  9. Knowl 9 — Failure Modes and Transformation Distribution Limits of EOT

    limitation

    The Expectation Over Transformation (EOT) optimization framework exhibits two structural failure modes:

    1. Over-constrained perturbation budget: If the distance constraint bound ϵ\epsilon is set too small, or the Lagrangian penalty λ\lambda is set too large, the optimizer cannot discover a perturbation pattern that remains adversarial over varying transformation samples.
    2. Excessively broad transformation distribution: If the chosen distribution of transformations T\mathcal{T} has support over operations that destroy discriminatory image features (e.g., replacing every pixel with independent uniform noise U[0,1]\mathcal{U}[0, 1]), no static perturbation x′x' can remain invariant and simultaneously adversarial across the distribution.

Coverage note — None was omitted; all key contributions including EOT mathematical formulation, differentiable rendering approach, experimental setups, simulated/physical quantitative tables, semantic findings, and failure modes are fully captured.

References

  1. 1.Athalye, A., Carlini, N., and Wagner, D. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. 2018. URL https://arxiv.org/abs/1802.00420.
  2. 2.Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., Giacinto, G., and Roli, F. Evasion attacks against machine learning at test time. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pp. 387–402. Springer, 2013.
  3. 3.Brown, T. B., Mané, D., Roy, A., Abadi, M., and Gilmer, J. Defensive distillation is not robust to adversarial examples. 2016. URL https://arxiv.org/abs/1607.04311.
  4. 4.Buckman, J., Roy, A., Raffel, C., and Goodfellow, I. Thermometer encoding: One hot way to resist adversarial examples. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=S18Su--CW. accepted as poster.
  5. 5.Carlini, N. and Wagner, D. Defensive distillation is not robust to adversarial examples. 2016. URL https://arxiv.org/abs/1607.04311.
  6. 6.Carlini, N. and Wagner, D. Adversarial examples are not easily detected: Bypassing ten detection methods. AISec, 2017a.
  7. 7.Carlini, N. and Wagner, D. Magnet and “efficient defenses against adversarial attacks” are not robust to adversarial examples. arXiv preprint arXiv:1711.08478, 2017b.
  8. 8.Carlini, N. and Wagner, D. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security & Privacy, 2017c.
  9. 9.Carlini, N., Mishra, P., Vaidya, T., Zhang, Y., Sherr, M., Shields, C., Wagner, D., and Zhou, W. Hidden voice commands. In 25th USENIX Security Symposium (USENIX Security 16), pp. 513–530, Austin, TX, 2016. USENIX Association. ISBN 978-1-931971-32-4. URL https://www.usenix.org/conference/usenixsecurity16/technical-sessions/presentation/carlini.
  10. 10.Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., and Hsieh, C.-J. Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, AISec ’17, pp. 15–26, New York, NY, USA, 2017. ACM. ISBN 978-1-4503-5202-4. doi: 10.1145/3128572.3140448. URL http://doi.acm.org/10.1145/3128572.3140448.
  11. 11.Dhillon, G. S., Azizzadenesheli, K., Bernstein, J. D., Kossaifi, J., Khanna, A., Lipton, Z. C., and Anandkumar, A. Stochastic activation pruning for robust adversarial defense. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=H1uR4GZRZ. accepted as poster.
  12. 12.Evtimov, I., Eykholt, K., Fernandes, E., Kohno, T., Li, B., Prakash, A., Rahmati, A., and Song, D. Robust Physical-World Attacks on Deep Learning Models. 2017. URL https://arxiv.org/abs/1707.08945.
  13. 13.Goodfellow, I. J., Shlens, J., and Szegedy, C. Explaining and harnessing adversarial examples. In Proceedings of the International Conference on Learning Representations (ICLR), 2015.
  14. 14.Guo, C., Rana, M., Cisse, M., and van der Maaten, L. Countering adversarial images using input transformations. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=SyJ7ClWCb. accepted as poster.
  15. 15.Hendrik Metzen, J., Genewein, T., Fischer, V., and Bischoff, B. On detecting adversarial perturbations. In International Conference on Learning Representations, 2017.
  16. 16.Hendrycks, D. and Gimpel, K. Early methods for detecting adversarial images. In International Conference on Learning Representations (Workshop Track), 2017.
  17. 17.Kurakin, A., Goodfellow, I., and Bengio, S. Adversarial examples in the physical world. 2016. URL https://arxiv.org/abs/1607.02533.
  18. 18.Lu, J., Sibai, H., Fabry, E., and Forsyth, D. No need to worry about adversarial examples in object detection in autonomous vehicles. 2017. URL https://arxiv.org/abs/1707.03501.
  19. 19.Luo, Y., Boix, X., Roig, G., Poggio, T., and Zhao, Q. Foveation-based mechanisms alleviate adversarial examples. 2016. URL https://arxiv.org/abs/1511.06292.
  20. 20.Ma, X., Li, B., Wang, Y., Erfani, S. M., Wijewickrema, S., Schoenebeck, G., Houle, M. E., Song, D., and Bailey, J. Characterizing adversarial subspaces using local intrinsic dimensionality. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=B1gJ1L2aW. accepted as oral presentation.
  21. 21.Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. Towards deep learning models resistant to adversarial attacks. 2017. URL https://arxiv.org/abs/1706.06083.
  22. 22.McLaren, K. Xiiithe development of the cie 1976 (l* a* b*) uniform colour space and colourdifference formula. Journal of the Society of Dyers and Colourists, 92(9):338–341, September 1976. doi: 10.1111/j.1478-4408.1976.tb03301.x. URL https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1478-4408.1976.tb03301.x.
  23. 23.Meng, D. and Chen, H. MagNet: a two-pronged defense against adversarial examples. In ACM Conference on Computer and Communications Security (CCS), 2017. arXiv preprint arXiv:1705.09064.
  24. 24.Moosavi-Dezfooli, S., Fawzi, A., and Frossard, P. Deepfool: a simple and accurate method to fool deep neural networks. CoRR, abs/1511.04599, 2015. URL http://arxiv.org/abs/1511.04599.
  25. 25.Moosavi-Dezfooli, S.-M., Fawzi, A., Fawzi, O., and Frossard, P. Universal adversarial perturbations. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  26. 26.Papernot, N., McDaniel, P., and Goodfellow, I. Transferability in machine learning: from phenomena to blackbox attacks using adversarial samples. 2016a. URL https://arxiv.org/abs/1605.07277.
  27. 27.Papernot, N., McDaniel, P., Jha, S., Fredrikson, M., Celik, Z. B., and Swami, A. The limitations of deep learning in adversarial settings. In IEEE European Symposium on Security & Privacy, 2016b.
  28. 28.Papernot, N., McDaniel, P., Wu, X., Jha, S., and Swami, A. Distillation as a defense to adversarial perturbations against deep neural networks. In Security and Privacy (SP), 2016 IEEE Symposium on, pp. 582–597. IEEE, 2016c.
  29. 29.Papernot, N., McDaniel, P., Goodfellow, I., Jha, S., Celik, Z. B., and Swami, A. Practical black-box attacks against machine learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, ASIA CCS ’17, pp. 506–519, New York, NY, USA, 2017. ACM. ISBN 978-1-4503-4944-4. doi: 10.1145/3052973.3053009. URL http://doi.acm.org/10.1145/3052973.3053009.
  30. 30.Samangouei, P., Kabkab, M., and Chellappa, R. Defensegan: Protecting classifiers against adversarial attacks using generative models. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=BkJ3ibb0-. accepted as poster.
  31. 31.Sharif, M., Bhagavatula, S., Bauer, L., and Reiter, M. K. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, CCS ’16, pp. 1528–1540, New York, NY, USA, 2016. ACM. ISBN 978-1-4503-4139-4. doi: 10.1145/2976749.2978392. URL http://doi.acm.org/10.1145/2976749.2978392.
  32. 32.Song, Y., Kim, T., Nowozin, S., Ermon, S., and Kushman, N. Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJUYGxbCW. accepted as poster.
  33. 33.Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. Intriguing properties of neural networks. 2013. URL https://arxiv.org/abs/1312.6199.
  34. 34.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the inception architecture for computer vision. 2015. URL https://arxiv.org/abs/1512.00567.
  35. 35.Xie, C., Wang, J., Zhang, Z., Ren, Z., and Yuille, A. Mitigating adversarial effects through randomization. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=Sk9yuql0Z. accepted as poster.
  36. 36.Zantedeschi, V., Nicolae, M.-I., and Rawat, A. Efficient defenses against adversarial attacks. arXiv preprint arXiv:1707.06728, 2017.

Citation

MLA
Athalye, A., et al. “Synthesizing Robust Adversarial Examples”. arXiv, 2017, http://arxiv.org/abs/1707.07397v3.
APA
Athalye, A., Engstrom, L., Ilyas, A., & Kwok, K. (2017). Synthesizing Robust Adversarial Examples. arXiv. http://arxiv.org/abs/1707.07397v3
Chicago
Athalye, A., L. Engstrom, A. Ilyas, and K. Kwok. 2017. “Synthesizing Robust Adversarial Examples”. arXiv. http://arxiv.org/abs/1707.07397v3.
Harvard
Athalye, A. et al. (2017) “Synthesizing Robust Adversarial Examples”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1707.07397v3.
Vancouver
1. Athalye A, Engstrom L, Ilyas A, Kwok K (2017) Synthesizing Robust Adversarial Examples. arXiv

BibTeX

@article{athalye2017synthesizing,
  title = {Synthesizing Robust Adversarial Examples},
  author = {Athalye, Anish and Engstrom, Logan and Ilyas, Andrew and Kwok, Kevin},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1707.07397v3},
  eprint = {1707.07397}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/