Adversarial Machine Learning at Scale

Alexey KurakinIan GoodfellowSamy Bengio

article2016ICLR3,436 citations

Establishes practical methods to scale adversarial training to large datasets like ImageNet, resolving the label leaking problem and uncovering critical differences in attack transferability.

Listen

Adversarial examples are small, often imperceptible changes to inputs that cause machine learning models to misclassify them, and they frequently transfer across different models to enable attacks without access to the target system’s parameters. This vulnerability creates security risks for deployed systems, yet prior defenses such as adversarial training had been demonstrated only on small datasets like MNIST and CIFAR-10.

The work set out to scale adversarial training to large models on the ImageNet dataset and to measure resulting robustness against both one-step and iterative attack methods. Researchers trained Inception v3 models using synchronous distributed training across 50 machines, mixing clean and adversarially perturbed examples in each minibatch, and evaluated performance on the 50,000-image ImageNet validation set.

Adversarial training with one-step methods raised top-1 accuracy on one-step adversarial examples from roughly 30 percent to 73–75 percent while reducing clean-image accuracy by less than 1 percent. Larger models, created by increasing the number of filters, gained more robustness from the same procedure and showed higher ratios of adversarial to clean accuracy. Iterative attacks remained effective against the trained models, but those same iterative examples transferred to other models at markedly lower rates than one-step examples. A label-leaking artifact appeared when training and testing used the fast gradient sign method, artificially inflating accuracy on adversarial inputs; the effect disappeared when label-free one-step methods were substituted.

These results indicate that adversarial training can be applied at ImageNet scale to protect against practical one-step attacks with only modest cost to normal accuracy, while larger models amplify the benefit. The lower transferability of iterative attacks offers partial protection against black-box threats even when direct robustness is limited. Practitioners should therefore adopt adversarial training with non-label-leaking one-step methods during both training and evaluation, pair it with increased model capacity, and avoid relying on the fast gradient sign method for robustness testing.

Further work is needed to develop defenses that also resist iterative attacks, possibly through still-larger models or new training regimes, and to confirm whether the observed scaling trends continue beyond the architectures tested here. The study’s main limitations are its focus on a single dataset and family of models, the computational expense of iterative training, and the absence of exhaustive hyperparameter sweeps for very large networks.

arXiv: 1611.01236
  • Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). Reading this foundational 2015 study on adversarial examples and the fast gradient sign method is essential to understanding the core attack mechanisms scaled up in the source paper.
Cover for Adversarial Machine Learning at Scale

Abstract

Adversarial examples are malicious inputs designed to fool machine learning models. They often transfer from one model to another, allowing attackers to mount black box attacks without knowledge of the target model's parameters. Adversarial training is the process of explicitly training a model on adversarial examples, in order to make it more robust to attack or to reduce its test error on clean inputs. So far, adversarial training has primarily been applied to small problems. In this research, we apply adversarial training to ImageNet. Our contributions include: (1) recommendations for how to succesfully scale adversarial training to large models and datasets, (2) the observation that adversarial training confers robustness to single-step attack methods, (3) the finding that multi-step attack methods are somewhat less transferable than single-step attack methods, so single-step attacks are the best for mounting black-box attacks, and (4) resolution of a "label leaking" effect that causes adversarially trained models to perform better on adversarial examples than on clean examples, because the adversarial example construction process uses the true label and the model can learn to exploit regularities in the construction process.

Table of Contents

  • 1 INTRODUCTION
  • 2 METHODS GENERATING ADVERSARIAL EXAMPLES
  • 2.1 TERMINOLOGY AND NOTATION
  • 2.2 ATTACK METHODS
  • 3 ADVERSARIAL TRAINING
  • 4 EXPERIMENTS
  • 4.1 RESULTS OF ADVERSARIAL TRAINING
  • 4.2 LABEL LEAKING
  • 4.3 INFLUENCE OF MODEL CAPACITY ON ADVERSARIAL ROBUSTNESS
  • 4.4 TRANSFERABILITY OF ADVERSARIAL EXAMPLES
  • 5 CONCLUSION
  • A COMPARISON OF ONE-STEP ADVERSARIAL METHODS
  • B ADDITIONAL RESULTS WITH SIZE OF THE MODEL
  • C ADDITIONAL RESULTS ON TRANSFERABILITY
  • D RESULTS WITH DIFFERENT ACTIVATION FUNCTIONS
  • E RESULTS WITH DIFFERENT NUMBER OF ADVERSARIAL EXAMPLES IN THE MINIBATCH

Knowls

  1. Knowl 1 — Adversarial Training Algorithm with Mixed Minibatches and Dynamic Perturbations

    algorithm

    To scale adversarial training to large datasets such as ImageNet and architectures utilizing batch normalization, each training minibatch is constructed to contain both clean and newly generated adversarial examples. Generating adversarial examples at each step using a randomized perturbation magnitude prevents the network from overfitting to a single attack scale.

    Input: Neural network NN with parameters θ\theta, training dataset, batch size mm, number of adversarial examples per minibatch kk, perturbation magnitude distribution Dϵ\mathcal{D}_{\epsilon}
    Output: Adversarially trained network NN
    Initialize network parameters θ\theta randomly
    repeat
        Sample clean minibatch B={X1,…,Xm}B = \{X_1, \ldots, X_m\} from training set with labels {y1,…,ym}\{y_1, \ldots, y_m\}
        for i=1i = 1 to kk do
            Sample ϵi∼Dϵ\epsilon_i \sim \mathcal{D}_{\epsilon} independently
            Generate adversarial example XiadvX_i^{\text{adv}} from XiX_i using network NN at magnitude ϵi\epsilon_i
        end for
        Form mixed minibatch B′={X1adv,…,Xkadv,Xk+1,…,Xm}B' = \{X_1^{\text{adv}}, \ldots, X_k^{\text{adv}}, X_{k+1}, \ldots, X_m\}
        Perform one optimization update step on θ\theta using minibatch B′B'
    until training converged

    For ImageNet training with Inception v3, typical hyperparameters are minibatch size m=32m = 32, adversarial sample count k=16k = 16 (replacing the first kk clean examples), and optimization via RMSProp with an initial learning rate of 0.0450.045. The perturbation magnitude ϵ\epsilon (in pixel intensity units on [0,255][0, 255]) is drawn independently per example from a truncated normal distribution on [0,16][0, 16] derived from N(μ=0,σ=8)\mathcal{N}(\mu = 0, \sigma = 8) (specifically ∣Z∣|Z| where Z∼N(0,82)Z \sim \mathcal{N}(0, 8^2) truncated to [0,16][0, 16]).

  2. Knowl 2 — Mixed Clean and Adversarial Training Loss Objective

    equation

    To provide independent control over the number and relative regularization weight of adversarial examples during large-scale network optimization, the training loss is formulated as:

    L=1(m−k)+λk(∑i∈CLEANL(Xi∣yi)+λ∑i∈ADVL(Xiadv∣yi))\mathcal{L} = \frac{1}{(m - k) + \lambda k} \left( \sum_{i \in \text{CLEAN}} L(X_i \mid y_i) + \lambda \sum_{i \in \text{ADV}} L(X_i^{\text{adv}} \mid y_i) \right)

    where:

    • m∈N+m \in \mathbb{N}^+ is the total number of examples in the minibatch.
    • k∈{0,1,…,m}k \in \{0, 1, \ldots, m\} is the number of adversarial examples included in the minibatch.
    • CLEAN\text{CLEAN} is the index set of unmodified clean examples, with cardinality ∣CLEAN∣=m−k|\text{CLEAN}| = m - k.
    • ADV\text{ADV} is the index set of adversarial examples, with cardinality ∣ADV∣=k|\text{ADV}| = k.
    • XiX_i denotes the ii-th clean input image with corresponding ground truth class label yiy_i.
    • XiadvX_i^{\text{adv}} is the candidate adversarial image generated from XiX_i.
    • L(X∣y)L(X \mid y) is the cross-entropy classification loss of the network evaluated on input XX given target label yy.
    • λ∈R+\lambda \in \mathbb{R}^+ is a scaling hyperparameter controlling the relative weight of adversarial examples in the loss. For ImageNet models, setting λ=0.3\lambda = 0.3, m=32m = 32, and k=16k = 16 balances clean classification performance with robustness against adversarial attacks.
  3. Knowl 3 — One-Step Least-Likely Class Adversarial Attack Method

    model/method

    The one-step least-likely class attack ("step l.l.") generates adversarial perturbations by driving network predictions toward the class assigned the lowest probability by the model on the clean input, without utilizing the ground truth class label.

    Given a clean image XX, the network's predicted class conditional distribution p(y∣X)p(y \mid X), and the model cost function J(X,y)J(X, y), the least likely target class is:

    yLL=arg⁡min⁡yp(y∣X)y_{\text{LL}} = \arg\min_y p(y \mid X)

    The adversarial image XadvX^{\text{adv}} is generated by taking a single step of magnitude ϵ\epsilon in the direction that minimizes loss with respect to yLLy_{\text{LL}} (maximizing p(yLL∣X)p(y_{\text{LL}} \mid X)):

    Xadv=Clip⁡X,ϵ(X−ϵsign⁡(∇XJ(X,yLL)))X^{\text{adv}} = \operatorname{Clip}_{X, \epsilon}\left( X - \epsilon \operatorname{sign}\left(\nabla_X J(X, y_{\text{LL}})\right) \right)

    where ϵ\epsilon represents the maximum allowable L∞L_\infty perturbation magnitude (measured in pixel values in [0,255][0, 255]) and Clip⁡X,ϵ(⋅)\operatorname{Clip}_{X, \epsilon}(\cdot) restricts pixel intensity values to the range [Xi,j−ϵ,Xi,j+ϵ]∩[0,255][X_{i,j} - \epsilon, X_{i,j} + \epsilon] \cap [0, 255]. Because yLLy_{\text{LL}} is derived entirely from the model output rather than ground truth data, this formulation prevents label leaking when used as a defense or evaluation mechanism.

  4. Knowl 4 — Label Leaking Phenomenon in Adversarial Training

    definition

    Label leaking is an artifact of adversarial training where a model trained on adversarial examples constructed using the true class label ytruey_{\text{true}} (such as the Fast Gradient Sign Method, FGSM) learns to detect regularities in the perturbation rather than learning robust invariant features, causing its accuracy on adversarial examples to become artificially higher than its accuracy on clean examples.

    A label for a given input XX is defined as leaked if and only if the model correctly classifies an adversarial example Xadv(ytrue)X^{\text{adv}}(y_{\text{true}}) constructed with knowledge of the true label, but misclassifies a corresponding adversarial example generated without using the true label (such as via least-likely target classes or model predictions).

    The effect occurs because one-step perturbations using ytruey_{\text{true}} apply a simple additive transformation X+ϵsign⁡(∇XJ(X,ytrue))X + \epsilon \operatorname{sign}(\nabla_X J(X, y_{\text{true}})) that directly encodes information about ytruey_{\text{true}}. The model learns to reverse this transformation to recover the ground truth class. Label leaking is eliminated by generating adversarial training examples using methods that do not access ytruey_{\text{true}} (such as the one-step least-likely method, random target class methods, or predicted classes), or by using multi-step iterative attacks whose outputs are more complex and less predictable.

  5. Knowl 5 — Non-Robustness of Single-Step Adversarially Trained Models to Iterative Attacks

    limitation

    Adversarial training using single-step methods (such as "step l.l." or FGSM) confers substantial robustness against all single-step attack variations, but does not provide resistance against multi-step iterative attacks.

    When an Inception v3 network adversarially trained on ImageNet using the one-step least-likely method is evaluated against iterative attacks across perturbation magnitudes ϵ∈[2,16]\epsilon \in [2, 16] (pixel scale in [0,255][0, 255]):

    • Against the iterative least-likely attack ("iter. l.l."), top-1 accuracy falls from 77.4%77.4\% (clean) to 29.1%29.1\% at ϵ=2\epsilon = 2, 7.5%7.5\% at ϵ=4\epsilon = 4, 3.0%3.0\% at ϵ=8\epsilon = 8, and 1.5%1.5\% at ϵ=16\epsilon = 16.
    • Against the basic iterative attack ("iter. basic"), top-1 accuracy falls from 77.4%77.4\% (clean) to 30.0%30.0\% at ϵ=2\epsilon = 2, 25.2%25.2\% at ϵ=4\epsilon = 4, 23.5%23.5\% at ϵ=8\epsilon = 8, and 23.2%23.2\% at ϵ=16\epsilon = 16.

    Directly generating and training on multi-step iterative adversarial examples on large-scale datasets like ImageNet is computationally prohibitive and fails to produce robustness or maintain acceptable clean image accuracy without substantially larger model capacity.

  6. Knowl 6 — Transferability Asymmetry Between Single-Step and Multi-Step Attacks

    empirical result

    In black-box attack scenarios where adversarial examples generated on a source network are evaluated against a separate target network, single-step attacks transfer at substantially higher rates than multi-step iterative attacks, demonstrating an inverse relationship between white-box attack strength and transferability.

    For perturbations of size ϵ=16\epsilon = 16 generated on an Inception v3 source model and evaluated across independently initialized Inception v3 models, Inception v3 with ELU activations, and Inception v4:

    1. Fast Gradient Sign Method (FGSM, single-step) achieves moderate white-box error (63%−69%63\% - 69\%) on the source model but transfers with high target top-1 error rates (47%−58%47\% - 58\%).
    2. Basic Iterative Method achieves high white-box error but shows intermediate transfer error rates (33%−46%33\% - 46\%).
    3. Iterative Least-Likely Class method ("iter. l.l.") achieves over 99%99\% top-1 white-box misclassification on the source model, yet transfers poorly, causing only 9%−13%9\% - 13\% top-1 error on target models.

    This occurs because multi-step iterative attacks overfit to the local gradient landscapes and specific parameters of the source network, whereas single-step attacks produce broad directional perturbations that generalize across diverse architectures and parameter initializations. Furthermore, across all attack types, transfer rates increase monotonically as ϵ\epsilon increases.

  7. Knowl 7 — Impact of Network Capacity on Adversarial Robustness

    empirical result

    Scaling the capacity of a deep convolutional network (by multiplying the number of convolutional filters across all layers by a factor ρ∈[0.5,2.0]\rho \in [0.5, 2.0]) leads to distinct robustness behaviors depending on the training regime:

    1. Standard Models (No Adversarial Training): Robustness (measured as the ratio of adversarial accuracy to clean accuracy) peaks at an intermediate filter scale ρ\rho. Models that are either under-parameterized or over-parameterized exhibit lower robustness ratios, indicating that excessive model capacity without adversarial training leads to overfitting on the input manifold.
    2. Adversarially Trained Models: Robustness against one-step attacks increases monotonically with model capacity ρ\rho. For an Inception v3 model with ρ=2.0\rho = 2.0 (doubled filter count), the ratio of adversarial top-1 accuracy to clean top-1 accuracy approaches 1.01.0 across perturbation sizes ϵ∈[2,16]\epsilon \in [2, 16].
    3. Mitigating Clean Accuracy Loss: Increasing model depth (e.g., adding two Inception blocks) reduces the clean accuracy penalty of adversarial training from 0.8%0.8\% to 0.6%0.6\% while improving top-1 adversarial accuracy by 1.1%−1.9%1.1\% - 1.9\% across ϵ∈[2,16]\epsilon \in [2, 16].
  8. Knowl 8 — Iterative Least-Likely Class Adversarial Attack Method

    model/method

    The iterative least-likely class attack ("iter. l.l.") is a multi-step targeted attack that applies small gradient updates to drive network predictions toward the least likely class yLLy_{\text{LL}} within an L∞L_\infty ball of radius ϵ\epsilon.

    The recurrence relation is defined as:

    X0adv=XX_0^{\text{adv}} = X

    XN+1adv=Clip⁡X,ϵ{XNadv−αsign⁡(∇XJ(XNadv,yLL))}X_{N+1}^{\text{adv}} = \operatorname{Clip}_{X, \epsilon}\left\{ X_N^{\text{adv}} - \alpha \operatorname{sign}\left( \nabla_X J(X_N^{\text{adv}}, y_{\text{LL}}) \right) \right\}

    where:

    • XX is the clean input image.
    • yLL=arg⁡min⁡yp(y∣X)y_{\text{LL}} = \arg\min_y p(y \mid X) is the least likely predicted class evaluated on the clean input XX and held constant across iterations.
    • J(X,y)J(X, y) is the cross-entropy loss.
    • α\alpha is the step size per iteration, set to α=1\alpha = 1 (changing each pixel value by 1 in the range [0,255][0, 255] per step).
    • Clip⁡X,ϵ{A}\operatorname{Clip}_{X, \epsilon}\{A\} clips each pixel Ai,jA_{i,j} to [Xi,j−ϵ,Xi,j+ϵ]∩[0,255][X_{i,j} - \epsilon, X_{i,j} + \epsilon] \cap [0, 255].
    • The total number of iterations is N=min⁡(ϵ+4,⌊1.25ϵ⌋)N = \min(\epsilon + 4, \lfloor 1.25\epsilon \rfloor).

    This method achieves a white-box misclassification top-1 error rate exceeding 99%99\% on standard ImageNet classification models.

  9. Knowl 9 — Accuracy of Adversarially Trained Inception v3 Under Step Least-Likely Attack

    data/table

    Adversarial training on ImageNet using the one-step least-likely ("step l.l.") method increases classification accuracy against one-step attacks across various test-time perturbation magnitudes ϵ∈{2,4,8,16}\epsilon \in \{2, 4, 8, 16\} (pixel values in [0,255][0, 255]), while causing a slight drop of under 1%1\% in clean accuracy. Adding two Inception blocks (deeper model) improves robustness and mitigates the drop in clean accuracy.

    Model Training Metric Clean ϵ=2\epsilon = 2 ϵ=4\epsilon = 4 ϵ=8\epsilon = 8 ϵ=16\epsilon = 16
    Baseline (standard training) Top-1 78.4% 30.8% 27.2% 27.2% 29.5%
    Baseline (standard training) Top-5 94.0% 60.0% 55.6% 55.1% 57.2%
    Adversarial training Top-1 77.6% 73.5% 74.0% 74.5% 73.9%
    Adversarial training Top-5 93.8% 91.7% 91.9% 92.0% 91.4%
    Deeper model (standard training) Top-1 78.7% 33.5% 30.0% 30.0% 31.6%
    Deeper model (standard training) Top-5 94.4% 63.3% 58.9% 58.1% 59.5%
    Deeper model (adversarial training) Top-1 78.1% 75.4% 75.7% 75.6% 74.4%
    Deeper model (adversarial training) Top-5 94.1% 92.6% 92.7% 92.5% 91.6%

    Evaluations were conducted on the 50,000-image ImageNet validation set with adversarial images generated using the "step l.l." method. Adversarial training raises adversarial top-1 accuracy from ∼27−31%\sim 27-31\% up to ∼74−76%\sim 74-76\%.

  10. Knowl 10 — Cross-Model Adversarial Transfer Rates Across Architectures

    data/table

    Transfer rates measure the percentage of adversarial examples that fool a target model among those that successfully fooled the source model (prefiltered for source misclassification). Evaluated at ϵ=16\epsilon = 16 on 1,000 test images, single-step FGSM transfers at high rates across architectures, whereas iterative least-likely attacks transfer poorly.

    Models evaluated:

    • A, B: Inception v3 models trained with different random initializations.
    • C: Inception v3 trained with Exponential Linear Unit (ELU) activations instead of ReLU.
    • D: Inception v4 architecture.
    FGSM Basic Iterative Iterative L.L.
    Metric Source A B C D A B C D A B C D
    Top-1 A (v3) 100% 56% 58% 47% 100% 46% 45% 33% 100% 13% 13% 9%
    Top-1 B (v3) 58% 100% 59% 51% 41% 100% 40% 30% 15% 100% 13% 10%
    Top-1 C (v3 ELU) 56% 58% 100% 52% 44% 44% 100% 32% 12% 11% 100% 9%
    Top-1 D (v4) 50% 54% 52% 100% 35% 39% 37% 100% 12% 13% 13% 100%
    Top-5 A (v3) 100% 50% 50% 36% 100% 15% 17% 11% 100% 8% 7% 5%
    Top-5 B (v3) 51% 100% 50% 37% 16% 100% 14% 10% 7% 100% 5% 4%
    Top-5 C (v3 ELU) 44% 45% 100% 37% 16% 18% 100% 13% 6% 6% 100% 4%
    Top-5 D (v4) 42% 38% 46% 100% 11% 15% 15% 100% 6% 6% 6% 100%

    While iterative least-likely ("iter l.l.") achieves 100%100\% top-1 white-box success on source models, its transfer rate to different architectures is only 9%−15%9\% - 15\%. Conversely, FGSM transfers with 47%−59%47\% - 59\% top-1 error, demonstrating higher transferability for black-box attacks.

  11. Knowl 11 — Comparison of One-Step Adversarial Methods for Training Robust Models

    empirical result

    Comparing different single-step perturbation methods for adversarial training indicates that methods avoiding the true class label produce the best trade-off between clean and adversarial accuracy without label leaking:

    1. Targeted Methods without True Labels ("Step L.L." and "Step Rnd."): Using either the one-step least-likely class ("step l.l.") or one-step random class ("step rnd.") achieves 76.3%−76.4%76.3\% - 76.4\% top-1 clean accuracy and 72.2%−76.5%72.2\% - 76.5\% top-1 adversarial accuracy across perturbation magnitudes ϵ∈[2,16]\epsilon \in [2, 16]. Adversarial training with a single one-step method confers robustness against all other one-step methods.
    2. True-Label Methods (FGSM): Training with FGSM leads to label leaking, causing artificially inflated accuracy on FGSM adversarial examples (79.3%−85.3%79.3\% - 85.3\%) exceeding clean accuracy (74.9%74.9\%), while offering lower general robustness.
    3. Predicted Label and Entropy Methods ("FGSM-pred" and "Fast Entropy"): Methods using predicted labels or entropy maximization prevent label leaking (76.4%76.4\% clean accuracy), but yield lower adversarial robustness (40.0%−43.2%40.0\% - 43.2\% for FGSM-pred and 54.8%−62.8%54.8\% - 62.8\% for Fast Entropy at ϵ∈[2,16]\epsilon \in [2, 16]) compared to step l.l.
    4. Random Noise Perturbations: Adding Gaussian or uniform random noise during training fails to increase adversarial robustness (adversarial accuracy remains 31.8%−38.8%31.8\% - 38.8\%, comparable to undefended models).
  12. Knowl 12 — Effect of Adversarial Batch Proportion on Robustness and Clean Accuracy

    empirical result

    In large-scale adversarial training with total minibatch size m=32m = 32, varying the count of adversarial examples k∈{0,4,8,16,24,32}k \in \{0, 4, 8, 16, 24, 32\} (where kk clean examples are replaced with adversarial examples generated via "step l.l.") demonstrates a direct trade-off between clean data performance and adversarial robustness:

    • As kk increases from 00 to 1616, top-1 accuracy on adversarial examples across ϵ∈[2,16]\epsilon \in [2, 16] rises sharply from ∼28−31%\sim 28-31\% (for k=0k = 0, un-defended) to 73.8%−76.1%73.8\% - 76.1\% (for k=16k = 16).
    • Clean top-1 accuracy declines gradually as kk increases: 78.2%78.2\% (k=0k=0), 78.3%78.3\% (k=4k=4), 78.1%78.1\% (k=8k=8), 77.6%77.6\% (k=16k=16), 77.1%77.1\% (k=24k=24), and 76.3%76.3\% (k=32k=32).
    • Increasing kk beyond 1616 (such that adversarial examples constitute more than half the batch, up to 100%100\% at k=32k = 32) provides no significant additional adversarial robustness gain (remaining 73.4%−76.0%73.4\% - 76.0\%) but incurs an additional 1.3%1.3\% drop in clean accuracy. Therefore, setting k=m/2=16k = m / 2 = 16 achieves the optimal operating trade-off.

Coverage note — Ablations on alternative activation functions (tanh, relu6, ReluDecay) and training delay schedules were omitted as secondary hyperparameter explorations.

References

  1. 1.Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pp. 387–402. Springer, 2013.
  2. 2.Djork-Arne Clevert, Thomas Unterthiner, and Sepp Hochreiter. Fast and accurate deep network learning by exponential linear units (elus). CoRR, abs/1511.07289, 2015. URL http://arxiv.org/abs/1511.07289.
  3. 3.Nilesh Dalvi, Pedro Domingos, Sumit Sanghai, Deepak Verma, et al. Adversarial classification. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 99–108. ACM, 2004.
  4. 4.Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. CoRR, abs/1412.6572, 2014. URL http://arxiv.org/abs/1412.6572.
  5. 5.Ruitong Huang, Bing Xu, Dale Schuurmans, and Csaba Szepesvari. Learning with a strong adversary. CoRR, abs/1511.03034, 2015. URL http://arxiv.org/abs/1511.03034.
  6. 6.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. 2015.
  7. 7.Alex Kurakin, Ian Goodfellow, and Samy Bengio. Adversarial examples in the physical world. Technical report, arXiv, 2016. URL https://arxiv.org/abs/1607.02533.
  8. 8.Takeru Miyato, Andrew M Dai, and Ian Goodfellow. Virtual adversarial training for semi-supervised text classification. arXiv preprint arXiv:1605.07725, 2016a.
  9. 9.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, Ken Nakae, and Shin Ishii. Distributional smoothing with virtual adversarial training. In International Conference on Learning Representations (ICLR2016), April 2016b.
  10. 10.N. Papernot, P. McDaniel, and I. Goodfellow. Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples. ArXiv e-prints, May 2016b. URL http://arxiv.org/abs/1605.07277.
  11. 11.Nicolas Papernot, Patrick Drew McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. Distillation as a defense to adversarial perturbations against deep neural networks. CoRR, abs/1511.04508, 2015. URL http://arxiv.org/abs/1511.04508.
  12. 12.Nicolas Papernot, Patrick Drew McDaniel, Ian J. Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami. Practical black-box attacks against deep learning systems using adversarial examples. CoRR, abs/1602.02697, 2016a. URL http://arxiv.org/abs/1602.02697.
  13. 13.Andras Rozsa, Manuel Gunther, and Terrance E Boult. Are accuracy and robustness correlated? arXiv preprint arXiv:1610.04563, 2016.
  14. 14.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. arXiv preprint arXiv:1409.0575, 2014.
  15. 15.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. Intriguing properties of neural networks. ICLR, abs/1312.6199, 2014. URL http://arxiv.org/abs/1312.6199.
  16. 16.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567, 2015. URL http://arxiv.org/abs/1512.00567.
  17. 17.Christian Szegedy, Sergey Ioffe, and Vincent Vanhoucke. Inception-v4, inception-resnet and the impact of residual connections on learning. CoRR, abs/1602.07261, 2016. URL http://arxiv.org/abs/1602.07261.

Citation

MLA
Kurakin, A., et al. “Adversarial Machine Learning at Scale”. arXiv, 2016, http://arxiv.org/abs/1611.01236v2.
APA
Kurakin, A., Goodfellow, I., & Bengio, S. (2016). Adversarial Machine Learning at Scale. arXiv. http://arxiv.org/abs/1611.01236v2
Chicago
Kurakin, A., I. Goodfellow, and S. Bengio. 2016. “Adversarial Machine Learning at Scale”. arXiv. http://arxiv.org/abs/1611.01236v2.
Harvard
Kurakin, A., Goodfellow, I. and Bengio, S. (2016) “Adversarial Machine Learning at Scale”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.01236v2.
Vancouver
1. Kurakin A, Goodfellow I, Bengio S (2016) Adversarial Machine Learning at Scale. arXiv

BibTeX

@article{kurakin2016adversarial,
  title = {Adversarial Machine Learning at Scale},
  author = {Kurakin, Alexey and Goodfellow, Ian and Bengio, Samy},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.01236v2},
  eprint = {1611.01236}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission