Improving Transferability of Adversarial Examples With Input Diversity

Cihang XieZhishuai ZhangJianyu WangYuyin ZhouZhou RenA. Yuille

article2018CVPR1,497 citations

Proposes the Diverse Inputs method, which applies random transformations during iterative gradient optimization to prevent overfitting and significantly improve the black-box transferability of adversarial examples across convolutional neural networks.

Listen

Deep learning models are increasingly deployed in mission-critical applications such as autonomous driving and medical diagnostics, yet they remain vulnerable to adversarial examples—images altered with imperceptible noise that trick models into making severe classification errors. While attackers can easily mislead a system when its internal architecture and parameters are fully known (the white-box setting), these manipulated images historically struggle to fool target systems when their underlying parameters are unknown (the black-box setting). This occurs because standard optimization techniques overfit to the source network, limiting their transferability across diverse, real-world deployment environments.

The article demonstrates an effective method to enhance the transferability of adversarial examples by introducing input diversity during image generation. The approach applies random, differentiable transformations—specifically random resizing and random padding—at each optimization step to prevent adversarial noise from overfitting to a specific network architecture.

To evaluate this technique, the researchers conducted extensive empirical evaluations on standard benchmark datasets, using 5,000 correctly classified images from ImageNet. They tested several standard and adversarially trained convolutional networks under single-model and multi-model ensemble scenarios. Furthermore, they tested their approach against top defense systems and baseline models from the NIPS 2017 Adversarial Competition.

The findings show that introducing input diversity substantially improves attack transferability while maintaining nearly 100% white-box success rates. When paired with momentum optimization and multi-network ensembles, the combined method attained an average attack success rate of 73.0% across leading competitive defense mechanisms, outperforming the competition-winning baseline by 6.6 percentage points. When targeting individual black-box models, input diversity alone doubled or tripled transfer success rates compared to standard iterative attacks, and it proved similarly effective when integrated into alternative attack frameworks.

These results demonstrate that many prevailing security defenses, including transformation-based inference mitigations and adversarial training, provide less protection in black-box environments than previously assumed. This indicates a heightened operational risk for systems relying on security through obscurity or simple input defenses, underscoring that current defenses can be bypassed if an adversary crafts inputs that generalize across varied network transformations.

Organizations developing or deploying safety-critical computer vision models should adopt input-diversity attacks as a standard benchmark to stress-test system robustness. Reliance solely on basic input transformations or single-model defenses should be reconsidered in favor of more robust defenses that account for transformation-invariant adversarial noise. Future work should further investigate the underlying mathematical properties of shared decision boundaries and validate the approach across non-vision tasks and newer network architectures.

Cover for Improving Transferability of Adversarial Examples With Input Diversity

Abstract

Though CNNs have achieved the state-of-the-art performance on various vision tasks, they are vulnerable to adversarial examples — crafted by adding human-imperceptible perturbations to clean images. However, most of the existing adversarial attacks only achieve relatively low success rates under the challenging black-box setting, where the attackers have no knowledge of the model structure and parameters. To this end, we propose to improve the transferability of adversarial examples by creating diverse input patterns. Instead of only using the original images to generate adversarial examples, our method applies random transformations to the input images at each iteration. Extensive experiments on ImageNet show that the proposed attack method can generate adversarial examples that transfer much better to different networks than existing baselines. By evaluating our method against top defense solutions and official baselines from NIPS 2017 adversarial competition, the enhanced attack reaches an average success rate of 73.0%, which outperforms the top-1 attack submission in the NIPS competition by a large margin of 6.6%. We hope that our proposed attack strategy can serve as a strong benchmark baseline for evaluating the robustness of networks to adversaries and the effectiveness of different defense methods in the future. Code is available at https://github.com/cihangxie/DI-2-FGSM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Generating Adversarial Examples
  • 2.2 Defending Against Adversarial Examples
  • 3 Methodology
  • 3.1 Family of Fast Gradient Sign Methods
  • 3.2 Motivation
  • 3.3 Diverse Input Patterns
  • 3.4 Relationships between Different Attacks
  • 3.5 Attacking an Ensemble of Networks
  • 4 Experiment
  • 4.1 Experiment Setup
  • 4.2 Attacking a Single Network
  • 4.3 Attacking an Ensemble of Networks
  • 4.4 Ablation Studies
  • 4.5 NIPS 2017 Adversarial Competition
  • 4.6 Discussion
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Diverse Inputs Iterative Fast Gradient Sign Method (DI2-FGSM)

    model/method

    The Diverse Inputs Iterative Fast Gradient Sign Method (DI2-FGSM\text{DI}^2\text{-FGSM}) improves the transferability of adversarial examples generated by iterative gradient attacks by applying stochastic, differentiable image transformations at each attack iteration. Standard iterative fast gradient sign methods (I-FGSM) greedily step in the gradient sign direction of the original input image XX, which causes the perturbation to overfit the specific loss landscape of the white-box model parameters θ\theta and fail to transfer to unknown black-box models θ^\hat{\theta}.

    To prevent this optimization overfitting, DI2-FGSM\text{DI}^2\text{-FGSM} applies a stochastic transformation operator T(X;p)T(X; p) with activation probability p∈[0,1]p \in [0, 1]:

    T(X;p)={T(X)with probability pXwith probability 1−pT(X; p) = \begin{cases} T(X) & \text{with probability } p \\ X & \text{with probability } 1 - p \end{cases}

    where T(⋅)T(\cdot) consists of randomly resizing the image to dimension rnd×rnd×3\text{rnd} \times \text{rnd} \times 3 with rnd∈[299,330)\text{rnd} \in [299, 330) and randomly padding zero margins to restore the image shape to 330×330×3330 \times 330 \times 3.

    Let L(X,ytrue;θ)=−1ytrue⋅log⁡(softmax(l(X;θ)))L(X, y^{\text{true}}; \theta) = - \mathbf{1}_{y^{\text{true}}} \cdot \log(\text{softmax}(l(X; \theta))) denote the cross-entropy loss with one-hot ground-truth label vector 1ytrue\mathbf{1}_{y^{\text{true}}} and logits l(X;θ)l(X; \theta). The iterative update rule of DI2-FGSM\text{DI}^2\text{-FGSM} is defined as:

    X0adv=XX_0^{\text{adv}} = X

    Xn+1adv=ClipXϵ{Xnadv+α⋅sign(∇XL(T(Xnadv;p),ytrue;θ))}X_{n+1}^{\text{adv}} = \text{Clip}_X^\epsilon \left\{ X_n^{\text{adv}} + \alpha \cdot \text{sign}\left( \nabla_X L\left(T(X_n^{\text{adv}}; p), y^{\text{true}}; \theta\right) \right) \right\}

    where n∈{0,…,N−1}n \in \{0, \dots, N-1\} is the iteration step index, α\alpha is the step size, ϵ\epsilon is the maximum L∞L_\infty perturbation magnitude, and ClipXϵ(⋅)\text{Clip}_X^\epsilon(\cdot) clips the updated image pixel-wise into both the valid image range [0,255][0, 255] and the L∞L_\infty-ball [X−ϵ,X+ϵ][X - \epsilon, X + \epsilon].

  2. Knowl 2 — Momentum Diverse Inputs Iterative Fast Gradient Sign Method (M-DI2-FGSM)

    model/method

    The Momentum Diverse Inputs Iterative Fast Gradient Sign Method (M-DI2-FGSM\text{M-DI}^2\text{-FGSM}) combines input diversity with momentum-based gradient accumulation to alleviate adversarial overfitting and escape poor local maxima during adversarial image generation.

    At iteration step nn, the loss gradient is computed with respect to the stochastically transformed input T(Xnadv;p)T(X_n^{\text{adv}}; p) and normalized by its L1L_1-norm before being accumulated into a momentum buffer gn+1g_{n+1} with decay factor μ∈[0,1]\mu \in [0, 1]:

    g0=0g_0 = 0

    gn+1=μ⋅gn+∇XL(T(Xnadv;p),ytrue;θ)∥∇XL(T(Xnadv;p),ytrue;θ)∥1g_{n+1} = \mu \cdot g_n + \frac{\nabla_X L\left(T(X_n^{\text{adv}}; p), y^{\text{true}}; \theta\right)}{\left\|\nabla_X L\left(T(X_n^{\text{adv}}; p), y^{\text{true}}; \theta\right)\right\|_1}

    Xn+1adv=ClipXϵ{Xnadv+α⋅sign(gn+1)}X_{n+1}^{\text{adv}} = \text{Clip}_X^\epsilon \left\{ X_n^{\text{adv}} + \alpha \cdot \text{sign}(g_{n+1}) \right\}

    where X0adv=XX_0^{\text{adv}} = X, L(X,ytrue;θ)L(X, y^{\text{true}}; \theta) is the classification loss function, α\alpha is the step size, NN is the total number of iterations, and ClipXϵ(⋅)\text{Clip}_X^\epsilon(\cdot) constrains the updated adversarial image to remain within an L∞L_\infty distance ϵ\epsilon of the clean original image XX while preserving valid pixel bounds.

  3. Knowl 3 — Algorithm for M-DI2-FGSM with Multi-Model Ensemble Attack

    algorithm

    The following algorithm details the execution of M-DI2-FGSM\text{M-DI}^2\text{-FGSM} when attacking an ensemble of KK distinct neural network models simultaneously. The network outputs are fused at the logit level using ensemble weights wk≥0w_k \ge 0 satisfying ∑k=1Kwk=1\sum_{k=1}^K w_k = 1:

    l(X;θ1,…,θK)=∑k=1Kwklk(X;θk)l(X; \theta_1, \dots, \theta_K) = \sum_{k=1}^K w_k l_k(X; \theta_k)

    Input: Clean image XX with ground-truth label ytruey^{\text{true}}, ensemble of KK models with parameters {θk}k=1K\{\theta_k\}_{k=1}^K and ensemble weights {wk}k=1K\{w_k\}_{k=1}^K
    Input: Loss function LL, perturbation budget ϵ\epsilon, step size α\alpha, iterations NN, momentum decay factor μ\mu, transformation probability pp
    Output: Adversarial example XadvX^{\text{adv}}
    X0adv←XX_0^{\text{adv}} \leftarrow X
    g0←0g_0 \leftarrow 0
    for n←0n \leftarrow 0 to N−1N - 1 do
        Draw r∼Uniform(0,1)r \sim \text{Uniform}(0, 1)
        if r<pr < p then
            Sample integer rnd∼UniformInteger(299,329)rnd \sim \text{UniformInteger}(299, 329)
            Xtrans←RandomResizeAndZeroPad(Xnadv,target_size=rnd,pad_size=330)X_{\text{trans}} \leftarrow \text{RandomResizeAndZeroPad}(X_n^{\text{adv}}, \text{target\_size}=rnd, \text{pad\_size}=330)
        else
            Xtrans←XnadvX_{\text{trans}} \leftarrow X_n^{\text{adv}}
        end if
        lens←∑k=1Kwklk(Xtrans;θk)l_{\text{ens}} \leftarrow \sum_{k=1}^K w_k l_k(X_{\text{trans}}; \theta_k)
        Gn+1←∇XtransL(lens,ytrue)\mathcal{G}_{n+1} \leftarrow \nabla_{X_{\text{trans}}} L(l_{\text{ens}}, y^{\text{true}})
        gn+1←μ⋅gn+Gn+1∥Gn+1∥1g_{n+1} \leftarrow \mu \cdot g_n + \frac{\mathcal{G}_{n+1}}{\|\mathcal{G}_{n+1}\|_1}
        Xn+1adv←ClipXϵ(Xnadv+α⋅sign(gn+1))X_{n+1}^{\text{adv}} \leftarrow \text{Clip}_X^\epsilon \left( X_n^{\text{adv}} + \alpha \cdot \text{sign}(g_{n+1}) \right)
    end for
    return XNadvX_N^{\text{adv}}

    In standard benchmark configurations, α=1\alpha = 1, N=min⁡(ϵ+4,1.25ϵ)N = \min(\epsilon + 4, 1.25\epsilon), ϵ=15\epsilon = 15, μ=1.0\mu = 1.0, and p=0.5p = 0.5 (or p=0.4p = 0.4 for NIPS competition setups).

  4. Knowl 4 — Unified Parametric Framework of Fast Gradient Sign Attack Family

    model/method

    The family of gradient-based attacks spanning FGSM, I-FGSM, MI-FGSM, DI2-FGSM\text{DI}^2\text{-FGSM}, and M-DI2-FGSM\text{M-DI}^2\text{-FGSM} can be unified into a single parametric formulation controlled by transformation probability p∈[0,1]p \in [0, 1], momentum decay factor μ≥0\mu \ge 0, and total iteration number N≥1N \ge 1:

    1. If p=0p = 0, M-DI2-FGSM\text{M-DI}^2\text{-FGSM} reduces to MI-FGSM, and DI2-FGSM\text{DI}^2\text{-FGSM} reduces to I-FGSM.
    2. If μ=0\mu = 0, M-DI2-FGSM\text{M-DI}^2\text{-FGSM} reduces to DI2-FGSM\text{DI}^2\text{-FGSM}, and MI-FGSM reduces to I-FGSM.
    3. If N=1N = 1, I-FGSM reduces to FGSM.
    4. If both p=0p = 0 and μ=0\mu = 0, M-DI2-FGSM\text{M-DI}^2\text{-FGSM} reduces to standard iterative FGSM (I-FGSM).
  5. Knowl 5 — Single-Model Attack Transferability Benchmark

    data/table

    Adversarial examples were generated on single normally trained networks (Inception-v3, Inception-v4, Inception-ResNet-v2, and ResNet-v2-152) and evaluated across 7 networks, including four normally trained models and three adversarially trained models (Inc-v3ens3_{\text{ens3}}, Inc-v3ens4_{\text{ens4}}, IncRes-v2ens_{\text{ens}}) on 5000 ImageNet images with ϵ=15\epsilon = 15, α=1\alpha = 1, and N=19N = 19.

    Source Model Attack Inc-v3 Inc-v4 IncRes-v2 Res-152 Inc-v3ens3_{\text{ens3}} Inc-v3ens4_{\text{ens4}} IncRes-v2ens_{\text{ens}}
    Inc-v3 FGSM 64.6% 23.5% 21.7% 21.7% 8.0% 7.5% 3.6%
    I-FGSM 99.9% 14.8% 11.6% 8.9% 3.3% 2.9% 1.5%
    DI2-FGSM\text{DI}^2\text{-FGSM} 99.9% 35.5% 27.8% 21.4% 5.5% 5.2% 2.8%
    MI-FGSM 99.9% 36.6% 34.5% 27.5% 8.9% 8.4% 4.7%
    M-DI2-FGSM\text{M-DI}^2\text{-FGSM} 99.9% 63.9% 59.4% 47.9% 14.3% 14.0% 7.0%
    Inc-v4 FGSM 26.4% 49.6% 19.7% 20.4% 8.4% 7.7% 4.1%
    I-FGSM 22.0% 99.9% 13.2% 10.9% 3.2% 3.0% 1.7%
    DI2-FGSM\text{DI}^2\text{-FGSM} 43.3% 99.7% 28.9% 23.1% 5.9% 5.5% 3.2%
    MI-FGSM 51.1% 99.9% 39.4% 33.7% 11.2% 10.7% 5.3%
    M-DI2-FGSM\text{M-DI}^2\text{-FGSM} 72.4% 99.5% 62.2% 52.1% 17.6% 15.6% 8.8%
    IncRes-v2 FGSM 24.3% 19.3% 39.6% 19.4% 8.5% 7.3% 4.8%
    I-FGSM 22.2% 17.7% 97.9% 12.6% 4.6% 3.7% 2.5%
    DI2-FGSM\text{DI}^2\text{-FGSM} 46.5% 40.5% 95.8% 28.6% 8.2% 6.6% 4.8%
    MI-FGSM 53.5% 45.9% 98.4% 37.8% 15.3% 13.0% 8.8%
    M-DI2-FGSM\text{M-DI}^2\text{-FGSM} 71.2% 67.4% 96.1% 57.4% 25.1% 20.7% 14.9%
    Res-152 FGSM 34.4% 28.5% 27.1% 75.2% 12.4% 11.0% 6.0%
    I-FGSM 20.8% 17.2% 14.9% 99.1% 5.4% 4.6% 2.8%
    DI2-FGSM\text{DI}^2\text{-FGSM} 53.8% 49.0% 44.8% 99.2% 13.0% 11.1% 6.9%
    MI-FGSM 50.1% 44.1% 42.2% 99.0% 18.2% 15.2% 9.0%
    M-DI2-FGSM\text{M-DI}^2\text{-FGSM} 78.9% 76.5% 74.8% 99.2% 35.2% 29.4% 19.0%

    The diagonal values represent white-box attack success rates, while off-diagonal values represent black-box attack transferability. DI2-FGSM\text{DI}^2\text{-FGSM} substantially improves black-box transferability compared to I-FGSM (e.g., from 14.8% to 35.5% on Inc-v4 when attacking Inc-v3). M-DI2-FGSM\text{M-DI}^2\text{-FGSM} achieves the highest black-box transferability across all tested architectures while maintaining near-perfect white-box success rates.

  6. Knowl 6 — Hold-Out Transferability of Multi-Model Ensemble Attacks

    data/table

    Adversarial attacks were generated against an ensemble of six networks (with uniform weights wk=1/6w_k = 1/6) and evaluated simultaneously on the ensembled models (white-box setting) and a hold-out target model (black-box setting) across seven ImageNet architectures with ϵ=15\epsilon = 15.

    Evaluation Attack -Inc-v3 -Inc-v4 -IncRes-v2 -Res-152 -Inc-v3ens3_{\text{ens3}} -Inc-v3ens4_{\text{ens4}} -IncRes-v2ens_{\text{ens}}
    Ensemble I-FGSM 96.6% 96.9% 98.7% 96.2% 97.0% 97.3% 94.3%
    (White-box) DI2-FGSM\text{DI}^2\text{-FGSM} 88.9% 89.6% 93.2% 87.7% 91.7% 91.7% 93.2%
    MI-FGSM 96.9% 96.9% 98.8% 96.8% 96.8% 97.0% 94.6%
    M-DI2-FGSM\text{M-DI}^2\text{-FGSM} 90.1% 91.1% 94.0% 89.3% 92.8% 92.7% 94.9%
    Hold-out I-FGSM 43.7% 36.4% 33.3% 25.4% 12.9% 15.1% 8.8%
    (Black-box) DI2-FGSM\text{DI}^2\text{-FGSM} 69.9% 67.9% 64.1% 51.7% 36.3% 35.0% 30.4%
    MI-FGSM 71.4% 65.9% 64.6% 55.6% 22.8% 26.1% 15.8%
    M-DI2-FGSM\text{M-DI}^2\text{-FGSM} 80.7% 80.6% 80.7% 70.9% 44.6% 44.5% 39.4%

    The prefix '-' denotes the hold-out black-box evaluation model. M-DI2-FGSM\text{M-DI}^2\text{-FGSM} outperforms all existing baseline attacks on every black-box hold-out model by a substantial margin. On adversarially trained hold-out models, input diversity alone (DI2-FGSM\text{DI}^2\text{-FGSM}) outperforms momentum alone (MI-FGSM) (e.g., 36.3% vs 22.8% on Inc-v3ens3_{\text{ens3}}).

  7. Knowl 7 — Evaluation on NIPS 2017 Adversarial Competition Defenses

    data/table

    Attacks were evaluated against the top-3 defense entries and 3 official baseline models from the NIPS 2017 Adversarial Competition using 5000 ImageNet images. Batches were attacked under randomly sampled maximum perturbation budgets ϵ∈{4/255,8/255,12/255,16/255}\epsilon \in \{4/255, 8/255, 12/255, 16/255\} against an ensemble of 8 networks (Inc-v3, Inc-v4, IncRes-v2, Res-152, Inc-v3ens3_{\text{ens3}}, Inc-v3ens4_{\text{ens4}}, IncRes-v2ens_{\text{ens}} weighted 1/7.251/7.25 each, and Inc-v3adv_{\text{adv}} weighted 0.25/7.250.25/7.25) with N=10N = 10, μ=1.0\mu = 1.0, and p=0.4p = 0.4.

    Attack TsAIL (1st) iyswim (2nd) Anil Thomas (3rd) Inc-v3adv_{\text{adv}} IncRes-v2ens_{\text{ens}} Inc-v3 Average
    I-FGSM 14.0% 35.6% 30.9% 98.2% 96.4% 99.0% 62.4%
    DI2-FGSM\text{DI}^2\text{-FGSM} (Ours) 22.7% 58.4% 48.0% 91.5% 90.7% 97.3% 68.1%
    MI-FGSM 14.9% 45.7% 46.6% 97.3% 95.4% 98.7% 66.4%
    MI-FGSM* 13.6% 43.2% 43.9% 94.4% 93.0% 97.3% 64.2%
    M-DI2-FGSM\text{M-DI}^2\text{-FGSM} (Ours) 20.0% 69.8% 64.4% 93.3% 92.4% 97.9% 73.0%

    Asterisk (*) marks the official reported competition result. M-DI2-FGSM\text{M-DI}^2\text{-FGSM} achieves an average success rate of 73.0%, surpassing the competition-winning MI-FGSM attack (66.4% replicated, 64.2% official) by 6.6% to 8.8% overall, with particularly large gains on defense models utilizing randomized transformations such as iyswim (+24.1% over MI-FGSM) and Anil Thomas (+17.8% over MI-FGSM).

  8. Knowl 8 — Generalization of Input Diversity to Optimization-Based Carlini-Wagner Attack (D-C&W)

    data/table

    Input diversity applies generally to optimization-based adversarial attacks beyond gradient sign methods. Integrating diverse input transformations into the Carlini & Wagner (C&W) attack creates D-C&W. Evaluated on 1000 correctly classified ImageNet images with maximum iterations 250, learning rate 0.01, and confidence parameter 10:

    Source Model Attack Inc-v3 Inc-v4 IncRes-v2 Res-152 Inc-v3ens3_{\text{ens3}} Inc-v3ens4_{\text{ens4}} IncRes-v2ens_{\text{ens}}
    Inc-v3 C W 100.0% 5.7% 5.3% 5.1% 3.0% 2.5% 1.1%
    D-C W (Ours) 100.0% 16.8% 13.0% 11.2% 5.8% 3.9% 2.1%
    Inc-v4 C W 15.1% 100.0% 9.2% 7.8% 4.4% 3.5% 1.9%
    D-C W (Ours) 29.3% 100.0% 20.1% 15.4% 7.1% 5.3% 3.1%
    IncRes-v2 C W 15.8% 11.2% 99.9% 8.6% 6.3% 3.6% 3.4%
    D-C W (Ours) 33.9% 25.6% 100.0% 19.4% 11.2% 7.3% 4.0%
    Res-152 C W 11.4% 6.9% 6.1% 100.0% 4.4% 4.1% 2.3%
    D-C W (Ours) 33.0% 27.7% 24.4% 100.0% 13.1% 9.3% 5.7%

    D-C&W substantially improves black-box transferability over standard C&W across all target models (e.g., from 6.9% to 27.7% on Inc-v4 when crafted on Res-152) while retaining 100% white-box attack success.

  9. Knowl 9 — Ablation on Attack Hyperparameters: Transformation Probability, Iterations, and Step Size

    empirical result

    Empirical ablation studies on ensemble attacks at perturbation budget ϵ=15\epsilon = 15 reveal the following operational properties:

    1. Transformation probability (p∈[0,1]p \in [0, 1]): As pp increases, black-box transferability improves monotonically while white-box success rate exhibits a mild decrease. A small probability (p≈0.1–0.2p \approx 0.1\text{--}0.2) yields large gains in black-box transfer with negligible white-box penalty. Setting p=1.0p = 1.0 maximizes pure black-box transfer when targeting completely unknown architectures, whereas p≈0.4–0.5p \approx 0.4\text{--}0.5 provides a strong balance when targeting mixed environments.
    2. Iteration count (N∈[15,31]N \in [15, 31] with α=1\alpha = 1, p=0.5p = 0.5): For DI2-FGSM\text{DI}^2\text{-FGSM}, both white-box and black-box success rates monotonically increase with NN. For M-DI2-FGSM\text{M-DI}^2\text{-FGSM}, higher NN improves white-box and normally trained black-box success rates but plateaus on adversarially trained models, narrowing the performance gap between M-DI2-FGSM\text{M-DI}^2\text{-FGSM} and DI2-FGSM\text{DI}^2\text{-FGSM}.
    3. Step size (α∈[1/30,1/5]\alpha \in [1/30, 1/5] with N=ϵ/αN = \epsilon / \alpha, p=0.5p = 0.5): Smaller step sizes consistently improve white-box success for both methods. Under black-box evaluation, DI2-FGSM\text{DI}^2\text{-FGSM} transferability is insensitive to step size, whereas M-DI2-FGSM\text{M-DI}^2\text{-FGSM} transferability further improves as α\alpha decreases.
  10. Knowl 10 — Decision Boundary Invariance Hypothesis for Adversarial Transferability

    theoretical result

    The enhanced transferability induced by input diversity is explained by a decision boundary geometry hypothesis: Deep neural networks trained on identical datasets (e.g., ImageNet) learn decision boundaries with similar macro-level geometric properties, as evidenced by consistent misclassification patterns across diverse architectures on transferable adversaries.

    Standard iterative attacks overfit to the local high-frequency perturbations of the source network's specific decision boundary wrinkles. By applying random affine and padding transformations T(⋅)T(\cdot) during gradient calculation, the optimization forces the generated perturbation to remain effective under geometric variance. The resulting adversarial examples occupy broader, shared adversarial regions across the decision boundary rather than narrow, architecture-specific peaks, yielding substantially higher cross-network transferability.

Coverage note — No substantial contributed material was omitted from the knowls.

References

  1. 1.A. Arnab, O. Miksik, and P. H. Torr. On the robustness of semantic segmentation models to adversarial attacks. arXiv preprint arXiv:1711.09856, 2017.
  2. 2.A. Athalye, L. Engstrom, A. Ilyas, and K. Kwok. Synthesizing robust adversarial examples. In International Conference on Machine Learning, pages 284–293, 2018.
  3. 3.B. Biggio, I. Corona, D. Maiorca, B. Nelson, N. Šrndić, P. Laskov, G. Giacinto, and F. Roli. Evasion attacks against machine learning at test time. In Joint European conference on machine learning and knowledge discovery in databases, pages 387–402, 2013.
  4. 4.N. Carlini and D. Wagner. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy, 2017.
  5. 5.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017.
  6. 6.M. Cisse, Y. Adi, N. Neverova, and J. Keshet. Houdini: Fooling deep structured prediction models. arXiv preprint arXiv:1707.05373, 2017.
  7. 7.N. Dalvi, P. Domingos, S. Sanghai, D. Verma, et al. Adversarial classification. In ACM SIGKDD international conference on Knowledge discovery and data mining, 2004.
  8. 8.G. S. Dhillon, K. Azizzadenesheli, J. D. Bernstein, J. Kossaifi, A. Khanna, Z. C. Lipton, and A. Anandkumar. Stochastic activation pruning for robust adversarial defense. In International Conference on Learning Representations, 2018.
  9. 9.Y. Dong, F. Liao, T. Pang, H. Su, X. Hu, J. Li, and J. Zhu. Boosting adversarial attacks with momentum. arXiv preprint arXiv:1710.06081, 2017.
  10. 10.R. Girshick. Fast r-cnn. In International Conference on Computer Vision, 2015.
  11. 11.I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. In International Conference on Learning Representations, 2015.
  12. 12.C. Guo, M. Rana, M. Cissé, and L. van der Maaten. Countering adversarial images using input transformations. In International Conference on Learning Representations, 2018.
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In European Conference on Computer Vision, 2016.
  14. 14.L. Huang, A. D. Joseph, B. Nelson, B. I. Rubinstein, and J. Tygar. Adversarial machine learning. In ACM workshop on Security and artificial intelligence, 2011.
  15. 15.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, 2012.
  16. 16.A. Kurakin, I. Goodfellow, and S. Bengio. Adversarial examples in the physical world. In International Conference on Learning Representations Workshop, 2017.
  17. 17.A. Kurakin, I. Goodfellow, and S. Bengio. Adversarial machine learning at scale. In International Conference on Learning Representations, 2017.
  18. 18.A. Kurakin, I. Goodfellow, S. Bengio, Y. Dong, F. Liao, M. Liang, T. Pang, J. Zhu, X. Hu, C. Xie, et al. Adversarial attacks and defences competition. arXiv preprint arXiv:1804.00097, 2018.
  19. 19.F. Liao, M. Liang, Y. Dong, T. Pang, X. Hu, and J. Zhu. Defense against adversarial attacks using high-level representation guided denoiser. In Computer Vision and Pattern Recognition, 2018.
  20. 20.Y.-C. Lin, Z.-W. Hong, Y.-H. Liao, M.-L. Shih, M.-Y. Liu, and M. Sun. Tactics of adversarial attack on deep reinforcement learning agents. In International Joint Conference on Artificial Intelligence, 2017.
  21. 21.Y. Liu, X. Chen, C. Liu, and D. Song. Delving into transferable adversarial examples and black-box attacks. In International Conference on Learning Representations, 2017.
  22. 22.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In Computer Vision and Pattern Recognition, 2015.
  23. 23.Y. Luo, X. Boix, G. Roig, T. Poggio, and Q. Zhao. Foveation-based mechanisms alleviate adversarial examples. arXiv preprint arXiv:1511.06292, 2015.
  24. 24.D. Meng and H. Chen. Magnet: a two-pronged defense against adversarial examples. arXiv preprint arXiv:1705.09064, 2017.
  25. 25.A. Nguyen, J. Yosinski, and J. Clune. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Computer Vision and Pattern Recognition, 2015.
  26. 26.N. Papernot, F. Faghri, N. Carlini, I. Goodfellow, R. Feinman, A. Kurakin, C. Xie, Y. Sharma, T. Brown, A. Roy, A. Matyasko, V. Behzadan, K. Hambardzumyan, Z. Zhang, Y.-L. Juang, Z. Li, R. Sheatsley, A. Garg, J. Uesato, W. Gierke, Y. Dong, D. Berthelot, P. Hendricks, J. Rauber, and R. Long. cleverhans v2.1.0: an adversarial machine learning library. arXiv preprint arXiv:1610.00768, 2018.
  27. 27.A. Prakash, N. Moran, S. Garber, A. DiLillo, and J. Storer. Deflecting adversarial attacks with pixel deflection. arXiv preprint arXiv:1801.08926, 2018.
  28. 28.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, 2015.
  29. 29.P. Samangouei, M. Kabkab, and R. Chellappa. Defense-GAN: Protecting classifiers against adversarial attacks using generative models. In International Conference on Learning Representations, 2018.
  30. 30.A. Shrivastava, A. Gupta, and R. Girshick. Training region-based object detectors with online hard example mining. In Computer Vision and Pattern Recognition, 2016.
  31. 31.E. Simo-Serra, E. Trulls, L. Ferraz, I. Kokkinos, P. Fua, and F. Moreno-Noguer. Discriminative learning of deep convolutional feature point descriptors. In International Conference on Computer Vision, 2015.
  32. 32.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 2015.
  33. 33.Y. Song, T. Kim, S. Nowozin, S. Ermon, and N. Kushman. Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. arXiv preprint arXiv:1710.10766, 2017.
  34. 34.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI, 2017.
  35. 35.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In Computer Vision and Pattern Recognition, 2016.
  36. 36.C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations, 2014.
  37. 37.F. Tramèr, A. Kurakin, N. Papernot, D. Boneh, and P. McDaniel. Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204, 2017.
  38. 38.C. Xie, J. Wang, Z. Zhang, Z. Ren, and A. Yuille. Mitigating adversarial effects through randomization. In International Conference on Learning Representations, 2018.
  39. 39.C. Xie, J. Wang, Z. Zhang, Y. Zhou, L. Xie, and A. Yuille. Adversarial Examples for Semantic Segmentation and Object Detection. In International Conference on Computer Vision, 2017.
  40. 40.Z. Zhang, S. Qiao, C. Xie, W. Shen, B. Wang, and A. L. Yuille. Single-shot object detection with enriched semantics. arXiv preprint arXiv:1712.00433, 2017.

Citation

MLA
Xie, C., et al. “Improving Transferability of Adversarial Examples with Input Diversity”. arXiv, 2018, http://arxiv.org/abs/1803.06978v4.
APA
Xie, C., Zhang, Z., Zhou, Y., Bai, S., Wang, J., Ren, Z., & Yuille, A. (2018). Improving Transferability of Adversarial Examples with Input Diversity. arXiv. http://arxiv.org/abs/1803.06978v4
Chicago
Xie, C., Z. Zhang, Y. Zhou, et al. 2018. “Improving Transferability of Adversarial Examples with Input Diversity”. arXiv. http://arxiv.org/abs/1803.06978v4.
Harvard
Xie, C. et al. (2018) “Improving Transferability of Adversarial Examples with Input Diversity”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.06978v4.
Vancouver
1. Xie C, Zhang Z, Zhou Y, Bai S, Wang J, Ren Z, Yuille A (2018) Improving Transferability of Adversarial Examples with Input Diversity. arXiv

BibTeX

@article{xie2018improving,
  title = {Improving Transferability of Adversarial Examples with Input Diversity},
  author = {Xie, Cihang and Zhang, Zhishuai and Zhou, Yuyin and Bai, Song and Wang, Jianyu and Ren, Zhou and Yuille, Alan},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.06978v4},
  eprint = {1803.06978}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE