Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation

Zeyu QinYanbo FanYi LiuLi ShenYong ZhangJue WangBaoyuan Wu

article2022NeurIPS119 citations

Proposes reverse adversarial perturbation, a min-max bi-level optimization technique that prevents attacks from overfitting to surrogate models by targeting flat loss regions, substantially improving black-box attack transferability across both standard networks and commercial vision systems.

Listen

Deep neural networks are central to safety-critical applications like automated vision and face recognition, but they remain susceptible to adversarial examples—minor, imperceptible modifications to input data that fool models into incorrect classifications. In practical black-box settings, attackers lack direct access to internal target models and instead generate adversarial examples using known substitute models. These attacks often fail to transfer across different systems because the generated perturbations overfit the specific geometry of the substitute model, landing in sharp decision regions that do not generalize.

The article introduces a method called Reverse Adversarial Perturbation (RAP) to systematically improve the transferability of black-box attacks. The primary objective is to evaluate whether driving adversarial optimization toward broader, flatter loss regions—rather than single sharp points—prevents substitute overfitting and improves cross-model attack success.

To accomplish this, the authors model adversarial generation as a two-level optimization problem. The inner step computes a worst-case shift within a defined neighborhood to identify vulnerability, while the outer step updates the attack sample to ensure low error across that entire local area. To enhance early computational efficiency, the authors also introduce a late-start variation (RAP-LS). The approach was evaluated across standard benchmarks, testing multiple model architectures (such as Inception, ResNet, DenseNet, VGG, and Vision Transformers), robust defense techniques, and a real-world commercial platform using 1,000 ImageNet samples and 500 real-world API queries.

The findings show that finding flatter loss regions significantly boosts attack success across unknown models. When integrated with foundational attack methods, RAP increases average untargeted attack success rates by roughly 6% to 16% and targeted success rates by up to 18.5%. Combining RAP with advanced input-transformation methods achieves untargeted transfer success rates between 95% and 98%, while targeted baseline attacks improved by approximately 9% to 14%. Furthermore, RAP-LS consistently outperformed competitive feature-based and generative attack approaches, improving targeted attack rates against defense-hardened models by 11% to 15%. In a real-world evaluation against the Google Cloud Vision API, RAP-LS achieved a 22.0% absolute increase in targeted attack success over existing baseline techniques.

These findings demonstrate that black-box adversarial transferability poses a severe and practical risk to commercial artificial intelligence deployments. Models previously considered resilient due to hidden internal parameters or defensive training remain vulnerable to neighborhood-optimized attacks. Consequently, system architects and security teams cannot rely on model obscurity or isolated input defenses. Organizations operating computer vision systems should proactively audit existing architectures against multi-model transfer attacks and prioritize structural defenses, such as diverse ensemble pipelines and robust loss regularization.

While the empirical results provide strong evidence across diverse model families, the study focuses predominantly on image classification benchmarks and standard perturbation limits. Readers should note that computational costs increase during neighborhood searches, and transferring attacks to substantially different architectures, such as Vision Transformers, remains inherently more challenging. Future work should focus on developing advanced defense mechanisms tailored against flat-loss attacks and evaluating transfer dynamics in broader domains such as object detection and natural language processing.

Cover for Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation

Abstract

Deep neural networks (DNNs) have been shown to be vulnerable to adversarial examples, which can produce erroneous predictions by injecting imperceptible perturbations. In this work, we study the transferability of adversarial examples, which is significant due to its threat to real-world applications where model architecture or parameters are usually unknown. Many existing works reveal that the adversarial examples are likely to overfit the surrogate model that they are generated from, limiting its transfer attack performance against different target models. To mitigate the overfitting of the surrogate model, we propose a novel attack method, dubbed reverse adversarial perturbation (RAP). Specifically, instead of minimizing the loss of a single adversarial point, we advocate seeking adversarial example located at a region with unified low loss value, by injecting the worst-case perturbation (i.e., the reverse adversarial perturbation) for each step of the optimization procedure. The adversarial attack with RAP is formulated as a min-max bi-level optimization problem. By integrating RAP into the iterative process for attacks, our method can find more stable adversarial examples which are less sensitive to the changes of decision boundary, mitigating the overfitting of the surrogate model. Comprehensive experimental comparisons demonstrate that RAP can significantly boost adversarial transferability. Furthermore, RAP can be naturally combined with many existing black-box attack techniques, to further boost the transferability. When attacking a real-world image recognition system, i.e., Google Cloud Vision API, we obtain 22% performance improvement of targeted attacks over the compared method. Our codes are available at: https://github.com/SCLBD/Transfer_attack_RAP.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Preliminaries of Transfer Adversarial Attack
  • 3.2 Reverse Adversarial Perturbation
  • 3.3 A Closer Look at RAP
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 The Evaluation of Untargeted Attacks
  • 4.3 The Evaluation of Targeted Attacks
  • 4.4 The Comparison with Other Types of Attacks
  • 4.5 The Evaluation on Diverse Network Architectures and Defense Models
  • 4.6 Ablation Study
  • 4.7 The Targeted Attack Against Google Cloud Vision API
  • 5 Conclusion
  • Acknowledgments
  • References
  • Checklist

Knowls

  1. Knowl 1 — Reverse Adversarial Perturbation Optimization Framework

    model/method

    To alleviate the overfitting of adversarial examples to a white-box surrogate model Ms(x;θ)M^s(x; \theta) and boost black-box transferability to unknown target models Mt(x;ϕ)M^t(x; \phi), Reverse Adversarial Perturbation (RAP) formulates adversarial attack generation as a min-max bi-level optimization problem. Rather than minimizing the loss of a single point (which frequently converges to sharp local minima sensitive to shifts in decision boundaries), RAP searches for an adversarial example xadvx^{adv} situated in a flat loss region where points within its local neighborhood also maintain low attack loss.

    For a targeted attack on a benign input sample xx with target label yty_t, perturbation budget Bϵ(x)={x′:∥x′−x∥∞≤ϵ}B_\epsilon(x) = \{x' : \|x' - x\|_\infty \le \epsilon\}, input transformation G(⋅)G(\cdot), and surrogate loss function L\mathcal{L}, the RAP objective is:

    min⁡xadv∈Bϵ(x)L(Ms(G(xadv+nrap);θ),yt)\min_{x^{adv} \in B_\epsilon(x)} \mathcal{L}\left(M^s\left(G(x^{adv} + n^{rap}); \theta\right), y_t\right)

    subject to the inner maximization:

    nrap=arg⁡max⁡∥nrap∥∞≤ϵnL(Ms(xadv+nrap;θ),yt)n^{rap} = \arg\max_{\|n^{rap}\|_\infty \le \epsilon_n} \mathcal{L}\left(M^s(x^{adv} + n^{rap}; \theta), y_t\right)

    where nrapn^{rap} is the worst-case reverse perturbation within an ℓ∞\ell_\infty-ball of radius ϵn\epsilon_n.

    For an untargeted attack with true label yy, the loss L\mathcal{L} and target label yty_t in both levels are replaced with −L-\mathcal{L} and yy, respectively.

    The inner maximization is solved via TT steps of projected gradient ascent:

    nrap←Clip[−ϵn,ϵn](nrap+αn⋅sign(∇nrapL(Ms(xadv+nrap;θ),yt)))n^{rap} \leftarrow \text{Clip}_{[-\epsilon_n, \epsilon_n]}\left(n^{rap} + \alpha_n \cdot \text{sign}\left(\nabla_{n^{rap}} \mathcal{L}\left(M^s(x^{adv} + n^{rap}; \theta), y_t\right)\right)\right)

    with inner step size αn=ϵn/T\alpha_n = \epsilon_n / T. The outer minimization updates xadvx^{adv} using projected gradient descent:

    xadv←ClipBϵ(x)(xadv−α⋅sign(∇xadvL(Ms(G(xadv+nrap);θ),yt)))x^{adv} \leftarrow \text{Clip}_{B_\epsilon(x)}\left(x^{adv} - \alpha \cdot \text{sign}\left(\nabla_{x^{adv}} \mathcal{L}\left(M^s\left(G(x^{adv} + n^{rap}); \theta\right), y_t\right)\right)\right)

    where α\alpha is the outer attack step size.

  2. Knowl 2 — Reverse Adversarial Perturbation Algorithm with Late-Start

    algorithm

    The Reverse Adversarial Perturbation (RAP) algorithm, incorporated with a late-start strategy (RAP-LS), iteratively generates transferable adversarial perturbations by calculating worst-case inner neighborhood perturbations after an optional warmup stage.

    Input: Surrogate model MsM^s, input (x,y)(x, y), target label yty_t (or true label yy for untargeted), loss L\mathcal{L}, transformation GG, total iterations KK, late-start threshold KLSK_{LS}, outer step size α\alpha, inner perturbation budget ϵn\epsilon_n, inner steps TT, inner step size αn=ϵn/T\alpha_n = \epsilon_n / T, perturbation budget ϵ\epsilon.
    Output: Adversarial example xadvx^{adv}.
    Initialize xadv←xx^{adv} \leftarrow x
    for k=1,…,Kk = 1, \dots, K do
        if k≥KLSk \ge K_{LS} then
            Initialize nrap←0n^{rap} \leftarrow 0
            for t=1,…,Tt = 1, \dots, T do
                gn←∇nrapL(Ms(xadv+nrap;θ),yt)g_n \leftarrow \nabla_{n^{rap}} \mathcal{L}(M^s(x^{adv} + n^{rap}; \theta), y_t)
                nrap←Clip[−ϵn,ϵn](nrap+αn⋅sign(gn))n^{rap} \leftarrow \text{Clip}_{[-\epsilon_n, \epsilon_n]}(n^{rap} + \alpha_n \cdot \text{sign}(g_n))
            end for
        else
            nrap←0n^{rap} \leftarrow 0
        end if
        gx←∇xadvL(Ms(G(xadv+nrap);θ),yt)g_x \leftarrow \nabla_{x^{adv}} \mathcal{L}(M^s(G(x^{adv} + n^{rap}); \theta), y_t)
        xadv←ClipBϵ(x)(xadv−α⋅sign(gx))x^{adv} \leftarrow \text{Clip}_{B_\epsilon(x)}(x^{adv} - \alpha \cdot \text{sign}(g_x))
    end for
    return xadvx^{adv}

    In standard RAP, KLS=1K_{LS} = 1. In the late-start variant (RAP-LS), setting KLS=100K_{LS} = 100 for a total budget of K=400K = 400 delays the min-max perturbation calculation until the attack point has traversed beyond the initial weak attack region.

  3. Knowl 3 — Late-Start Optimization Mechanism for RAP

    model/method

    During the early iterations of adversarial example optimization on a surrogate model MsM^s, the current iterate xadvx^{adv} resides in a region of high classification confidence for the original class and very weak adversarial efficacy. Solving the full bi-level min-max problem at this stage consumes inner optimization steps unnecessarily.

    The Late-Start variant (RAP-LS) introduces a warm-up phase defined by an iteration threshold parameter KLS∈[0,K]K_{LS} \in [0, K]:

    1. For iterations k<KLSk < K_{LS}, the reverse perturbation is held at zero (nrap=0n^{rap} = 0), reducing the update to standard single-level minimization of the surrogate attack loss L(Ms(G(xadv);θ),yt)\mathcal{L}(M^s(G(x^{adv}); \theta), y_t). This enables the iterate to rapidly reach a region of high adversarial loss.
    2. For iterations k≥KLSk \ge K_{LS}, the inner maximization for nrapn^{rap} is activated, computing the worst-case neighborhood perturbation over TT steps before updating xadvx^{adv}.

    When KLS=0K_{LS} = 0, RAP-LS reduces to standard RAP. In empirical comparisons with K=400K = 400 total iterations, setting KLS=100K_{LS} = 100 accelerates convergence and improves targeted transfer attack success rate across evaluated models by an average of 2.6%2.6\% over standard RAP.

  4. Knowl 4 — Untargeted Black-Box Attack Transferability of RAP and RAP-LS

    data/table

    Untargeted attack success rates (ASR, in %) measure the percentage of adversarial examples generated from a surrogate model that fool black-box target models. Evaluated on 1,000 ImageNet-compatible images with ℓ∞\ell_\infty budget ϵ=16/255\epsilon = 16/255, total outer iterations K=400K = 400, step size α=2/255\alpha = 2/255, cross-entropy loss, and KLS=100K_{LS} = 100 for RAP-LS.

    The tables report untargeted ASRs comparing baseline gradient/transformation methods (I-FGSM, MI-FGSM, TI, DI, SI, Admix) and combinational attacks (MI-TI-DI / MTDI, MI-TI-DI-SI / MTDSI, MI-TI-DI-Admix / MTDAI) against their RAP and RAP-LS augmented counterparts:

    Attack ResNet-50 ⟹\Longrightarrow DenseNet-121 ⟹\Longrightarrow
    Dense-121 VGG-16 Inc-v3 Res-50 VGG-16 Inc-v3
    I / +RAP / +RAP-LS 79.2 / 91.5 / 91.9 78.0 / 91.1 / 92.9 34.6 / 57.0 / 57.2 87.4 / 94.2 / 94.3 85.1 / 91.7 / 92.8 46.5 / 60.2 / 61.1
    MI / +RAP / +RAP-LS 85.8 / 95.0 / 96.1 82.4 / 93.9 / 94.5 50.3 / 75.9 / 77.4 90.3 / 97.6 / 97.9 87.5 / 96.0 / 97.6 59.3 / 80.4 / 82.8
    TI / +RAP / +RAP-LS 82.0 / 94.1 / 95.1 81.0 / 93.1 / 93.3 45.5 / 66.1 / 67.0 89.6 / 94.2 / 94.8 87.0 / 92.1 / 93.3 54.2 / 66.7 / 70.0
    DI / +RAP / +RAP-LS 99.0 / 99.6 / 99.7 99.0 / 99.6 / 99.7 57.7 / 82.9 / 85.0 98.2 / 99.6 / 99.7 98.1 / 99.4 / 99.4 67.6 / 86.6 / 86.9
    SI / +RAP / +RAP-LS 94.9 / 98.9 / 99.7 88.6 / 95.7 / 97.2 65.9 / 79.7 / 84.4 95.1 / 96.9 / 98.8 91.9 / 95.0 / 97.5 71.6 / 83.2 / 87.4
    Admix / +RAP / +RAP-LS 97.9 / 99.6 / 99.9 95.8 / 97.7 / 99.0 77.7 / 87.4 / 92.6 97.0 / 99.0 / 99.2 95.6 / 97.7 / 98.6 82.0 / 89.8 / 93.8
    Attack VGG-16 ⟹\Longrightarrow Inc-v3 ⟹\Longrightarrow
    Res-50 Dense-121 Inc-v3 Res-50 Dense-121 VGG-16
    I / +RAP / +RAP-LS 53.7 / 53.0 / 54.2 49.1 / 50.6 / 51.4 22.0 / 24.7 / 24.9 51.5 / 62.1 / 62.0 48.7 / 60.8 / 60.0 55.1 / 65.9 / 68.0
    MI / +RAP / +RAP-LS 62.5 / 76.2 / 76.4 60.5 / 73.0 / 73.9 30.0 / 42.7 / 42.2 62.0 / 85.8 / 84.8 56.7 / 84.6 / 84.6 63.1 / 84.9 / 84.6
    TI / +RAP / +RAP-LS 62.8 / 64.8 / 65.8 55.9 / 63.7 / 62.1 29.1 / 36.2 / 37.1 49.3 / 63.4 / 61.6 49.4 / 63.4 / 63.8 58.1 / 68.6 / 69.5
    DI / +RAP / +RAP-LS 72.2 / 86.0 / 88.8 68.8 / 85.0 / 87.4 29.9 / 46.6 / 51.6 68.4 / 81.7 / 81.8 71.9 / 85.0 / 84.0 76.1 / 85.2 / 86.4
    SI / +RAP / +RAP-LS 80.0 / 92.7 / 94.7 82.1 / 94.8 / 95.7 45.8 / 74.0 / 74.7 66.2 / 69.8 / 72.8 65.9 / 74.9 / 77.2 66.0 / 69.2 / 73.0
    Admix / +RAP / +RAP-LS 87.3 / 94.6 / 96.8 88.2 / 96.4 / 97.2 55.5 / 77.6 / 80.8 75.9 / 80.2 / 84.9 78.5 / 83.7 / 87.4 74.5 / 77.2 / 83.5
    Combinational Attack ResNet-50 ⟹\Longrightarrow DenseNet-121 ⟹\Longrightarrow
    Dense-121 VGG-16 Inc-v3 Res-50 VGG-16 Inc-v3
    MTDI / +RAP / +RAP-LS 99.8 / 100 / 100 99.8 / 100 / 99.9 85.7 / 96.0 / 96.9 99.4 / 99.8 / 100 99.2 / 99.5 / 100 89.1 / 97.1 / 97.1
    MTDSI / +RAP / +RAP-LS 100 / 100 / 100 99.7 / 99.9 / 99.8 97.0 / 99.1 / 99.1 99.8 / 99.9 / 99.9 99.2 / 99.3 / 99.7 95.1 / 98.3 / 98.4
    MTDAI / +RAP / +RAP-LS 100 / 100 / 100 99.8 / 99.9 / 99.9 98.3 / 99.2 / 99.8 99.8 / 99.8 / 99.9 99.4 / 99.6 / 99.8 97.9 / 98.8 / 98.9
    Combinational Attack VGG-16 ⟹\Longrightarrow Inc-v3 ⟹\Longrightarrow
    Res-50 Dense-121 Inc-v3 Res-50 Dense-121 VGG-16
    MTDI / +RAP / +RAP-LS 90.0 / 97.2 / 97.7 88.8 / 97.0 / 97.3 56.8 / 82.6 / 81.4 82.9 / 91.8 / 90.6 85.7 / 94.2 / 93.3 85.1 / 92.7 / 91.0
    MTDSI / +RAP / +RAP-LS 97.6 / 98.8 / 99.4 98.1 / 99.2 / 99.4 85.0 / 94.1 / 95.2 89.0 / 91.2 / 92.3 92.0 / 95.2 / 95.6 87.6 / 90.3 / 92.2
    MTDAI / +RAP / +RAP-LS 97.8 / 99.2 / 99.6 98.9 / 99.5 / 99.6 89.3 / 95.0 / 95.5 91.5 / 94.1 / 94.7 95.4 / 96.2 / 97.6 91.4 / 93.2 / 94.1

    RAP improves the average untargeted transfer success rate across all models (by 9.6% over I-FGSM, 16.3% over MI-FGSM, 10.2% over TI, 10.9% over DI, 9.3% over SI, and 6.3% over Admix). When combined with MTDI, MTDSI, and MTDAI, RAP-LS achieves average transfer success rates of 95.4%, 97.6%, and 98.3%, respectively.

  5. Knowl 5 — Targeted Black-Box Attack Transferability of RAP and RAP-LS

    data/table

    Targeted transfer attack success rate (ASR, in %) measures the percentage of adversarial examples generated against a surrogate model that force a black-box target model to predict a pre-specified target label yty_t. Evaluated with logit loss, ℓ∞\ell_\infty budget ϵ=16/255\epsilon = 16/255, outer iterations K=400K = 400, outer step size α=2/255\alpha = 2/255, and KLS=100K_{LS} = 100:

    Attack ResNet-50 ⟹\Longrightarrow DenseNet-121 ⟹\Longrightarrow
    Dense-121 VGG-16 Inc-v3 Res-50 VGG-16 Inc-v3
    I / +RAP / +RAP-LS 4.5 / 9.5 / 14.3 2.4 / 9.8 / 11.8 0.1 / 0.1 / 0.7 5.0 / 12.8 / 17.9 2.9 / 10.1 / 15.9 0.0 / 0.8 / 1.2
    MI / +RAP / +RAP-LS 6.3 / 17.5 / 29.6 2.2 / 14.5 / 20.6 0.1 / 1.1 / 2.4 4.6 / 16.2 / 26.5 3.1 / 13.4 / 23.2 0.3 / 2.0 / 3.4
    TI / +RAP / +RAP-LS 7.2 / 11.0 / 17.3 4.0 / 12.9 / 15.3 0.1 / 0.8 / 1.2 8.4 / 13.5 / 20.8 5.2 / 12.4 / 16.4 0.2 / 2.1 / 3.0
    DI / +RAP / +RAP-LS 62.6 / 64.9 / 73.9 57.2 / 63.4 / 69.3 1.5 / 7.9 / 10.1 30.2 / 52.6 / 60.4 32.1 / 49.5 / 58.9 1.4 / 8.8 / 10.0
    SI / +RAP / +RAP-LS 30.0 / 53.2 / 61.1 9.5 / 32.8 / 36.0 1.8 / 9.3 / 10.5 14.2 / 41.5 / 43.4 8.4 / 31.0 / 35.2 1.6 / 8.5 / 10.4
    Admix / +RAP / +RAP-LS 54.6 / 68.0 / 74.6 26.0 / 45.4 / 51.6 5.8 / 17.1 / 19.6 29.3 / 53.0 / 58.2 21.5 / 42.7 / 48.2 5.0 / 17.1 / 17.6
    Attack VGG-16 ⟹\Longrightarrow Inc-v3 ⟹\Longrightarrow
    Res-50 Dense-121 Inc-v3 Res-50 Dense-121 VGG-16
    I / +RAP / +RAP-LS 0.1 / 0.7 / 1.4 0.2 / 1.4 / 1.7 0.0 / 0.1 / 0.2 0.2 / 0.9 / 0.5 0.2 / 0.6 / 0.3 0.1 / 0.5 / 0.5
    MI / +RAP / +RAP-LS 0.5 / 1.3 / 1.9 0.5 / 2.3 / 3.0 0.0 / 0.0 / 0.3 0.2 / 1.7 / 1.5 0.1 / 1.6 / 1.5 0.2 / 1.3 / 1.0
    TI / +RAP / +RAP-LS 0.7 / 1.2 / 3.2 0.8 / 1.7 / 2.9 0.0 / 0.1 / 0.4 0.2 / 0.5 / 0.7 0.1 / 0.7 / 0.6 0.2 / 0.8 / 0.6
    DI / +RAP / +RAP-LS 2.8 / 7.3 / 9.7 3.8 / 8.4 / 12.7 0.0 / 0.4 / 1.1 1.6 / 4.6 / 6.4 2.8 / 5.8 / 7.5 2.6 / 6.3 / 8.1
    SI / +RAP / +RAP-LS 3.3 / 9.8 / 9.8 7.2 / 16.8 / 17.8 0.2 / 1.7 / 1.8 0.6 / 2.9 / 2.5 0.9 / 2.7 / 3.2 0.5 / 1.5 / 2.3
    Admix / +RAP / +RAP-LS 5.6 / 11.1 / 11.9 13.0 / 20.2 / 23.6 0.7 / 2.4 / 2.8 1.5 / 4.9 / 5.2 2.0 / 6.9 / 7.5 1.3 / 3.3 / 4.4
    Combinational Attack ResNet-50 ⟹\Longrightarrow DenseNet-121 ⟹\Longrightarrow
    Dense-121 VGG-16 Inc-v3 Res-50 VGG-16 Inc-v3
    MTDI / +RAP / +RAP-LS 74.9 / 78.2 / 88.5 62.8 / 72.9 / 81.5 10.9 / 28.3 / 33.2 44.9 / 64.3 / 74.5 38.5 / 55.0 / 65.5 7.7 / 23.0 / 26.5
    MTDSI / +RAP / +RAP-LS 86.3 / 88.4 / 93.3 70.1 / 77.7 / 84.7 38.1 / 51.8 / 58.0 55.0 / 71.2 / 75.8 42.0 / 58.4 / 62.3 19.8 / 39.0 / 39.2
    MTDAI / +RAP / +RAP-LS 91.4 / 89.4 / 93.6 79.9 / 79.0 / 86.3 50.8 / 57.1 / 64.1 69.1 / 74.2 / 82.1 54.7 / 63.1 / 69.3 32.0 / 43.5 / 49.3
    Combinational Attack VGG-16 ⟹\Longrightarrow Inc-v3 ⟹\Longrightarrow
    Res-50 Dense-121 Inc-v3 Res-50 Dense-121 VGG-16
    MTDI / +RAP / +RAP-LS 11.8 / 16.7 / 22.9 13.7 / 19.4 / 27.4 0.7 / 3.4 / 4.6 1.8 / 8.3 / 7.5 4.1 / 14.8 / 13.4 2.9 / 8.0 / 9.8
    MTDSI / +RAP / +RAP-LS 31.0 / 35.3 / 38.7 41.7 / 44.4 / 49.6 9.6 / 15.2 / 13.7 5.6 / 11.9 / 10.7 10.4 / 21.2 / 20.9 4.2 / 8.9 / 8.6
    MTDAI / +RAP / +RAP-LS 36.2 / 39.0 / 43.1 48.0 / 45.1 / 55.2 11.6 / 17.1 / 17.6 9.6 / 13.6 / 16.7 17.9 / 27.5 / 31.6 8.4 / 12.0 / 12.1

    On ResNet-50 and DenseNet-121 surrogates, RAP enhances baseline targeted transferability by an average of 5.0% for I-FGSM, 8.1% for MI-FGSM, 4.6% for TI, 10.4% for DI, 18.5% for SI, and 15.1% for Admix. Across all combinational attacks, RAP-LS achieves average improvements of 14.2%, 11.8%, and 9.3% over MTDI, MTDSI, and MTDAI, respectively.

  6. Knowl 6 — Transfer Attack Performance on Diverse Architectures and Adversarially Trained Defense Models

    data/table

    Transferability across disparate model families—including Vision Transformers (ViT-Base/16), NASNet-Large, Inception-ResNet-v2 (IncRes-v2), and ensemble adversarially trained (AT) models (adv-Inception-v3 / Inc-v3adv_{adv} and ens-adv-Inception-ResNet-v2 / IncRes-v2ens_{ens})—demonstrates the generalizability of RAP-LS.

    When attacking diverse target architectures using a single ResNet-50 surrogate, and when attacking defense models using an ensemble surrogate (averaging logits across ResNet-50, ResNet-101, Inception-v3, and Inception-ResNet-v2), the attack success rates (%) under ϵ=16/255\epsilon = 16/255 and K=400K = 400 are:

    Attack Standard Diverse Target Models (ResNet-50) Defense Models (Ensemble Surrogate)
    Untargeted Targeted Untargeted Targeted
    IncRes-v2 NASNet-L ViT-B/16 IncRes-v2 NASNet-L ViT-B/16 Inc-v3adv_{adv} IncRes-v2ens_{ens} Inc-v3adv_{adv} IncRes-v2ens_{ens}
    MTDI 83.4 89.0 27.9 14.8 32.1 0.4 68.1 50.9 0.8 0.0
    MTDI+RAP-LS 95.6 97.5 42.7 43.0 62.5 1.7 86.5 72.3 9.7 4.1
    MTDSI 95.7 98.0 43.0 45.5 67.9 2.6 90.0 79.6 12.7 6.7
    MTDSI+RAP-LS 98.6 99.7 57.4 64.0 80.4 5.3 96.5 91.5 31.0 22.0
    MTDAI 97.3 98.8 45.5 58.4 75.3 3.3 92.1 82.7 17.2 12.2
    MTDAI+RAP-LS 99.2 99.8 60.2 70.4 82.6 7.4 96.7 91.6 34.4 26.0

    For the Vision Transformer target (ViT-B/16), RAP-LS increases the untargeted transferability of MTDAI from 45.5% to 60.2% and targeted transferability from 3.3% to 7.4%. Against ensemble-AT defense models, RAP-LS provides average improvements of 9.8% (Inc-v3adv_{adv}) and 14.1% (IncRes-v2ens_{ens}) for untargeted attacks, and 14.8% and 11.1% for targeted attacks.

  7. Knowl 7 — Transfer Attack Comparison against LinBP, ILA, and Generative Attacks

    data/table

    RAP was compared against the model-specific skip-connection backpropagation method LinBP (with LinBP-ILA, LinBP-ILA-SGM, LinBP-MI-DI, and LinBP-MI-DI-SGM variants), the intermediate-level feature attack ILA, and the targeted generative perturbation method TTP (Trainable Targeted Perturbations). All adversarial examples were generated using a ResNet-50 surrogate with ϵ=16/255\epsilon = 16/255.

    Comparison with feature-level and model-specific methods (ILA and LinBP variants):

    Attack Untargeted ASR (%) Targeted ASR (%)
    Dense-121 VGG-16 Inc-v3 Dense-121 VGG-16 Inc-v3
    ILA 95.0 94.2 77.7 2.8 1.5 0.5
    LinBP-ILA 99.5 99.2 89.8 9.4 4.9 2.0
    LinBP-ILA-SGM 99.7 99.3 91.1 13.3 7.2 2.8
    LinBP-MI-DI 99.5 99.2 89.3 26.1 16.5 3.2
    LinBP-MI-DI-SGM 99.8 99.3 90.2 32.6 22.1 4.6
    MI-DI+RAP 99.9 100 93.7 75.1 69.7 13.9

    MI-DI+RAP outperforms the strongest LinBP baseline (LinBP-MI-DI-SGM) by an average of 33.5% across the three target models in targeted attack success rate.

    Comparison with generative targeted attack (TTP, evaluated in the 10-Targets all-source protocol):

    Attack Dense-121 VGG-16 Inc-v3
    TTP 79.6 78.6 40.3
    MTDI 78.6 74.6 12.7
    MTDI+RAP-LS 90.8 87.2 35.4
    MTDSI 93.2 80.0 41.3
    MTDSI+RAP-LS 95.7 88.1 59.3

    MTDSI+RAP-LS surpasses TTP by 16.1% on Dense-121, 9.5% on VGG-16, and 19.0% on Inception-v3 (an average margin of 14.9%).

  8. Knowl 8 — Targeted Transfer Attack against the Google Cloud Vision API

    empirical result

    To evaluate transferability against a real-world black-box image recognition system, targeted adversarial examples were generated using a ResNet-50 surrogate model on 500 randomly sampled ImageNet images and submitted to the commercial Google Cloud Vision API.

    Because the API returns the top 10 predicted labels without strictly mirroring ImageNet-1k category semantics, an attack is counted as successful if the target class or a semantically equivalent concept appears within the returned top-10 predictions.

    Under this evaluation:

    • Baseline MTDAI (MI-TI-DI-Admix) successfully attacked 232 out of 500 images (46.4%46.4\% success rate).
    • MTDAI combined with RAP-LS (MTDAI-RAP-LS) successfully attacked 342 out of 500 images (68.4%68.4\% success rate).

    This represents an absolute improvement of 22.0%22.0\% (110 additional successfully transferred targeted attacks) over the state-of-the-art baseline on a production cloud recognition system.

  9. Knowl 9 — Hyperparameter Sensitivity of Inner Perturbation Size, Steps, and Late-Start Threshold

    empirical result

    Ablation experiments on a ResNet-50 surrogate model for targeted transfer attacks characterize the behavior of the three main hyperparameters in RAP and RAP-LS:

    1. Neighborhood radius (ϵn\epsilon_n): Evaluated across ϵn∈{2/255,4/255,8/255,12/255,16/255,20/255}\epsilon_n \in \{2/255, 4/255, 8/255, 12/255, 16/255, 20/255\}. Increasing ϵn\epsilon_n monotonically boosts the transfer attack success rate across all evaluated target models until performance stabilizes around ϵn=12/255\epsilon_n = 12/255 to 16/25516/255.
    2. Inner iteration steps (TT) and step size (αn=ϵn/T\alpha_n = \epsilon_n / T): For a given ϵn\epsilon_n, larger inner iteration counts TT (smaller inner step sizes αn\alpha_n) yield higher transferability. Fixing αn=2/255\alpha_n = 2/255 (T=ϵn/αn=8T = \epsilon_n / \alpha_n = 8 for ϵn=16/255\epsilon_n = 16/255) provides an effective trade-off between performance and computation.
    3. Late-start iteration threshold (KLSK_{LS}): Evaluated across KLS∈{0,25,50,100,150,200}K_{LS} \in \{0, 25, 50, 100, 150, 200\} with total iterations K=400K = 400. Late-start (KLS>0K_{LS} > 0) consistently outperforms standard RAP (KLS=0K_{LS} = 0). Success rates rise up to KLS=100K_{LS} = 100 and remain steady for KLS∈[100,200]K_{LS} \in [100, 200]. Setting KLS=100K_{LS} = 100 demonstrates optimal and consistent transfer performance across model architectures.
  10. Knowl 10 — Empirical Verification of Loss Landscape Flatness via Directional Slicing

    empirical result

    To empirically test whether Reverse Adversarial Perturbation (RAP) guides adversarial examples into flatter loss basins on the surrogate model MsM^s, directional loss slicing was conducted around adversarial examples generated by I-FGSM, MI-FGSM, DI, and MTDI, with and without RAP, using ResNet-50 as the surrogate.

    The loss landscape was measured by computing surrogate loss variations L(Ms(xadv+a⋅v;θ),y)\mathcal{L}(M^s(x^{adv} + a \cdot v; \theta), y) along random unit vectors vv across perturbation magnitudes a∈[−20/255,20/255]a \in [-20/255, 20/255].

    Adversarial examples created by baseline attacks reside in sharp local minima, where slight perturbations a⋅va \cdot v away from xadvx^{adv} cause rapid, steep increases in attack loss. In contrast, adversarial examples generated with RAP exhibit a significantly wider, flatter basin where the loss remains consistently low across the entire neighborhood a∈[−20/255,20/255]a \in [-20/255, 20/255], explaining why they remain effective under parameter perturbations and decision boundary variations across different target models.

Coverage note — Omitted additional supplementary appendix experiments (such as evaluations on NRP, feature denoising, multi-step AT defense models, and comparisons with VT, EMI, and Ghost Net) as these are supplementary validation extensions of the main paper benchmarks already fully captured in the knowls.

References

  1. 1.Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. Synthesizing robust adversarial examples. In International conference on machine learning, pages 284–293. PMLR, 2018. 2, 4.1, C, C.3
  2. 2.Minhao Cheng, Thong Le, Pin-Yu Chen, Huan Zhang, JinFeng Yi, and Cho-Jui Hsieh. Query-efficient hard-label black-box attack: An optimization-based approach. In International Conference on Learning Representations, 2019. 1
  3. 3.Minhao Cheng, Simranjit Singh, Patrick H. Chen, Pin-Yu Chen, Sijia Liu, and Cho-Jui Hsieh. Sign-opt: A query-efficient hard-label adversarial attack. In International Conference on Learning Representations, 2020. 1
  4. 4.Ambra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski, Battista Biggio, Alina Oprea, Cristina Nita-Rotaru, and Fabio Roli. Why do adversarial attacks transfer? explaining transferability of evasion and poisoning attacks. In 28th USENIX security symposium (USENIX security 19), pages 321–338, 2019. 3.2
  5. 5.Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 9185–9193, 2018. 1, 2, 3.1, 3.2, 3.2, 3.3, 4.1
  6. 6.Yinpeng Dong, Tianyu Pang, Hang Su, and Jun Zhu. Evading defenses to transferable adversarial examples by translation-invariant attacks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4312–4321, 2019. 1, 2, 3.1, 3.2, 3.2, 4.1
  7. 7.Yinpeng Dong, Qi-An Fu, Xiao Yang, Tianyu Pang, Hang Su, Zihao Xiao, and Jun Zhu. Benchmarking adversarial robustness on image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 321–331, 2020. D
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. 4.1
  9. 9.Yanbo Fan, Baoyuan Wu, Tuanhui Li, Yong Zhang, Mingyang Li, Zhifeng Li, and Yujiu Yang. Sparse adversarial attack via perturbation factorization. In European conference on computer vision, pages 35–50. Springer, 2020. 1
  10. 10.Yan Feng, Baoyuan Wu, Yanbo Fan, Li Liu, Zhifeng Li, and Shu-Tao Xia. Boosting black-box attack with partially transferred conditional adversarial distribution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15095–15104, June 2022. 1
  11. 11.Lianli Gao, Qilong Zhang, Jingkuan Song, Xianglong Liu, and Heng Tao Shen. Patch-wise attack for fooling deep neural network. In European Conference on Computer Vision, pages 307–322. Springer, 2020. B
  12. 12.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014. 1, 2
  13. 13.Martin Gubri, Maxime Cordy, Mike Papadakis, Yves Le Traon, and Koushik Sen. Lgv: Boosting adversarial example transferability from large geometric vicinity. arXiv preprint arXiv:2207.13129, 2022. 2
  14. 14.Yiwen Guo, Qizhang Li, and Hao Chen. Backpropagating linearly improves transferability of adversarial examples. Advances in Neural Information Processing Systems, 33:85–95, 2020. 1, 2, 4.1, 4.4, B
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 4.1
  16. 16.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017. 4.1
  17. 17.Qian Huang, Isay Katsman, Horace He, Zeqi Gu, Serge Belongie, and Ser-Nam Lim. Enhancing adversarial example transferability with an intermediate level attack. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4733–4742, 2019. 1, 2, 4.1, 4.4
  18. 18.Andrew Ilyas, Logan Engstrom, Anish Athalye, and Jessy Lin. Black-box adversarial attacks with limited queries and information. In International Conference on Machine Learning, pages 2137–2146. PMLR, 2018. 1
  19. 19.Nathan Inkawhich, Kevin Liang, Lawrence Carin, and Yiran Chen. Transferable perturbations of deep feature distributions. In International Conference on Learning Representations, 2020. 2
  20. 20.Nathan Inkawhich, Kevin Liang, Binghui Wang, Matthew Inkawhich, Lawrence Carin, and Yiran Chen. Perturbing across the feature hierarchy to improve standard and strict blackbox attack transferability. Advances in Neural Information Processing Systems, 33, 2020. 2
  21. 21.Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. Adversarial examples in the physical world. Artificial Intelligence Safety and Security, page 99–112, Jul 2018. doi: 10.1201/9781351251389-8. 2, 3.1, 3.3, 4.1
  22. 22.Yingwei Li, Song Bai, Yuyin Zhou, Cihang Xie, Zhishuai Zhang, and Alan Yuille. Learning transferable adversarial examples via ghost networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11458–11465, 2020. 4.1, C
  23. 23.Siyuan Liang, Baoyuan Wu, Yanbo Fan, Xingxing Wei, and Xiaochun Cao. Parallel rectangle flip attack: A query-based black-box attack against object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7697–7707, 2021. 1
  24. 24.Jiadong Lin, Chuanbiao Song, Kun He, Liwei Wang, and John E. Hopcroft. Nesterov accelerated gradient and scale invariance for adversarial attacks. In International Conference on Learning Representations, 2020. 1, 2, 3.1, 3.2, 4.1
  25. 25.Risheng Liu, Jiaxin Gao, Jin Zhang, Deyu Meng, and Zhouchen Lin. Investigating bi-level optimization for learning and vision from a unified perspective: A survey and beyond. arXiv preprint arXiv:2101.11517, 2021. 3.2
  26. 26.Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. Delving into transferable adversarial examples and black-box attacks. arXiv preprint arXiv:1611.02770, 2016. 1
  27. 27.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations, 2018. 1
  28. 28.Muhammad Muzammal Naseer, Salman H Khan, Muhammad Haris Khan, Fahad Shahbaz Khan, and Fatih Porikli. Cross-domain transferability of adversarial perturbations. Advances in Neural Information Processing Systems, 32:12905–12915, 2019. 2
  29. 29.Muzammal Naseer, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Fatih Porikli. A self-supervised approach for adversarial robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 262–271, 2020. 4.1, 4.5, D
  30. 30.Muzammal Naseer, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Fatih Porikli. On generating transferable targeted perturbations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7708–7717, 2021. 1, 2, 4.1, 4.4, 4.4, B
  31. 31.Omid Poursaeed, Isay Katsman, Bicheng Gao, and Serge Belongie. Generative adversarial perturbations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4422–4431, 2018. 2
  32. 32.Zeyu Qin, Yanbo Fan, Hongyuan Zha, and Baoyuan Wu. Random noise defense against query-based black-box attacks. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021. 1
  33. 33.Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor, and Aleksander Madry. Do adversarially robust imagenet models transfer better? In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 3533–3545. Curran Associates, Inc., 2020. 4.1, 4.5, D
  34. 34.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 4.1
  35. 35.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013. 1
  36. 36.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016. 4.1
  37. 37.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In Thirty-first AAAI conference on artificial intelligence, 2017. 4.1
  38. 38.Florian Tramèr, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. Ensemble adversarial training: Attacks and defenses. In International Conference on Learning Representations, 2018. 1, 3.2, 4.1, 4.5
  39. 39.Jingkang Wang, Tianyun Zhang, Sijia Liu, Pin-Yu Chen, Jiacen Xu, Makan Fardad, and Bo Li. Adversarial attack generation empowered by min-max optimization. Advances in Neural Information Processing Systems, 34, 2021. 2
  40. 40.Xiaosen Wang and Kun He. Enhancing the transferability of adversarial attacks through variance tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1924–1933, June 2021. 2, 4.1, C
  41. 41.Xiaosen Wang, Xuanran He, Jingdong Wang, and Kun He. Admix: Enhancing the transferability of adversarial attacks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16158–16167, 2021. 2, 3.1, 3.2, 3.2, 4.1, 4.2, 4.5, B
  42. 42.Xiaosen Wang, Jiadong Lin, Han Hu, Jingdong Wang, and Kun He. Boosting adversarial transferability through enhanced momentum. arXiv preprint arXiv:2103.10609, 2021. 4.1, C
  43. 43.Dongxian Wu, Yisen Wang, Shu-Tao Xia, James Bailey, and Xingjun Ma. Skip connections matter: On the transferability of adversarial examples generated with resnets. In International Conference on Learning Representations, 2020. 2
  44. 44.Jiancong Xiao, Yanbo Fan, Ruoyu Sun, and Zhi-Quan Luo. Adversarial rademacher complexity of deep neural networks, 2022. 1
  45. 45.Jiancong Xiao, Yanbo Fan, Ruoyu Sun, Jue Wang, and Zhi-Quan Luo. Stability analysis and generalization bounds of adversarial training. arXiv preprint arXiv:2210.00960, 2022. 1
  46. 46.Cihang Xie, Jianyu Wang, Zhishuai Zhang, Zhou Ren, and Alan Yuille. Mitigating adversarial effects through randomization. In International Conference on Learning Representations, 2018. 3.2, 4.1, 4.5, D
  47. 47.Cihang Xie, Yuxin Wu, Laurens van der Maaten, Alan L Yuille, and Kaiming He. Feature denoising for improving adversarial robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 501–509, 2019. 4.1, 4.5, D, D
  48. 48.Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai, Jianyu Wang, Zhou Ren, and Alan L Yuille. Improving transferability of adversarial examples with input diversity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2730–2739, 2019. 1, 2, 3.1, 3.2, 3.3, 4.1, 4.3, 4.5
  49. 49.Zhengyu Zhao, Zhuoran Liu, and Martha Larson. On success and simplicity: A second look at transferable targeted attacks. arXiv preprint arXiv:2012.11207, 2020. 2, 3.1, 4.1, 4.2, 4.3, 4.4, B, E.3
  50. 50.Xin Zheng, Yanbo Fan, Baoyuan Wu, Yong Zhang, Jue Wang, and Shirui Pan. Robust physical-world attacks on face recognition. Pattern Recognition, page 109009, 2022. 1
  51. 51.Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le. Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8697–8710, 2018. 4.1

Citation

MLA
Qin, Z., et al. “Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 29845–58, https://proceedings.neurips.cc/paper_files/paper/2022/file/c0f9419caa85d7062c7e6d621a335726-Paper-Conference.pdf.
APA
Qin, Z., Fan, Y., Liu, Y., Shen, L., Zhang, Y., Wang, J., & Wu, B. (2022). Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation. Advances in Neural Information Processing Systems, 35, 29845–29858. https://proceedings.neurips.cc/paper_files/paper/2022/file/c0f9419caa85d7062c7e6d621a335726-Paper-Conference.pdf
Chicago
Qin, Z., Y. Fan, Y. Liu, et al. 2022. “Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation”. Advances in Neural Information Processing Systems 35: 29845–58. https://proceedings.neurips.cc/paper_files/paper/2022/file/c0f9419caa85d7062c7e6d621a335726-Paper-Conference.pdf.
Harvard
Qin, Z. et al. (2022) “Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 29845–29858. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/c0f9419caa85d7062c7e6d621a335726-Paper-Conference.pdf.
Vancouver
1. Qin Z, Fan Y, Liu Y, Shen L, Zhang Y, Wang J, Wu B (2022) Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 29845–29858

BibTeX

@inproceedings{qin2022boosting,
  title = {Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation},
  author = {Qin, Zeyu and Fan, Yanbo and Liu, Yi and Shen, Li and Zhang, Yong and Wang, Jue and Wu, Baoyuan},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {29845-29858},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/c0f9419caa85d7062c7e6d621a335726-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors