Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation
Zeyu QinYanbo FanYi LiuLi ShenYong ZhangJue WangBaoyuan Wu
Proposes reverse adversarial perturbation, a min-max bi-level optimization technique that prevents attacks from overfitting to surrogate models by targeting flat loss regions, substantially improving black-box attack transferability across both standard networks and commercial vision systems.
Deep neural networks are central to safety-critical applications like automated vision and face recognition, but they remain susceptible to adversarial examples—minor, imperceptible modifications to input data that fool models into incorrect classifications. In practical black-box settings, attackers lack direct access to internal target models and instead generate adversarial examples using known substitute models. These attacks often fail to transfer across different systems because the generated perturbations overfit the specific geometry of the substitute model, landing in sharp decision regions that do not generalize.
The article introduces a method called Reverse Adversarial Perturbation (RAP) to systematically improve the transferability of black-box attacks. The primary objective is to evaluate whether driving adversarial optimization toward broader, flatter loss regions—rather than single sharp points—prevents substitute overfitting and improves cross-model attack success.
To accomplish this, the authors model adversarial generation as a two-level optimization problem. The inner step computes a worst-case shift within a defined neighborhood to identify vulnerability, while the outer step updates the attack sample to ensure low error across that entire local area. To enhance early computational efficiency, the authors also introduce a late-start variation (RAP-LS). The approach was evaluated across standard benchmarks, testing multiple model architectures (such as Inception, ResNet, DenseNet, VGG, and Vision Transformers), robust defense techniques, and a real-world commercial platform using 1,000 ImageNet samples and 500 real-world API queries.
The findings show that finding flatter loss regions significantly boosts attack success across unknown models. When integrated with foundational attack methods, RAP increases average untargeted attack success rates by roughly 6% to 16% and targeted success rates by up to 18.5%. Combining RAP with advanced input-transformation methods achieves untargeted transfer success rates between 95% and 98%, while targeted baseline attacks improved by approximately 9% to 14%. Furthermore, RAP-LS consistently outperformed competitive feature-based and generative attack approaches, improving targeted attack rates against defense-hardened models by 11% to 15%. In a real-world evaluation against the Google Cloud Vision API, RAP-LS achieved a 22.0% absolute increase in targeted attack success over existing baseline techniques.
These findings demonstrate that black-box adversarial transferability poses a severe and practical risk to commercial artificial intelligence deployments. Models previously considered resilient due to hidden internal parameters or defensive training remain vulnerable to neighborhood-optimized attacks. Consequently, system architects and security teams cannot rely on model obscurity or isolated input defenses. Organizations operating computer vision systems should proactively audit existing architectures against multi-model transfer attacks and prioritize structural defenses, such as diverse ensemble pipelines and robust loss regularization.
While the empirical results provide strong evidence across diverse model families, the study focuses predominantly on image classification benchmarks and standard perturbation limits. Readers should note that computational costs increase during neighborhood searches, and transferring attacks to substantially different architectures, such as Vision Transformers, remains inherently more challenging. Future work should focus on developing advanced defense mechanisms tailored against flat-loss attacks and evaluating transfer dynamics in broader domains such as object detection and natural language processing.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). This foundational study examines the transferability of adversarial perturbations across vision models and introduces ensemble-based black-box attacks, establishing the exact problem setting that reverse adversarial perturbation seeks to improve.
- Paper: Improving Transferability of Adversarial Examples With Input Diversity, Cihang Xie et al. (2018). This paper presents input diversity as an optimization technique to prevent substitute model overfitting, serving as a primary baseline and complementary method evaluated alongside reverse adversarial perturbation.
- Paper: Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, Nicolas Papernot et al. (2016). This work formalizes the concept of black-box adversarial transferability using local substitute models, which constitutes the foundational threat model analyzed in the source paper.
- Paper: Ensemble Adversarial Training: Attacks and Defenses, Florian Tramèr et al. (2018). This article analyzes how sharp local curvature in the loss landscape leads to substitute model overfitting and attack failure, directly motivating the flat-loss neighborhood optimization strategy.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This seminal paper introduces the Fast Gradient Sign Method and the fundamental principles of adversarial sample generation upon which modern gradient-based transfer attacks are built.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). This work establishes the min-max robust optimization formulation for adversarial perturbations, providing the mathematical framing adapted by the two-level optimization in reverse adversarial perturbation.
- Paper: Countering Adversarial Images using Input Transformations, Chuan Guo et al. (2018). This paper evaluates input transformation defenses against adversarial images, providing key defensive benchmark baselines that the source paper tests its transferable attacks against.
- Paper: Practical Black-Box Attacks against Machine Learning, Nicolas Papernot et al. (2017). This study demonstrates practical black-box attacks against commercial cloud vision APIs via substitute models, establishing the real-world evaluation paradigm used in the source paper.
- Paper: Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition, Zexin Li et al. (2023). Extends the study of black-box adversarial transferability from general image classification to commercial face recognition APIs using multi-task gradient stabilization.
- Paper: Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains, Qilong Zhang et al. (2022). Generalizes black-box adversarial transfer attacks by disrupting domain-agnostic intermediate features to fool unknown classifiers across completely different visual domains.
- Paper: RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts, Han Liu et al. (2023). Expands black-box adversarial perturbation techniques from traditional discriminative classifiers to modern generative text-to-image foundation models.
