Improving Transferability of Adversarial Examples With Input Diversity
Cihang XieZhishuai ZhangJianyu WangYuyin ZhouZhou RenA. Yuille
Proposes the Diverse Inputs method, which applies random transformations during iterative gradient optimization to prevent overfitting and significantly improve the black-box transferability of adversarial examples across convolutional neural networks.
Deep learning models are increasingly deployed in mission-critical applications such as autonomous driving and medical diagnostics, yet they remain vulnerable to adversarial examples—images altered with imperceptible noise that trick models into making severe classification errors. While attackers can easily mislead a system when its internal architecture and parameters are fully known (the white-box setting), these manipulated images historically struggle to fool target systems when their underlying parameters are unknown (the black-box setting). This occurs because standard optimization techniques overfit to the source network, limiting their transferability across diverse, real-world deployment environments.
The article demonstrates an effective method to enhance the transferability of adversarial examples by introducing input diversity during image generation. The approach applies random, differentiable transformations—specifically random resizing and random padding—at each optimization step to prevent adversarial noise from overfitting to a specific network architecture.
To evaluate this technique, the researchers conducted extensive empirical evaluations on standard benchmark datasets, using 5,000 correctly classified images from ImageNet. They tested several standard and adversarially trained convolutional networks under single-model and multi-model ensemble scenarios. Furthermore, they tested their approach against top defense systems and baseline models from the NIPS 2017 Adversarial Competition.
The findings show that introducing input diversity substantially improves attack transferability while maintaining nearly 100% white-box success rates. When paired with momentum optimization and multi-network ensembles, the combined method attained an average attack success rate of 73.0% across leading competitive defense mechanisms, outperforming the competition-winning baseline by 6.6 percentage points. When targeting individual black-box models, input diversity alone doubled or tripled transfer success rates compared to standard iterative attacks, and it proved similarly effective when integrated into alternative attack frameworks.
These results demonstrate that many prevailing security defenses, including transformation-based inference mitigations and adversarial training, provide less protection in black-box environments than previously assumed. This indicates a heightened operational risk for systems relying on security through obscurity or simple input defenses, underscoring that current defenses can be bypassed if an adversary crafts inputs that generalize across varied network transformations.
Organizations developing or deploying safety-critical computer vision models should adopt input-diversity attacks as a standard benchmark to stress-test system robustness. Reliance solely on basic input transformations or single-model defenses should be reconsidered in favor of more robust defenses that account for transformation-invariant adversarial noise. Future work should further investigate the underlying mathematical properties of shared decision boundaries and validate the approach across non-vision tasks and newer network architectures.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). This paper establishes foundational methodologies for generating transferable adversarial examples across deep architectures in black-box ImageNet settings, which the source directly aims to improve via input diversity.
- Paper: Adversarial Machine Learning at Scale, Alexey Kurakin et al. (2016). It introduces iterative fast gradient methods at scale on ImageNet and highlights the issue of transferability under iterative attacks, forming the algorithmic baseline for DI-2-FGSM.
- Paper: Synthesizing Robust Adversarial Examples, Anish Athalye et al. (2017). It introduces the Expectation Over Transformation framework for applying random transformations during optimization, which inspires the random input transformations used in the source attack.
- Paper: Countering Adversarial Images using Input Transformations, Chuan Guo et al. (2018). It analyzes using image transformations as defensive countermeasures on ImageNet, providing the key defense paradigm that input diversity attacks are designed to circumvent.
- Paper: Ensemble Adversarial Training: Attacks and Defenses, Florian Tramèr et al. (2018). It details ensemble adversarial training and random single-step perturbations to resist transferable attacks, serving as a primary defense baseline evaluated in the source.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). It defines the foundational Fast Gradient Sign Method (FGSM) upon which iterative and diverse-input gradient-based attacks are directly built.
- Paper: Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, Nicolas Papernot et al. (2016). It introduces the practical black-box threat model relying on adversarial transferability between distinct classifiers.
- Paper: Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples, Anish Athalye et al. (2018). It systematically examines gradient obfuscation in defenses, showing how transformation-based defenses and randomization can be circumvented via techniques closely tied to input transformation handling.
- Paper: Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks, Francesco Croce et al. (2020). It builds an ensemble of parameter-free attacks to reliably assess robustness, advancing standardized benchmarking beyond iterative transfer attacks.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). It provides theoretical and empirical insights into why transferable adversarial perturbations exist by framing them as non-robust dataset features.
- Paper: AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty, Dan Hendrycks et al. (2019). It adapts the principle of randomized input transformations and mixing as a defensive training augmentation to improve general robustness against corruptions and perturbations.
- Paper: Manifold Mixup: Better Representations by Interpolating Hidden States, Vikas Verma et al. (2018). It extends data diversity and interpolation beyond input pixel space into hidden internal states to smooth decision boundaries against adversarial perturbations.
