Adversarial Machine Learning at Scale
Alexey KurakinIan GoodfellowSamy Bengio
Establishes practical methods to scale adversarial training to large datasets like ImageNet, resolving the label leaking problem and uncovering critical differences in attack transferability.
Adversarial examples are small, often imperceptible changes to inputs that cause machine learning models to misclassify them, and they frequently transfer across different models to enable attacks without access to the target system’s parameters. This vulnerability creates security risks for deployed systems, yet prior defenses such as adversarial training had been demonstrated only on small datasets like MNIST and CIFAR-10.
The work set out to scale adversarial training to large models on the ImageNet dataset and to measure resulting robustness against both one-step and iterative attack methods. Researchers trained Inception v3 models using synchronous distributed training across 50 machines, mixing clean and adversarially perturbed examples in each minibatch, and evaluated performance on the 50,000-image ImageNet validation set.
Adversarial training with one-step methods raised top-1 accuracy on one-step adversarial examples from roughly 30 percent to 73–75 percent while reducing clean-image accuracy by less than 1 percent. Larger models, created by increasing the number of filters, gained more robustness from the same procedure and showed higher ratios of adversarial to clean accuracy. Iterative attacks remained effective against the trained models, but those same iterative examples transferred to other models at markedly lower rates than one-step examples. A label-leaking artifact appeared when training and testing used the fast gradient sign method, artificially inflating accuracy on adversarial inputs; the effect disappeared when label-free one-step methods were substituted.
These results indicate that adversarial training can be applied at ImageNet scale to protect against practical one-step attacks with only modest cost to normal accuracy, while larger models amplify the benefit. The lower transferability of iterative attacks offers partial protection against black-box threats even when direct robustness is limited. Practitioners should therefore adopt adversarial training with non-label-leaking one-step methods during both training and evaluation, pair it with increased model capacity, and avoid relying on the fast gradient sign method for robustness testing.
Further work is needed to develop defenses that also resist iterative attacks, possibly through still-larger models or new training regimes, and to confirm whether the observed scaling trends continue beyond the architectures tested here. The study’s main limitations are its focus on a single dataset and family of models, the computational expense of iterative training, and the absence of exhaustive hyperparameter sweeps for very large networks.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). Reading this foundational 2015 study on adversarial examples and the fast gradient sign method is essential to understanding the core attack mechanisms scaled up in the source paper.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, A. Ma̧dry et al. (2017). This paper extends the source's adversarial training concepts by formulating robust optimization as a saddle-point problem and utilizing projected gradient descent for defense.
