Fast is better than free: Revisiting adversarial training

Eric WongL. RiceJ. Kolter

article2020ICLR1,408 citations

Demonstrates that combining the single-step Fast Gradient Sign Method with random initialization matches the defensive strength of expensive multi-step adversarial training while cutting computation times from hours to minutes.

Listen

Deep neural networks are vulnerable to adversarial attacks, which are small, intentionally designed input perturbations that deceive machine learning models into making incorrect predictions. Developing robust networks is critical for security-sensitive applications, but standard defensive training methods—specifically multi-step Projected Gradient Descent (PGD)—require substantial computational resources, frequently inflating training times by up to an order of magnitude. This high computational burden has hindered rapid experimentation and scalability across large networks and datasets.

The article evaluates whether single-step Fast Gradient Sign Method (FGSM) adversarial training—previously considered ineffective against strong attacks—can produce robust models when paired with proper initialization and efficient training techniques. The authors demonstrate that this simplified approach achieves defensive performance comparable to multi-step methods while requiring a fraction of the computational time and cost.

To establish these results, the authors conducted empirical experiments across standard computer vision benchmarks (MNIST, CIFAR-10, and ImageNet) using residual neural networks. They combined uniform random initialization with FGSM and integrated established training optimizations, including cyclical learning rates, mixed-precision arithmetic, and progressive image resizing. Exact formal verification via mixed-integer linear programming on MNIST was also utilized to confirm that the observed robustness was genuine and mathematically provable.

The primary findings show that FGSM with uniform random initialization produces robust accuracy equivalent to multi-step PGD defenses, disproving long-standing assumptions about the weakness of single-step training. On CIFAR-10, the method trained a robust model (45% robust accuracy against strong PGD attacks) in just 6 minutes, compared to 10 hours for previous "free" adversarial training and 80 hours for standard PGD. On ImageNet, the approach achieved 43% robust accuracy in 12 hours, versus roughly 50 hours required by prior methods. Additionally, the authors identified and characterized "catastrophic overfitting"—a failure mode where robust accuracy drops to 0% during training due to a lack of perturbation diversity or excessive step size—and showed that it can be monitored and mitigated using lightweight early stopping.

These findings indicate that achieving adversarial robustness does not inherently require prohibitive computational budgets. Engineering teams can dramatically cut the hardware costs, energy consumption, and turnaround times associated with training secure models. Furthermore, the findings show that defense mechanisms during training do not require complex, multi-step adversaries, provided the single-step perturbations adequately span the threat model space.

Organizations training robust models should replace costly multi-step inner optimization loops with randomly initialized FGSM, paired with cyclical learning rates and mixed-precision computation. Teams adopting this method should set the perturbation step size slightly above the threat radius (e.g., 1.25 times epsilon) and incorporate periodic multi-step evaluation on a single training minibatch to guard against catastrophic overfitting.

A key limitation is that exact mathematical verification of robustness remains computationally restricted to small networks and datasets, meaning that robustness on larger benchmarks like CIFAR-10 and ImageNet must rely on empirical evaluations. In addition, hyperparameters such as step sizes and learning rate schedules require careful management to prevent sudden overfitting failures. However, across the empirical and verified evaluations tested, confidence in the method's ability to match traditional multi-step defenses is high.

Cover for Fast is better than free: Revisiting adversarial training

Abstract

Adversarial training, a method for learning robust deep networks, is typically assumed to be more expensive than traditional training due to the necessity of constructing adversarial examples via a first-order method like projected gradient decent (PGD). In this paper, we make the surprising discovery that it is possible to train empirically robust models using a much weaker and cheaper adversary, an approach that was previously believed to be ineffective, rendering the method no more costly than standard training in practice. Specifically, we show that adversarial training with the fast gradient sign method (FGSM), when combined with random initialization, is as effective as PGD-based training but has significantly lower cost. Furthermore we show that FGSM adversarial training can be further accelerated by using standard techniques for efficient training of deep networks, allowing us to learn a robust CIFAR10 classifier with 45% robust accuracy to PGD attacks with ϵ=8/255\epsilon=8/255 in 6 minutes, and a robust ImageNet classifier with 43% robust accuracy at ϵ=2/255\epsilon=2/255 in 12 hours, in comparison to past work based on "free" adversarial training which took 10 and 50 hours to reach the same respective thresholds. Finally, we identify a failure mode referred to as "catastrophic overfitting" which may have caused previous attempts to use FGSM adversarial training to fail. All code for reproducing the experiments in this paper as well as pretrained model weights are at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Adversarial training overview
  • 3.1 “Free” adversarial training
  • 4 Fast adversarial training
  • 4.1 Revisiting FGSM adversarial training
  • 4.2 DAWNBench improvements
  • 5 Experiments
  • 5.1 Verified performance on MNIST
  • 5.2 Fast CIFAR10
  • 5.3 Fast ImageNet
  • 5.4 Catastrophic overfitting
  • 5.5 Takeaways from FGSM adversarial training
  • 6 Conclusion
  • References
  • A A direct comparison to R+FGSM from Tramèr et al. (2017)
  • B Training parameters for Table
  • C Optimal step size for FGSM adversarial training
  • D Catastrophic overfitting and the effect of early stopping
  • E Training parameters for Figure
  • F Combining free adversarial training with DAWNBench improvements on ImageNet

Knowls

  1. Knowl 1 — Fast FGSM Adversarial Training with Random Initialization

    algorithm

    Fast adversarial training uses a single-step Fast Gradient Sign Method (FGSM) adversary initialized with a uniform random perturbation to achieve robustness against multi-step projected gradient descent (PGD) attacks at a computational cost equivalent to two standard training epochs per pass.

    Input: Dataset D={(xi,yi)}i=1MD = \{(x_i, y_i)\}_{i=1}^M, network fθf_\theta parameterized by θ\theta, loss function ℓ\ell, perturbation radius ϵ>0\epsilon > 0, adversarial step size α\alpha (default α=1.25ϵ\alpha = 1.25\epsilon), total training epochs TT, optimizer update rule OptimizerUpdate\text{OptimizerUpdate}
    Output: Robust network parameters θ\theta
    for t=1t = 1 to TT do
        for each sample or minibatch (xi,yi)(x_i, y_i) in DD do
            Sample initial perturbation δ∼Uniform(−ϵ,ϵ)\delta \sim \text{Uniform}(-\epsilon, \epsilon)
            Compute input gradient: gδ=∇δℓ(fθ(xi+δ),yi)g_\delta = \nabla_\delta \ell(f_\theta(x_i + \delta), y_i)
            Update perturbation: δ←δ+α⋅sign(gδ)\delta \leftarrow \delta + \alpha \cdot \text{sign}(g_\delta)
            Project perturbation: δ←max⁡(min⁡(δ,ϵ),−ϵ)\delta \leftarrow \max(\min(\delta, \epsilon), -\epsilon)
            Compute parameter gradient: gθ=∇θℓ(fθ(xi+δ),yi)g_\theta = \nabla_\theta \ell(f_\theta(x_i + \delta), y_i)
            Update model parameters: θ←OptimizerUpdate(θ,gθ)\theta \leftarrow \text{OptimizerUpdate}(\theta, g_\theta)
        end for
    end for
    return θ\theta

    Unlike standard FGSM training which initializes from δ=0\delta = 0, uniform random initialization forces generated adversarial samples to populate the interior of the ℓ∞\ell_\infty ball rather than solely discrete boundary points {−ϵ,0,ϵ}\{-\epsilon, 0, \epsilon\}, preventing the model from fitting to degenerate single-step attacks.

  2. Knowl 2 — Catastrophic Overfitting in Adversarial Training

    definition

    Catastrophic overfitting is a training failure mode in single-step (FGSM-based) adversarial training where a model's robust accuracy against iterative adversaries (such as multi-step PGD) abruptly collapses to 0%0\% within very few training epochs, even as the model continues to exhibit high accuracy and low loss against the single-step FGSM training adversary.

    This phenomenon occurs when training perturbations lack sufficient geometric diversity across the threat model Δ={δ:∥δ∥∞≤ϵ}\Delta = \{\delta : \|\delta\|_\infty \le \epsilon\}, such as when perturbations are initialized at zero or when the step size is set too large (e.g., α=2ϵ\alpha = 2\epsilon, forcing all perturbations onto the boundary). In response, the network learns a highly distorted local decision boundary that fools single-step gradient attacks but leaves the model vulnerable to multi-step optimization.

  3. Knowl 3 — Minibatch PGD Early Stopping for Catastrophic Overfitting Detection

    model/method

    To monitor and mitigate catastrophic overfitting during FGSM adversarial training, a lightweight early stopping check is evaluated at the conclusion of each training epoch.

    The check computes robust accuracy on a single training minibatch using a 5-step PGD adversary with 1 random restart within the threat model ∥δ∥∞≤ϵ\|\delta\|_\infty \le \epsilon. If the robust classification accuracy on this single minibatch suddenly drops toward 0%0\%, catastrophic overfitting is triggered, and training is terminated to preserve the model parameters from the peak-performance epoch immediately preceding the collapse.

  4. Knowl 4 — CIFAR-10 Benchmark Comparison of Fast Adversarial Training

    data/table

    Models using the PreAct ResNet-18 architecture were trained on CIFAR-10 to evaluate ℓ∞\ell_\infty robustness at radius ϵ=8/255\epsilon = 8/255. Robust accuracy was evaluated using a 50-step PGD attack with step size α=2/255\alpha = 2/255 and 10 random restarts. Training times were measured on a single NVIDIA GeForce RTX 2080 Ti GPU.

    Method Standard Accuracy PGD Accuracy (ϵ=8/255\epsilon = 8/255) Total Time (min)
    FGSM + DAWNBench (+ zero init) 85.18% 0.00% 12.37
    FGSM + DAWNBench (+ early stopping) 71.14% 38.86% 7.89
    FGSM + DAWNBench (+ previous init) 86.02% 42.37% 12.21
    FGSM + DAWNBench (+ random init, α=8/255\alpha = 8/255) 85.32% 44.01% 12.33
    FGSM + DAWNBench (+ random init, α=10/255\alpha = 10/255) 83.81% 46.06% 12.17
    FGSM + DAWNBench (+ random init, α=16/255\alpha = 16/255) 86.05% 0.00% 12.06
    FGSM + DAWNBench (+ α=16/255\alpha = 16/255 + early stop) 70.93% 40.38% 8.81
    Free (m=8m = 8) 85.96% 46.33% 785.00
    Free (m=8m = 8) + DAWNBench 78.38% 46.18% 20.91
    PGD-7 87.30% 45.80% 4965.71
    PGD-7 + DAWNBench 82.46% 50.69% 68.80

    When combined with cyclic learning rates and mixed-precision arithmetic, FGSM with uniform random initialization and step size α=10/255\alpha = 10/255 achieves 46.06%46.06\% robust accuracy, matching or exceeding Free (m=8m=8) and PGD-7 training while requiring only 12.1712.17 minutes (or 6.346.34 minutes over 15 epochs to reach 45%45\% robust accuracy).

  5. Knowl 5 — ImageNet Robust Classification with Fast FGSM and Progressive Resizing

    data/table

    ResNet-50 classifiers were trained on ImageNet against ℓ∞\ell_\infty perturbations of radii ϵ=2/255\epsilon = 2/255 and ϵ=4/255\epsilon = 4/255. Fast FGSM training used 15 epochs total divided across three progressive resizing phases (6 epochs at 160×160160 \times 160, 6 epochs at 352×352352 \times 352, and 3 epochs at full resolution) with mixed-precision arithmetic (Apex AMP O1). Evaluation was conducted against a 50-step PGD attack with 1 restart (PGD+1) and 10 random restarts (PGD+10) on a machine with 4 NVIDIA GeForce RTX 2080 Ti GPUs.

    Method ϵ\epsilon Standard Accuracy PGD+1 Restart PGD+10 Restarts Total Time (hrs)
    Fast FGSM 2/255 60.90% 43.46% 43.43% 12.14
    Free (m=4m = 4) 2/255 64.37% 43.31% 43.28% 52.20
    Fast FGSM 4/255 55.45% 30.28% 30.18% 12.14
    Free (m=4m = 4) 4/255 60.42% 31.22% 31.08% 52.20

    Fast FGSM training reaches equivalent robust accuracy (43.43%43.43\% at ϵ=2/255\epsilon = 2/255) in 12.1412.14 hours, compared to 52.2052.20 hours for Free (m=4m=4) adversarial training.

  6. Knowl 6 — Exact Verification of FGSM Adversarial Robustness on MNIST

    empirical result

    To establish whether FGSM adversarial training confers genuine robustness or merely causes gradient masking/obfuscation, small convolutional networks (two convolutional layers with 16 and 32 filters, followed by a 100-unit fully connected layer) trained on MNIST at ϵ=0.3\epsilon = 0.3 were verified using Mixed-Integer Linear Programming (MILP).

    Method Standard Accuracy PGD (ϵ=0.1\epsilon = 0.1) PGD (ϵ=0.3\epsilon = 0.3) Verified (ϵ=0.1\epsilon = 0.1)
    PGD-40 Adversarial Training 99.20% 97.66% 89.90% 96.7%
    Fast FGSM Adversarial Training 99.20% 97.53% 88.77% 96.8%

    Fast FGSM adversarial training achieves a certified exact robustness of 96.8%96.8\% at ϵ=0.1\epsilon = 0.1, matching the 96.7%96.7\% certified robustness of PGD-40 adversarial training, verifying that the empirical defense is non-trivial and mathematically sound.

  7. Knowl 7 — Ablation of Random Initialization Schemes for FGSM Training

    data/table

    An ablation study on MNIST (with ℓ∞\ell_\infty radius ϵ=0.3\epsilon = 0.3) compares uniform initialization against the R+FGSM strategy of initializing on the surface of a hypercube of radius ϵ/2\epsilon/2, evaluated across 10 random seeds against a PGD adversary.

    Method Step Size α\alpha Initialization Distribution Robust Accuracy
    R+FGSM 0.15 Hypercube(0.15)\text{Hypercube}(0.15) 34.58±36.06%34.58 \pm 36.06\%
    R+FGSM (+ full step size) 0.30 Hypercube(0.15)\text{Hypercube}(0.15) 26.53±32.48%26.53 \pm 32.48\%
    R+FGSM (+ uniform init.) 0.15 Uniform(0.30)\text{Uniform}(0.30) 72.92±10.40%72.92 \pm 10.40\%
    Fast FGSM (Uniform + full step size) 0.30 Uniform(0.30)\text{Uniform}(0.30) 86.21±0.75%86.21 \pm 0.75\%

    Uniform random initialization across the full interval [−ϵ,ϵ][-\epsilon, \epsilon] provides the dominant improvement in robust accuracy and dramatically stabilizes training, reducing variance across seeds from ±36.06%\pm 36.06\% to ±0.75%\pm 0.75\%.

  8. Knowl 8 — Accelerated Training Pipeline for Adversarial Optimization

    model/method

    Adversarial training convergence is accelerated by combining three algorithmic techniques adapted from efficient deep learning pipelines:

    1. Cyclic Learning Rates: The learning rate schedules linearly from 00 up to a maximum learning rate λmax⁡\lambda_{\max} over the first N/2N/2 epochs, and decreases linearly from λmax⁡\lambda_{\max} back down to 00 over the remaining N/2N/2 epochs. This enables models to reach peak robust accuracy in 1515 to 3030 epochs rather than hundreds of epochs.
    2. Mixed-Precision Arithmetic: Computations utilize 16-bit floating point (FP16) tensor cores (via Apex AMP at optimization level O2 for CIFAR-10 without loss scaling, and level O1 for ImageNet) to decrease GPU memory bandwidth demands and increase throughput.
    3. Progressive Resizing and Regularization Adjustment: For large-scale datasets such as ImageNet, image resolutions are scaled across three sequential training phases (160×160160 \times 160, 352×352352 \times 352, and full resolution), paired with the complete removal of weight decay regularization from all batch normalization layers.
  9. Knowl 9 — Adversarial Step Size Selection and Boundary Overfitting

    empirical result

    The choice of step size α\alpha in single-step FGSM adversarial training controls the tradeoff between adversarial strength and catastrophic overfitting. For CIFAR-10 training target radius ϵ=8/255\epsilon = 8/255:

    1. Small step sizes (α<ϵ\alpha < \epsilon) generate perturbations that are too weak, yielding substandard robust performance (e.g., <10%< 10\% PGD accuracy for α=1/255\alpha = 1/255).
    2. Increasing step size up to α=1.25ϵ=10/255\alpha = 1.25\epsilon = 10/255 improves robust accuracy monotonically to its maximum (46.06%46.06\%).
    3. Setting α≥2ϵ\alpha \ge 2\epsilon (16/25516/255) forces every component of the perturbation onto the ℓ∞\ell_\infty boundary [−ϵ,ϵ][-\epsilon, \epsilon], triggering catastrophic overfitting where the model exhibits 0.00%0.00\% robust accuracy against PGD.

Coverage note — No substantial contributed material was omitted. All primary algorithmic contributions, failure mode analyses, experimental tables, and verification benchmarks are covered.

References

  1. 1.Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. Synthesizing robust adversarial examples. arXiv preprint arXiv:1707.07397, 2017.
  2. 2.Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. arXiv preprint arXiv:1802.00420, 2018.
  3. 3.Jacob Buckman, Aurko Roy, Colin Raffel, and Ian Goodfellow. Thermometer encoding: One hot way to resist adversarial examples. 2018.
  4. 4.Nicholas Carlini. Is ami (attacks meet interpretability) robust to adversarial examples? arXiv preprint arXiv:1902.02322, 2019.
  5. 5.Nicholas Carlini and David Wagner. Defensive distillation is not robust to adversarial examples. arXiv preprint arXiv:1607.04311, 2016.
  6. 6.Nicholas Carlini and David Wagner. Adversarial examples are not easily detected: Bypassing ten detection methods. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, pp. 3–14. ACM, 2017a.
  7. 7.Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In 2017 IEEE Symposium on Security and Privacy (SP), pp. 39–57. IEEE, 2017b.
  8. 8.Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, and Aleksander Madry. On evaluating adversarial robustness. arXiv preprint arXiv:1902.06705, 2019.
  9. 9.Jeremy M Cohen, Elan Rosenfeld, and J Zico Kolter. Certified adversarial robustness via randomized smoothing. arXiv preprint arXiv:1902.02918, 2019.
  10. 10.Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia. Dawnbench: An end-to-end deep learning benchmark and competition. Training, 100(101):102, 2017.
  11. 11.Francesco Croce, Maksym Andriushchenko, and Matthias Hein. Provable robustness of relu networks via maximization of linear regions. arXiv preprint arXiv:1810.07481, 2018.
  12. 12.Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 9185–9193, 2018.
  13. 13.Logan Engstrom, Andrew Ilyas, and Anish Athalye. Evaluating and understanding the robustness of adversarial logit pairing. arXiv preprint arXiv:1807.10272, 2018.
  14. 14.Reuben Feinman, Ryan R Curtin, Saurabh Shintre, and Andrew B Gardner. Detecting adversarial samples from artifacts. arXiv preprint arXiv:1703.00410, 2017.
  15. 15.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.
  16. 16.Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Timothy Mann, and Pushmeet Kohli. On the effectiveness of interval bound propagation for training verifiably robust models. arXiv preprint arXiv:1810.12715, 2018.
  17. 17.Chuan Guo, Mayank Rana, Moustapha Cisse, and Laurens Van Der Maaten. Countering adversarial images using input transformations. arXiv preprint arXiv:1711.00117, 2017.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  19. 19.Harini Kannan, Alexey Kurakin, and Ian Goodfellow. Adversarial logit pairing. arXiv preprint arXiv:1803.06373, 2018.
  20. 20.Guy Katz, Clark Barrett, David L Dill, Kyle Julian, and Mykel J Kochenderfer. Reluplex: An efficient smt solver for verifying deep neural networks. In International Conference on Computer Aided Verification, pp. 97–117. Springer, 2017.
  21. 21.Alexey Kurakin, Ian Goodfellow, and Samy Bengio. Adversarial examples in the physical world. arXiv preprint arXiv:1607.02533, 2016.
  22. 22.Jiajun Lu, Hussein Sibai, Evan Fabry, and David Forsyth. No need to worry about adversarial examples in object detection in autonomous vehicles. arXiv preprint arXiv:1707.03501, 2017.
  23. 23.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017.
  24. 24.Pratyush Maini, Eric Wong, and J Zico Kolter. Adversarial robustness against the union of multiple perturbation models. arXiv preprint arXiv:1909.04068, 2019.
  25. 25.Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff. On detecting adversarial perturbations. arXiv preprint arXiv:1702.04267, 2017.
  26. 26.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. Mixed precision training. arXiv preprint arXiv:1710.03740, 2017.
  27. 27.Matthew Mirman, Timon Gehr, and Martin Vechev. Differentiable abstract interpretation for provably robust neural networks. In International Conference on Machine Learning, pp. 3575–3583, 2018.
  28. 28.Marius Mosbach, Maksym Andriushchenko, Thomas Trost, Matthias Hein, and Dietrich Klakow. Logit pairing methods can fool gradient-based attacks. arXiv preprint arXiv:1810.12042, 2018.
  29. 29.Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. Distillation as a defense to adversarial perturbations against deep neural networks. In 2016 IEEE Symposium on Security and Privacy (SP), pp. 582–597. IEEE, 2016.
  30. 30.Aditi Raghunathan, Jacob Steinhardt, and Percy S Liang. Semidefinite relaxations for certifying robustness to adversarial examples. In Advances in Neural Information Processing Systems, pp. 10877–10887, 2018.
  31. 31.Hadi Salman, Greg Yang, Jerry Li, Pengchuan Zhang, Huan Zhang, Ilya Razenshteyn, and Sebastien Bubeck. Provably robust deep learning via adversarially trained smoothed classifiers. arXiv preprint arXiv:1906.04584, 2019.
  32. 32.Ali Shafahi, Mahyar Najibi, Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! arXiv preprint arXiv:1904.12843, 2019.
  33. 33.Aman Sinha, Hongseok Namkoong, and John Duchi. Certifying some distributional robustness with principled adversarial training. arXiv preprint arXiv:1710.10571, 2017.
  34. 34.Leslie N Smith. Cyclical learning rates for training neural networks. In 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 464–472. IEEE, 2017.
  35. 35.Leslie N Smith and Nicholay Topin. Super-convergence: Very fast training of residual networks using large learning rates. 2018.
  36. 36.Yang Song, Taesup Kim, Sebastian Nowozin, Stefano Ermon, and Nate Kushman. Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. arXiv preprint arXiv:1710.10766, 2017.
  37. 37.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013.
  38. 38.Guanhong Tao, Shiqing Ma, Yingqi Liu, and Xiangyu Zhang. Attacks meet interpretability: Attribute-steered detection of adversarial samples. In Advances in Neural Information Processing Systems, pp. 7717–7728, 2018.
  39. 39.Vincent Tjeng, Kai Xiao, and Russ Tedrake. Evaluating robustness of neural networks with mixed integer programming. arXiv preprint arXiv:1711.07356, 2017.
  40. 40.Florian Tramèr and Dan Boneh. Adversarial training and robustness for multiple perturbations. arXiv preprint arXiv:1904.13000, 2019.
  41. 41.Florian Tramèr, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204, 2017.
  42. 42.Jonathan Uesato, Brendan O’Donoghue, Aaron van den Oord, and Pushmeet Kohli. Adversarial risk and the dangers of evaluating against weak attacks. arXiv preprint arXiv:1802.05666, 2018.
  43. 43.Jianyu Wang. Bilateral adversarial training: Towards fast training of more robust models against adversarial attacks. arXiv preprint arXiv:1811.10716, 2018.
  44. 44.Eric Wong and J Zico Kolter. Provable defenses against adversarial examples via the convex outer adversarial polytope. arXiv preprint arXiv:1711.00851, 2017.
  45. 45.Eric Wong, Frank Schmidt, Jan Hendrik Metzen, and J Zico Kolter. Scaling provable adversarial defenses. In Advances in Neural Information Processing Systems, pp. 8400–8409, 2018.
  46. 46.Kai Y Xiao, Vincent Tjeng, Nur Muhammad Shafiullah, and Aleksander Madry. Training for faster adversarial robustness verification via inducing relu stability. arXiv preprint arXiv:1809.03008, 2018.
  47. 47.Yuzhe Yang, Guo Zhang, Dina Katabi, and Zhi Xu. Me-net: Towards effective adversarial robustness with matrix estimation. arXiv preprint arXiv:1905.11971, 2019.
  48. 48.Dinghuai Zhang, Tianyuan Zhang, Yiping Lu, Zhanxing Zhu, and Bin Dong. You only propagate once: Painless adversarial training using maximal principle. arXiv preprint arXiv:1905.00877, 2019.

Citation

MLA
Wong, E., et al. “Fast Is Better Than Free: Revisiting Adversarial Training”. arXiv, 2020, http://arxiv.org/abs/2001.03994v1.
APA
Wong, E., Rice, L., & Kolter, J. Z. (2020). Fast is better than free: Revisiting adversarial training. arXiv. http://arxiv.org/abs/2001.03994v1
Chicago
Wong, E., L. Rice, and J. Z. Kolter. 2020. “Fast Is Better Than Free: Revisiting Adversarial Training”. arXiv. http://arxiv.org/abs/2001.03994v1.
Harvard
Wong, E., Rice, L. and Kolter, J.Z. (2020) “Fast is better than free: Revisiting adversarial training”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2001.03994v1.
Vancouver
1. Wong E, Rice L, Kolter JZ (2020) Fast is better than free: Revisiting adversarial training. arXiv

BibTeX

@article{wong2020fast,
  title = {Fast is better than free: Revisiting adversarial training},
  author = {Wong, Eric and Rice, Leslie and Kolter, J. Zico},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2001.03994v1},
  eprint = {2001.03994}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors