Built independently by an author, for readers. Read the story and support ChapterPal

keyword

PGD adversary

A PGD adversary is an attack mechanism in machine learning that generates adversarial examples by using projected gradient descent to maximize a model prediction error within predefined perturbation boundaries. As an iterative, first-order white-box threat model, it computes the gradient of the loss function with respect to the input features, repeatedly updates the input in the direction of steepest loss ascent, and projects the resulting values back into a constrained set, typically defined by an Lp-norm bound around the original input. This adversary often incorporates random initialization within the constraint space to explore multiple local maxima. In robust deep learning, the PGD adversary serves as a standard benchmark for evaluating empirical defense mechanisms and functions as the inner maximization component in adversarial training to train networks that resist worst-case input perturbations.

2 items

Fast is better than free: Revisiting adversarial training

Fast is better than free: Revisiting adversarial training

Eric Wong, L. Rice, J. Kolter

OrganizationsBosch Center for AICarnegie Mellon University

Why you should read this

Demonstrates that combining the single-step Fast Gradient Sign Method with random initialization matches the defensive strength of expensive multi-step adversarial training while cutting computation times from hours to minutes.

Adversarial training, a method for learning robust deep networks, is typically assumed to be more expensive than traditional training due to the necessity of constructing adversarial examples via a first-order method like projected gradient decent (PGD). In this paper, we make the surprising discovery that it is possible to train empirically robust models using a much weaker and cheaper adversary, an approach that was previously believed to be ineffective, rendering the method no more costly than standard training in practice. Specifically, we show that adversarial training with the fast gradient sign method (FGSM), when combined with random initialization, is as effective as PGD-based training but has significantly lower cost. Furthermore we show that FGSM adversarial training can be further accelerated by using standard techniques for efficient training of deep networks, allowing us to learn a robust CIFAR10 classifier with 45% robust accuracy to PGD attacks with ϵ=8/255\epsilon=8/255 in 6 minutes, and a robust ImageNet classifier with 43% robust accuracy at ϵ=2/255\epsilon=2/255 in 12 hours, in comparison to past work based on "free" adversarial training which took 10 and 50 hours to reach the same respective thresholds. Finally, we identify a failure mode referred to as "catastrophic overfitting" which may have caused previous attempts to use FGSM adversarial training to fail. All code for reproducing the experiments in this paper as well as pretrained model weights are at this https URL.

Added

2026-09-25

Towards Deep Learning Models Resistant to Adversarial Attacks

Towards Deep Learning Models Resistant to Adversarial Attacks

Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, Adrian Vladu

OrganizationsMassachusetts Institute of Technology

Why you should read this

Establishes a min-max optimization framework for adversarial training, demonstrating that models trained against projected gradient descent attacks achieve reliable defense against first-order adversarial examples.

Recent work has demonstrated that deep neural networks are vulnerable to adversarial examples---inputs that are almost indistinguishable from natural data and yet classified incorrectly by the network. In fact, some of the latest findings suggest that the existence of adversarial attacks may be an inherent weakness of deep learning models. To address this problem, we study the adversarial robustness of neural networks through the lens of robust optimization. This approach provides us with a broad and unifying view on much of the prior work on this topic. Its principled nature also enables us to identify methods for both training and attacking neural networks that are reliable and, in a certain sense, universal. In particular, they specify a concrete security guarantee that would protect against any adversary. These methods let us train networks with significantly improved resistance to a wide range of adversarial attacks. They also suggest the notion of security against a first-order adversary as a natural and broad security guarantee. We believe that robustness against such well-defined classes of adversaries is an important stepping stone towards fully resistant deep learning models. Code and pre-trained models are available at this https URL and this https URL.

Added

2026-09-06