Understanding Robust Overfitting of Adversarial Training and Beyond

Chaojian YuBo HanLi ShenJun YuChen GongMingming GongTongliang Liu

article2022ICML89 citations

Shows that overfitting in adversarial training is driven by easily fitted, small-loss data under strong attacks and introduces minimum loss constrained adversarial training to prevent test degradation by actively increasing the loss on easy examples.

Listen

Deep neural networks are vulnerable to adversarial attacks, where subtle, human-imperceptible input modifications mislead models into making incorrect predictions. Adversarial training serves as one of the primary defenses by training models directly on attacked examples. However, standard adversarial training routinely suffers from robust overfitting, a phenomenon where a model's defense performance against attacks peaks during training and then significantly degrades. Understanding and fixing this issue without relying on expensive, newly collected datasets is crucial for developing reliable, secure artificial intelligence systems.

The article aims to uncover the exact root causes of robust overfitting and evaluate a new training prototype designed to eliminate this degradation while boosting overall adversarial defense performance.

To investigate the mechanism, the authors compared data loss distributions across weak and strong adversarial training regimes. Through data ablation experiments, they selectively removed samples from specific loss ranges across multiple benchmark image datasets (CIFAR10, CIFAR100, and SVHN), network architectures, and threat conditions. Building on these empirical insights, the article introduces Minimum Loss Constrained Adversarial Training (MLCAT), a framework that identifies easy-to-learn (small-loss) adversarial samples within each training batch and applies targeted interventions to artificially increase their loss, evaluated via two orthogonal implementations: loss scaling (MLCATLS) and weight perturbation (MLCATWP).

Key findings show that robust overfitting under strong attack settings is caused specifically by small-loss adversarial data—particularly samples that transition from hard to easy as the network learns—rather than large-loss samples. Standard adversarial training exhibited substantial robust accuracy drops of roughly 3 to 8 percentage points between peak and final training checkpoints. In contrast, both MLCAT implementations virtually eliminated robust overfitting, compressing this accuracy drop to under 1 percentage point across all test settings. Furthermore, MLCAT using weight perturbation consistently improved final robust accuracy (e.g., reaching approximately 50.3% to 54.6% Auto Attack accuracy on CIFAR10, compared to 42.1% to 46.1% for standard adversarial training) while maintaining standard, unattacked classification performance.

These findings challenge the assumption that large-loss data induce robust overfitting, demonstrating instead that easily fitted adversarial samples undermine defense generalization. Practically, this framework delivers higher security and model stability without the substantial operational and financial costs of acquiring extra training data or the operational risks of premature early stopping. However, the results show an important trade-off: while loss scaling eliminates the training gap, it creates vulnerability to logit-scaling attacks, making weight perturbation the superior practical implementation.

Organizations developing robust deep learning models should adopt parameter-perturbation-based minimum loss constraints within their adversarial training pipelines to prevent performance decay. Implementation teams must carefully calibrate the minimum loss threshold to the specific dataset and threat level, as setting the constraint too high can destabilize training and cause model collapse. Confidence in these conclusions is high given consistent validation across diverse datasets, network backbones, and standard attack benchmarks, though practical deployment requires standard tuning of loss thresholds for novel data domains.

Yu et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Understanding Robust Overfitting of Adversarial Training and Beyond

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Adversarial Training
  • 2.2. Robust Overfitting
  • 3. Understanding Robust Overfitting in AT
  • 3.1. Non-overfit AT vs. Overfitted AT: Data Distribution Perspective
  • 3.2. Causes of Robust Overfitting
  • 3.3. A Prototype of MLCAT
  • 4. Two Realizations of MLCAT
  • 5. Experiment
  • 5.1. Experimental Settings
  • 5.2. Performance Evaluation
  • 5.3. Ablation Studies
  • 6. Conclusion
  • Acknowledgement
  • References
  • A. More Evidences for Robust Overfitting and Data Distribution
  • B. More Evidences for the Causes of Robust Overfitting
  • C. More Experimental Results
  • C.1. Performance Evaluation
  • C.2. Ablation Studies

Knowls

  1. Knowl 1 — Minimum-loss-constrained adversarial training

    algorithm

    Minimum loss constrained adversarial training (MLCAT) modifies a base adversarial-training procedure so that each mini-batch retains its large-loss examples and applies a loss-adjustment strategy to examples whose losses fall below a threshold. The adjustment is intended to stop already-easy adversarial examples from being fitted too readily, without discarding them from the training set. Let AA be a base procedure with an inner adversarial-example generation step and an outer parameter-optimization step; let ℓi\ell_i be the loss of example ii after adversarial-example generation; let ℓmin⁡\ell_{\min} be the threshold; and let SS be a strategy that returns an adjusted loss for a small-loss example. The mini-batch objective is the mean of the original losses for examples at or above the threshold and the adjusted losses for examples below it.

    Input: Base adversarial-training procedure A, model parameters w, optimizer O, training set D, batch size m, threshold ell_min, adjustment strategy S
    Sample a mini-batch B of m labeled examples from D
    Generate adversarial examples B' using A's inner maximization step
    Compute the per-example losses ell_i on B'
    Set the accumulated loss L to 0
    For each example i in the mini-batch:
        If ell_i >= ell_min:
            Add ell_i to L
        Otherwise:
            Compute adjusted loss ell_i_S = S(w, example i in B', ell_min)
            Add ell_i_S to L
    Divide L by m
    Compute the outer-minimization gradient using the mean loss L
    Update w using optimizer O

    With an identity adjustment, MLCAT reduces to standard adversarial training; with an adjustment that returns zero for every small-loss example, it reduces to data-ablation adversarial training. The threshold determines which examples are adjusted: a nonpositive threshold leaves ordinary nonnegative losses unadjusted, while an increasingly large threshold causes more examples to be adjusted. The authors caution that adjusting all examples can harm optimization, so the threshold must be selected for the perturbation strength and task.

  2. Knowl 2 — Small-loss examples are implicated in robust overfitting

    empirical result

    Data-ablation experiments identify small-loss adversarial examples—not large-loss examples—as the main cause of robust overfitting in the tested strong-adversary settings. In the primary experiment, a PreAct ResNet-18 was adversarially trained on CIFAR10 under an L∞L_{\infty} threat model with perturbation size ϵ=8\epsilon=8. Removing examples in selected large-loss ranges before overfitting began (for example, at epoch 100) did not eliminate the subsequent decline in test robustness. By contrast, removing small-loss examples from the start of training eliminated the observed robust-overfitting pattern. The authors report similar outcomes across additional datasets, architectures, and threat models.

    The small-loss group has two observed sources: examples already in low-loss ranges before the learning-rate decay, and examples that move into those ranges later from other loss ranges. Separating these groups in the CIFAR10 experiment indicated that robust overfitting was mainly associated with the latter, transformed examples. The authors suggest that these examples become less aggressive relative to the increasingly robust model as training proceeds; this explanation is offered as a possible account, rather than established as a separate mechanism. Their proposed remedy adjusts small losses instead of discarding these examples, thereby retaining the training sample size.

  3. Knowl 3 — Loss distributions differ between weak- and strong-adversary training

    empirical result

    In experiments training PreAct ResNet-18 on CIFAR10 under an L∞L_{\infty} threat model, the authors varied the perturbation size ϵ\epsilon across 0,1,2,4,6,8,0, 1, 2, 4, 6, 8, and 1010, and tracked both test robustness and the training examples’ loss ranges over 200 epochs. Weak-adversary settings, which did not show robust overfitting, were dominated by small-loss examples. Strong-adversary settings, in which test robustness degraded during continued training, had a more dispersed loss distribution containing substantial proportions of both small- and large-loss examples. The authors report analogous qualitative patterns across further datasets, architectures, and threat models. These observations motivate their hypothesis that some low-loss examples in strong-adversary training are not challenging enough to justify the attack strength.

  4. Knowl 4 — Loss scaling as an MLCAT realization

    equation

    MLCAT with loss scaling, denoted MLCATLS_{\mathrm{LS}}, raises every below-threshold per-example loss directly to the minimum-loss threshold. For an adversarial example with original loss ℓi\ell_i satisfying 0<ℓi<ℓmin⁡0<\ell_i<\ell_{\min}, the adjusted loss is

    ℓiS=ℓmin⁡ℓi ℓi=ℓmin⁡.\ell_i^{S}=\frac{\ell_{\min}}{\ell_i}\,\ell_i=\ell_{\min}.

    Here, ℓi\ell_i is the scalar classification loss for example ii, ℓmin⁡\ell_{\min} is the threshold, and ℓiS\ell_i^{S} is the adjusted scalar loss used for optimization. Examples not below the threshold retain their original loss. The scaling coefficient is larger for smaller original losses; equivalently, scaling the loss scales its gradient and acts like using a larger learning rate for those examples. The paper notes that this realization can make models more sensitive to logit-scaling attacks, which is reflected in its AutoAttack results.

  5. Knowl 5 — Targeted weight perturbation as an MLCAT realization

    equation

    MLCAT with weight perturbation, denoted MLCATWP_{\mathrm{WP}}, increases the losses of small-loss adversarial examples by perturbing the model parameters in a direction targeted at those examples. Let ww be the current vector of model parameters, ℓi\ell_i the current loss of adversarial example ii, and ℓmin⁡\ell_{\min} the threshold. First form the gradient over the selected mini-batch examples, then scale that gradient to obtain the weight perturbation:

    g=∇w∑i1(ℓi≤ℓmin⁡)ℓi,v=γ∥w∥2∥g∥2g.g=\nabla_w\sum_i \mathbf{1}(\ell_i\leq\ell_{\min})\ell_i, \qquad v=\gamma\frac{\lVert w\rVert_2}{\lVert g\rVert_2}g.

    The indicator 1(ℓi≤ℓmin⁡)\mathbf{1}(\ell_i\leq\ell_{\min}) selects examples at or below the threshold, gg is their aggregate loss gradient, and γ\gamma is the weight-perturbation size. For each selected adversarial input xi′x'_i with label yiy_i, the adjusted loss is computed using perturbed parameters:

    ℓiS=ℓ(fw+v(xi′),yi),\ell_i^{S}=\ell(f_{w+v}(x'_i),y_i),

    where fw+vf_{w+v} is the classifier with parameters w+vw+v and ℓ\ell is the classification loss. The perturbation raises the selected examples’ losses collectively but does not guarantee that every individual adjusted loss reaches ℓmin⁡\ell_{\min}.

  6. Knowl 6 — Evaluation protocol for the MLCAT experiments

    experimental setup

    The experiments evaluated MLCAT loss scaling and weight perturbation on CIFAR10, SVHN, and CIFAR100, using PreAct ResNet-18 and Wide ResNet-34-10 under L∞L_{\infty} and L2L_2 threat models. Models were trained for 200 epochs with SGD, momentum 0.90.9, weight decay 5×10−45\times10^{-4}, and initial learning rate 0.10.1, divided by 10 at epochs 100 and 150. CIFAR10 and CIFAR100 used random crops with four pixels of padding and random horizontal flips; SVHN used no data augmentation.

    Training adversarial examples were generated with 10-step PGD. For the L∞L_{\infty} threat model, ϵ=8/255\epsilon=8/255 and the step size was 1/2551/255 for SVHN and 2/2552/255 for CIFAR10 and CIFAR100. For the L2L_2 threat model, ϵ=128/255\epsilon=128/255 and the step size was 15/25515/255 on all three datasets. Test robustness was measured using PGD-20 and AutoAttack (AA). Reported main-table results are means over five runs; the paper states that omitted standard deviations were below 0.6 percentage points. The chosen threshold was ℓmin⁡=1.5\ell_{\min}=1.5 for CIFAR10 and SVHN and ℓmin⁡=4.0\ell_{\min}=4.0 for CIFAR100.

  7. Knowl 7 — CIFAR10 robustness and overfitting results

    data/table

    On CIFAR10, both MLCAT realizations greatly reduced the gap between the best test robustness observed during training and the last-epoch robustness, relative to standard adversarial training. The table gives test accuracy in percent for PGD-20 and AutoAttack; “Diff” is last-epoch accuracy minus best accuracy, so a value closer to zero indicates less robust overfitting. MLCATWP_{\mathrm{WP}} improved AA robustness over adversarial training in all four architecture–threat-model settings. MLCATLS_{\mathrm{LS}} also reduced the gap, but its AA accuracy was much lower than the baseline in these settings.

    Network Threat Method PGD-20 AA
    Best Last Diff Best Last Diff
    PreAct ResNet-18 L∞L_{\infty} AT 52.29 44.43 -7.86 47.99 42.08 -5.91
    MLCATLS_{\mathrm{LS}} 56.90 56.87 -0.03 28.12 26.93 -1.19
    MLCATWP_{\mathrm{WP}} 58.48 57.65 -0.83 50.70 50.32 -0.38
    PreAct ResNet-18 L2L_2 AT 69.27 65.86 -3.41 67.70 64.64 -3.06
    MLCATLS_{\mathrm{LS}} 73.16 72.48 -0.68 49.70 48.94 -0.76
    MLCATWP_{\mathrm{WP}} 74.38 73.86 -0.52 70.46 70.15 -0.31
    Wide ResNet-34-10 L∞L_{\infty} AT 55.57 47.37 -8.20 52.13 46.09 -6.04
    MLCATLS_{\mathrm{LS}} 64.73 63.94 -0.79 35.00 34.51 -0.49
    MLCATWP_{\mathrm{WP}} 62.50 61.91 -0.59 54.65 54.56 -0.09
    Wide ResNet-34-10 L2L_2 AT 71.57 69.99 -1.58 70.44 68.92 -1.52
    MLCATLS_{\mathrm{LS}} 75.05 74.97 -0.08 55.31 55.11 -0.20
    MLCATWP_{\mathrm{WP}} 76.92 76.55 -0.37 74.35 73.97 -0.38

    The experiment used both network architectures and both threat models. The results show that reducing robust overfitting does not by itself ensure high robustness against every evaluation attack: in particular, loss scaling’s strong PGD-20 scores coexist with substantially lower AA scores.

  8. Knowl 8 — Robustness results on SVHN and CIFAR100

    data/table

    The following results give the PreAct ResNet-18, L∞L_{\infty} settings on SVHN and CIFAR100. For each attack, entries are best test accuracy, last-epoch test accuracy, and last-minus-best difference, in percent. They illustrate the main cross-dataset pattern: weight perturbation improves robustness and sharply reduces the gap; loss scaling reduces the gap but does not consistently improve AA accuracy over standard adversarial training. The full experiments also include Wide ResNet-34-10 and L2L_2 settings, for which the authors report the same general ability of MLCAT to reduce robust overfitting.

    Dataset Method PGD-20: Best, Last, Diff AA: Best, Last, Diff
    SVHN AT 52.88 45.29 -7.59 45.09 40.36 -4.73
    SVHN MLCATLS_{\mathrm{LS}} 64.28 62.30 -1.98 34.48 32.33 -2.15
    SVHN MLCATWP_{\mathrm{WP}} 60.34 57.79 -2.55 51.90 49.76 -2.14
    CIFAR100 AT 28.01 20.39 -7.62 23.61 18.41 -5.20
    CIFAR100 MLCATLS_{\mathrm{LS}} 20.09 18.14 -1.95 13.41 11.35 -2.06
    CIFAR100 MLCATWP_{\mathrm{WP}} 31.27 30.57 -0.70 25.66 25.28 -0.38

    For example, on CIFAR100, MLCATWP_{\mathrm{WP}} raises last-epoch AA accuracy from 18.41% to 25.28% and changes the robustness gap from -5.20 to -0.38 percentage points. On SVHN, MLCATWP_{\mathrm{WP}} raises last-epoch AA accuracy from 40.36% to 49.76%, while shrinking the gap from -4.73 to -2.14 points.

  9. Knowl 9 — Threshold and loss-adjustment ablations

    empirical result

    Ablations on CIFAR10 with PreAct ResNet-18 under an L∞L_{\infty} threat model tested thresholds ℓmin⁡\ell_{\min} from 0 to 3. Raising the threshold generally reduced the best-to-last robust-accuracy gap. At comparatively small thresholds, increasing ℓmin⁡\ell_{\min} could also improve robustness over standard adversarial training; above about 1.5, further increases reduced robustness and could cause training collapse. The authors therefore treat the threshold as a task- and perturbation-dependent control, rather than a value to maximize. They used 1.51.5 for CIFAR10 and SVHN and 4.04.0 for CIFAR100 in their main experiments.

    The adjustment-direction ablation compared increasing, preserving, and decreasing the losses of small-loss examples. Increasing those losses with either loss scaling or weight perturbation reduced robust overfitting and improved robustness; leaving the losses unchanged was equivalent to standard adversarial training; decreasing them failed to suppress robust overfitting and produced worse robustness. The authors report similar threshold trends on SVHN and CIFAR100.

  10. Knowl 10 — Transfer to TRADES and comparison with AWP

    empirical result

    The authors also applied the MLCAT idea to TRADES, retaining TRADES for adversarial-example generation and outer minimization and using ℓmin⁡=1.5\ell_{\min}=1.5. On CIFAR10 under the reported PGD-20 evaluation, weight-perturbation MLCAT for TRADES raised best accuracy from 52.56 ±\pm 0.43% to 55.28 ±\pm 0.21% and last-epoch accuracy from 49.12 ±\pm 0.39% to 54.99 ±\pm 0.19%; the best-to-last difference changed from -3.53 to -0.29 percentage points. Loss-scaling MLCAT for TRADES yielded 42.82 ±\pm 0.25% best and 41.4 ±\pm 0.38% last accuracy, with a -1.42-point difference.

    In a separate CIFAR10 comparison under the L∞L_{\infty} setup, MLCATWP_{\mathrm{WP}} was compared with adversarial weight perturbation (AWP). For PGD-20, AWP gave 55.54 ±\pm 0.20% best, 54.64 ±\pm 0.25% last, and -0.9 points difference; MLCATWP_{\mathrm{WP}} gave 58.48 ±\pm 0.39% best, 57.65 ±\pm 0.19% last, and -0.83 points difference. For AA, AWP gave 49.94 ±\pm 0.08% best, 49.69 ±\pm 0.10% last, and -0.25 points difference; MLCATWP_{\mathrm{WP}} gave 50.70 ±\pm 0.11% best, 50.32 ±\pm 0.09% last, and -0.38 points difference. Thus MLCATWP_{\mathrm{WP}} improved the reported best and last accuracies over AWP in both evaluations, although its AA best-to-last gap was slightly larger in magnitude.

Coverage note — The detailed appendix plots for additional dataset, architecture, and threat-model combinations, along with the natural-accuracy table, are omitted because they corroborate the reported trends without adding a distinct core method or conclusion.

References

  1. 1.Andriushchenko, M., Croce, F., Flammarion, N., and Hein, M. Square attack: a query-efficient black-box adversarial attack via random search. In European Conference on Computer Vision, pp. 484–501. Springer, 2020.
  2. 2.Athalye, A., Carlini, N., and Wagner, D. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International conference on machine learning, pp. 274–283. PMLR, 2018.
  3. 3.Bai, Y., Zeng, Y., Jiang, Y., Xia, S.-T., Ma, X., and Wang, Y. Improving adversarial robustness via channel-wise activation suppressing. arXiv preprint arXiv:2103.08307, 2021.
  4. 4.Cai, Q.-Z., Du, M., Liu, C., and Song, D. Curriculum adversarial training. arXiv preprint arXiv:1805.04807, 2018.
  5. 5.Carmon, Y., Raghunathan, A., Schmidt, L., Liang, P., and Duchi, J. C. Unlabeled data improves adversarial robustness. arXiv preprint arXiv:1905.13736, 2019.
  6. 6.Chen, C., Zhang, J., Xu, X., Hu, T., Niu, G., Chen, G., and Sugiyama, M. Guided interpolation for adversarial training. arXiv preprint arXiv:2102.07327, 2021.
  7. 7.Chen, T., Zhang, Z., Liu, S., Chang, S., and Wang, Z. Robust overfitting may be mitigated by properly learned smoothening. In International Conference on Learning Representations, 2020.
  8. 8.Croce, F. and Hein, M. Minimally distorted adversarial examples with a fast adaptive boundary attack. In International Conference on Machine Learning, pp. 2196–2205. PMLR, 2020a.
  9. 9.Croce, F. and Hein, M. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In International conference on machine learning, pp. 2206–2216. PMLR, 2020b.
  10. 10.Croce, F., Gowal, S., Brunner, T., Shelhamer, E., Hein, M., and Cemgil, T. Evaluating the adversarial robustness of adaptive test-time defenses. arXiv preprint arXiv:2202.13711, 2022.
  11. 11.Dong, Y., Xu, K., Yang, X., Pang, T., Deng, Z., Su, H., and Zhu, J. Exploring memorization in adversarial training. In International Conference on Learning Representations, 2022.
  12. 12.Goodfellow, I. J., Shlens, J., and Szegedy, C. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.
  13. 13.Gowal, S., Qin, C., Uesato, J., Mann, T., and Kohli, P. Uncovering the limits of adversarial training against norm-bounded adversarial examples. arXiv preprint arXiv:2010.03593, 2020.
  14. 14.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  15. 15.Hitaj, D., Pagnotta, G., Masi, I., and Mancini, L. V. Evaluating the robustness of geometry-aware instance-reweighted adversarial training. arXiv preprint arXiv:2103.01914, 2021.
  16. 16.Kannan, H., Kurakin, A., and Goodfellow, I. Adversarial logit pairing. arXiv preprint arXiv:1803.06373, 2018.
  17. 17.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  18. 18.Lee, S., Lee, H., and Yoon, S. Adversarial vertex mixup: Toward better adversarially robust generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 272–281, 2020.
  19. 19.Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017.
  20. 20.Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I. Deep double descent: Where bigger models and more data hurt. arXiv preprint arXiv:1912.02292, 2019.
  21. 21.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. Reading digits in natural images with unsupervised feature learning. 2011.
  22. 22.Pang, T., Yang, X., Dong, Y., Su, H., and Zhu, J. Bag of tricks for adversarial training. arXiv preprint arXiv:2010.00467, 2020.
  23. 23.Rice, L., Wong, E., and Kolter, Z. Overfitting in adversarially robust deep learning. In International Conference on Machine Learning, pp. 8093–8104. PMLR, 2020.
  24. 24.Schmidt, L., Santurkar, S., Tsipras, D., Talwar, K., and Madry, A. Adversarially robust generalization requires more data. Advances in Neural Information Processing Systems, 31:5014–5026, 2018.
  25. 25.Song, C., He, K., Lin, J., Wang, L., and Hopcroft, J. E. Robust local features for improving the generalization of adversarial training. In International Conference on Learning Representations, 2020.
  26. 26.Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013.
  27. 27.Uesato, J., Alayrac, J.-B., Huang, P.-S., Stanforth, R., Fawzi, A., and Kohli, P. Are labels required for improving adversarial robustness? arXiv preprint arXiv:1905.13725, 2019.
  28. 28.Wang, Y., Ma, X., Bailey, J., Yi, J., Zhou, B., and Gu, Q. On the convergence and robustness of adversarial training. In ICML, volume 1, pp. 2, 2019a.
  29. 29.Wang, Y., Zou, D., Yi, J., Bailey, J., Ma, X., and Gu, Q. Improving adversarial robustness requires revisiting misclassified examples. In International Conference on Learning Representations, 2019b.
  30. 30.Wu, B., Pan, H., Shen, L., Gu, J., Zhao, S., Li, Z., Cai, D., He, X., and Liu, W. Attacking adversarial attacks as a defense. arXiv preprint arXiv:2106.04938, 2021.
  31. 31.Wu, D., Xia, S.-T., and Wang, Y. Adversarial weight perturbation helps robust generalization. arXiv preprint arXiv:2004.05884, 2020.
  32. 32.Xia, X., Liu, T., Wang, N., Han, B., Gong, C., Niu, G., and Sugiyama, M. Are anchor points really indispensable in label-noise learning? In NeurIPS, 2019.
  33. 33.Xia, X., Liu, T., Han, B., Wang, N., Gong, M., Liu, H., Niu, G., Tao, D., and Sugiyama, M. Part-dependent label noise: Towards instance-dependent label noise. In NeurIPS, 2020.
  34. 34.Yan, H., Zhang, J., Niu, G., Feng, J., Tan, V. Y., and Sugiyama, M. Cifs: Improving adversarial robustness of cnns via channel-wise importance-based feature selection. arXiv preprint arXiv:2102.05311, 2021.
  35. 35.Yu, C., Han, B., Gong, M., Shen, L., Ge, S., Du, B., and Liu, T. Robust weight perturbation for adversarial training. arXiv preprint arXiv:2205.14826, 2022.
  36. 36.Zagoruyko, S. and Komodakis, N. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  37. 37.Zhai, R., Cai, T., He, D., Dan, C., He, K., Hopcroft, J., and Wang, L. Adversarially robust generalization just requires more unlabeled data. arXiv preprint arXiv:1906.00555, 2019.
  38. 38.Zhang, H. and Xu, W. Adversarial interpolation training: A simple approach for improving model robustness. 2019.
  39. 39.Zhang, H., Yu, Y., Jiao, J., Xing, E., El Ghaoui, L., and Jordan, M. Theoretically principled trade-off between robustness and accuracy. In International Conference on Machine Learning, pp. 7472–7482. PMLR, 2019.
  40. 40.Zhang, J., Xu, X., Han, B., Niu, G., Cui, L., Sugiyama, M., and Kankanhalli, M. Attacks which do not kill training make adversarial learning stronger. In International Conference on Machine Learning, pp. 11278–11287. PMLR, 2020a.
  41. 41.Zhang, J., Zhu, J., Niu, G., Han, B., Sugiyama, M., and Kankanhalli, M. Geometry-aware instance-reweighted adversarial training. arXiv preprint arXiv:2010.01736, 2020b.
  42. 42.Zhou, D., Wang, N., Han, B., and Liu, T. Modeling adversarial noise for adversarial defense. arXiv preprint arXiv:2109.09901, 2021.

Citation

MLA
Yu, C., et al. “Understanding Robust Overfitting of Adversarial Training and Beyond”. International Conference on Machine Learning, vol. 162, 2022, pp. 25595–610, https://proceedings.mlr.press/v162/yu22b.html.
APA
Yu, C., Han, B., Shen, L., Yu, J., Gong, C., Gong, M., & Liu, T. (2022). Understanding Robust Overfitting of Adversarial Training and Beyond. International Conference on Machine Learning, 162, 25595–25610. https://proceedings.mlr.press/v162/yu22b.html
Chicago
Yu, C., B. Han, L. Shen, et al. 2022. “Understanding Robust Overfitting of Adversarial Training and Beyond”. International Conference on Machine Learning 162: 25595–610. https://proceedings.mlr.press/v162/yu22b.html.
Harvard
Yu, C. et al. (2022) “Understanding Robust Overfitting of Adversarial Training and Beyond”, International Conference on Machine Learning. PMLR, pp. 25595–25610. Available at: https://proceedings.mlr.press/v162/yu22b.html.
Vancouver
1. Yu C, Han B, Shen L, Yu J, Gong C, Gong M, Liu T (2022) Understanding Robust Overfitting of Adversarial Training and Beyond. In: International Conference on Machine Learning. PMLR, pp 25595–25610

BibTeX

@InProceedings{pmlr-v162-yu22b,
  title = 	 {Understanding Robust Overfitting of Adversarial Training and Beyond},
  author =       {Yu, Chaojian and Han, Bo and Shen, Li and Yu, Jun and Gong, Chen and Gong, Mingming and Liu, Tongliang},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {25595--25610},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/yu22b/yu22b.pdf},
  url = 	 {https://proceedings.mlr.press/v162/yu22b.html},
  abstract = 	 {Robust overfitting widely exists in adversarial training of deep networks. The exact underlying reasons for this are still not completely understood. Here, we explore the causes of robust overfitting by comparing the data distribution of non-overfit (weak adversary) and overfitted (strong adversary) adversarial training, and observe that the distribution of the adversarial data generated by weak adversary mainly contain small-loss data. However, the adversarial data generated by strong adversary is more diversely distributed on the large-loss data and the small-loss data. Given these observations, we further designed data ablation adversarial training and identify that some small-loss data which are not worthy of the adversary strength cause robust overfitting in the strong adversary mode. To relieve this issue, we propose minimum loss constrained adversarial training (MLCAT): in a minibatch, we learn large-loss data as usual, and adopt additional measures to increase the loss of the small-loss data. Technically, MLCAT hinders data fitting when they become easy to learn to prevent robust overfitting; philosophically, MLCAT reflects the spirit of turning waste into treasure and making the best use of each adversarial data; algorithmically, we designed two realizations of MLCAT, and extensive experiments demonstrate that MLCAT can eliminate robust overfitting and further boost adversarial robustness.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/