Robustness and Accuracy Could Be Reconcilable by (Proper) Definition

Tianyu PangMin LinXiao YangJun ZhuShuicheng Yan

article2022ICML179 citations

Proposes a self-consistent error objective termed SCORE that replaces conventional local invariance with local equivariance, eliminating the artificial conflict between clean accuracy and adversarial resilience to achieve top-ranked defense performance under AutoAttack.

Listen

Modern machine learning models deployed in safety-critical applications are highly vulnerable to adversarial attacks, where small, imperceptible input perturbations cause critical misclassifications. Defending against these attacks through adversarial training has historically introduced a persistent drawback: a significant drop in baseline accuracy on unperturbed, standard data. The prevailing view in the artificial intelligence community has long held that this trade-off between robustness and accuracy is an unavoidable, inherent cost. The article investigates the origin of this tension and evaluates whether it is actually a mathematical artifact resulting from an improper definition of robust error.

The article demonstrates that the standard formulation of adversarial training forces the model toward excessive local invariance—an overcorrection that demands model predictions remain entirely constant around an input, ignoring natural data shifts. To resolve this, the article introduces the Self-Consistent Robust Error (SCORE), a theoretical objective that promotes local equivariance, allowing the model to smoothly match the underlying true data distribution point by point. Because directly calculating SCORE requires inaccessible gradient information from the true data distribution, the authors mathematically derive bounds using distance metrics. This reveals that substituting conventional information-divergence losses with monotonically increasing convex distance variants—specifically squared error—enables efficient optimization without extra computational overhead.

The empirical evaluation was conducted across benchmark image datasets (CIFAR-10, CIFAR-100, and ImageNet) using standard residual network architectures against standard adversarial benchmarks, including the rigorous AutoAttack framework. Key findings demonstrate that substituting squared error into standard adversarial training pipelines consistently improves accuracy on clean data while maintaining or increasing robustness against attacks. On large-scale benchmarks utilizing synthetic data, the proposed method raised clean accuracy by approximately 2.0 to 5.1 percentage points over state-of-the-art baselines while matching or surpassing top-tier robust accuracy (achieving up to 63.74% robust accuracy on CIFAR-10 and 33.65% on CIFAR-100). The theoretical framework also successfully explains why adversarial training frequently suffers from robust overfitting and why robust models inherently learn semantically meaningful, shape-based visual gradients.

These findings provide immediate practical implications for engineering, safety, and compliance teams deploying machine learning systems in security-sensitive environments. Organizations do not need to accept substantial performance degradation on standard operations to achieve adversarial resilience. Because the proposed modification simply replaces the loss function during model training, it incurs zero additional computational overhead or inference latency, effectively lowering the cost and technical risk of robust machine learning deployments.

Organizations developing or deploying safety-critical vision systems should adopt distance-based formulations, such as squared error losses, within their current adversarial training pipelines. Furthermore, practitioners should tune regularization parameters toward lower values when leveraging synthetic or augmented data to maximize baseline performance. Future work should expand these principles beyond standard vision threat models to evaluate non-vision modalities and test generative score-matching implementations as data distribution estimators mature.

Confidence in these findings is high regarding standard image classification benchmarks, supported by verified sanity checks confirming the absence of false robustness. However, decision-makers should note that empirical trade-offs can still arise in practical settings with limited training samples, as the theoretical reconciliation relies on sufficient data coverage. Caution is also advised when transferring these methods to distinct data modalities or specialized hardware setups where data augmentation dynamics may vary.

No sufficiently relevant recommendations were found.

Cover for Robustness and Accuracy Could Be Reconcilable by (Proper) Definition

Abstract

The trade-off between robustness and accuracy has been widely studied in the adversarial literature. Although still controversial, the prevailing view is that this trade-off is inherent, either empirically or theoretically. Thus, we dig for the origin of this trade-off in adversarial training and find that it may stem from the improperly defined robust error, which imposes an inductive bias of local invariance — an overcorrection towards smoothness. Given this, we advocate employing local equivalence to describe the ideal behavior of a robust model, leading to a self-consistent robust error named SCORE. By definition, SCORE facilitates the reconciliation between robustness and accuracy, while still handling the worst-case uncertainty via robust optimization. By simply substituting KL divergence with variants of distance metrics, SCORE can be efficiently minimized. Empirically, our models achieve top-rank performance on RobustBench under AutoAttack. Besides, SCORE provides instructive insights for explaining the overfitting phenomenon and semantic input gradients observed on robust models.

Table of Contents

  • 1. Introduction
  • 2. Self-Consistent Robust Error
  • 2.1. Preliminaries
  • 2.2. Definitions of Robustness
  • 2.3. A Self-Consistent Robust Error
  • 3. How to Practically Optimize SCORE?
  • 3.1. Substituting KL Divergence with Distance Metrics
  • 3.2. Monotonically Increasing Convex Variants
  • 3.3. Equivalent Relation Induced by Distance Metrics
  • 4. New Insights Brought by SCORE
  • 4.1. Overfitting and Early-Stopping
  • 4.2. Semantic Gradients: Adversarial Training
  • 4.3. Semantic Gradients: Randomized Smoothing
  • 5. Experiments
  • 5.1. The 0-1 Version of SCORE for Evaluation
  • 5.2. Basic Setting without Extra or Generated Data
  • 5.3. Advanced Setting with DDPM Generated Data
  • 6. Conclusion and Discussion
  • Acknowledgements
  • References
  • A. Proofs
  • A.1. Proof of Theorem 1
  • A.2. Proof of Theorem 2
  • A.3. Proof of Theorem 3
  • A.4. Proof of Theorem 4
  • A.5. Proof of Theorem 5
  • A.6. Proof of Corollary 1
  • B. Detailed Derivations
  • B.1. The Connection between RMadry ( θ ) and the Original Objective in Madry et al. (2018)
  • B.2. Using Score-based Learning to Optimize RSCORE ( θ )
  • C. Detailed Discussion on the Effects of Randomized Smoothing
  • D. Additional Experiments
  • D.1. Visualization of (KL-based) Overfitting
  • D.2. Visualization of Semantic Gradients
  • D.3. Checking for Gradient Obfuscation
  • E. Clarification on Backgrounds
  • E.1. Overfitting; Catastrophic Overfitting; Benign Overfitting
  • E.2. Other Trade-offs in the Adversarial Literature
  • E.3. AutoAttack for Evaluation

Knowls

  1. Knowl 1 — SCORE replaces distributional invariance with local equivariance

    model/method

    Let pd(x,y)p_d(x,y) be the data distribution, pd(y∣x)p_d(y\mid x) its conditional label distribution, pθ(y∣x)p_\theta(y\mid x) a classifier, and B(x)B(x) a set of inputs associated with xx. The conventional robust objective is RMadry(θ)=Ex∼pd(x)[max⁡x′∈B(x)KL(pd(y∣x)∥pθ(y∣x′))]R_{\mathrm{Madry}}(\theta)=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}\mathrm{KL}(p_d(y\mid x)\|p_\theta(y\mid x'))]: it compares predictions at every x′x' with the conditional distribution at the unperturbed input xx, encouraging local invariance. The paper proposes the Self-COnsistent Robust Error (SCORE), RSCORE(θ)=Ex∼pd(x)[max⁡x′∈B(x)KL(pd(y∣x′)∥pθ(y∣x′))]R_{\mathrm{SCORE}}(\theta)=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}\mathrm{KL}(p_d(y\mid x')\|p_\theta(y\mid x'))]. SCORE instead asks the model to match the data conditional at each challenged input, an equivariant behavior. At the population optimum, the paper states that pθ(y∣x)=pd(y∣x)p_\theta(y\mid x)=p_d(y\mid x), reconciling accuracy with this definition of robustness; the maximization still handles worst-case uncertainty. Unlike the invariance-based objective, SCORE permits B(x)B(x) to be nonlocal or otherwise arbitrary.

  2. Knowl 2 — Distance-based SCORE is bounded by the conventional robust objective

    theoretical result

    Let DD be any distance metric on label distributions, let xx be drawn from pd(x)p_d(x), and let x′∈B(x)x'\in B(x). Define RMadryD(θ)=Ex∼pd(x)[max⁡x′∈B(x)D(pd(y∣x),pθ(y∣x′))]R^D_{\mathrm{Madry}}(\theta)=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}D(p_d(y\mid x),p_\theta(y\mid x'))] and RSCORED(θ)=Ex∼pd(x)[max⁡x′∈B(x)D(pd(y∣x′),pθ(y∣x′))]R^D_{\mathrm{SCORE}}(\theta)=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}D(p_d(y\mid x'),p_\theta(y\mid x'))]. The quantity CD=Ex∼pd(x)[max⁡x′∈B(x)D(pd(y∣x),pd(y∣x′))]C^D=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}D(p_d(y\mid x),p_d(y\mid x'))] measures variation in the data conditional over the allowed set and is independent of model parameters. The paper proves ∣RMadryD(θ)−CD∣≤RSCORED(θ)≤RMadryD(θ)+CD|R^D_{\mathrm{Madry}}(\theta)-C^D|\le R^D_{\mathrm{SCORE}}(\theta)\le R^D_{\mathrm{Madry}}(\theta)+C^D. In particular, zero distance-based SCORE requires RMadryD(θ)=CDR^D_{\mathrm{Madry}}(\theta)=C^D; reducing the conventional objective far below this data-variation level can therefore be counterproductive for SCORE.

  3. Knowl 3 — Convex variants of distance metrics also bound SCORE

    theoretical result

    Let DD be a distance metric on label distributions, and let ϕ\phi be a monotonically increasing convex function with an inverse on the relevant range. Define RMadryϕ∘D(θ)=Ex∼pd(x)[max⁡x′∈B(x)ϕ(D(pd(y∣x),pθ(y∣x′)))]R^{\phi\circ D}_{\mathrm{Madry}}(\theta)=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}\phi(D(p_d(y\mid x),p_\theta(y\mid x')))] and CD=Ex∼pd(x)[max⁡x′∈B(x)D(pd(y∣x),pd(y∣x′))]C^D=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}D(p_d(y\mid x),p_d(y\mid x'))]. For the distance-based SCORE RSCORED(θ)=Ex∼pd(x)[max⁡x′∈B(x)D(pd(y∣x′),pθ(y∣x′))]R^D_{\mathrm{SCORE}}(\theta)=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}D(p_d(y\mid x'),p_\theta(y\mid x'))], the paper proves ∣RSCORED(θ)−CD∣≤ϕ−1(RMadryϕ∘D(θ))|R^D_{\mathrm{SCORE}}(\theta)-C^D|\le\phi^{-1}(R^{\phi\circ D}_{\mathrm{Madry}}(\theta)). This permits optimization with convex transformations that avoid the practical weakness the authors observed for unsquared, sublinear distance losses. Squared error is one such useful variant; the paper also identifies Jensen–Shannon divergence as a possible variant through its square-root distance.

  4. Knowl 4 — SE-based TRADES improves results with generated data

    empirical result

    The paper evaluates TRADES with squared error (SE, ∥P−Q∥22\|P-Q\|_2^2) replacing KL divergence, using 1M DDPM-generated examples, WideResNet architectures with SiLU activations, and AutoAttack. The authors trained for 400 epochs with batch size 512 and no loss clipping; their matched Rebuffi et al. comparisons used 1M generated examples, batch size 1024, and 800 epochs. The values below are clean accuracy / AutoAttack accuracy, in percent. On CIFAR-10 under ℓ∞\ell_\infty perturbations of 8/2558/255, WRN-28-10 achieved 88.61 / 61.04 at β=3\beta=3 and 88.10 / 61.51 at β=4\beta=4, compared with 85.97 / 60.73 for the matched baseline; WRN-70-16 achieved 89.01 / 63.35 and 88.57 / 63.74, respectively, compared with 86.94 / 63.58. On CIFAR-10 under ℓ2\ell_2 perturbations of 128/255128/255, WRN-28-10 achieved 91.52 / 77.89 at β=3\beta=3 and 90.83 / 78.10 at β=4\beta=4, compared with 90.24 / 77.37. On CIFAR-100 under ℓ∞\ell_\infty perturbations of 8/2558/255, WRN-28-10 achieved 63.66 / 31.08 and 62.08 / 31.40, compared with 59.18 / 30.81; WRN-70-16 achieved 65.56 / 33.05 and 63.99 / 33.65, compared with 60.46 / 33.49. Thus the reported SE-based models generally raised clean accuracy while maintaining comparable or slightly improving robust accuracy, with some individual robust-accuracy results below the matched baseline.

  5. Knowl 5 — SE improves PGD-AT and TRADES without extra data

    empirical result

    On CIFAR-10 with ResNet-18, the authors replaced KL divergence with squared error (SE, ∥P−Q∥22\|P-Q\|_2^2) in PGD-AT and TRADES, and evaluated clean and AutoAttack accuracy under an ℓ∞\ell_\infty threat model with ϵ=8/255\epsilon=8/255. Results are mean ±\pm standard deviation over five runs, in percent. For PGD-AT, KL achieved 82.46 ±\pm 0.41 clean and 48.39 ±\pm 0.14 AutoAttack accuracy; SE without clipping achieved 82.13 ±\pm 0.14 and 49.41 ±\pm 0.27; SE with clipping achieved 82.80 ±\pm 0.16 and 49.63 ±\pm 0.17. For TRADES, KL achieved 81.47 ±\pm 0.12 and 49.14 ±\pm 0.16; SE without clipping achieved 83.50 ±\pm 0.05 and 49.44 ±\pm 0.35; SE with clipping achieved 83.75 ±\pm 0.14 and 49.57 ±\pm 0.28. Clipping was applied at each training step, with thresholds 0.4 for PGD-AT and 0.3 for TRADES. The SE models used an initial learning rate of 0.05, versus 0.1 for KL baselines. The authors selected checkpoints using PGD-10 accuracy on a separate validation set. These results show improved robust accuracy for PGD-AT and improved clean and robust accuracy for TRADES under the tested setup.

  6. Knowl 6 — Pinsker’s inequality connects robust overfitting to data variation

    theoretical result

    Let RMadry(θ)R_{\mathrm{Madry}}(\theta) be the KL-based robust objective, and let RSCOREℓ1(θ)R^{\ell_1}_{\mathrm{SCORE}}(\theta) be SCORE using the ℓ1\ell_1 distance between label distributions. Define Cℓ1=Ex∼pd(x)[max⁡x′∈B(x)∥pd(y∣x)−pd(y∣x′)∥1]C^{\ell_1}=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}\|p_d(y\mid x)-p_d(y\mid x')\|_1]. The paper derives ∣RSCOREℓ1(θ)−Cℓ1∣≤2RMadry(θ)|R^{\ell_1}_{\mathrm{SCORE}}(\theta)-C^{\ell_1}|\le\sqrt{2R_{\mathrm{Madry}}(\theta)} using Pinsker’s inequality. Consequently, if RSCOREℓ1(θ)=0R^{\ell_1}_{\mathrm{SCORE}}(\theta)=0, then RMadry(θ)≥(Cℓ1)2/2R_{\mathrm{Madry}}(\theta)\ge (C^{\ell_1})^2/2. The authors interpret the bound as explaining why continued minimization of the KL robust objective can lead to over-smoothing and overfitting once it falls below the natural variation level. In their toy experiment, SCORE began to overfit when the KL robust objective was about 0.051, near the calculated CKL≈0.054C^{\mathrm{KL}}\approx0.054; this observed threshold is an example, not a universal equality established by the bound.

  7. Knowl 7 — Adversarial training induces an input-gradient alignment term

    theoretical result

    Consider the ℓp\ell_p threat set B(x)={x′:∥x′−x∥p≤ϵ}B(x)=\{x':\|x'-x\|_p\le\epsilon\}, and let qq satisfy 1/p+1/q=11/p+1/q=1. Write Yd(x)=arg⁡max⁡ypd(y∣x)Y_d(x)=\arg\max_y p_d(y\mid x), and let RStandardℓ1R^{\ell_1}_{\mathrm{Standard}} and RSCOREℓ1R^{\ell_1}_{\mathrm{SCORE}} be the standard and SCORE risks using ℓ1\ell_1 distance between label distributions. Under the paper’s condition that, for the data-predicted class, pd(Yd(x)∣x)p_d(Y_d(x)\mid x) is above pθ(Yd(x)∣x)p_\theta(Y_d(x)\mid x) or below it in the stated respective cases, a first-order expansion gives RSCOREℓ1(θ)=RStandardℓ1(θ)+2ϵ Ex∼pd(x)[∥∇xpd(Yd(x)∣x)−∇xpθ(Yd(x)∣x)∥q]+o(ϵ)R^{\ell_1}_{\mathrm{SCORE}}(\theta)=R^{\ell_1}_{\mathrm{Standard}}(\theta)+2\epsilon\,\mathbb{E}_{x\sim p_d(x)}[\|\nabla_xp_d(Y_d(x)\mid x)-\nabla_xp_\theta(Y_d(x)\mid x)\|_q]+o(\epsilon). Thus the first-order contribution beyond standard error penalizes disagreement between the data and model input gradients for the data-predicted class. The authors use this relation to interpret semantic or perceptually aligned gradients in adversarially trained models; they describe the interpretation as suggestive rather than conclusive.

  8. Knowl 8 — Gaussian augmentation has a gradient-alignment derivative

    theoretical result

    For σ≥0\sigma\ge0, let ω∼N(0,I)\omega\sim\mathcal{N}(0,I) and define the Gaussian-augmented data distribution by pdσ(x,y)=Eω[pd(x−σ ω,y)]p_d^\sigma(x,y)=\mathbb{E}_\omega[p_d(x-\sqrt{\sigma}\,\omega,y)], with conditional pdσ(x∣y)=Eω[pd(x−σ ω∣y)]p_d^\sigma(x\mid y)=\mathbb{E}_\omega[p_d(x-\sqrt{\sigma}\,\omega\mid y)]. Define the augmented cross-entropy RG(θ;σ)=E(x,y)∼pdσ(x,y)[−log⁡pθ(y∣x)]R_G(\theta;\sigma)=\mathbb{E}_{(x,y)\sim p_d^\sigma(x,y)}[-\log p_\theta(y\mid x)]. For any fixed model parameters θ\theta, the paper proves ddσRG(θ;σ)=12E(x,y)∼pdσ(x,y)[∇xlog⁡pθ(y∣x)⊤∇xlog⁡pdσ(x∣y)]\frac{d}{d\sigma}R_G(\theta;\sigma)=\frac12\mathbb{E}_{(x,y)\sim p_d^\sigma(x,y)}[\nabla_x\log p_\theta(y\mid x)^\top\nabla_x\log p_d^\sigma(x\mid y)]. For small σ\sigma, the loss is its value at zero plus σ\sigma times this derivative at zero, up to o(σ)o(\sigma). The authors interpret the added term as encouraging the model gradient toward the negative conditional data score, providing a gradient-alignment account of semantic gradients associated with Gaussian-augmented training.

  9. Knowl 9 — Distance-based PGD-AT and TRADES objectives are comparable

    theoretical result

    Let DD be a distance metric on label distributions, let β≥1\beta\ge1, and define RMadryD(θ)=Ex∼pd(x)[max⁡x′∈B(x)D(pd(y∣x),pθ(y∣x′))]R^D_{\mathrm{Madry}}(\theta)=\mathbb{E}_{x\sim p_d(x)}[\max_{x'\in B(x)}D(p_d(y\mid x),p_\theta(y\mid x'))]. Define the distance-based TRADES objective as RTRADESD(θ;β)=Ex∼pd(x)[D(pd(y∣x),pθ(y∣x))+βmax⁡x′∈B(x)D(pθ(y∣x),pθ(y∣x′))]R^D_{\mathrm{TRADES}}(\theta;\beta)=\mathbb{E}_{x\sim p_d(x)}[D(p_d(y\mid x),p_\theta(y\mid x))+\beta\max_{x'\in B(x)}D(p_\theta(y\mid x),p_\theta(y\mid x'))]. For every model parameter setting θ\theta, the paper proves RMadryD(θ)≤RTRADESD(θ;β)≤(1+2β)RMadryD(θ)R^D_{\mathrm{Madry}}(\theta)\le R^D_{\mathrm{TRADES}}(\theta;\beta)\le(1+2\beta)R^D_{\mathrm{Madry}}(\theta). This bounds the two objectives by constant factors and supports transferring qualitative conclusions about distance-based Madry training to distance-based TRADES.

  10. Knowl 10 — Finite data and unavailable data scores limit practical guarantees

    limitation

    SCORE’s population-level self-consistency does not eliminate empirical robustness–accuracy trade-offs when training data are finite; the paper attributes these residual effects to insufficient sample size and states that convergence to a self-consistent solution is expected as more data are collected. Direct first-order optimization of KL-based SCORE also requires ∇xlog⁡pd(y∣x)\nabla_x\log p_d(y\mid x), which is unavailable from samples alone. The authors tried estimating the required data scores with score-matching methods, including NCSN++ with denoising score matching, but found the estimates too high-variance for reliable use in discriminative learning and did not pursue that pipeline. Their practical experiments therefore optimize distance-based variants such as squared error rather than directly optimizing KL-based SCORE. The proposed explanations of overfitting and semantic gradients are presented as insights for further investigation, not as conclusive accounts.

Coverage note — Omitted the toy-plot demonstrations, the auxiliary randomized-smoothing coefficient controls, and detailed gradient visualizations because the principal findings they illustrate are captured in the stated bounds, gradient results, and experiments.

References

  1. 1.Jean-Baptiste Alayrac, Jonathan Uesato, Po-Sen Huang, Alhussein Fawzi, Robert Stanforth, and Pushmeet Kohli. Are labels required for improving adversarial robustness? In Advances in Neural Information Processing Systems (NeurIPS), pages 12192–12202, 2019.
  2. 2.Maksym Andriushchenko and Nicolas Flammarion. Understanding and improving fast adversarial training. In Advances in neural information processing systems (NeurIPS), 2020.
  3. 3.Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International Conference on Machine Learning (ICML), 2018.
  4. 4.Yogesh Balaji, Tom Goldstein, and Judy Hoffman. Instance adaptive adversarial training: Improved accuracy trade-offs in neural nets. arXiv preprint arXiv:1910.08051, 2019.
  5. 5.Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 387–402. Springer, 2013.
  6. 6.Wieland Brendel, Jonas Rauber, Matthias Kummerer, Ivan Ustyuzhaninov, and Matthias Bethge. Accurate, reliable and fast robustness evaluation. Advances in Neural Information Processing Systems, 32:12861–12871, 2019.
  7. 7.Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy (S&P), 2017.
  8. 8.Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin. On evaluating adversarial robustness. arXiv preprint arXiv:1902.06705, 2019.
  9. 9.Yair Carmon, Aditi Raghunathan, Ludwig Schmidt, Percy Liang, and John C Duchi. Unlabeled data improves adversarial robustness. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  10. 10.Alvin Chan, Yi Tay, Yew Soon Ong, and Jie Fu. Jacobian adversarially regularized networks for robustness. In International Conference on Learning Representations (ICLR), 2020.
  11. 11.Jinghui Chen and Quanquan Gu. Rays: A ray searching method for hard-label adversarial attack. In International Conference on Knowledge Discovery & Data Mining (KDD), 2020.
  12. 12.Jinghui Chen, Yuan Cao, and Quanquan Gu. Benign overfitting in adversarially robust linear classification. arXiv preprint arXiv:2112.15250, 2021a.
  13. 13.Tianlong Chen, Zhenyu Zhang, Sijia Liu, Shiyu Chang, and Zhangyang Wang. Robust overfitting may be mitigated by properly learned smoothening. In International Conference on Learning Representations (ICLR), 2021b.
  14. 14.Jeremy M Cohen, Elan Rosenfeld, and J Zico Kolter. Certified adversarial robustness via randomized smoothing. In International Conference on Machine Learning (ICML), 2019.
  15. 15.Keith Conrad. Equivalence of norms. Expository Paper, University Of Connecticut, 2018.
  16. 16.Francesco Croce and Matthias Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In International Conference on Machine Learning (ICML), 2020.
  17. 17.Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. Robustbench: a standardized adversarial robustness benchmark. arXiv preprint arXiv:2010.09670, 2020.
  18. 18.Imre Csiszar and János Körner. Information theory: coding theorems for discrete memoryless systems. Cambridge University Press, 2011.
  19. 19.Gavin Weiguang Ding, Yash Sharma, Kry Yik Chau Lui, and Ruitong Huang. Mma training: Direct input space margin maximization through adversarial training. In International Conference on Learning Representations (ICLR), 2020.
  20. 20.Yinpeng Dong, Qi-An Fu, Xiao Yang, Tianyu Pang, Hang Su, Zihao Xiao, and Jun Zhu. Benchmarking adversarial robustness. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  21. 21.Dominik Maria Endres and Johannes E Schindelin. A new metric for probability distributions. IEEE Transactions on Information theory, 49(7):1858–1860, 2003.
  22. 22.Logan Engstrom, Brandon Tran, Dimitris Tsipras, Ludwig Schmidt, and Aleksander Madry. Exploring the landscape of spatial robustness. In International Conference on Machine Learning (ICML), 2019.
  23. 23.Christian Etmann, Sebastian Lunz, Peter Maass, and Carola-Bibiane Schönlieb. On the connection between adversarial robustness and saliency map interpretability. In International Conference on Machine Learning (ICML), 2019.
  24. 24.Jerome Friedman, Trevor Hastie, and Robert Tibshirani. The elements of statistical learning. Springer series in statistics New York, 2001.
  25. 25.Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations (ICLR), 2019.
  26. 26.Justin Gilmer, Luke Metz, Fartash Faghri, Samuel S Schoenholz, Maithra Raghu, Martin Wattenberg, and Ian Goodfellow. Adversarial spheres. arXiv preprint arXiv:1801.02774, 2018.
  27. 27.Zeinab Golgooni, Mehrdad Saberi, Masih Eskandar, and Mohammad Hossein Rohban. Zerograd: Mitigating and explaining catastrophic overfitting in fgsm adversarial training. arXiv preprint arXiv:2103.15476, 2021.
  28. 28.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR), 2015.
  29. 29.Sven Gowal, Chongli Qin, Jonathan Uesato, Timothy Mann, and Pushmeet Kohli. Uncovering the limits of adversarial training against norm-bounded adversarial examples. arXiv preprint arXiv:2010.03593, 2020.
  30. 30.Sven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg, Dan Andrei Calian, and Timothy A Mann. Improving robustness using generated data. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  31. 31.Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky. Your classifier is secretly an energy based model and you should treat it like one. In International Conference on Learning Representations (ICLR), 2020.
  32. 32.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016.
  33. 33.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  34. 34.Dan Hendrycks, Kimin Lee, and Mantas Mazeika. Using pre-training can improve model robustness and uncertainty. In International Conference on Machine Learning (ICML), 2019.
  35. 35.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  36. 36.Aapo Hyvarinen. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research (JMLR), 6(Apr):695–709, 2005.
  37. 37.Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Anish Athalye, Tran, and Aleksander Madry. Adversarial examples are not bugs, they are features. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  38. 38.P Izmailov, AG Wilson, D Podoprikhin, D Vetrov, and T Garipov. Averaging weights leads to wider optima and better generalization. In Conference on Uncertainty in Artificial Intelligence (UAI), 2018.
  39. 39.Adel Javanmard, Mahdi Soltanolkotabi, and Hamed Hassani. Precise tradeoffs in adversarial training for linear regression. In Conference on Learning Theory (COLT), pages 2034–2078. PMLR, 2020.
  40. 40.Peilin Kang and Seyed-Mohsen Moosavi-Dezfooli. Understanding catastrophic overfitting in adversarial training. arXiv preprint arXiv:2105.02942, 2021.
  41. 41.Harini Kannan, Alexey Kurakin, and Ian Goodfellow. Adversarial logit pairing. arXiv preprint arXiv:1803.06373, 2018.
  42. 42.Simran Kaur, Jeremy Cohen, and Zachary C Lipton. Are perceptually-aligned gradients a general property of robust classifiers? arXiv preprint arXiv:1910.08640, 2019.
  43. 43.Hoki Kim, Woojin Lee, and Jaewook Lee. Understanding catastrophic overfitting in single-step adversarial training. In AAAI Conference on Artificial Intelligence (AAAI), 2021.
  44. 44.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In International Conference on Learning Representations (ICLR), 2014.
  45. 45.Bai Li, Shiqi Wang, Suman Jana, and Lawrence Carin. Towards understanding fast adversarial training. arXiv preprint arXiv:2006.03089, 2020a.
  46. 46.Zichao Li, Liyuan Liu, Chengyu Dong, and Jingbo Shang. Overfitting or underfitting? understand robustness drop in adversarial training. arXiv preprint arXiv:2010.08034, 2020b.
  47. 47.Chen Liu, Zhichao Huang, Mathieu Salzmann, Tong Zhang, and Sabine Susstrunk. On the impact of hard adversarial instances on overfitting in adversarial training. arXiv preprint arXiv:2112.07324, 2021.
  48. 48.Peter Lorenz, Dominik Strassel, Margret Keuper, and Janis Keuper. Is robustbench/autoattack a suitable benchmark for adversarial robustness? In The AAAI Workshop on Adversarial Machine Learning and Beyond, 2022.
  49. 49.Siwei Lyu. Interpretation and generalization of score matching. In Conference on Uncertainty in Artificial Intelligence (UAI), 2009.
  50. 50.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations (ICLR), 2018.
  51. 51.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. Deepfool: a simple and accurate method to fool deep neural networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2574–2582, 2016.
  52. 52.Preetum Nakkiran. Adversarial robustness may be at odds with simplicity. arXiv preprint arXiv:1901.00532, 2019.
  53. 53.Yurii E Nesterov. A method for solving the convex programming problem with convergence rate o (1/k2). In Dokl. akad. nauk Sssr, volume 269, pages 543–547, 1983.
  54. 54.Tianyu Pang, Chao Du, and Jun Zhu. Max-mahalanobis linear discriminant analysis networks. In International Conference on Machine Learning (ICML), 2018.
  55. 55.Tianyu Pang, Kun Xu, Chongxuan Li, Yang Song, Stefano Ermon, and Jun Zhu. Efficient learning of generative models via finite-difference score matching. In Annual Conference on Neural Information Processing Systems (NeurIPS), 2020a.
  56. 56.Tianyu Pang, Xiao Yang, Yinpeng Dong, Kun Xu, Hang Su, and Jun Zhu. Boosting adversarial training with hypersphere embedding. In Advances in Neural Information Processing Systems (NeurIPS), 2020b.
  57. 57.Tianyu Pang, Xiao Yang, Yinpeng Dong, Hang Su, and Jun Zhu. Bag of tricks for adversarial training. In International Conference on Learning Representations (ICLR), 2021.
  58. 58.Maura Pintor, Fabio Roli, Wieland Brendel, and Battista Biggio. Fast minimum-norm adversarial attacks through adaptive norm constraints. In Annual Conference on Neural Information Processing Systems (NeurIPS), 2021.
  59. 59.Rahul Rade and Seyed-Mohsen Moosavi-Dezfooli. Helper-based adversarial training: Reducing excessive margin to achieve a better accuracy vs. robustness trade-off. In ICML 2021 Workshop on Adversarial Machine Learning, 2021.
  60. 60.Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John Duchi, and Percy Liang. Understanding and mitigating the tradeoff between robustness and accuracy. In International Conference on Machine Learning (ICML), 2020.
  61. 61.Sylvestre-Alvise Rebuffi, Sven Gowal, Dan A Calian, Florian Stimberg, Olivia Wiles, and Timothy Mann. Fixing data augmentation to improve adversarial robustness. In Advances in neural information processing systems (NeurIPS), 2021.
  62. 62.Leslie Rice, Eric Wong, and J Zico Kolter. Overfitting in adversarially robust deep learning. In International Conference on Machine Learning (ICML), 2020.
  63. 63.Jérôme Rony, Luiz G Hafemann, Luiz S Oliveira, Ismail Ben Ayed, Robert Sabourin, and Eric Granger. Decoupling direction and norm for efficient gradient-based l2 adversarial attacks and defenses. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  64. 64.Andras Rozsa, Manuel Gunther, and Terrance E Boult. Are accuracy and robustness correlated. In IEEE international conference on machine learning and applications (ICMLA), pages 227–232. IEEE, 2016.
  65. 65.Shibani Santurkar, Andrew Ilyas, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. Image synthesis with a single (robust) classifier. In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  66. 66.Amartya Sanyal, Puneet K. Dokania, Varun Kanade, and Philip Torr. How benign is benign overfitting ? In International Conference on Learning Representations (ICLR), 2021.
  67. 67.Ludwig Schmidt, Shibani Santurkar, Dimitris Tsipras, Kunal Talwar, and Aleksander Madry. Adversarially robust generalization requires more data. In Advances in Neural Information Processing Systems (NeurIPS), pages 5019–5031, 2018.
  68. 68.Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina, Sihui Dai, Chong Xiang, Mung Chiang, and Prateek Mittal. Improving adversarial robustness using proxy distributions. arXiv preprint arXiv:2104.09425, 2021.
  69. 69.Ali Shafahi, Mahyar Najibi, Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! In Advances in Neural Information Processing Systems (NeurIPS), 2019.
  70. 70.Carl-Johann Simon-Gabriel, Yann Ollivier, Leon Bottou, Bernhard Scholkopf, and David Lopez-Paz. First-order adversarial vulnerability of neural networks and input dimension. In International Conference on Machine Learning (ICML), 2019.
  71. 71.Vasu Singla, Sahil Singla, Soheil Feizi, and David Jacobs. Low curvature activations reduce overfitting in adversarial training. In IEEE International Conference on Computer Vision (ICCV), pages 16423–16433, 2021.
  72. 72.Leslie N Smith and Nicholay Topin. Super-convergence: Very fast training of neural networks using large learning rates. In Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications, volume 11006, page 1100612. International Society for Optics and Photonics, 2019.
  73. 73.Chuanbiao Song, Kun He, Liwei Wang, and John E Hopcroft. Improving the generalization of adversarial training with domain adaptation. In International Conference on Learning Representations (ICLR), 2019a.
  74. 74.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems (NeurIPS), pages 11895–11907, 2019.
  75. 75.Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. Sliced score matching: A scalable approach to density and score estimation. In Conference on Uncertainty in Artificial Intelligence (UAI), 2019b.
  76. 76.Gaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, et al. Guided adversarial attack for evaluating and enhancing adversarial defenses. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  77. 77.David Stutz, Matthias Hein, and Bernt Schiele. Disentangling adversarial robustness and generalization. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  78. 78.David Stutz, Matthias Hein, and Bernt Schiele. Confidence-calibrated adversarial training: Generalizing to unseen attacks. In International Conference on Machine Learning (ICML), 2020.
  79. 79.Dong Su, Huan Zhang, Hongge Chen, Jinfeng Yi, Pin-Yu Chen, and Yupeng Gao. Is robustness the cost of accuracy? – a comprehensive study on the robustness of 18 deep image classification models. In The European Conference on Computer Vision (ECCV), 2018.
  80. 80.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations (ICLR), 2014.
  81. 81.Shiyu Tang, Ruihao Gong, Yan Wang, Aishan Liu, Jiakai Wang, Xinyun Chen, Fengwei Yu, Xianglong Liu, Dawn Song, Alan Yuille, et al. Robustart: Benchmarking robustness on architecture design and training techniques. arXiv preprint arXiv:2109.05211, 2021.
  82. 82.Guanhong Tao, Shiqing Ma, Yingqi Liu, and Xiangyu Zhang. Attacks meet interpretability: Attribute-steered detection of adversarial samples. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  83. 83.Florian Tramer and Dan Boneh. Adversarial training and robustness for multiple perturbations. In Advances in Neural Information Processing Systems (NeurIPS), pages 5858–5868, 2019.
  84. 84.Florian Tramer, Jens Behrmann, Nicholas Carlini, Nicolas Papernot, and Jorn-Henrik Jacobsen. Fundamental tradeoffs between invariance and sensitivity to adversarial perturbations. In International Conference on Machine Learning (ICML). PMLR, 2020.
  85. 85.François Treves. Topological Vector Spaces, Distributions and Kernels: Pure and Applied Mathematics, Vol. 25, volume 25. Elsevier, 2016.
  86. 86.Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. Robustness may be at odds with accuracy. In International Conference on Learning Representations (ICLR), 2019.
  87. 87.Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama. Lipschitz-margin training: Scalable certification of perturbation invariance for deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  88. 88.Pascal Vincent. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661–1674, 2011.
  89. 89.S Vivek B and R Venkatesh Babu. Single-step adversarial training with dropout scheduling. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  90. 90.Abraham Wald. Statistical decision functions which minimize the maximum risk. Annals of Mathematics, pages 265–280, 1945.
  91. 91.Haotao Wang, Tianlong Chen, Shupeng Gui, Ting-Kuei Hu, Ji Liu, and Zhangyang Wang. Once-for-all adversarial training: In-situ tradeoff between robustness and accuracy for free. In Advances in Neural Information Processing Systems (NeurIPS), 2020a.
  92. 92.Yisen Wang, Difan Zou, Jinfeng Yi, James Bailey, Xingjun Ma, and Quanquan Gu. Improving adversarial robustness requires revisiting misclassified examples. In International Conference on Learning Representations (ICLR), 2020b.
  93. 93.Eric Wong, Leslie Rice, and J. Zico Kolter. Fast is better than free: Revisiting adversarial training. In International Conference on Learning Representations (ICLR), 2020.
  94. 94.Boxi Wu, Jinghui Chen, Deng Cai, Xiaofei He, and Quanquan Gu. Do wider neural networks really help adversarial robustness? In Annual Conference on Neural Information Processing Systems (NeurIPS), 2021.
  95. 95.Dongxian Wu, Shu-Tao Xia, and Yisen Wang. Adversarial weight perturbation helps robust generalization. Advances in Neural Information Processing Systems (NeurIPS), 33, 2020.
  96. 96.Cihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang, Alan Yuille, and Quoc V Le. Adversarial examples improve image recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  97. 97.Yao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov, and Kamalika Chaudhuri. A closer look at accuracy vs. robustness. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  98. 98.Yaodong Yu, Zitong Yang, Edgar Dobriban, Jacob Steinhardt, and Yi Ma. Understanding generalization in adversarial training via the bias-variance decomposition. arXiv preprint arXiv:2103.09947, 2021.
  99. 99.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In The British Machine Vision Conference (BMVC), 2016.
  100. 100.Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric P Xing, Laurent El Ghaoui, and Michael I Jordan. Theoretically principled trade-off between robustness and accuracy. In International Conference on Machine Learning (ICML), 2019.
  101. 101.Jingfeng Zhang, Xilie Xu, Bo Han, Gang Niu, Lizhen Cui, Masashi Sugiyama, and Mohan Kankanhalli. Attacks which do not kill training make adversarial learning stronger. In International Conference on Machine Learning (ICML), 2020.
  102. 102.Jingfeng Zhang, Jianing Zhu, Gang Niu, Bo Han, Masashi Sugiyama, and Mohan Kankanhalli. Geometry-aware instance-reweighted adversarial training. In International Conference on Learning Representations (ICLR), 2021a.
  103. 103.Yihua Zhang, Guanhuan Zhang, Prashant Khanduri, Mingyi Hong, Shiyu Chang, and Sijia Liu. Revisiting and advancing fast adversarial training through the lens of bi-level optimization. arXiv preprint arXiv:2112.12376, 2021b.

Citation

MLA
Pang, T., et al. “Robustness and Accuracy Could Be Reconcilable by (Proper) Definition”. International Conference on Machine Learning, vol. 162, 2022, pp. 17258–77, https://proceedings.mlr.press/v162/pang22a.html.
APA
Pang, T., Lin, M., Yang, X., Zhu, J., & Yan, S. (2022). Robustness and Accuracy Could Be Reconcilable by (Proper) Definition. International Conference on Machine Learning, 162, 17258–17277. https://proceedings.mlr.press/v162/pang22a.html
Chicago
Pang, T., M. Lin, X. Yang, J. Zhu, and S. Yan. 2022. “Robustness and Accuracy Could Be Reconcilable by (Proper) Definition”. International Conference on Machine Learning 162: 17258–77. https://proceedings.mlr.press/v162/pang22a.html.
Harvard
Pang, T. et al. (2022) “Robustness and Accuracy Could Be Reconcilable by (Proper) Definition”, International Conference on Machine Learning. PMLR, pp. 17258–17277. Available at: https://proceedings.mlr.press/v162/pang22a.html.
Vancouver
1. Pang T, Lin M, Yang X, Zhu J, Yan S (2022) Robustness and Accuracy Could Be Reconcilable by (Proper) Definition. In: International Conference on Machine Learning. PMLR, pp 17258–17277

BibTeX

@InProceedings{pmlr-v162-pang22a,
  title = 	 {Robustness and Accuracy Could Be Reconcilable by ({P}roper) Definition},
  author =       {Pang, Tianyu and Lin, Min and Yang, Xiao and Zhu, Jun and Yan, Shuicheng},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {17258--17277},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/pang22a/pang22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/pang22a.html},
  abstract = 	 {The trade-off between robustness and accuracy has been widely studied in the adversarial literature. Although still controversial, the prevailing view is that this trade-off is inherent, either empirically or theoretically. Thus, we dig for the origin of this trade-off in adversarial training and find that it may stem from the improperly defined robust error, which imposes an inductive bias of local invariance — an overcorrection towards smoothness. Given this, we advocate employing local equivariance to describe the ideal behavior of a robust model, leading to a self-consistent robust error named SCORE. By definition, SCORE facilitates the reconciliation between robustness and accuracy, while still handling the worst-case uncertainty via robust optimization. By simply substituting KL divergence with variants of distance metrics, SCORE can be efficiently minimized. Empirically, our models achieve top-rank performance on RobustBench under AutoAttack. Besides, SCORE provides instructive insights for explaining the overfitting phenomenon and semantic input gradients observed on robust models.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/