Robustness and Accuracy Could Be Reconcilable by (Proper) Definition
Tianyu PangMin LinXiao YangJun ZhuShuicheng Yan
Proposes a self-consistent error objective termed SCORE that replaces conventional local invariance with local equivariance, eliminating the artificial conflict between clean accuracy and adversarial resilience to achieve top-ranked defense performance under AutoAttack.
Modern machine learning models deployed in safety-critical applications are highly vulnerable to adversarial attacks, where small, imperceptible input perturbations cause critical misclassifications. Defending against these attacks through adversarial training has historically introduced a persistent drawback: a significant drop in baseline accuracy on unperturbed, standard data. The prevailing view in the artificial intelligence community has long held that this trade-off between robustness and accuracy is an unavoidable, inherent cost. The article investigates the origin of this tension and evaluates whether it is actually a mathematical artifact resulting from an improper definition of robust error.
The article demonstrates that the standard formulation of adversarial training forces the model toward excessive local invariance—an overcorrection that demands model predictions remain entirely constant around an input, ignoring natural data shifts. To resolve this, the article introduces the Self-Consistent Robust Error (SCORE), a theoretical objective that promotes local equivariance, allowing the model to smoothly match the underlying true data distribution point by point. Because directly calculating SCORE requires inaccessible gradient information from the true data distribution, the authors mathematically derive bounds using distance metrics. This reveals that substituting conventional information-divergence losses with monotonically increasing convex distance variants—specifically squared error—enables efficient optimization without extra computational overhead.
The empirical evaluation was conducted across benchmark image datasets (CIFAR-10, CIFAR-100, and ImageNet) using standard residual network architectures against standard adversarial benchmarks, including the rigorous AutoAttack framework. Key findings demonstrate that substituting squared error into standard adversarial training pipelines consistently improves accuracy on clean data while maintaining or increasing robustness against attacks. On large-scale benchmarks utilizing synthetic data, the proposed method raised clean accuracy by approximately 2.0 to 5.1 percentage points over state-of-the-art baselines while matching or surpassing top-tier robust accuracy (achieving up to 63.74% robust accuracy on CIFAR-10 and 33.65% on CIFAR-100). The theoretical framework also successfully explains why adversarial training frequently suffers from robust overfitting and why robust models inherently learn semantically meaningful, shape-based visual gradients.
These findings provide immediate practical implications for engineering, safety, and compliance teams deploying machine learning systems in security-sensitive environments. Organizations do not need to accept substantial performance degradation on standard operations to achieve adversarial resilience. Because the proposed modification simply replaces the loss function during model training, it incurs zero additional computational overhead or inference latency, effectively lowering the cost and technical risk of robust machine learning deployments.
Organizations developing or deploying safety-critical vision systems should adopt distance-based formulations, such as squared error losses, within their current adversarial training pipelines. Furthermore, practitioners should tune regularization parameters toward lower values when leveraging synthetic or augmented data to maximize baseline performance. Future work should expand these principles beyond standard vision threat models to evaluate non-vision modalities and test generative score-matching implementations as data distribution estimators mature.
Confidence in these findings is high regarding standard image classification benchmarks, supported by verified sanity checks confirming the absence of false robustness. However, decision-makers should note that empirical trade-offs can still arise in practical settings with limited training samples, as the theoretical reconciliation relies on sufficient data coverage. Caution is also advised when transferring these methods to distinct data modalities or specialized hardware setups where data augmentation dynamics may vary.
- Paper: Theoretically Principled Trade-off between Robustness and Accuracy, Hongyang Zhang et al. (2019). Its decomposition of robust error and the TRADES objective provide the framework whose definition of robustness this paper revisits.
- Paper: Robustness May Be at Odds with Accuracy, Dimitris Tsipras et al. (2018). This influential account of an inherent accuracy–robustness trade-off supplies the prevailing claim that the source challenges.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). Its robust-optimization formulation establishes the worst-case adversarial training setup that the source retains while redefining robust error.
- Paper: Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks, Francesco Croce et al. (2020). Its AutoAttack ensemble underpins the robustness evaluation used to assess the source’s models.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). Its account of adversarial examples and gradient-based attacks supplies foundational context for the adversarial robustness problem the source addresses.
No sufficiently relevant recommendations were found.
