Understanding Robust Overfitting of Adversarial Training and Beyond
Chaojian YuBo HanLi ShenJun YuChen GongMingming GongTongliang Liu
Shows that overfitting in adversarial training is driven by easily fitted, small-loss data under strong attacks and introduces minimum loss constrained adversarial training to prevent test degradation by actively increasing the loss on easy examples.
Deep neural networks are vulnerable to adversarial attacks, where subtle, human-imperceptible input modifications mislead models into making incorrect predictions. Adversarial training serves as one of the primary defenses by training models directly on attacked examples. However, standard adversarial training routinely suffers from robust overfitting, a phenomenon where a model's defense performance against attacks peaks during training and then significantly degrades. Understanding and fixing this issue without relying on expensive, newly collected datasets is crucial for developing reliable, secure artificial intelligence systems.
The article aims to uncover the exact root causes of robust overfitting and evaluate a new training prototype designed to eliminate this degradation while boosting overall adversarial defense performance.
To investigate the mechanism, the authors compared data loss distributions across weak and strong adversarial training regimes. Through data ablation experiments, they selectively removed samples from specific loss ranges across multiple benchmark image datasets (CIFAR10, CIFAR100, and SVHN), network architectures, and threat conditions. Building on these empirical insights, the article introduces Minimum Loss Constrained Adversarial Training (MLCAT), a framework that identifies easy-to-learn (small-loss) adversarial samples within each training batch and applies targeted interventions to artificially increase their loss, evaluated via two orthogonal implementations: loss scaling (MLCATLS) and weight perturbation (MLCATWP).
Key findings show that robust overfitting under strong attack settings is caused specifically by small-loss adversarial data—particularly samples that transition from hard to easy as the network learns—rather than large-loss samples. Standard adversarial training exhibited substantial robust accuracy drops of roughly 3 to 8 percentage points between peak and final training checkpoints. In contrast, both MLCAT implementations virtually eliminated robust overfitting, compressing this accuracy drop to under 1 percentage point across all test settings. Furthermore, MLCAT using weight perturbation consistently improved final robust accuracy (e.g., reaching approximately 50.3% to 54.6% Auto Attack accuracy on CIFAR10, compared to 42.1% to 46.1% for standard adversarial training) while maintaining standard, unattacked classification performance.
These findings challenge the assumption that large-loss data induce robust overfitting, demonstrating instead that easily fitted adversarial samples undermine defense generalization. Practically, this framework delivers higher security and model stability without the substantial operational and financial costs of acquiring extra training data or the operational risks of premature early stopping. However, the results show an important trade-off: while loss scaling eliminates the training gap, it creates vulnerability to logit-scaling attacks, making weight perturbation the superior practical implementation.
Organizations developing robust deep learning models should adopt parameter-perturbation-based minimum loss constraints within their adversarial training pipelines to prevent performance decay. Implementation teams must carefully calibrate the minimum loss threshold to the specific dataset and threat level, as setting the constraint too high can destabilize training and cause model collapse. Confidence in these conclusions is high given consistent validation across diverse datasets, network backbones, and standard attack benchmarks, though practical deployment requires standard tuning of loss thresholds for novel data domains.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). Madry et al. establish the robust-optimization formulation of adversarial training that this paper uses as the foundation for analyzing and modifying training.
- Paper: Theoretically Principled Trade-off between Robustness and Accuracy, Hongyang Zhang et al. (2019). TRADES develops a widely used adversarial-training objective, giving useful context for how the paper’s minimum-loss intervention relates to established robustness-training methods.
No sufficiently relevant recommendations were found.
