keyword
robust generalization
Robust generalization is the capacity of a machine learning model to maintain high predictive accuracy and reliable performance on unseen data that has been subjected to adversarial attacks, noise, or distribution shifts. Unlike standard generalization, which measures how well a model performs on clean test samples drawn from the same distribution as the training data, robust generalization specifically evaluates performance under perturbed or worst-case inputs. Achieving robust generalization is a central goal in trustworthy machine learning, requiring models to learn invariant representations while overcoming challenges such as robust overfitting, where a model successfully defends against training perturbations but fails to generalize to unseen perturbed test samples.
3 items

Understanding Robust Overfitting of Adversarial Training and Beyond
Chaojian Yu, Bo Han, Li Shen, Jun Yu, Chen Gong, Mingming Gong, Tongliang Liu
Why you should read this
Shows that overfitting in adversarial training is driven by easily fitted, small-loss data under strong attacks and introduces minimum loss constrained adversarial training to prevent test degradation by actively increasing the loss on easy examples.
Added
2026-10-02

Understanding The Robustness in Vision Transformers
Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Animashree Anandkumar, Jiashi Feng, José M. Álvarez
Why you should read this
Explains how visual grouping in self-attention reduces corruption sensitivity through an information bottleneck perspective, introducing fully attentional networks that integrate dynamic channel selection to achieve superior corruption error on ImageNet-C.
Recent studies show that Vision Transformers (ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code will be available at https://github.com/NVlabs/FAN.
Added
2026-09-26

Theoretically Principled Trade-off between Robustness and Accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric P. Xing, Laurent El Ghaoui, Michael I. Jordan
Why you should read this
Introduces TRADES, a defense algorithm derived from a tight theoretical bound that decomposes prediction error into classification and boundary terms to systematically balance clean accuracy against adversarial attacks.
We identify a trade-off between robustness and accuracy that serves as a guiding principle in the design of defenses against adversarial examples. Although this problem has been widely studied empirically, much remains unknown concerning the theory underlying this trade-off. In this work, we decompose the prediction error for adversarial examples (robust error) as the sum of the natural (classification) error and boundary error, and provide a differentiable upper bound using the theory of classification-calibrated loss, which is shown to be the tightest possible upper bound uniform over all probability distributions and measurable predictors. Inspired by our theoretical analysis, we also design a new defense method, TRADES, to trade adversarial robustness off against accuracy. Our proposed algorithm performs well experimentally in real-world datasets. The methodology is the foundation of our entry to the NeurIPS 2018 Adversarial Vision Challenge in which we won the 1st place out of ~2,000 submissions, surpassing the runner-up approach by in terms of mean perturbation distance.
Added
2026-09-12

