A Modern Look at the Relationship between Sharpness and Generalization
Maksym AndriushchenkoFrancesco CroceMaximilian MüllerMatthias HeinNicolas Flammarion
Reveals that reparametrization-invariant sharpness measures fail to reliably predict generalization in modern deep architectures like vision transformers and language models, showing that sharpness often correlates primarily with training hyperparameters rather than test performance and can even inversely correlate with out-of-distribution error.
In modern machine learning, understanding why certain neural networks generalize well to unseen data remains a fundamental challenge. A long-standing intuition holds that flatter loss minima produce better generalization and higher robustness than sharp minima, motivating widely used optimization techniques designed to explicitly minimize sharpness. However, prior empirical support was largely restricted to older convolutional architectures on small datasets, while standard mathematical measures of sharpness fail to remain consistent when models undergo basic parameter transformations. The article systematically evaluates whether modern, scale-invariant adaptive sharpness measures can reliably predict model generalization across modern machine learning benchmarks.
To conduct this evaluation, the authors performed extensive empirical analyses across a range of modern deep learning workflows. They examined hundreds of models across vision and natural language processing tasks, evaluating both standard in-distribution test error and out-of-distribution performance under data corruptions and domain shifts. The empirical suite included vision transformers and ResNets trained from scratch on CIFAR-10 and ImageNet, as well as CLIP and BERT transformer architectures fine-tuned on ImageNet and the Multi-genre Natural Language Inference benchmark. The study evaluated twelve mathematical formulations of sharpness—varying optimization bounds, perturbation types, radii, and logit normalization—using Auto-PGD optimization to ensure stable, hyperparameter-free convergence.
The investigation produced several key findings that challenge conventional assumptions. First, reparametrization-invariant sharpness does not consistently correlate with test accuracy across broad sets of models, showing rank correlation coefficients near zero across ImageNet and language tasks. Second, when models are evaluated on out-of-distribution benchmarks, sharpness frequently exhibits a statistically significant negative correlation with error (often reaching rank correlations between negative 0.30 and negative 0.60), demonstrating that sharper minima can generalize significantly better than flatter ones. Third, measured sharpness primarily acts as a proxy for specific training hyperparameters, notably the learning rate, rather than reflecting generalization capability. Finally, theoretical and empirical analyses of simplified diagonal linear models confirm that the optimal definition of sharpness is strictly dependent on the underlying data distribution, showing that universal geometric measures of generalization do not exist.
These findings have critical implications for research and engineering practices. Practitioners and leaders should no longer rely on sharpness metrics as reliable indicators of model quality, validation proxies, or guarantees of out-of-distribution safety. The results demonstrate that blanket assertions such as "flatter minima generalize better" are factually incorrect in modern transformer and transfer-learning regimes. While optimization algorithms targeting sharpness can sometimes provide regularizing benefits, their performance gains stem from stochastic dynamics and learning rate interactions rather than a universal geometric principle.
Organizations developing machine learning models should avoid using sharpness as a model-selection criterion or diagnostic metric for deployment readiness. Future research should prioritize developing data-dependent complexity measures and investigating the implicit regularization mechanics of training algorithms directly on target data distributions. While the empirical findings are supported by a rigorous experimental protocol across standard benchmarks, readers should note that evaluations were conducted primarily within fixed model families and standard vision and language tasks; further verification may be required when assessing novel architectures or unconventional training objectives.
- Paper: Sharpness-Aware Minimization for Efficiently Improving Generalization, Pierre Foret et al. (2020). Read the paper that introduced SAM first to understand the sharpness-minimization approach and theoretical motivation this study evaluates.
- Paper: On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, Nitish Shirish Keskar et al. (2016). This influential study established the empirical link between sharp minima and poorer generalization that the source re-examines across modern settings.
- Paper: Exploring Generalization in Deep Learning, Behnam Neyshabur et al. (2017). Its analysis of sharpness’s sensitivity to parameter rescaling provides essential context for the source’s focus on reparametrization-invariant measures.
- Paper: Towards Understanding Sharpness-Aware Minimization, Maksym Andriushchenko et al. (2022). Its theoretical and empirical critique of flat-minima explanations for SAM prepares readers for the source’s broader challenge to sharpness as a generalization indicator.
- Paper: Visualizing the Loss Landscape of Neural Nets, Hao Li et al. (2017). Its filter-wise normalization addresses scale distortions in loss-landscape visualizations, clarifying why sharpness measurements must account for parameter transformations.
- Paper: Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves Generalization, Xingxuan Zhang et al. (2023). Building on the debate over whether sharpness predicts generalization, this paper proposes first-order flatness as an alternative criterion and develops an optimizer around it.
