A Modern Look at the Relationship between Sharpness and Generalization

Maksym AndriushchenkoFrancesco CroceMaximilian MüllerMatthias HeinNicolas Flammarion

article2023ICML105 citations

Reveals that reparametrization-invariant sharpness measures fail to reliably predict generalization in modern deep architectures like vision transformers and language models, showing that sharpness often correlates primarily with training hyperparameters rather than test performance and can even inversely correlate with out-of-distribution error.

Listen

In modern machine learning, understanding why certain neural networks generalize well to unseen data remains a fundamental challenge. A long-standing intuition holds that flatter loss minima produce better generalization and higher robustness than sharp minima, motivating widely used optimization techniques designed to explicitly minimize sharpness. However, prior empirical support was largely restricted to older convolutional architectures on small datasets, while standard mathematical measures of sharpness fail to remain consistent when models undergo basic parameter transformations. The article systematically evaluates whether modern, scale-invariant adaptive sharpness measures can reliably predict model generalization across modern machine learning benchmarks.

To conduct this evaluation, the authors performed extensive empirical analyses across a range of modern deep learning workflows. They examined hundreds of models across vision and natural language processing tasks, evaluating both standard in-distribution test error and out-of-distribution performance under data corruptions and domain shifts. The empirical suite included vision transformers and ResNets trained from scratch on CIFAR-10 and ImageNet, as well as CLIP and BERT transformer architectures fine-tuned on ImageNet and the Multi-genre Natural Language Inference benchmark. The study evaluated twelve mathematical formulations of sharpness—varying optimization bounds, perturbation types, radii, and logit normalization—using Auto-PGD optimization to ensure stable, hyperparameter-free convergence.

The investigation produced several key findings that challenge conventional assumptions. First, reparametrization-invariant sharpness does not consistently correlate with test accuracy across broad sets of models, showing rank correlation coefficients near zero across ImageNet and language tasks. Second, when models are evaluated on out-of-distribution benchmarks, sharpness frequently exhibits a statistically significant negative correlation with error (often reaching rank correlations between negative 0.30 and negative 0.60), demonstrating that sharper minima can generalize significantly better than flatter ones. Third, measured sharpness primarily acts as a proxy for specific training hyperparameters, notably the learning rate, rather than reflecting generalization capability. Finally, theoretical and empirical analyses of simplified diagonal linear models confirm that the optimal definition of sharpness is strictly dependent on the underlying data distribution, showing that universal geometric measures of generalization do not exist.

These findings have critical implications for research and engineering practices. Practitioners and leaders should no longer rely on sharpness metrics as reliable indicators of model quality, validation proxies, or guarantees of out-of-distribution safety. The results demonstrate that blanket assertions such as "flatter minima generalize better" are factually incorrect in modern transformer and transfer-learning regimes. While optimization algorithms targeting sharpness can sometimes provide regularizing benefits, their performance gains stem from stochastic dynamics and learning rate interactions rather than a universal geometric principle.

Organizations developing machine learning models should avoid using sharpness as a model-selection criterion or diagnostic metric for deployment readiness. Future research should prioritize developing data-dependent complexity measures and investigating the implicit regularization mechanics of training algorithms directly on target data distributions. While the empirical findings are supported by a rigorous experimental protocol across standard benchmarks, readers should note that evaluations were conducted primarily within fixed model families and standard vision and language tasks; further verification may be required when assessing novel architectures or unconventional training objectives.

Cover for A Modern Look at the Relationship between Sharpness and Generalization

Citation

MLA
Andriushchenko, M., et al. “A Modern Look at the Relationship Between Sharpness and Generalization”. International Conference on Machine Learning, vol. 202, 2023, pp. 840–902, https://proceedings.mlr.press/v202/andriushchenko23a.html.
APA
Andriushchenko, M., Croce, F., Müller, M., Hein, M., & Flammarion, N. (2023). A Modern Look at the Relationship between Sharpness and Generalization. International Conference on Machine Learning, 202, 840–902. https://proceedings.mlr.press/v202/andriushchenko23a.html
Chicago
Andriushchenko, M., F. Croce, M. Müller, M. Hein, and N. Flammarion. 2023. “A Modern Look at the Relationship Between Sharpness and Generalization”. International Conference on Machine Learning 202: 840–902. https://proceedings.mlr.press/v202/andriushchenko23a.html.
Harvard
Andriushchenko, M. et al. (2023) “A Modern Look at the Relationship between Sharpness and Generalization”, International Conference on Machine Learning. PMLR, pp. 840–902. Available at: https://proceedings.mlr.press/v202/andriushchenko23a.html.
Vancouver
1. Andriushchenko M, Croce F, Müller M, Hein M, Flammarion N (2023) A Modern Look at the Relationship between Sharpness and Generalization. In: International Conference on Machine Learning. PMLR, pp 840–902

BibTeX

@InProceedings{pmlr-v202-andriushchenko23a,
  title = 	 {A Modern Look at the Relationship between Sharpness and Generalization},
  author =       {Andriushchenko, Maksym and Croce, Francesco and M\"{u}ller, Maximilian and Hein, Matthias and Flammarion, Nicolas},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {840--902},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/andriushchenko23a/andriushchenko23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/andriushchenko23a.html},
  abstract = 	 {Sharpness of minima is a promising quantity that can correlate with generalization in deep networks and, when optimized during training, can improve generalization. However, standard sharpness is not invariant under reparametrizations of neural networks, and, to fix this, reparametrization-invariant sharpness definitions have been proposed, most prominently adaptive sharpness (Kwon et al., 2021). But does it really capture generalization in modern practical settings? We comprehensively explore this question in a detailed study of various definitions of adaptive sharpness in settings ranging from training from scratch on ImageNet and CIFAR-10 to fine-tuning CLIP on ImageNet and BERT on MNLI. We focus mostly on transformers for which little is known in terms of sharpness despite their widespread usage. Overall, we observe that sharpness does not correlate well with generalization but rather with some training parameters like the learning rate that can be positively or negatively correlated with generalization depending on the setup. Interestingly, in multiple cases, we observe a consistent negative correlation of sharpness with OOD generalization implying that sharper minima can generalize better. Finally, we illustrate on a simple model that the right sharpness measure is highly data-dependent, and that we do not understand well this aspect for realistic data distributions.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/