SGD with Large Step Sizes Learns Sparse Features
Maksym AndriushchenkoAditya Vardhan VarreLoucas Pillaud-VivienNicolas Flammarion
Explains how large initial step sizes in stochastic gradient descent cause loss stabilization that amplifies gradient noise, implicitly biasing deep networks toward learning sparse features and improving generalization.
Deep neural networks are central to modern artificial intelligence, yet explaining why they generalize effectively to unseen data remains an open challenge. In practice, training routines rely on stochastic gradient descent using large initial step sizes that are later reduced. These large step sizes often cause the training error to temporarily level off—a state termed loss stabilization—before eventually dropping. The article investigates the underlying mechanics of this phase to understand how optimization schedules drive neural networks toward simpler, more generalizable solutions.
The main objective of the article is to demonstrate that stochastic gradient descent with large step sizes induces an implicit mathematical regularization that actively discovers sparse features. The authors evaluate this mechanism to show how the interplay between gradient updates and stochastic noise naturally simplifies the internal representations of deep networks without requiring explicit penalties like weight decay.
To conduct this evaluation, the authors combine theoretical modeling with extensive empirical testing across models of increasing complexity. The theoretical framework maps the training process during loss stabilization to a continuous-time stochastic differential equation with multiplicative noise. Empirically, the authors evaluate diagonal linear networks, shallow and deep networks with piecewise-linear activations, and modern deep convolutional architectures (ResNet-18, ResNet-34, and DenseNet-100) trained on image benchmark datasets including CIFAR-10, CIFAR-100, and Tiny ImageNet under both basic and state-of-the-art training configurations.
The investigation yields three primary findings. First, maintaining large step sizes keeps the optimization bouncing across the walls of loss valleys, sustaining a non-vanishing noise level that drives a hidden simplification process. Second, longer loss stabilization periods produce substantially sparser feature representations; for example, the fraction of active, distinct features in deep DenseNet layers dropped from 50–60% under small step sizes down to 10–20% under large step size schedules. Third, this representation sparsity directly translates into superior predictive performance: on CIFAR-100 image classification, models trained with extended large step sizes achieved test error rates around 35%, compared to 60% for runs using small step sizes.
These findings provide practical insights for machine learning operations and model design. They show that step size schedules act as primary regularizers rather than purely computational speed controls. This explains why standard engineering practices, such as learning rate warmup and batch normalization, improve performance: they prevent early numerical divergence and allow models to operate safely in large-step regimes that promote feature selection. The results also clarify that the training lifecycle naturally separates into an initial representation-simplification phase followed by a data-fitting phase once the step size is decayed.
Practitioners should design learning rate schedules that prolong the large step size phase before decay, using gradual warmup schedules to avoid instability. Teams tuning neural network training pipelines can leverage smaller batch sizes or higher learning rates to enhance implicit regularization instead of relying entirely on explicit tuning penalties. Further research is recommended to mathematically establish whether this feature sparsity is strictly causal to general predictive gains on complex, real-world data.
The study's primary limitations stem from the theoretical derivations relying on simplified models and continuous-time stochastic approximations, which do not fully capture every discrete interaction in complex architectures. Additionally, setting step sizes excessively high can lead to overregularization where models fail to fit the training data entirely. Nonetheless, the high consistency of empirical results across diverse model scales provides strong confidence in the core conclusion that large step sizes drive sparse feature learning.
- Paper: Understanding the unstable convergence of gradient descent, Kwangjun Ahn et al. (2022). Its analysis of unstable convergence clarifies how gradient descent can remain bounded and make progress beyond conventional step-size limits, setting up the large-step training regime studied here.
- Paper: The alignment property of SGD noise and how it helps select flat minima: A stability analysis, Lei Wu et al. (2022). Its account of how SGD noise aligns with sharp loss directions provides useful groundwork for understanding the stochastic dynamics behind large-step implicit regularization.
- Paper: How Does Batch Normalization Help Optimization?, Shibani Santurkar et al. (2018). Its analysis of Batch Normalization’s optimization effects helps contextualize the source’s claim that normalization enables stable training at large step sizes.
- Paper: High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation, Jimmy Ba et al. (2022). Its study of large-step feature learning gives a theoretical point of comparison for the source’s account of how step-size choices reshape learned representations.
No sufficiently relevant recommendations were found.
