SGD with Large Step Sizes Learns Sparse Features

Maksym AndriushchenkoAditya Vardhan VarreLoucas Pillaud-VivienNicolas Flammarion

article2023ICML83 citations

Explains how large initial step sizes in stochastic gradient descent cause loss stabilization that amplifies gradient noise, implicitly biasing deep networks toward learning sparse features and improving generalization.

Listen

Deep neural networks are central to modern artificial intelligence, yet explaining why they generalize effectively to unseen data remains an open challenge. In practice, training routines rely on stochastic gradient descent using large initial step sizes that are later reduced. These large step sizes often cause the training error to temporarily level off—a state termed loss stabilization—before eventually dropping. The article investigates the underlying mechanics of this phase to understand how optimization schedules drive neural networks toward simpler, more generalizable solutions.

The main objective of the article is to demonstrate that stochastic gradient descent with large step sizes induces an implicit mathematical regularization that actively discovers sparse features. The authors evaluate this mechanism to show how the interplay between gradient updates and stochastic noise naturally simplifies the internal representations of deep networks without requiring explicit penalties like weight decay.

To conduct this evaluation, the authors combine theoretical modeling with extensive empirical testing across models of increasing complexity. The theoretical framework maps the training process during loss stabilization to a continuous-time stochastic differential equation with multiplicative noise. Empirically, the authors evaluate diagonal linear networks, shallow and deep networks with piecewise-linear activations, and modern deep convolutional architectures (ResNet-18, ResNet-34, and DenseNet-100) trained on image benchmark datasets including CIFAR-10, CIFAR-100, and Tiny ImageNet under both basic and state-of-the-art training configurations.

The investigation yields three primary findings. First, maintaining large step sizes keeps the optimization bouncing across the walls of loss valleys, sustaining a non-vanishing noise level that drives a hidden simplification process. Second, longer loss stabilization periods produce substantially sparser feature representations; for example, the fraction of active, distinct features in deep DenseNet layers dropped from 50–60% under small step sizes down to 10–20% under large step size schedules. Third, this representation sparsity directly translates into superior predictive performance: on CIFAR-100 image classification, models trained with extended large step sizes achieved test error rates around 35%, compared to 60% for runs using small step sizes.

These findings provide practical insights for machine learning operations and model design. They show that step size schedules act as primary regularizers rather than purely computational speed controls. This explains why standard engineering practices, such as learning rate warmup and batch normalization, improve performance: they prevent early numerical divergence and allow models to operate safely in large-step regimes that promote feature selection. The results also clarify that the training lifecycle naturally separates into an initial representation-simplification phase followed by a data-fitting phase once the step size is decayed.

Practitioners should design learning rate schedules that prolong the large step size phase before decay, using gradual warmup schedules to avoid instability. Teams tuning neural network training pipelines can leverage smaller batch sizes or higher learning rates to enhance implicit regularization instead of relying entirely on explicit tuning penalties. Further research is recommended to mathematically establish whether this feature sparsity is strictly causal to general predictive gains on complex, real-world data.

The study's primary limitations stem from the theoretical derivations relying on simplified models and continuous-time stochastic approximations, which do not fully capture every discrete interaction in complex architectures. Additionally, setting step sizes excessively high can lead to overregularization where models fail to fit the training data entirely. Nonetheless, the high consistency of empirical results across diverse model scales provides strong confidence in the core conclusion that large step sizes drive sparse feature learning.

Andriushchenko et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for SGD with Large Step Sizes Learns Sparse Features

Abstract

Deep neural networks trained with SGD learn features as data representations which form the basis of predictions made during test time.

In over-parameterized models where sparse solutions exist, stochastic gradient descent (SGD) learns dense-to-sparse conversion

that reliably identifies good sparse predictors whose performance matches original dense networks across several models

to possibly ease interpretability. However, using large learning rates for (stochastic) gradient descent induces

difficulty due to labels while also requiring at least two feature types through different training stages.

Table of Contents

  • 1. Introduction
  • 1.1. Our Contributions
  • 1.2. Related Work
  • 2. The Effective Dynamics of SGD with Large Step Size: Sparse Feature Learning
  • 2.1. Background: SGD is GD with Specific Label Noise
  • 2.2. The Effective Dynamics Behind Loss Stabilization
  • 2.3. Sparse Feature Learning
  • 2.3.1. A WARM-UP: DIAGONAL LINEAR NETWORKS
  • 2.3.2. THE SPARSE FEATURE LEARNING CONJECTURE FOR MORE GENERAL MODELS
  • 3. Empirical Evidence of Sparse Feature Learning Driven by SGD
  • 3.1. Sparse Feature Learning in Diagonal Linear Networks
  • 3.2. Sparse Feature Learning in Simple ReLU Networks
  • 3.3. Sparse Feature Learning in Deep ReLU Networks
  • 4. Conclusions and Insights from our Understanding of the Training Dynamics
  • Acknowledgements
  • References
  • APPENDIX
  • A. SGD and Label Noise GD
  • B. Quadratic Parameterization in One Dimension
  • C. Empirical Validation of the SDE Modeling
  • D. Additional Experimental Results

Knowls

  1. Knowl 1 — Single-sample SGD is full-batch descent with state-dependent label noise

    theoretical result

    For squared-error training on nn examples (xi,yi)(x_i,y_i), let hθ(xi)h_\theta(x_i) be the prediction of a model with parameter vector θ∈Rp\theta\in\mathbb{R}^p, and define L(θ)=12n∑i=1n(hθ(xi)−yi)2L(\theta)=\frac{1}{2n}\sum_{i=1}^n(h_\theta(x_i)-y_i)^2. At iteration tt, single-sample SGD draws iti_t uniformly from {1,…,n}\{1,\ldots,n\}. Define a random perturbation of each label by

    ξt,i=(hθt(xi)−yi)(1−n1{it=i}),yt,i=yi+ξt,i,\xi_{t,i}=(h_{\theta_t}(x_i)-y_i)\bigl(1-n\mathbf{1}_{\{i_t=i\}}\bigr),\qquad y_{t,i}=y_i+\xi_{t,i},

    where 1{it=i}\mathbf{1}_{\{i_t=i\}} is one if sample ii was drawn and zero otherwise. Then the resulting full-batch update is exactly the SGD update:

    θt+1=θt−ηn∑i=1n(hθt(xi)−yt,i)∇θhθt(xi),\theta_{t+1}=\theta_t-\frac{\eta}{n}\sum_{i=1}^n(h_{\theta_t}(x_i)-y_{t,i})\nabla_\theta h_{\theta_t}(x_i),

    with step size η>0\eta>0. Conditional on θt\theta_t, the label-noise vector has mean zero and satisfies E∥ξt∥22=2n(n−1)L(θt)\mathbb{E}\|\xi_t\|_2^2=2n(n-1)L(\theta_t). Thus this exact rewriting ties the noise scale to the current training loss, and the noisy update is confined to the span of the sample prediction gradients.

  2. Knowl 2 — A quadratic one-dimensional model exhibits almost-sure loss stabilization and bouncing

    theoretical result

    Consider scalar inputs x>0x>0 drawn from a distribution supported on [xmin⁡,xmax⁡][x_{\min},x_{\max}], targets y=xθ∗2y=x\theta_*^2 for a fixed θ∗>0\theta_*>0, prediction xθ2x\theta^2, and loss F(θ)=14E[(xθ∗2−xθ2)2]F(\theta)=\frac14\mathbb{E}[(x\theta_*^2-x\theta^2)^2]. For SGD initialized at 0<θ0<θ∗0<\theta_0<\theta_*, if

    (θ∗xmin⁡)−2<η<1.25(θ∗xmax⁡)−2,(\theta_*x_{\min})^{-2}<\eta<1.25(\theta_*x_{\max})^{-2},

    then, almost surely, the loss remains bounded away from zero and below a fixed positive level: F(θt)∈(ϵ02θ∗2, 0.17θ∗2)F(\theta_t)\in(\epsilon_0^2\theta_*^2,\,0.17\theta_*^2), where ϵ0=min⁡{(η(θ∗xmin⁡)2−1)/3, 0.02}\epsilon_0=\min\{(\eta(\theta_*x_{\min})^2-1)/3,\,0.02\}. After a transient, successive iterates lie on opposite sides of θ∗\theta_*, with the lower-side iterates in (0.65θ∗,(1−ϵ0)θ∗)(0.65\theta_*,(1-\epsilon_0)\theta_*) and the upper-side iterates in ((1+ϵ0)θ∗,1.162θ∗)((1+\epsilon_0)\theta_*,1.162\theta_*). This stochastic nonconvex example shows how a range of large step sizes can sustain a high-loss oscillation rather than produce either convergence or divergence.

  3. Knowl 3 — Proposed SDE for the slow dynamics during loss stabilization

    model/method

    For squared-error training, let θ∈Rp\theta\in\mathbb{R}^p be the network parameters, nn the number of training examples, L(θ)=12n∑i=1n(hθ(xi)−yi)2L(\theta)=\frac{1}{2n}\sum_{i=1}^n(h_\theta(x_i)-y_i)^2, and let ϕθ(X)∈Rn×p\phi_\theta(X)\in\mathbb{R}^{n\times p} have row ii equal to ∇θhθ(xi)⊤\nabla_\theta h_\theta(x_i)^\top. The paper proposes the following stochastic differential equation as a model of the effective, slower parameter dynamics while large-step-size SGD rapidly bounces in the loss valley:

    dθt=−∇θL(θt) dt+ηδ ϕθt(X)⊤dBt.d\theta_t=-\nabla_\theta L(\theta_t)\,dt+\sqrt{\eta\delta}\,\phi_{\theta_t}(X)^\top dB_t.

    Here η>0\eta>0 is the SGD step size, δ>0\delta>0 is the approximately stabilized loss level that sets the noise intensity, and BtB_t is standard Brownian motion in Rn\mathbb{R}^n. The drift is full-batch gradient descent; the multiplicative noise acts through the sample-gradient (neural tangent kernel feature) matrix. This is a proposed effective model, not a claim that the discrete SGD process is exactly an SDE.

  4. Knowl 4 — Diagonal linear networks connect multiplicative noise to sparse predictors

    theoretical result

    A diagonal linear network predicts hu,v(x)=⟨u⊙v,x⟩h_{u,v}(x)=\langle u\odot v,x\rangle with parameter vectors u,v∈Rdu,v\in\mathbb{R}^d and elementwise product ⊙\odot; its effective linear predictor is β=u⊙v\beta=u\odot v. The paper uses existing analysis of the corresponding noise-driven dynamics under sparse-recovery assumptions: if β∗\beta^* is the sparsest predictor interpolating the data, then coordinates of βt=ut⊙vt\beta_t=u_t\odot v_t outside supp⁡(β∗)\operatorname{supp}(\beta^*) converge exponentially fast to zero, while the coordinates on the support lie with high probability in an O(ηδ)O(\sqrt{\eta\delta}) neighborhood of β∗\beta^* after time O(δ−1)O(\delta^{-1}). Here η\eta is the SGD step size and δ\delta is the stabilized noise-intensity level. This result supplies a concrete sparse-predictor mechanism for the paper's proposed effective dynamics; it is invoked from prior analysis rather than established as a new theorem in this paper.

  5. Knowl 5 — Conjectured Jacobian simplification from multiplicative SGD noise

    model/method

    The paper conjectures that the multiplicative noise in its loss-stabilization SDE favors smaller ℓ2\ell_2 norms of the columns of the Jacobian feature matrix ϕθ(X)\phi_\theta(X). The fitting drift prevents every feature from collapsing, but features not needed to fit the data can shrink. In diagonal linear networks, the sample-gradient structure is determined by v⊙xv\odot x and u⊙xu\odot x, so this pressure can promote zero coordinates in the effective predictor u⊙vu\odot v. In a one-hidden-layer ReLU network ha,W(x)=⟨a,σ(Wx)⟩h_{a,W}(x)=\langle a,\sigma(Wx)\rangle, where aa contains output weights, WW contains hidden-unit weights, and σ\sigma is ReLU, a hidden unit's Jacobian contribution is reduced when it is active on fewer training examples. The paper further argues that fitting can align neurons with similar activations, reducing the effective rank of the Jacobian. These are proposed mechanisms; the paper does not prove the general Jacobian-minimization conjecture.

  6. Knowl 6 — Operational measures of feature sparsity and Jacobian rank

    definition

    For a layer's activation vectors evaluated on the training set, the paper defines the feature sparsity coefficient as the average fraction of distinct, nonzero activations. Activations with Pearson correlation coefficient at least 0.950.95 are treated as one feature, so aligned neurons that implement similar activations are not counted repeatedly. A value of 100%100\% denotes dense features and 0%0\% denotes all-zero features. As a complementary measure, the paper tracks the rank of the Jacobian feature matrix using a fixed threshold on singular values normalized by the largest singular value. For rank measurements, the Jacobian is evaluated on as many fresh samples as there are model parameters, avoiding an artificially low rank caused simply by having fewer samples than parameters. The feature sparsity coefficient is also used for deep networks where full Jacobian-rank computation is too costly.

  7. Knowl 7 — SDE discretization reproduces sparsification patterns absent from gradient flow

    empirical result

    The proposed SDE was tested on the paper's diagonal-linear and ReLU models, excluding deep networks because repeatedly computing their Jacobians was too costly. For an SGD step size ηt\eta_t, the authors used SDE discretization step γt=ηt/10\gamma_t=\eta_t/10, ran the discretization for ten times as many steps as the corresponding SGD run, and set the noise level to δt=cL(θ⌊t/10⌋SGD)\delta_t=cL(\theta^{\mathrm{SGD}}_{\lfloor t/10\rfloor}), with the constant cc chosen separately to match each run. The resulting SDE trajectories qualitatively reproduced the corresponding SGD trajectories, including comparable decreases in Jacobian rank and feature sparsity; exact curve matching was not expected because of stochastic variability. A discretized gradient-flow control, with the noise term removed, showed no corresponding rank minimization or feature sparsification. These comparisons support a role for noise in the observed representation simplification.

  8. Knowl 8 — Diagonal-network experiments distinguish SGD sparsification from loss stabilization alone

    empirical result

    The diagonal-network experiment used n=80n=80 inputs drawn from N(0,Id)\mathcal{N}(0,I_d) with d=200d=200, and targets generated by a predictor β∗∈R200\beta^*\in\mathbb{R}^{200} having 2020 nonzero coordinates. The network parameters were initialized with ui=0.1u_i=0.1 and vi=0v_i=0. A small-step run with η=0.25\eta=0.25 was compared with η=0.28\eta=0.28 runs decayed after 10%10\%, 30%30\%, or 50%50\% of the iterations; no explicit regularizer was used. In the large-step runs, training loss stabilized near 10−1.510^{-1.5}, while longer stabilization schedules produced better test loss and decreases in both rank⁡(ϕθ(X))\operatorname{rank}(\phi_\theta(X)) and the number of nonzero entries of u⊙vu\odot v. A full-batch gradient-descent control could also exhibit loss stabilization, but showed much weaker sparsification and no corresponding test-loss improvement. The comparison indicates that stabilization alone is insufficient to account for the SGD effect.

  9. Knowl 9 — ReLU experiments show that longer stabilization simplifies features before fitting

    empirical result

    The paper tested ReLU networks in two settings without weight decay. In one-dimensional regression with 12 data points, a two-layer network with 100 hidden units required a long linear step-size warmup to stabilize its loss; runs transitioned from warmup to decay at 2%2\% or 50%50\% of iterations. The training loss stabilized near 10−0.510^{-0.5}, and the longer stabilization run developed a simpler predictor with fewer distinct ReLU kinks and substantially lower Jacobian rank and feature sparsity coefficient. In a teacher–student experiment, a three-layer teacher had two hidden units per layer, while a student had ten per layer and trained on 50 examples. Large-step runs with warmup and decay at 10%10\%, 30%30\%, or 50%50\% showed stabilization near 10−1.510^{-1.5} and reductions in Jacobian rank and feature sparsity on both hidden layers. All schedules reached training loss around 10−310^{-3} after 10410^4 iterations, yet longer stabilization yielded lower test loss. Together these experiments exhibit a progression in which representation simplification during the plateau precedes the final data fitting enabled by step-size decay.

  10. Knowl 10 — Large-step schedules yield sparse features in deep convolutional networks

    empirical result

    The deep-network experiments trained DenseNet-100-12 on CIFAR-10, CIFAR-100, and Tiny ImageNet with batch size 256. An exponentially increasing warmup schedule with exponent 1.051.05 was used to reach loss stabilization; the authors measured feature sparsity at the ends of DenseNet super-blocks 3 and 4 because computing the full Jacobian rank was infeasible. They compared a basic setting without momentum or data augmentation and a practice-oriented setting with both, and used no weight decay. Across datasets and settings, larger-step schedules showed approximately stabilized training loss, progressively lower test error for longer stabilization, and decreasing feature sparsity at both measured blocks before step-size decay. In the basic CIFAR-100 experiment, the reported test error was approximately 60%60\% for a small-step schedule versus 35%35\% for a large-step schedule; at block 4, feature sparsity was typically about 50%50\%–60%60\% for small steps versus 10%10\%–20%20\% for larger steps. The authors caution that the concurrent decrease in test error and feature sparsity does not establish that sparsity causes the generalization improvement.

Coverage note — The paper's informal discussion of weight decay, normalization, SAM, and batch size as possible related routes to stronger implicit regularization is omitted because these are speculative extensions rather than separately tested contributions; additional ResNet plots are omitted as they repeat the reported deep-network pattern.

References

  1. 1.Agarwal, N., Goel, S., and Zhang, C. Acceleration via fractal learning rate schedules. In International Conference on Machine Learning, pp. 87–99. PMLR, 2021.
  2. 2.Beugnot, G., Mairal, J., and Rudi, A. On the benefits of large learning rates for kernel methods. arXiv preprint arXiv:2202.13733, 2022.
  3. 3.Bjorck, N., Gomes, C. P., Selman, B., and Weinberger, K. Q. Understanding batch normalization. Advances in neural information processing systems, 31, 2018.
  4. 4.Blanc, G., Gupta, N., Valiant, G., and Valiant, P. Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process. In Conference on learning theory, pp. 483–513. PMLR, 2020.
  5. 5.Boursier, E., Pillaud-Vivien, L., and Flammarion, N. Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs. arXiv preprint arXiv:2206.00939, 2022.
  6. 6.Chen, L. and Bruna, J. On gradient descent convergence beyond the edge of stability. arXiv preprint arXiv:2206.04172, 2022.
  7. 7.Chizat, L. and Bach, F. Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss. In Conference on Learning Theory, pp. 1305–1338. PMLR, 2020.
  8. 8.Chizat, L., Oyallon, E., and Bach, F. On lazy training in differentiable programming. Advances in Neural Information Processing Systems, 32, 2019.
  9. 9.Cohen, J. M., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A. Gradient descent on neural networks typically occurs at the edge of stability. arXiv preprint arXiv:2103.00065, 2021.
  10. 10.Damian, A., Ma, T., and Lee, J. D. Label noise sgd provably prefers flat global minimizers. Advances in Neural Information Processing Systems, 34:27449–27461, 2021.
  11. 11.Denton, E. L., Zaremba, W., Bruna, J., LeCun, Y., and Fergus, R. Exploiting linear structure within convolutional networks for efficient evaluation. Advances in neural information processing systems, 27, 2014.
  12. 12.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  13. 13.Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. Sharp minima can generalize for deep nets. In International Conference on Machine Learning, pp. 1019–1028. PMLR, 2017.
  14. 14.Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. Sharpness-aware minimization for efficiently improving generalization. In International Conference on Learning Representations, 2021.
  15. 15.Frankle, J. and Carbin, M. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635, 2018.
  16. 16.Geiping, J., Goldblum, M., Pope, P. E., Moeller, M., and Goldstein, T. Stochastic training is not necessary for generalization. In International Conference on Learning Representations, 2022.
  17. 17.HaoChen, J. Z., Wei, C., Lee, J. D., and Ma, T. Shape matters: Understanding the implicit bias of the noise covariance. In Conference on Learning Theory, 2021.
  18. 18.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  19. 19.Hinton, G., Vinyals, O., Dean, J., et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  20. 20.Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. JMLR, 22(241):1–124, 2021.
  21. 21.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. Advances in neural information processing systems, 31, 2018.
  22. 22.Jacot, A., Ged, F., S¸ims¸ek, B., Hongler, C., and Gabriel, F. Saddle-to-saddle dynamics in deep linear networks: Small initialization training, symmetry, and sparsity. In International Conference on Machine Learning, 2021.
  23. 23.Jastrzebski, S., Arpit, D., Astrand, O., Kerg, G. B., Wang, H., Xiong, C., Socher, R., Cho, K., and Geras, K. J. Catastrophic fisher explosion: Early phase fisher matrix impacts generalization. In Proceedings of the 38th International Conference on Machine Learning. PMLR, 2021.
  24. 24.Kaur, S., Cohen, J., and Lipton, Z. C. On the maximum Hessian eigenvalue and generalization. arXiv preprint arXiv:2206.10654, 2022.
  25. 25.Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint arXiv:1609.04836, 2016.
  26. 26.Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G. The large learning rate phase of deep learning: the catapult mechanism. arXiv preprint arXiv:2003.02218, 2020.
  27. 27.Li, Q., Tai, C., and Weinan, E. Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations. The Journal of Machine Learning Research, 20(1):1474–1520, 2019a.
  28. 28.Li, Y., Ma, T., and Zhang, H. Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations. In Conference On Learning Theory, pp. 2–47. PMLR, 2018.
  29. 29.Li, Y., Wei, C., and Ma, T. Towards explaining the regularization effect of initial large learning rate in training neural networks. In NeurIPS, 2019b.
  30. 30.Li, Z. and Arora, S. An exponential learning rate schedule for deep learning. arXiv preprint arXiv:1910.07454, 2019.
  31. 31.Li, Z., Lyu, K., and Arora, S. Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate. In Advances in Neural Information Processing Systems, volume 33, pp. 14544–14555, 2020.
  32. 32.Li, Z., Wang, T., and Arora, S. What happens after sgd reaches zero loss?–a mathematical framework. In International Conference on Learning Representations, 2022.
  33. 33.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  34. 34.Lyu, K. and Li, J. Gradient descent maximizes the margin of homogeneous neural networks. In International Conference on Learning Representations, 2020.
  35. 35.Ma, C. and Ying, L. The Sobolev regularization effect of stochastic gradient descent. arXiv preprint arXiv:2105.13462, 2021.
  36. 36.Ma, C., Wu, L., and Ying, L. The multiscale structure of neural network loss functions: The effect on optimization and origin. arXiv preprint arXiv:2204.11326, 2022.
  37. 37.Moroshko, E., Woodworth, B. E., Gunasekar, S., Lee, J. D., Srebro, N., and Soudry, D. Implicit bias in deep linear classification: Initialization scale vs training accuracy. In Advances in Neural Information Processing Systems, volume 33, pp. 22182–22193, 2020.
  38. 38.Mulayoff, R., Michaeli, T., and Soudry, D. The implicit bias of minima stability: A view from function space. Advances in Neural Information Processing Systems, 34:17749–17761, 2021.
  39. 39.Nacson, M. S., Ravichandran, K., Srebro, N., and Soudry, D. Implicit bias of the step size in linear diagonal neural networks. In International Conference on Machine Learning, pp. 16270–16295. PMLR, 2022.
  40. 40.Nakkiran, P. Learning rate annealing can provably help generalization, even for convex problems. arXiv preprint arXiv:2005.07360, 2020.
  41. 41.Oksendal, B. Stochastic differential equations: an introduction with applications. Springer Science & Business Media, 2013.
  42. 42.Pesme, S., Pillaud-Vivien, L., and Flammarion, N. Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity. In Advances in Neural Information Processing Systems, 2021.
  43. 43.Pillaud-Vivien, L., Reygner, J., and Flammarion, N. Label noise (stochastic) gradient descent implicitly solves the lasso for quadratic parametrisation. In Conference on Learning Theory, 2022.
  44. 44.Smith, S. L. and Le, Q. V. A Bayesian perspective on generalization and stochastic gradient descent. In International Conference on Learning Representations, 2018.
  45. 45.Smith, S. L., Dherin, B., Barrett, D. G., and De, S. On the origin of implicit regularization in stochastic gradient descent. arXiv preprint arXiv:2101.12176, 2021.
  46. 46.Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. The implicit bias of gradient descent on separable data. The Journal of Machine Learning Research, 19(1):2822–2878, 2018.
  47. 47.Vaskevicius, T., Kanade, V., and Rebeschini, P. Implicit regularization for optimal sparse recovery. Advances in Neural Information Processing Systems, 32, 2019.
  48. 48.Wang, Y., Chen, M., Zhao, T., and Tao, M. Large learning rate tames homogeneity: Convergence and balancing effect. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=3tbDrs77LJ5.
  49. 49.Wojtowytsch, S. Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis. arXiv preprint arXiv:2106.02588, 2021a.
  50. 50.Wojtowytsch, S. Stochastic gradient descent with noise of machine learning type. Part I: Discrete time analysis. arXiv preprint arXiv:2105.01650, 2021b.
  51. 51.Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N. Kernel and rich regimes in overparametrized models. In Proceedings of Thirty Third Conference on Learning Theory, volume 125 of Proceedings of Machine Learning Research, pp. 3635–3673. PMLR, 09–12 Jul 2020.
  52. 52.Wu, J., Zou, D., Braverman, V., and Gu, Q. Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate. International Conference on Learning Representations, 2021.
  53. 53.Wu, L., Ma, C., et al. How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective. Advances in Neural Information Processing Systems, 31, 2018.
  54. 54.Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y. A walk with sgd. arXiv preprint arXiv:1802.08770, 2018.
  55. 55.Yang, N., Tang, C., and Tu, Y. Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions. arXiv preprint arXiv:2206.01246, 2022.
  56. 56.Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. Understanding deep learning requires rethinking generalization. In International Conference on Learning Representations, 2017.
  57. 57.Zhang, G., Wang, C., Xu, B., and Grosse, R. Three mechanisms of weight decay regularization. In International Conference on Learning Representations, 2018.
  58. 58.Ziyin, L., Liu, K., Mori, T., and Ueda, M. Strength of minibatch noise in sgd. In International Conference on Learning Representations, 2022.

Citation

MLA
Andriushchenko, M., et al. “SGD with Large Step Sizes Learns Sparse Features”. International Conference on Machine Learning, vol. 202, 2023, pp. 903–25, https://proceedings.mlr.press/v202/andriushchenko23b.html.
APA
Andriushchenko, M., Varre, A. V., Pillaud-Vivien, L., & Flammarion, N. (2023). SGD with Large Step Sizes Learns Sparse Features. International Conference on Machine Learning, 202, 903–925. https://proceedings.mlr.press/v202/andriushchenko23b.html
Chicago
Andriushchenko, M., A. V. Varre, L. Pillaud-Vivien, and N. Flammarion. 2023. “SGD with Large Step Sizes Learns Sparse Features”. International Conference on Machine Learning 202: 903–25. https://proceedings.mlr.press/v202/andriushchenko23b.html.
Harvard
Andriushchenko, M. et al. (2023) “SGD with Large Step Sizes Learns Sparse Features”, International Conference on Machine Learning. PMLR, pp. 903–925. Available at: https://proceedings.mlr.press/v202/andriushchenko23b.html.
Vancouver
1. Andriushchenko M, Varre AV, Pillaud-Vivien L, Flammarion N (2023) SGD with Large Step Sizes Learns Sparse Features. In: International Conference on Machine Learning. PMLR, pp 903–925

BibTeX

@InProceedings{pmlr-v202-andriushchenko23b,
  title = 	 {{SGD} with Large Step Sizes Learns Sparse Features},
  author =       {Andriushchenko, Maksym and Varre, Aditya Vardhan and Pillaud-Vivien, Loucas and Flammarion, Nicolas},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {903--925},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/andriushchenko23b/andriushchenko23b.pdf},
  url = 	 {https://proceedings.mlr.press/v202/andriushchenko23b.html},
  abstract = 	 {We showcase important features of the dynamics of the Stochastic Gradient Descent (SGD) in the training of neural networks. We present empirical observations that commonly used large step sizes (i) may lead the iterates to jump from one side of a valley to the other causing loss stabilization, and (ii) this stabilization induces a hidden stochastic dynamics that biases it implicitly toward simple predictors. Furthermore, we show empirically that the longer large step sizes keep SGD high in the loss landscape valleys, the better the implicit regularization can operate and find sparse representations. Notably, no explicit regularization is used: the regularization effect comes solely from the SGD dynamics influenced by the large step sizes schedule. Therefore, these observations unveil how, through the step size schedules, both gradient and noise drive together the SGD dynamics through the loss landscape of neural networks. We justify these findings theoretically through the study of simple neural network models as well as qualitative arguments inspired from stochastic processes. This analysis allows us to shed new light on some common practices and observed phenomena when training deep networks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/