Omnigrok: Grokking Beyond Algorithmic Data

Ziming LiuEric J. MichaudMax Tegmark

article2023ICLR128 citations

Demonstrates how delayed generalization in neural networks arises from training and test loss mismatches, providing actionable methods to intentionally induce or eliminate grokking across image, language, and molecular domains.

Listen

The phenomenon known as grokking—where a machine learning model achieves strong test performance long after it has completely memorized and overfit its training data—has raised fundamental questions about how neural networks learn. Understanding this behavior is critical for artificial intelligence development, as unexpected delays in generalization can cause practitioners to stop training prematurely, waste computational resources, or misunderstand model reliability.

The article aims to explain the underlying mechanism of grokking by analyzing neural network loss landscapes. It evaluates whether this delayed learning is uniquely tied to algorithmic math problems or represents a general property of neural networks that can be reproduced and controlled across diverse domains.

To investigate this, the researchers evaluated simplified models alongside complex real-world datasets, including MNIST handwritten digit classification, IMDb movie review sentiment analysis, and QM9 molecular property prediction. They used multi-layer networks, long short-term memory models, graph convolutional neural networks, and single-layer transformer architectures. The study examined model behavior across variations in initialization weight scale, dataset size, and weight decay regularization by projecting training and test performance onto lower-dimensional visualizations termed reduced loss landscapes.

The article established four primary findings. First, grokking is driven by the mismatch between training and test loss landscapes as a function of total model parameter magnitude, termed the LU mechanism: training loss forms an L-shape that flattens at zero, while test loss forms a U-shape with an optimal generalizing zone. Second, when models are initialized with large parameter scales, they quickly slide into non-generalizing overfitting regions; subsequent movement toward generalizing regions requires regularization, causing generalization time to scale inversely with the strength of weight decay across multiple orders of magnitude. Third, grokking is not unique to algorithmic data; the researchers successfully induced delayed generalization on images, text, and molecular data by artificially inflating initialization scales and reducing dataset sizes. Fourth, grokking is particularly dramatic in algorithmic tasks because success requires learning representations from scratch, which causes parameter norms to surge during initial memorization before shrinking upon generalization.

These findings provide practical implications for model training and resource allocation. They demonstrate that apparent memorization does not necessarily indicate a failed model, but allowing unconstrained parameter growth introduces severe training delays and computational costs. Crucially, the authors demonstrated that grokking can be nearly eliminated on algorithmic transformer tasks by constraining the model's parameter magnitude during training, aligning training and test accuracy improvements concurrently.

Engineering teams training neural networks should carefully configure parameter initialization scales and apply appropriate regularization to prevent models from settling into slow overfitting regimes. Where training delays occur on representation-heavy tasks, practitioners should consider constraining optimization to fixed parameter magnitudes rather than relying solely on extended training runs. Further work is recommended to evaluate how these landscape dynamics interact with large-scale pre-trained foundation models, where complex language representations are already established rather than learned from scratch.

Confidence in these findings is strong across the evaluated synthetic, algorithmic, and benchmark classification tasks. However, users should exercise caution when extrapolating directly to massive industrial models, as the article's empirical validations relied on small-to-medium network architectures, reduced datasets, and specific architectural assumptions.

arXiv: 2210.01117
  • Paper: Understanding deep learning requires rethinking generalization, Chiyuan Zhang et al. (2017). It provides the foundational empirical demonstration that overparameterized neural networks can easily achieve zero training error by memorizing random noise or data before generalizing.
  • Paper: Visualizing the Loss Landscape of Neural Nets, Hao Li et al. (2017). It introduces the normalized loss landscape visualization methods that the source directly adapts to project and analyze the LU mechanism of delayed generalization.
  • Paper: Decoupled Weight Decay Regularization, Ilya Loshchilov et al. (2019). It establishes the mechanics of decoupled weight decay regularization, which the source relies on to control parameter norm growth and accelerate transitions from overfitting to generalizing regimes.
  • Paper: A Closer Look at Memorization in Deep Networks, Devansh Arpit et al. (2017). It formalizes the separation between pattern learning and brute-force memorization dynamics during neural network optimization, providing necessary background for analyzing grokking.
  • Paper: Exploring Generalization in Deep Learning, Behnam Neyshabur et al. (2017). It evaluates how norm-based complexity measures and landscape geometry govern generalization in overparameterized networks, which underpins the source's LU parameter-magnitude framework.
  • Paper: On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, Nitish Shirish Keskar et al. (2016). It characterizes the distinction between sharp and flat minimizers in neural loss landscapes, providing essential context for understanding how regularization shifts optimization into generalizing basins.
Cover for Omnigrok: Grokking Beyond Algorithmic Data

Abstract

Grokking, the unusual phenomenon for algorithmic datasets where generalization happens long after overfitting the training data, has remained elusive. We aim to understand grokking by analyzing the loss landscapes of neural networks, identifying the mismatch between training and test losses as the cause for grokking. We refer to this as the "LU mechanism" because training and test losses (against model weight norm) typically resemble "L" and "U", respectively. This simple mechanism can nicely explain many aspects of grokking: data size dependence, weight decay dependence, the emergence of representations, etc. Guided by the intuitive picture, we are able to induce grokking on tasks involving images, language and molecules. In the reverse direction, we are able to eliminate grokking for algorithmic datasets. We attribute the dramatic nature of grokking for algorithmic datasets to representation learning.

Table of Contents

  • 1 Introduction
  • 2 The LU mechanism for grokking
  • 3 Grokking for a teacher-student setup
  • 4 Omnigrok: Grokking for more interesting tasks
  • 5 Representation is key to grokking
  • 5.1 Algorithmic datasets
  • 5.2 MNIST
  • 6 Relation to Related works
  • 7 Conclusions
  • References
  • A Experiment details
  • B Reduced loss for modular addition with transformers
  • C time to generalize versus weight decay
  • D Section setup
  • E MNIST experiments with cross entropy loss

Knowls

  1. Knowl 1 — The LU Mechanism for Grokking

    theoretical result

    Grokking—the phenomenon where generalization occurs long after a neural network overfits the training data—originates from a structural mismatch between the reduced training loss and reduced test loss plotted against the parameter L2L_2 weight norm w=∥w∥2w = \|\mathbf{w}\|_2.

    • The reduced training loss l~train(w)\tilde{l}_{\text{train}}(w) exhibits an "L-shape": for weight norms below a critical threshold wcw_c, training loss is high; for w≥wcw \ge w_c, training loss rapidly drops to near zero and remains flat, forming an extensive valley of overfitting solutions.
    • The reduced test loss l~test(w)\tilde{l}_{\text{test}}(w) exhibits a "U-shape": test loss is minimal in a finite spherical shell ("Goldilocks zone") around w≈wcw \approx w_c, but rises for both w<wcw < w_c (underfitting) and w>wcw > w_c (overfitting).

    When a network is initialized with a large weight norm w0>wcw_0 > w_c, optimization rapidly fits the training set by moving angularly to an overfitting solution in the w>wcw > w_c region, where training loss gradients vanish. Regularization (such as weight decay with parameter γ\gamma) then slowly decreases the weight norm w(t)w(t) radially toward wcw_c. Because the radial contraction velocity is proportional to γ\gamma, a small regularization magnitude causes a substantial time delay between training loss convergence and generalization. Conversely, small initializations (w0<wcw_0 < w_c) enter the Goldilocks zone directly, generalizing rapidly without grokking.

  2. Knowl 2 — Reduced 1D Loss Landscape

    definition

    For a neural network parameterized by weight vector w∈RD\mathbf{w} \in \mathbb{R}^D with L2L_2 norm w≡∥w∥2w \equiv \|\mathbf{w}\|_2 and angular direction w^≡w/w\hat{\mathbf{w}} \equiv \mathbf{w}/w, the reduced function f~(w)\tilde{f}(w) corresponding to any performance metric f(w)f(\mathbf{w}) (such as training loss, test loss, or classification error) is defined by evaluating ff at the angular direction that minimizes training loss on the sphere of radius ww:

    f~(w)≡f(w∗(w)),wherew∗(w)≡argmin⁡∥w∥2=wltrain(w)\tilde{f}(w) \equiv f(\mathbf{w}^*(w)), \quad \text{where} \quad \mathbf{w}^*(w) \equiv \operatorname{argmin}_{\|\mathbf{w}\|_2 = w} l_{\text{train}}(\mathbf{w})

    Computationally, w∗(w)\mathbf{w}^*(w) is obtained by training the network with standard gradient optimization while constraining the parameter norm to remain exactly ww via sphere projection (rescaling w←w⋅w/∥w∥2\mathbf{w} \leftarrow w \cdot \mathbf{w}/\|\mathbf{w}\|_2 after each unconstrained gradient step) until convergence. Initialization scales are parameterized relative to standard default initialization w0\mathbf{w}_0 by a constant scaling factor α≡w/∥w0∥2\alpha \equiv w/\|\mathbf{w}_0\|_2.

  3. Knowl 3 — Generalization Time Scaling with Weight Decay

    equation

    In the regime where the initial weight norm w0w_0 exceeds the critical generalizing radius wcw_c (w0>wcw_0 > w_c) and the training loss has converged to near zero, the weight norm dynamics under gradient descent with explicit L2L_2 weight decay regularization 12γ∥w∥22\frac{1}{2}\gamma \|\mathbf{w}\|_2^2 are governed by:

    dwdt≈−γw  ⟹  w(t)≈w0e−γt\frac{dw}{dt} \approx -\gamma w \implies w(t) \approx w_0 e^{-\gamma t}

    The number of optimization steps tgent_{\text{gen}} required to contract the model weight vector from w0w_0 to the generalizing Goldilocks radius wcw_c scales inversely with the weight decay parameter γ\gamma:

    tgen≈ln⁡(w0/wc)γ∝γ−1t_{\text{gen}} \approx \frac{\ln(w_0 / w_c)}{\gamma} \propto \gamma^{-1}

    Consequently, the number of steps required to reach training set overfitting is independent of γ\gamma, while the delay between overfitting and test generalization is proportional to γ−1\gamma^{-1}.

  4. Knowl 4 — Recipe for Inducing Grokking on Non-Algorithmic Datasets

    model/method

    Standard machine learning tasks (including vision, natural language, and molecular property prediction) do not typically display grokking under default training setups because standard initializations set w≤wcw \le w_c and datasets are sufficiently large that the Goldilocks zone is wide. Grokking can be induced in standard architectures by applying two modifications guided by the LU mechanism:

    1. Subsample the training dataset size NN: Reducing NN narrows the Goldilocks zone (the range of weight norms where test error is low), which steepens the U-shaped test error landscape relative to the flat, overfitting basin of the L-shaped training error landscape.
    2. Increase the weight initialization scale α\alpha: Multiplying initial weights (e.g., standard Kaiming uniform initialization) by a factor α=w/w0>1\alpha = w/w_0 > 1 places the starting parameter vector in the w>wcw > w_c overfitting basin, forcing the optimizer to first reach an overfitting state before slow radial decay brings the weights into the generalizing regime.
  5. Knowl 5 — Grokking Dynamics on the 2D Representation-Weight Landscape

    theoretical result

    When analyzing grokking over the 2D plane of decoder weight norm ww and representation messiness mm, the loss landscape is partitioned into:

    • Region I: Small ww, high training loss, and steep training loss gradients.
    • Region II: Large ww, training loss l~train≈0\tilde{l}_{\text{train}} \approx 0, and vanishing training loss gradients.

    Under gradient flow with decoder learning rate ηD\eta_D, representation learning rate ηR=ηD\eta_R = \eta_D, and weight decay γ\gamma:

    dwdt=−ηD(∂l~train∂w+γw),dmdt=−ηR∂l~train∂m\frac{dw}{dt} = -\eta_D \left(\frac{\partial \tilde{l}_{\text{train}}}{\partial w} + \gamma w\right), \quad \frac{dm}{dt} = -\eta_R \frac{\partial \tilde{l}_{\text{train}}}{\partial m}

    Starting from a large norm and a random representation (m≈1m \approx 1) in Region II, dmdt≈0\frac{dm}{dt} \approx 0, causing the parameters to drift purely in the −w-w direction with constant velocity v=ηDγv = \eta_D \gamma over distance L1=L−hcot⁡θL_1 = L - h \cot\theta to the boundary of Region I. At the boundary, training gradients balance the radial drift, forcing parameters to slide along the contour with velocity v′=vcos⁡θv' = v \cos\theta over distance L2=h/sin⁡θL_2 = h / \sin\theta (where θ\theta is the contour angle) toward the generalizing solution (wc,m=0)(w_c, m=0). The total grokking time is:

    t=L1v+L2v′=L+htan⁡θηDγt = \frac{L_1}{v} + \frac{L_2}{v'} = \frac{L + h \tan\theta}{\eta_D \gamma}

    Larger training datasets decrease the boundary angle θ\theta, thereby shortening tt and accelerating grokking.

  6. Knowl 6 — Representation Quality Discrepancy as the Origin of Algorithmic Grokking

    theoretical result

    The pronounced occurrence of grokking on algorithmic tasks compared to natural vision datasets (such as MNIST) under default initializations is caused by how critical weight norm wcw_c depends on representation quality:

    • Algorithmic Datasets: In modular arithmetic tasks, classifying with an unstructured/random representation requires learning O(p2)O(p^2) sample-target mappings, whereas a structured linear representation requires learning only O(p)O(p) mappings. This creates a large gap in required decoder capacity:

    wc(bad representation)>wc(good representation)w_c(\text{bad representation}) > w_c(\text{good representation})

    Standard initializations are scaled to allow parameter fitting under the initial poor representation, making w0w_0 effectively too large (w0>wc(good)w_0 > w_c(\text{good})). Overfitting causes weight norm to inflate, and generalization cannot occur until weight decay drives representation refinement and subsequent weight norm contraction.

    • Natural Datasets (e.g., MNIST): Raw input features Rraw\mathbf{R}_{\text{raw}} are already close to linearly separable. Consequently, training and test landscapes depend very weakly on representation messiness, wcw_c is essentially constant across representation qualities, and grokking does not emerge under standard initializations.
  7. Knowl 7 — Empirical Demonstration of Grokking Across Modalities

    empirical result

    Grokking was systematically induced across three standard domains using the reduced dataset and elevated initialization scale protocol:

    • Image Classification (MNIST): A depth-3, width-200 ReLU MLP trained with AdamW on N=1,000N = 1{,}000 samples with α∈[2,10]\alpha \in [2, 10] and weight decay γ=10−4\gamma = 10^{-4} (or α=100\alpha = 100 with cross-entropy loss) fits the training set within 10210^2 steps, while test accuracy plateaus and does not surge toward completion until 10410^4 to 10510^5 steps.
    • Sentiment Analysis (IMDb): A 2-layer LSTM (embedding dimension 64, hidden dimension 128) trained with Adam on N=1,000N = 1{,}000 reviews with α=6\alpha = 6 and weight decay γ=1\gamma = 1 reaches full training fit within 10210^2 steps, whereas generalization is delayed until approximately 10310^3 steps. Standard initialization (α=1,γ=0\alpha = 1, \gamma = 0) generalizes immediately.
    • Molecular Property Prediction (QM9): A 2-layer Graph Convolutional Neural Network (GCNN) predicting isotropic polarizability trained with Adam on N=200N = 200 molecules with α=3\alpha = 3 and γ=0\gamma = 0 displays a sharp test MSE decrease at 10410^4 steps (driven by implicit optimizer regularization), whereas α=1\alpha = 1 exhibits monotonic convergence without grokking.
  8. Knowl 8 — Grokking Elimination via Constrained Weight Norm Optimization

    empirical result

    In algorithmic tasks where grokking is standard (e.g., a 1-layer Transformer with dmodel=128d_{\text{model}} = 128, 4 attention heads, dmlp=512d_{\text{mlp}} = 512 trained on modular addition with p=113p = 113 and AdamW, γ=1\gamma = 1), grokking can be eliminated by:

    1. Initializing the network at a small weight scale (e.g., α=0.8\alpha = 0.8).
    2. Enforcing a hard weight norm constraint by projecting model parameters back to the fixed sphere ∥w∥2=0.8⋅∥w0∥2\|\mathbf{w}\|_2 = 0.8 \cdot \|\mathbf{w}_0\|_2 after each gradient step.

    Under this constrained norm optimization, train accuracy and test accuracy rise simultaneously without an intermediate plateau of overfitted training accuracy, proving that unconstrained parameter norm growth into the overparameterized w>wcw > w_c regime is necessary for delayed generalization.

  9. Knowl 9 — 2D Reduced Landscape Parameterization with Representation Messiness

    definition

    For an algorithmic task with input symbols i,j∈{0,…,p−1}i, j \in \{0, \dots, p-1\} and output labels k=i+jk = i+j, inputs are mapped to embeddings Ei\mathbf{E}_i and decoded by an MLP Decw(Ei+Ej)\text{Dec}_{\mathbf{w}}(\mathbf{E}_i + \mathbf{E}_j). The full representation matrix R={E0,…,Ep−1}\mathbf{R} = \{\mathbf{E}_0, \dots, \mathbf{E}_{p-1}\} is parameterized by a scalar m∈[0,1]m \in [0, 1] via linear interpolation:

    R(m)=mRrandom+(1−m)Rlinear\mathbf{R}(m) = m \mathbf{R}_{\text{random}} + (1 - m) \mathbf{R}_{\text{linear}}

    where Rlinear\mathbf{R}_{\text{linear}} embeds integer kk as [k,0,…,0]T[k, 0, \dots, 0]^T, Rrandom\mathbf{R}_{\text{random}} draws each embedding independently from N(0,I)\mathcal{N}(\mathbf{0}, \mathbf{I}), and mm represents representation messiness (m=0m=0 is fully structured; m=1m=1 is random Gaussian noise).

    The 2D reduced training and test losses are computed by minimizing training loss over weight directions w^\hat{\mathbf{w}} on constant-norm spheres at fixed (w,m)(w, m):

    l~train(w,m)≡ltrain(w∗(w,m),m),l~test(w,m)≡ltest(w∗(w,m),m)\tilde{l}_{\text{train}}(w, m) \equiv l_{\text{train}}(\mathbf{w}^*(w, m), m), \quad \tilde{l}_{\text{test}}(w, m) \equiv l_{\text{test}}(\mathbf{w}^*(w, m), m)

    w∗(w,m)≡argmin⁡∥w∥2=wltrain(w,m)\mathbf{w}^*(w, m) \equiv \operatorname{argmin}_{\|\mathbf{w}\|_2 = w} l_{\text{train}}(\mathbf{w}, m)

  10. Knowl 10 — Trainability Configurations for Loss Landscape and Trajectory Analysis

    data/table

    To decouple parameter dynamics and isolate the role of weight norm versus representation evolution in grokking, three distinct parameter training regimes are defined:

    Setup Decoder Norm w=∥w∥2w = \|\mathbf{w}\|_2 Decoder Direction w^=w/w\hat{\mathbf{w}} = \mathbf{w}/w Messiness mm Other Representation Params
    Landscape analysis Fixed (ww) Trainable Fixed (mm) Fixed (00)
    Reduced trajectory Trainable Fixed to w^∗(w,m)\hat{\mathbf{w}}^*(w, m) Trainable Fixed (00)
    Full trajectory Trainable Trainable Trainable Trainable
    • Landscape analysis: Freezes weight norm ww and messiness mm on a 2D grid, training only the angular direction w^\hat{\mathbf{w}} to minimize training loss with Adam (10410^4 steps, learning rate 10−310^{-3}) to compute static 2D loss surfaces l~(w,m)\tilde{l}(w, m).
    • Reduced trajectory analysis: Tracks 2D trajectory dynamics (w(t),m(t))(w(t), m(t)) assuming instant angular equilibration w^(t)=w^∗(w(t),m(t))\hat{\mathbf{w}}(t) = \hat{\mathbf{w}}^*(w(t), m(t)) (scale separation) and 1D linear representation evolution.
    • Full trajectory analysis: Standard end-to-end gradient descent where all decoder weights and multi-dimensional representation embeddings vary simultaneously.

Coverage note — None was omitted; the top 10 knowls fully capture the paper's core contributions, including the LU mechanism, 1D and 2D reduced loss landscapes, weight decay scaling relations, cross-modal induction and elimination of grokking, and the representation-based explanation of algorithmic vs. natural tasks.

References

  1. 1.Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. Grokking: Generalization beyond overfitting on small algorithmic datasets. arXiv preprint arXiv:2201.02177, 2022.
  2. 2.Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J Michaud, Max Tegmark, and Mike Williams. Towards understanding grokking: An effective theory of representation learning. arXiv preprint arXiv:2205.10343, 2022.
  3. 3.Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind. The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon. arXiv preprint arXiv:2206.04817, 2022.
  4. 4.Boaz Barak, Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang. Hidden progress in deep learning: Sgd learns parities near the computational limit. arXiv preprint arXiv:2207.08799, 2022.
  5. 5.Stanislav Fort and Adam Scherlis. The goldilocks zone: Towards better understanding of neural network loss landscapes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 3574–3581, 2019.
  6. 6.Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: Where bigger models and more data hurt. Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003, 2021.
  7. 7.Andrew Ng and Tengyu Ma. Cs229 lecture notes. https://cs229.stanford.edu/lectures-spring2022/main_notes.pdf, page 115, 2022.
  8. 8.Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein. Deep information propagation. arXiv preprint arXiv:1611.01232, 2016.
  9. 9.Ge Yang and Samuel Schoenholz. Mean field residual networks: On the edge of chaos. Advances in neural information processing systems, 30, 2017.
  10. 10.Li Deng. The mnist database of handwritten digit images for machine learning research. IEEE Signal Processing Magazine, 29(6):141–142, 2012.
  11. 11.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8): 1735–1780, 1997.
  12. 12.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA, June 2011. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/P11-1015.
  13. 13.Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole von Lilienfeld. Quantum chemistry structures and properties of 134 kilo molecules. Scientific Data, 1, 2014.
  14. 14.Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. arXiv preprint arXiv:2301.05217, 2023.
  15. 15.Rohin Shah. Alignment Newsletter #159. https://www.alignmentforum.org/posts/zvWqPmQasssaAWkrj/an-159-building-agents-that-know-how-to-experiment-by#DEEP_LEARNING_, 2021.
  16. 16.Beren Millidge. Grokking ’grokking’. https://beren.io/2022-01-11-Grokking-Grokking/, 2022.
  17. 17.Fatih Furkan Yilmaz and Reinhard Heckel. Regularization-wise double descent: Why it occurs and how to eliminate it. In 2022 IEEE International Symposium on Information Theory (ISIT), pages 426–431. IEEE, 2022.
  18. 18.Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi. Multiple descent: Design your own generalization curve. Advances in Neural Information Processing Systems, 34:8898–8912, 2021.
  19. 19.Stéphane d’Ascoli, Levent Sagun, and Giulio Biroli. Triple descent and the two kinds of overfitting: Where & why do they appear? Advances in Neural Information Processing Systems, 33:3058–3069, 2020.
  20. 20.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 249–256. JMLR Workshop and Conference Proceedings, 2010.
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026–1034, 2015.
  22. 22.Yasaman Bahri, Jonathan Kadmon, Jeffrey Pennington, Sam S Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli. Statistical mechanics of deep learning. Annual Review of Condensed Matter Physics, 11(1), 2020.
  23. 23.Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, and Marin Soljačić. Tunable efficient unitary neural networks (eunn) and their application to rnns. In International Conference on Machine Learning, pages 1733–1741. PMLR, 2017.
  24. 24.Maciej Skorski, Alessandro Temperoni, and Martin Theobald. Revisiting initialization of neural networks. arXiv preprint arXiv:2004.09506, 2020.
  25. 25.Zhi-Qin John Xu, Yaoyu Zhang, and Yanyang Xiao. Training behavior of deep neural network in frequency domain. In International Conference on Neural Information Processing, pages 264–274. Springer, 2019.
  26. 26.Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma. A type of generalization error induced by initialization in deep neural networks. In Mathematical and Scientific Machine Learning, pages 144–164. PMLR, 2020.
  27. 27.Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse. Three mechanisms of weight decay regularization. arXiv preprint arXiv:1810.12281, 2018.
  28. 28.Twan Van Laarhoven. L2 regularization versus batch and weight normalization. arXiv preprint arXiv:1706.05350, 2017.
  29. 29.Aitor Lewkowycz and Guy Gur-Ari. On the training dynamics of deep networks with l_2 regularization. Advances in Neural Information Processing Systems, 33:4790–4799, 2020.
  30. 30.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.

Citation

MLA
Liu, Z., et al. “Omnigrok: Grokking Beyond Algorithmic Data”. arXiv, 2022, http://arxiv.org/abs/2210.01117v2.
APA
Liu, Z., Michaud, E. J., & Tegmark, M. (2022). Omnigrok: Grokking Beyond Algorithmic Data. arXiv. http://arxiv.org/abs/2210.01117v2
Chicago
Liu, Z., E. J. Michaud, and M. Tegmark. 2022. “Omnigrok: Grokking Beyond Algorithmic Data”. arXiv. http://arxiv.org/abs/2210.01117v2.
Harvard
Liu, Z., Michaud, E.J. and Tegmark, M. (2022) “Omnigrok: Grokking Beyond Algorithmic Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.01117v2.
Vancouver
1. Liu Z, Michaud EJ, Tegmark M (2022) Omnigrok: Grokking Beyond Algorithmic Data. arXiv

BibTeX

@article{liu2022omnigrok,
  title = {Omnigrok: Grokking Beyond Algorithmic Data},
  author = {Liu, Ziming and Michaud, Eric J. and Tegmark, Max},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.01117v2},
  eprint = {2210.01117}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/