Progress measures for grokking via mechanistic interpretability
Neel NandaLawrence ChanTom LieberumJess SmithJacob Steinhardt
Explains the delayed generalization of grokking by reverse-engineering the trigonometric algorithms learned by transformers on modular addition, establishing continuous progress measures that track gradual circuit formation and the removal of memorization.
Modern artificial intelligence models frequently display emergent behaviors, where sudden and unexpected capabilities or safety risks arise abruptly after extensive training or scaling. Standard performance metrics often fail to anticipate these rapid shifts, making deployment timelines unpredictable and raising safety concerns. The article aims to evaluate whether reverse-engineering the internal mechanisms of neural networks can uncover continuous, hidden progress measures that explain and anticipate these sudden transitions.
To investigate this, the researchers conducted an in-depth mechanistic interpretability study on small transformer networks trained on modular addition tasks across tens of thousands of optimization epochs. By analyzing the internal weights, activations, and Fourier representations across different random seeds, prime moduli, and data fractions, the team mapped the exact algorithms learned by the models and constructed specific tracking metrics.
First, the article finds that instead of memorizing look-up tables indefinitely, the models learn a generalizable algorithm that maps inputs onto circles using discrete Fourier transforms and combines them via trigonometric identities. Second, training divides into three continuous phases: initial memorization of the training set, internal circuit formation where the generalizing mechanism builds up silently, and a cleanup phase where regularization eliminates memorization weights. Third, the apparent sudden jump in test performance (known as grokking) occurs entirely during the cleanup phase, well after the generalizing circuit has formed. Finally, both limited training data and explicit regularization—such as weight decay or dropout—are strictly necessary to drive this transition, whereas non-regularized models fail to generalize.
These findings demonstrate that sudden performance leaps are not random or instantaneous events, but rather the visible result of smooth, underlying structural changes within model weights. This insight offers a credible path toward anticipating capability jumps and safety failures before they appear on standard evaluation benchmarks, potentially lowering operational risks and preventing costly training overruns.
Organizations developing complex artificial intelligence systems should invest in automated mechanistic interpretability tools and track internal representation metrics rather than relying solely on high-level accuracy curves. For near-term applications, decision-makers should support pilot projects that extend internal progress tracking to larger models and multi-step tasks, while establishing quantitative thresholds to predict when transitions will occur.
Readers should exercise caution when extrapolating these findings directly to frontier-scale systems. The experiments rely primarily on small models trained on idealized mathematical and algorithmic tasks, and the manual reverse-engineering methods used here do not yet scale automatically to massive, complex models. Nevertheless, confidence in the demonstrated mechanism for algorithmic tasks remains very high across the studied settings.
- Paper: Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small, Kevin Wang et al. (2022). Its causal circuit analysis of GPT-2 provides an earlier example of tracing model behavior through internal mechanisms, preparing you for the source’s reverse-engineering approach.
- Paper: Omnigrok: Grokking Beyond Algorithmic Data, Ziming Liu et al. (2023). It carries the grokking investigation beyond algorithmic tasks, testing whether the delayed-generalization mechanisms identified in the source also explain learning on real-world data.
- Paper: A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations, Bilal Chughtai et al. (2023). It extends mechanistic study of algorithm learning across groups and model settings, using progress measures and circuit analysis to probe how learned solutions vary.
- Paper: Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory, Hari K. Prakash et al. (2026). It continues the weight-level study of grokking by testing whether random-matrix signatures can detect overfitting across training phases without access to the data.
