A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations

Bilal ChughtaiLawrence ChanNeel Nanda

article2023ICML172 citations

Reverse-engineers how neural networks learn finite group composition using mathematical representation theory, revealing that while networks share an underlying algorithmic family, the specific circuits they develop remain arbitrary.

Listen

As deep learning systems are increasingly deployed in high-stakes settings, understanding their internal mechanisms is critical for safety, auditing, and alignment. A foundational question in mechanistic interpretability is the universality hypothesis: whether different neural networks independently converge on the same internal features and algorithmic circuits when trained on similar tasks. If strong universality holds, reverse-engineering one model yields insights that directly generalize across other architectures and scales. Conversely, if models implement arbitrary solutions, analyzing individual systems in isolation cannot reliably predict the behavior of others.

The article evaluates the universality hypothesis by examining how neural networks learn group composition across various finite groups. The primary objective is to reverse-engineer the computational mechanisms models use to perform this algebraic reasoning and determine whether these learned circuits are identical across different random initializations, groups, and architectures.

To conduct this evaluation, the authors trained one-hidden-layer multilayer perceptrons and single-layer transformer models across seven distinct mathematical groups, including cyclic, dihedral, alternating, and symmetric permutation groups. The primary benchmark involved training over 50 random initializations on the symmetric group of order five (S5). Using mathematical representation theory, the authors reverse-engineered model weights, internal activations, and output predictions, validating their findings through targeted ablation experiments and continuous training-progress measures.

The analysis produced several key findings. First, networks universally implement a specific mathematical procedure named group composition via representations: input embeddings encode representation matrices, hidden neurons multiply these matrices using activation functions, and output layers compute matrix traces (characters) to generate predictions. Second, models achieve high accuracy while utilizing only a very sparse subset of possible mathematical representations; in the primary symmetric group model, just two out of six possible representations explained 84.8% of the output logit variance. Third, the specific representations learned, the total number of representations utilized, and the chronological order in which they developed varied significantly across random initializations. For instance, single-dimensional parity representations were learned rapidly but generalized poorly, whereas higher-dimensional representations developed later, often alongside distinct delayed-generalization ("grokking") phases.

These findings provide strong evidence for weak universality but refute strong universality. While all evaluated models converge on the same overarching algorithmic family, individual networks choose arbitrary subsets of mathematical representations to execute it. Consequently, fully reverse-engineering a single neural network is insufficient for characterizing model behavior broadly. In practical settings such as safety auditing or compliance verification, relying on mechanistic insights derived from a single trained instance introduces considerable risk, as parallel instances may rely on entirely different sub-circuits.

For practitioners and researchers, the article recommends conducting systematic multi-model robustness checks across varied random initializations before drawing broad conclusions about learned mechanisms. Interpretability workflows should aim to establish comprehensive catalogs or "periodic tables" of possible circuit implementations rather than analyzing solitary models. Furthermore, future efforts should explore whether feature emergence can be anticipated prior to training (such as through initialized sub-networks) and evaluate whether these representation-theoretic properties extend to complex real-world systems, such as large language models.

The conclusions should be interpreted within the context of the study's scope. The evaluation relied on small, synthetic algorithmic tasks and relatively compact neural architectures optimized with weight decay. While confidence in the mechanistic explanations for these specific algebraic tasks is exceptionally high due to rigorous mathematical ablations, further empirical validation is necessary to confirm how directly these principles scale to large, foundation-scale models operating on natural language and vision data.

  • Paper: Group Equivariant Convolutional Networks, Taco S. Cohen et al. (2016). Its treatment of group actions and equivariant neural-network operations provides useful mathematical groundwork for following how finite-group structure can shape a network’s learned composition algorithm.

No sufficiently relevant recommendations were found.

Cover for A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations

Abstract

Universality is a key hypothesis in mechanistic interpretability – that different models learn similar features and circuits when trained on similar tasks. In this work, we study the universality hypothesis by examining how small neural networks learn to implement group composition. We present a novel algorithm by which neural networks may implement composition for any finite group via mathematical representation theory. We then show that networks consistently learn this algorithm by reverse engineering model logits and weights, and confirm our understanding using ablations. By studying networks of differing architectures trained on various groups, we find mixed evidence for universality: using our algorithm, we can completely characterize the family of circuits and features that networks learn on this task, but for a given network the precise circuits learned – as well as the order they develop – are arbitrary.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Setup and Background
  • 3.1. Task Description
  • 3.2. Mathematical Representation Theory
  • 4. An Algorithm for Group Composition
  • 5. Reverse Engineering Permutation Group Composition in a One Layer ReLU MLP
  • 5.1. Logit Attribution
  • 5.2. Embeddings and Unembeddings
  • 5.3. MLP Neurons
  • 5.4. Logit Computation
  • 5.5. Correctness Checks: Ablations
  • 5.6. Understanding Training Dynamics using Progress Measures
  • 6. Universality
  • 7. Conclusion and Discussion
  • Author Contributions
  • Acknowledgments
  • References
  • A. Relevance for Mechanistic Interpretability
  • B. Similarities and differences with prior work on reverse engineering modular addition
  • C. Architecture Details
  • C.1. MLP
  • C.1.1. CHOICE OF NETWORK SIZE
  • C.2. Transformer
  • D. Mathematical Representation Theory
  • D.1. Explicit Groups and Representations
  • D.1.1. IRREDUCIBLE REPRESENTATIONS OF THE CYCLIC GROUP
  • D.1.2. IRREDUCIBLE REPRESENTATIONS OF THE DIHEDRAL GROUP
  • D.1.3. IRREDUCIBLE REPRESENTATIONS OF THE SYMMETRIC GROUP
  • D.1.4. IRREDUCIBLE REPRESENTATIONS OF A₅
  • E. Additional Reverse Engineering of Mainline Model
  • E.1. Progress Measures
  • E.2. Full Circuit Analysis: Sign Representation
  • E.3. Implementing Multiplication via ReLUs
  • E.4. Visualizing the Embeddings and Unembeddings.
  • E.5. Further Reverse Engineering Details
  • F. Further Future Work
  • G. Further Discussion on Inductive Biases
  • H. Universality Results

Knowls

  1. Knowl 1 — Group composition via representations

    model/method

    For a finite group GG, choose a real representation ρ:G→GLd(R)\rho:G\to GL_d(\mathbb{R}), where dd is the matrix dimension. Given input elements a,b∈Ga,b\in G, the group-composition-via-representations (GCR) procedure maps them to ρ(a)\rho(a) and ρ(b)\rho(b), multiplies the matrices to obtain ρ(ab)\rho(ab), and scores each candidate output c∈Gc\in G by

    sc=tr⁡ ⁣(ρ(ab)ρ(c−1))=χρ(abc−1),s_c=\operatorname{tr}\!\left(\rho(ab)\rho(c^{-1})\right)=\chi_\rho(abc^{-1}),

    where tr⁡\operatorname{tr} is matrix trace and χρ(g)=tr⁡(ρ(g))\chi_\rho(g)=\operatorname{tr}(\rho(g)) is the representation character. The score is maximized when c=abc=ab; if ρ\rho is faithful, this maximizer is unique. Neural implementations can map the inputs to representation matrices in embeddings, compute their product in hidden activations, and use output weights to score candidates. Multiple learned representations can contribute separate character terms to the logits. For a cyclic group, the two-dimensional rotation representations yield the familiar Fourier/trigonometric score 2cos⁡(2πk(a+b−c)/n)2\cos(2\pi k(a+b-c)/n), making modular addition a special case.

  2. Knowl 2 — Character scores identify the correct group product

    theoretical result

    Let GG be a finite group and let ρ:G→GLd(R)\rho:G\to GL_d(\mathbb{R}) be a real representation. For every g∈Gg\in G, its character satisfies χρ(g)=tr⁡(ρ(g))≤d\chi_\rho(g)=\operatorname{tr}(\rho(g))\le d, with equality exactly when ρ(g)=Id\rho(g)=I_d, the d×dd\times d identity matrix. Consequently, the GCR score χρ(abc−1)\chi_\rho(abc^{-1}) is maximal for every candidate cc such that abc−1abc^{-1} is in the kernel of ρ\rho. If ρ\rho is faithful, its kernel contains only the identity, so c=abc=ab is the unique maximizer; for a non-faithful representation, other candidates may tie.

  3. Knowl 3 — Neural-network setup for the composition experiments

    experimental setup

    The mainline experiment trains a one-hidden-layer ReLU MLP on composition in S5S_5, the permutation group of order 120120. Each input is an ordered pair of group elements, encoded as separate one-hot vectors; left and right embeddings are untied. Each embedding has dimension 256256, their concatenation feeds a bias-free hidden linear layer of width 128128, and an untied unembedding produces 120120 output logits. Training uses 40%40\% of the multiplication table, full-batch AdamW with weight decay λ=1\lambda=1, learning rate 0.0010.001, β1=0.9\beta_1=0.9, and β2=0.98\beta_2=0.98, for up to 250,000250{,}000 epochs; evaluation uses all input pairs not used for training. The broader comparison covers MLPs and one-layer ReLU Transformers on C113,C118,D59,D61,S5,S6C_{113}, C_{118}, D_{59}, D_{61}, S_5, S_6, and A5A_5, with four random seeds per group and architecture.

  4. Knowl 4 — Two representations explain most mainline logits

    empirical result

    For the trained S5S_5 MLP, the authors compare the logits over all input pairs and candidate outputs with character scores for each nontrivial irreducible representation. The sign and standard representations have logit cosine similarities of 0.5090.509 and 0.7670.767, respectively; the other tested representations have zero similarity. These two character directions explain 84.8%84.8\% of logit variance. Evaluating with the two-direction logit approximation reduces test loss by 70%70\% relative to the original, while ablating the remaining logit directions does not change loss. Thus, although the model has 120120 output logits, its useful output scores are largely accounted for by two representation-character patterns.

  5. Knowl 5 — Representations are localized in embeddings, hidden units, and output weights

    empirical result

    In the S5S_5 MLP, the left embedding, right embedding, and unembedding are well approximated by the same two key representation spaces, sign and standard. The sign representation explains 6.95%6.95\%, 6.95%6.95\%, and 9.58%9.58\% of their variance, respectively; the standard representation explains 93.0%93.0\%, 93.0%93.0\%, and 84.5%84.5\%. The embedding matrices therefore use only the 16+116+1 dimensions associated with the 44-dimensional standard and 11-dimensional sign representations, rather than their possible rank of 120120.

    The 128128 hidden neurons form representation-specific clusters: 77 sign neurons, 119119 standard neurons, and 22 neurons that are always off. Directions corresponding to ρ(a)\rho(a), ρ(b)\rho(b), and ρ(ab)\rho(ab) explain 99.9%99.9\% of sign-cluster variance and 88.0%88.0\% of standard-cluster variance. The model’s hidden activations contain recoverable ρ(ab)\rho(ab) matrices: the recovered sign matrices match exactly, and the recovered standard matrices have mean squared error below 10−810^{-8}. The maps from the neuron clusters to logits are also localized to their corresponding output representation spaces, explaining 99.9%99.9\% of the sign-map variance and 93.4%93.4\% of the standard-map variance.

  6. Knowl 6 — Ablations validate the representation-based circuit

    empirical result

    Ablations in the S5S_5 MLP support the claim that the ρ(ab)\rho(ab) components of the key representations drive the prediction. The baseline loss is 2.38×10−62.38\times10^{-6}. Removing the standard ρ(ab)\rho(ab) directions raises loss to 7.557.55, while removing the sign ρ(ab)\rho(ab) directions raises it to 0.00090.0009; removing both key representations gives loss 7.607.60. By contrast, ablating hidden directions corresponding to ρ(a)\rho(a) or ρ(b)\rho(b) does not change loss, consistent with the final character calculation using the product representation rather than those input terms directly. Replacing the identified hidden neurons with the corresponding representation-matrix elements lowers loss by 70%70\%, to 7.00×10−77.00\times10^{-7}. Removing the 5.96%5.96\% unembedding residual lowers loss by 12%12\%, whereas retaining only that residual raises loss to 4.804.80.

  7. Knowl 7 — GCR structure appears across groups and architectures

    data/table

    Across seven groups and both model architectures, four-seed averages show that embedding and unembedding weights are largely explained by key-representation subspaces, while the hidden layer contains the predicted representation terms. The table reports the percentage of variance explained (FVE) and the final test, excluded, and restricted losses. For MLPs, “MLP FVE” is the FVE for the hidden layer and “ρ(ab)\rho(ab) FVE” is the FVE specifically for the product term. For Transformers, the corresponding embedding and output-weight columns are WEW_E and WLW_L. Loss columns are reported in the paper’s order: test, excluded, restricted.

    GroupMLP WaW_a FVEMLP WbW_b FVEMLP WUW_U FVEMLP FVEMLP ρ(ab)\rho(ab) FVEMLP test lossMLP excluded lossMLP restricted lossTransformer WEW_E FVETransformer WLW_L FVETransformer MLP FVETransformer ρ(ab)\rho(ab) FVETransformer test lossTransformer excluded lossTransformer restricted loss
    C113C_{113}99.53%99.39%98.05%90.25%12.03%1.63e-055.956.88e-0395.18%99.52%92.12%16.77%2.67e-079.422.12e-02
    C118C_{118}99.75%99.74%98.43%95.84%13.26%5.39e-068.723.60e-0394.05%99.64%94.63%17.11%1.73e-0715.932.55e-01
    D59D_{59}99.71%99.73%98.52%87.68%12.44%6.34e-0612.371.60e-0698.58%98.53%85.01%10.85%3.20e-0646.422.82e-05
    D61D_{61}99.26%99.45%98.26%87.61%12.48%1.79e-0512.001.69e-0698.33%97.40%85.59%11.11%1.63e-0241.649.60e-02
    S5S_5100.00%99.99%94.14%88.91%12.13%1.02e-0511.722.21e-0799.84%99.97%85.28%10.23%1.43e-0717.774.44e-09
    S6S_699.65%99.78%93.67%86.38%8.98%4.95e-0512.172.66e-0699.94%99.93%86.32%9.35%2.21e-06291.671.05e-06
    A5A_599.04%99.31%93.27%86.69%10.26%1.94e-059.825.28e-0797.53%97.40%83.56%8.22%4.88e-0219.767.70e-04

    The high FVE values across the groups and architectures support a shared representation-theoretic mechanism, rather than a solution specific to the S5S_5 MLP. Some Transformer test losses remain comparatively high, so the structural evidence does not imply equally strong task performance in every run.

  8. Knowl 8 — Learned features and their learning order are not strongly universal

    empirical result

    The paper distinguishes weak universality—shared computational principles with potentially different implementations—from strong universality, which would require the same features and circuits to arise consistently. Across the tested groups and architectures, models exhibit GCR-like circuits, supporting weak universality. The particular irreducible representations selected, the number selected, and the order in which they appear vary across random seeds, including runs with the same architecture and training task. In 50 MLP runs on S5S_5, the sign and standard representations were learned most often, other representations were learned less often, and the six-dimensional representation was not observed; the number of key representations also varied. Transformers generally learned fewer representations than MLPs despite having more parameters. The one-dimensional sign representation tended to appear early, but representations were not learned in a strict order of dimension. These findings argue against strong universality: interpreting one trained network does not reveal all mechanisms used across the model family.

  9. Knowl 9 — Progress measures separate memorization, circuit formation, and cleanup

    empirical result

    For the mainline S5S_5 MLP, the authors define restricted loss by retaining only hidden ρ(ab)\rho(ab) directions in the key representations, thereby measuring the GCR circuit’s performance. Excluded loss removes those directions and is evaluated on training data to track the remaining, memorizing solution. In the analyzed training run, three partially overlapping phases were identified: memorization (epochs 00–2,0002{,}000), when training and excluded losses fall while test and restricted losses remain high; circuit formation (about epochs 2,2002{,}200–87,00087{,}000), when excluded loss and total squared weight fall and restricted loss begins to improve while test performance remains relatively flat; and cleanup (about epochs 87,00087{,}000–120,000120{,}000), when restricted loss continues to fall and test loss drops sharply. The authors interpret cleanup as weight decay favoring the lower-weight generalizing circuit over the memorizing solution. In this account, the GCR circuit begins forming before grokking becomes visible in test loss.

  10. Knowl 10 — Scope and implementation limitations

    limitation

    The experiments establish representation-based implementations for small networks trained on finite-group composition; they do not establish that the same mechanism explains behavior in large models or practical tasks. In the studied ReLU architecture, matrix multiplication of representation activations is only approximate. The authors associate this with extra hidden and logit components and with incomplete variance explanation; they suggest that the ReLU multiplication step can leave ρ(a)\rho(a) and ρ(b)\rho(b) terms that are not orthogonal to ρ(ab)\rho(ab). Consequently, the empirical FVE need not reach 100%100\% even when the identified circuit accounts for task performance. The broader universality conclusion is therefore evidence from the tested toy tasks and architectures, not a demonstration of universality in general.

Coverage note — No substantial contributed material was deliberately omitted; individual-run tables and ancillary representation visualizations were left out because the averaged cross-group results and the main mechanistic findings capture their contribution.

References

  1. 1.Alperin, J. L. and Bell, R. B. Groups and Representations, volume 162 of Graduate Texts in Mathematics. Springer, New York, NY, 1995. ISBN 978-0-387-94526-2 978-1-4612-0799-3. doi: 10.1007/978-1-4612-0799-3.
  2. 2.Bansal, Y., Nakkiran, P., and Barak, B. Revisiting Model Stitching to Compare Neural Representations, June 2021.
  3. 3.Barak, B., Edelman, B. L., Goel, S., Kakade, S., Malach, E., and Zhang, C. Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit, January 2023.
  4. 4.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language Models are Few-Shot Learners, July 2020.
  5. 5.Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K. Thread: Circuits. Distill, 5(3):e24, March 2020. ISSN 2476-0757. doi: 10.23915/distill.00024.
  6. 6.Davies, X., Langosco, L., and Krueger, D. Unifying Grokking and Double Descent. In NeurIPS ML Safety Workshop, December 2022.
  7. 7.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. A mathematical framework for transformer circuits, 2021.
  8. 8.Frankle, J. and Carbin, M. The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks, March 2019.
  9. 9.Ganguli, D., Hernandez, D., Lovitt, L., DasSarma, N., Henighan, T., Jones, A., Joseph, N., Kernion, J., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Elhage, N., Showk, S. E., Fort, S., Hatfield-Dodds, Z., Johnston, S., Kravec, S., Nanda, N., Ndousse, K., Olsson, C., Amodei, D., Amodei, D., Brown, T., Kaplan, J., McCandlish, S., Olah, C., and Clark, J. Predictability and Surprise in Large Generative Models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 1747–1764, June 2022. doi: 10.1145/3531146.3533229.
  10. 10.Goh, G., †, N. C., †, C. V., Carter, S., Petrov, M., Schubert, L., Radford, A., and Olah, C. Multimodal neurons in artificial neural networks. Distill, 2021. doi: 10.23915/distill.00030.
  11. 11.Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gerard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E. Array programming with NumPy. Nature, 585(7825):357–362, September 2020. ISSN 1476-4687. doi: 10.1038/s41586-020-2649-2.
  12. 12.Inc., P. T. Collaborative data science. https://plot.ly, 2015.
  13. 13.Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. Similarity of Neural Network Representations Revisited, July 2019.
  14. 14.Li, K., Hopkins, A. K., Bau, D., Viegas, F., Pfister, H., and Wattenberg, M. Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task, February 2023.
  15. 15.Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. Convergent Learning: Do different neural networks learn the same representations?, February 2016.
  16. 16.Lindner, D., Kramar, J., Rahtz, M., McGrath, T., and Mikulík, V. Tracr: Compiled Transformers as a Laboratory for Interpretability, January 2023.
  17. 17.Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C. Transformers Learn Shortcuts to Automata, October 2022a.
  18. 18.Liu, Z., Kitouni, O., Nolte, N., Michaud, E. J., Tegmark, M., and Williams, M. Towards Understanding Grokking: An Effective Theory of Representation Learning, October 2022b.
  19. 19.Liu, Z., Michaud, E. J., and Tegmark, M. Omnigrok: Grokking Beyond Algorithmic Data, October 2022c.
  20. 20.McGrath, T., Kapishnikov, A., Tomasev, N., Pearce, A., Hassabis, D., Kim, B., Paquet, U., and Kramnik, V. Acquisition of Chess Knowledge in AlphaZero. Proceedings of the National Academy of Sciences, 119(47):e2206625119, November 2022. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.2206625119.
  21. 21.McKinney, W. Data Structures for Statistical Computing in Python. Proceedings of the 9th Python in Science Conference, pp. 56–61, 2010. doi: 10.25080/Majora-92bf1922-00a.
  22. 22.Mehrer, J., Spoerer, C. J., Kriegeskorte, N., and Kietzmann, T. C. Individual differences among deep neural network models. Nature Communications, 11(1):5725, November 2020. ISSN 2041-1723. doi: 10.1038/s41467-020-19632-w.
  23. 23.Meurer, A., Smith, C. P., Paprocki, M., Certík, O., Kirpichev, S. B., Rocklin, M., Kumar, Am., Ivanov, S., Moore, J. K., Singh, S., Rathnayake, T., Vig, S., Granger, B. E., Muller, R. P., Bonazzi, F., Gupta, H., Vats, S., Johansson, F., Pedregosa, F., Curry, M. J., Terrel, A. R., Roucka, S., Saboo, A., Fernando, I., Kulal, S., Cimrman, R., and Scopatz, A. SymPy: Symbolic computing in Python. PeerJ Computer Science, 3:e103, January 2017. ISSN 2376-5992. doi: 10.7717/peerj-cs.103.
  24. 24.Morcos, A. S., Raghu, M., and Bengio, S. Insights on representational similarity in neural networks with canonical correlation, October 2018.
  25. 25.Nanda, N. A Comprehensive Mechanistic Interpretability Explainer & Glossary. https://www.neelnanda.io/glossary, December 2022.
  26. 26.Nanda, N. TransformerLens, January 2023.
  27. 27.Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J. Progress measures for grokking via mechanistic interpretability, January 2023.
  28. 28.Neyshabur, B., Tomioka, R., and Srebro, N. In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning, April 2015.
  29. 29.Olah, C., Mordvintsev, A., and Schubert, L. Feature Visualization. Distill, 2(11):e7, November 2017. ISSN 2476-0757. doi: 10.23915/distill.00007.
  30. 30.Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. Zoom In: An Introduction to Circuits. Distill, 5(3):e00024.001, March 2020. ISSN 2476-0757. doi: 10.23915/distill.00024.001.
  31. 31.Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. In-context Learning and Induction Heads, September 2022.
  32. 32.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raiṡon, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019.
  33. 33.Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V. Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets, January 2022.
  34. 34.Sellam, T., Yadlowsky, S., Wei, J., Saphra, N., D’Amour, A., Linzen, T., Bastings, J., Turc, I., Eisenstein, J., Das, D., Tenney, I., and Pavlick, E. The MultiBERTs: BERT Reproductions for Robustness Analysis, March 2022.
  35. 35.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention Is All You Need, December 2017.
  36. 36.Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 small, November 2022.
  37. 37.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. Emergent Abilities of Large Language Models, October 2022.
  38. 38.Weiss, G., Goldberg, Y., and Yahav, E. Thinking Like Transformers, July 2021.
  39. 39.Zhang, Y., Backurs, A., Bubeck, S., Eldan, R., Gunasekar, S., and Wagner, T. Unveiling Transformers with LEGO: A synthetic reasoning task, July 2022.

Citation

MLA
Chughtai, B., et al. “A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations”. International Conference on Machine Learning, vol. 202, 2023, pp. 6243–67, https://proceedings.mlr.press/v202/chughtai23a.html.
APA
Chughtai, B., Chan, L., & Nanda, N. (2023). A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations. International Conference on Machine Learning, 202, 6243–6267. https://proceedings.mlr.press/v202/chughtai23a.html
Chicago
Chughtai, B., L. Chan, and N. Nanda. 2023. “A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations”. International Conference on Machine Learning 202: 6243–67. https://proceedings.mlr.press/v202/chughtai23a.html.
Harvard
Chughtai, B., Chan, L. and Nanda, N. (2023) “A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations”, International Conference on Machine Learning. PMLR, pp. 6243–6267. Available at: https://proceedings.mlr.press/v202/chughtai23a.html.
Vancouver
1. Chughtai B, Chan L, Nanda N (2023) A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations. In: International Conference on Machine Learning. PMLR, pp 6243–6267

BibTeX

@InProceedings{pmlr-v202-chughtai23a,
  title = 	 {A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations},
  author =       {Chughtai, Bilal and Chan, Lawrence and Nanda, Neel},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {6243--6267},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/chughtai23a/chughtai23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/chughtai23a.html},
  abstract = 	 {Universality is a key hypothesis in mechanistic interpretability – that different models learn similar features and circuits when trained on similar tasks. In this work, we study the universality hypothesis by examining how small networks learn to implement group compositions. We present a novel algorithm by which neural networks may implement composition for any finite group via mathematical representation theory. We then show that these networks consistently learn this algorithm by reverse engineering model logits and weights, and confirm our understanding using ablations. By studying networks trained on various groups and architectures, we find mixed evidence for universality: using our algorithm, we can completely characterize the family of circuits and features that networks learn on this task, but for a given network the precise circuits learned – as well as the order they develop – are arbitrary.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/