Progress measures for grokking via mechanistic interpretability

Neel NandaLawrence ChanTom LieberumJess SmithJacob Steinhardt

article2023ICLR739 citations

Explains the delayed generalization of grokking by reverse-engineering the trigonometric algorithms learned by transformers on modular addition, establishing continuous progress measures that track gradual circuit formation and the removal of memorization.

Listen

Modern artificial intelligence models frequently display emergent behaviors, where sudden and unexpected capabilities or safety risks arise abruptly after extensive training or scaling. Standard performance metrics often fail to anticipate these rapid shifts, making deployment timelines unpredictable and raising safety concerns. The article aims to evaluate whether reverse-engineering the internal mechanisms of neural networks can uncover continuous, hidden progress measures that explain and anticipate these sudden transitions.

To investigate this, the researchers conducted an in-depth mechanistic interpretability study on small transformer networks trained on modular addition tasks across tens of thousands of optimization epochs. By analyzing the internal weights, activations, and Fourier representations across different random seeds, prime moduli, and data fractions, the team mapped the exact algorithms learned by the models and constructed specific tracking metrics.

First, the article finds that instead of memorizing look-up tables indefinitely, the models learn a generalizable algorithm that maps inputs onto circles using discrete Fourier transforms and combines them via trigonometric identities. Second, training divides into three continuous phases: initial memorization of the training set, internal circuit formation where the generalizing mechanism builds up silently, and a cleanup phase where regularization eliminates memorization weights. Third, the apparent sudden jump in test performance (known as grokking) occurs entirely during the cleanup phase, well after the generalizing circuit has formed. Finally, both limited training data and explicit regularization—such as weight decay or dropout—are strictly necessary to drive this transition, whereas non-regularized models fail to generalize.

These findings demonstrate that sudden performance leaps are not random or instantaneous events, but rather the visible result of smooth, underlying structural changes within model weights. This insight offers a credible path toward anticipating capability jumps and safety failures before they appear on standard evaluation benchmarks, potentially lowering operational risks and preventing costly training overruns.

Organizations developing complex artificial intelligence systems should invest in automated mechanistic interpretability tools and track internal representation metrics rather than relying solely on high-level accuracy curves. For near-term applications, decision-makers should support pilot projects that extend internal progress tracking to larger models and multi-step tasks, while establishing quantitative thresholds to predict when transitions will occur.

Readers should exercise caution when extrapolating these findings directly to frontier-scale systems. The experiments rely primarily on small models trained on idealized mathematical and algorithmic tasks, and the manual reverse-engineering methods used here do not yet scale automatically to massive, complex models. Nevertheless, confidence in the demonstrated mechanism for algorithmic tasks remains very high across the studied settings.

arXiv: 2301.05217
Cover for Progress measures for grokking via mechanistic interpretability

Abstract

Neural networks often exhibit emergent behavior, where qualitatively new capabilities arise from scaling up the amount of parameters, training data, or training steps. One approach to understanding emergence is to find continuous \textit{progress measures} that underlie the seemingly discontinuous qualitative changes. We argue that progress measures can be found via mechanistic interpretability: reverse-engineering learned behaviors into their individual components. As a case study, we investigate the recently-discovered phenomenon of ``grokking'' exhibited by small transformers trained on modular addition tasks. We fully reverse engineer the algorithm learned by these networks, which uses discrete Fourier transforms and trigonometric identities to convert addition to rotation about a circle. We confirm the algorithm by analyzing the activations and weights and by performing ablations in Fourier space. Based on this understanding, we define progress measures that allow us to study the dynamics of training and split training into three continuous phases: memorization, circuit formation, and cleanup. Our results show that grokking, rather than being a sudden shift, arises from the gradual amplification of structured mechanisms encoded in the weights, followed by the later removal of memorizing components.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Setup and Background
  • 3.1 The Fourier multiplication algorithm
  • 4 Reverse engineering a one-layer transformer
  • 4.1 Suggestive evidence: surprising periodicity
  • 4.2 Mechanistic Evidence: Composing Model Weights
  • 4.3 Zooming In: Approximating Neurons with Sines and Cosines
  • 4.4 Correctness checks: ablations
  • 5 Understanding grokking behavior using progress measures
  • 5.1 Progress measures
  • 5.2 Phases of grokking: memorization, circuit formation, and cleanup
  • 5.3 Grokking and Weight Decay
  • 6 Conclusion and discussion
  • References
  • A Mathematical Structure of the Transformer
  • A.1 Empirical Model Simplifications
  • B Why use constructive intereference?
  • C Supporting evidence for mechanistic analysis of modular arithmetic networks
  • C.1 Further analysis of the specific training run discussed in the paper
  • C.1.1 Periodicity in the activations of other attention heads
  • C.1.2 Approximating attention heads with Sines and Cosines
  • C.1.3 The attention pattern weights are well approximated by differences of sines and cosines of a single frequency.
  • C.1.4 Periodicity in the activations of additional neurons
  • C.1.5 Additional grokking figures for mainline run
  • C.2 Additional results from different runs
  • C.2.1 Additional results for different runs with the same architecture
  • C.2.2 Results for other experimental setups
  • C.2.3 Generalizing models consistently use the Fourier Multiplication Algorithm
  • D Additional results on grokking
  • D.1 Both regularization and limited data are necessary for grokking
  • D.2 The slingshot mechanism often occurs, but is unnecessary for grokking
  • D.3 Additional evidence from other algorithmic tasks
  • E Further speculations on grokking
  • E.1 An intuitive explanation of grokking
  • E.2 Hypothesis: Phase Transitions are inherent to composition
  • F Further discussion on using mechanistic interpretability and progress measures for studying emergent phenomena

Knowls

  1. Knowl 1 — Fourier multiplication is the learned modular-addition algorithm

    model/method

    For modular addition with prime modulus PP, the studied transformers represent each input number using sine and cosine features at selected Fourier frequencies, then combine those features to compute the sum. For a selected integer frequency index kk, define wk=2πk/Pw_k=2\pi k/P. The embedding maps each input aa or bb to features including sin⁡(wka)\sin(w_k a) and cos⁡(wka)\cos(w_k a), or the corresponding features of bb. The attention and MLP layers form the sine and cosine of the sum using

    cos⁡(wk(a+b))=cos⁡(wka)cos⁡(wkb)−sin⁡(wka)sin⁡(wkb),\cos(w_k(a+b))=\cos(w_k a)\cos(w_k b)-\sin(w_k a)\sin(w_k b), sin⁡(wk(a+b))=sin⁡(wka)cos⁡(wkb)+cos⁡(wka)sin⁡(wkb).\sin(w_k(a+b))=\sin(w_k a)\cos(w_k b)+\cos(w_k a)\sin(w_k b).

    For each candidate output c∈{0,…,P−1}c\in\{0,\ldots,P-1\}, the output weights combine these features to form terms proportional to cos⁡(wk(a+b−c))\cos(w_k(a+b-c)). Adding terms from several selected frequencies makes the correct output c=a+b mod Pc=a+b\bmod P have a high logit through constructive interference, while reducing logits for other outputs. In the main P=113P=113 model, the selected indices were k∈{14,35,41,42,52}k\in\{14,35,41,42,52\}. The paper calls this procedure Fourier multiplication.

  2. Knowl 2 — The output weights read Fourier addition features from the MLP

    empirical result

    For the main one-layer transformer with P=113P=113, the combined map from MLP activations to output logits, WL=WUWoutW_L=W_UW_{\mathrm{out}}, is well approximated by ten directions: one sine and one cosine direction for each of the five selected frequency indices k∈{14,35,41,42,52}k\in\{14,35,41,42,52\}. Here WUW_U is the unembedding matrix and WoutW_{\mathrm{out}} is the MLP output matrix. The resulting approximation has residual Frobenius norm below 0.55%0.55\% of ∥WL∥F\|W_L\|_F. Projecting MLP activations onto each corresponding direction yields an approximate multiple of cos⁡(wk(a+b))\cos(w_k(a+b)) or sin⁡(wk(a+b))\sin(w_k(a+b)); a single such function explains more than 90%90\% of the variance in each projection. The full logits are also well approximated by a weighted sum of cos⁡(wk(a+b−c))\cos(w_k(a+b-c)) over the five selected frequencies, explaining 95%95\% of logit variance.

  3. Knowl 3 — Attention heads and MLP neurons localize computation by frequency

    empirical result

    In the main model, 433433 of 512512 MLP neurons (84.6%84.6\%) have more than 85%85\% of their activation variance explained by a degree-two polynomial in sine and cosine features of one selected frequency. For each frequency-clustered group of neurons, the corresponding columns of the neuron-to-logit map have substantial components only in the sine and cosine directions at that frequency. The four attention heads also have periodic attention patterns: their attention from the final '=' token to the token containing aa is well approximated by 0.5+α[cos⁡(wka)−cos⁡(wkb)]+β[sin⁡(wka)−sin⁡(wkb)]0.5+\alpha[\cos(w_k a)-\cos(w_k b)]+\beta[\sin(w_k a)-\sin(w_k b)]. The variance explained for heads 00 through 33 is respectively 99.03%99.03\%, 98.49%98.49\%, 99.07%99.07\%, and 97.91%97.91\%; their fitted frequency indices are respectively 3535, 4242, 5252, and 4242. The authors’ mechanistic interpretation is that heads 00 and 22 approximately compute degree-two terms at their respective frequencies, while heads 11 and 33 amplify the selected Fourier components in the residual stream.

  4. Knowl 4 — Fourier-space ablations validate which components implement the solution

    empirical result

    Ablations of the main P=113P=113 model support the Fourier multiplication interpretation. Removing any one of the five selected logit frequencies substantially worsens performance, while removing individual non-selected frequencies does not. Removing all logit Fourier components except the constant and the twenty sine/cosine components associated with the five selected frequencies reduces loss by 70%70\%, to 7.24×10−87.24\times10^{-8}. Replacing the activations of the 433433 MLP neurons that admit good degree-two approximations with their fitted polynomials increases loss only from 2.41×10−72.41\times10^{-7} to 2.48×10−72.48\times10^{-7} and does not change accuracy. Restricting the MLP activations to the terms corresponding to cos⁡(wk(a+b))\cos(w_k(a+b)) and sin⁡(wk(a+b))\sin(w_k(a+b)) instead improves loss by 77%77\%, to 5.54×10−85.54\times10^{-8}. Projecting MLP activations onto the ten output directions associated with the selected frequencies reduces loss by 50%50\% to 1.19×10−71.19\times10^{-7}; projecting onto the nullspace of those directions raises loss to 5.275.27, worse than uniform prediction.

  5. Knowl 5 — Restricted and excluded loss measure progress in the learned circuit

    definition

    The paper defines two training-time progress measures from Fourier ablations of the logits. For each input pair (a,b)(a,b), take the two-dimensional discrete Fourier transform of the model’s logits as a function of the inputs. Restricted loss is the loss after retaining only the constant component and the twenty sine/cosine components associated with the model’s selected frequencies, corresponding to functions of a+ba+b; it measures how well the model performs using the putative generalizing circuit alone. Excluded loss is measured on the training examples after removing those selected-frequency components while retaining the other Fourier components; it tracks the performance remaining after the proposed generalizing circuit is removed, which the authors use as a proxy for memorization. The paper also tracks the Gini coefficient of the Fourier-component norms of the embedding matrix and WLW_L as a measure of spectral sparsity, and the total squared weight magnitude as a measure related to weight decay.

  6. Knowl 6 — Grokking separates into memorization, circuit formation, and cleanup

    empirical result

    For the main P=113P=113 run, the authors identify three phases in training. During memorization (approximately epochs 00–1,4001{,}400), training and excluded losses decline, while test and restricted losses remain high; the selected frequencies used by the eventual solution are not yet substantially used. During circuit formation (approximately epochs 1,4001{,}400–9,4009{,}400), excluded loss rises and restricted loss begins to fall, while train and test losses remain nearly flat; the total squared weight magnitude falls as the Fourier multiplication mechanism forms. During cleanup (approximately epochs 9,4009{,}400–14,00014{,}000), restricted loss continues to fall, test loss drops sharply, and squared weight magnitude drops more sharply; the embedding and output maps become more sparse in the Fourier basis. Thus, in this run, the generalizing mechanism develops before the abrupt increase in test accuracy: the apparent grokking transition occurs during cleanup, as memorizing components are removed.

  7. Knowl 7 — Main experiment uses a one-layer transformer on a 30% modular-addition split

    experimental setup

    The principal experiment trains a transformer to predict c=a+b mod Pc=a+b\bmod P from the input sequence “a b =a\ b\ =,” with a,b∈{0,…,P−1}a,b\in\{0,\ldots,P-1\} represented as one-hot tokens and the output read at the '=' token. The main modulus is P=113P=113, and training uses 30%30\% of the 1132113^2 possible input pairs; loss and accuracy are evaluated on every pair not used for training. The model is a one-layer ReLU transformer with token-embedding dimension 128128, learned positional embeddings, four attention heads of dimension 3232, and an MLP with 512512 hidden units. It has no LayerNorm, and its embedding and unembedding matrices are untied. Training uses full-batch gradient descent with AdamW, learning rate 0.0010.001, weight-decay parameter 11, and 40,00040{,}000 epochs. In this setup, the model first overfits the training set and later generalizes, with test accuracy rising to nearly 100%100\% after roughly 10,00010{,}000 epochs.

  8. Knowl 8 — Weight decay and limited data are associated with delayed generalization

    empirical result

    In the modular-addition experiments, the authors report that grokking does not occur without weight decay or another effective form of regularization, and that removing weight decay prevents the excluded loss from rising, consistent with failure to form the generalizing circuit. With weight decay λ=1\lambda=1 and P=113P=113, training on at least 60%60\% of all input pairs leads to immediate generalization rather than a delayed train–test gap; training on 10%10\% or 20%20\% does not produce grokking even after 40,00040{,}000 epochs, while smaller data fractions generally delay generalization. In three additional algorithmic tasks—five-digit addition, repeated-subsequence prediction, and skip-trigram prediction—the authors likewise observe grokking with restricted datasets but not when training uses newly sampled data at every step. These results support the paper’s claim that limited data and regularization are important conditions for the observed grokking behavior, without establishing that they are universal requirements.

  9. Knowl 9 — The Fourier mechanism recurs across seeds and model settings

    empirical result

    The Fourier multiplication account is not limited to the main random seed. Four additional runs with the same one-layer architecture use sparse Fourier components in their embeddings and neuron-to-logit maps, represent sine and cosine functions of a+ba+b in MLP directions, and lose performance when their selected frequencies are ablated. The particular selected frequencies vary across seeds. The authors also report Fourier-multiplication variants in experiments with different training fractions, two-layer transformers, and prime moduli including 5353 and 109109. The grokking outcome depends on the setting: for P=53P=53, the original weight decay did not yield generalization, whereas λ=5\lambda=5 did; for P=401P=401, the tested models generalized immediately across weight-decay values. Models trained with dropout could generalize, although their embedding and unembedding matrices were not as sparse in the Fourier basis as in the main weight-decay models.

  10. Knowl 10 — The approach is task-specific and does not predict transition timing

    limitation

    The study fully reverse engineers small transformers on a simple modular-addition task, where a single interpretable circuit implements the solution. The analysis requires substantial manual effort, and the resulting progress measures depend on identifying task-specific Fourier frequencies; the paper does not establish that the method scales to larger models or realistic tasks. Although restricted and excluded loss vary smoothly before grokking in the studied setting, the authors do not provide a general notion of criticality that predicts in advance when the phase transition will occur.

Coverage note — The paper’s speculative explanations of circuit formation, including lottery-ticket, random-walk, and evolutionary analogies, are omitted because the authors present them as hypotheses rather than established results; detailed auxiliary plots and per-seed coefficient tables are also omitted because their findings are captured by the broader empirical knowls.

References

  1. 1.Boaz Barak, Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang. Hidden progress in deep learning: Sgd learns parities near the computational limit. arXiv preprint arXiv:2207.08799, 2022.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  3. 3.Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim. Thread: Circuits. Distill, 2020. doi: 10.23915/distill.00024. https://distill.pub/2020/circuits.
  4. 4.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021. https://transformer-circuits.pub/2021/framework/index.html.
  5. 5.Jonathan Frankle and Michael Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635, 2018.
  6. 6.Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al. Predictability and surprise in large generative models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 1747–1764, 2022.
  7. 7.Charles R. Harris, K. Jarrod Millman, Stefan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernandez del Río, Mark Wiebe, Pearu Peterson, Pierre Gerard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant. Array programming with NumPy. Nature, 585(7825):357–362, September 2020. doi: 10.1038/s41586-020-2649-2. URL https://doi.org/10.1038/s41586-020-2649-2.
  8. 8.Niall Hurley and Scott Rickard. Comparing measures of sparsity. IEEE Transactions on Information Theory, 55(10):4723–4741, 2009.
  9. 9.Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J Michaud, Max Tegmark, and Mike Williams. Towards understanding grokking: An effective theory of representation learning. arXiv preprint arXiv:2205.10343, 2022.
  10. 10.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  11. 11.Thomas McGrath, Andrei Kapishnikov, Nenad Tomasev, Adam Pearce, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kramnik. Acquisition of chess knowledge in alphazero. arXiv preprint arXiv:2111.09259, 2021.
  12. 12.Beren Millidge. Grokking ’grokking’, 2022. URL https://www.beren.io/2022-01-11-Grokking-Grokking/.
  13. 13.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. In-context learning and induction heads. Transformer Circuits Thread, 2022. https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html.
  14. 14.Alexander Pan, Kush Bhatia, and Jacob Steinhardt. The effects of reward misspecification: Mapping and mitigating misaligned models. arXiv preprint arXiv:2201.03544, 2022.
  15. 15.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  16. 16.Plotly Technologies Inc. Collaborative data science, 2015. URL https://plot.ly.
  17. 17.Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. Grokking: Generalization beyond overfitting on small algorithmic datasets. arXiv preprint arXiv:2201.02177, 2022.
  18. 18.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  19. 19.Alex Rogozhnikov. Einops: Clear and reliable tensor manipulations with einstein-like notation. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=oapKSVM2bcj.
  20. 20.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  21. 21.Jacob Steinhardt. More is different for ai, Feb 2022. URL https://bounded-regret.ghost.io/more-is-different-for-ai/.
  22. 22.Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind. The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon. arXiv preprint arXiv:2206.04817, 2022.
  23. 23.Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. arXiv preprint arXiv:2211.00593, 2022.
  24. 24.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022a.
  25. 25.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022b.
  26. 26.Wes McKinney. Data Structures for Statistical Computing in Python. In Stefan van der Walt and Jarrod Millman (eds.), Proceedings of the 9th Python in Science Conference, pp. 56 – 61, 2010. doi: 10.25080/Majora-92bf1922-00a.

Citation

MLA
Nanda, N., et al. “Progress Measures for Grokking via Mechanistic Interpretability”. arXiv, 2023, http://arxiv.org/abs/2301.05217v3.
APA
Nanda, N., Chan, L., Lieberum, T., Smith, J., & Steinhardt, J. (2023). Progress measures for grokking via mechanistic interpretability. arXiv. http://arxiv.org/abs/2301.05217v3
Chicago
Nanda, N., L. Chan, T. Lieberum, J. Smith, and J. Steinhardt. 2023. “Progress Measures for Grokking via Mechanistic Interpretability”. arXiv. http://arxiv.org/abs/2301.05217v3.
Harvard
Nanda, N. et al. (2023) “Progress measures for grokking via mechanistic interpretability”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.05217v3.
Vancouver
1. Nanda N, Chan L, Lieberum T, Smith J, Steinhardt J (2023) Progress measures for grokking via mechanistic interpretability. arXiv

BibTeX

@article{nanda2023progress,
  title = {Progress measures for grokking via mechanistic interpretability},
  author = {Nanda, Neel and Chan, Lawrence and Lieberum, Tom and Smith, Jess and Steinhardt, Jacob},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.05217v3},
  eprint = {2301.05217}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors