Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Kevin WangAlexandre VariengienArthur ConmyBuck ShlegerisJacob Steinhardt

article2022ICLR1,276 citationsSpotlight (notable-top-25%)

Reverse-engineers the 26-attention-head circuit that GPT-2 small uses for indirect object identification, establishing quantitative criteria to validate mechanistic interpretability in transformer language models.

Listen

Modern transformer language models are deployed across critical commercial and public systems, yet their internal decision-making processes remain largely opaque black boxes. This opacity introduces operational, safety, and compliance risks because users cannot reliably anticipate or audit out-of-distribution behaviors and systematic failures. The article aims to demonstrate that mechanistic interpretability can reverse-engineer how a model performs a specific, natural language task by isolating the human-understandable algorithmic "circuit" embedded within its parameters.

To achieve this, the authors investigated GPT-2 small on an indirect object identification task, which requires predicting the indirect object in sentences with duplicated subjects (such as predicting "Mary" in "When Mary and John went to the store, John gave a drink to"). The analysis used causal intervention techniques, primarily "path patching" and mean-activation ablations, across 100,000 synthetic template samples and repeated token sequences. By systematically tracing pathways backward from the output logits, the authors evaluated how individual attention heads read, process, and write information across residual streams.

Key findings reveal that a sparse subnetwork of 26 attention heads—just 1.1% of the model's total head-position pairs—performs the core task across seven functional roles. Primary "Name Mover Heads" copy the target name to the output, while "S-Inhibition Heads" suppress attention to duplicated subjects using both positional and token signals. Furthermore, the analysis uncovered unexpected internal dynamics: "Backup Name Mover Heads" automatically compensate and restore performance when primary heads are knocked out (resulting in only a 5% drop in logit difference), while "Negative Name Mover Heads" systematically suppress confidence, likely to hedge against high loss on incorrect predictions. In quantitative evaluations, the isolated circuit achieved 87% of the full model's task performance.

These findings prove that complex linguistic capabilities in language models decompose into modular, understandable algorithms rather than uninterpretable statistical noise. However, the presence of redundant backup pathways and hedging heads highlights that standard ablation tests can be misleading. Using these mechanistic insights, the authors designed targeted adversarial examples: introducing an extra duplicate of the indirect object caused the model to fail and choose the incorrect subject 23.4% of the time, compared to an error rate of just 0.7% on standard inputs.

For practitioners and researchers, the authors recommend incorporating structured circuit-evaluation criteria—specifically faithfulness, completeness, and minimality—when validating internal model explanations. Organizations building safety and assurance audits should leverage circuit discovery to craft targeted adversarial tests rather than relying solely on surface-level evaluations. Next steps require extending these causal techniques to larger, modern models and analyzing the role of multilayer perceptrons and layer normalization, which were excluded from this head-level analysis.

No sufficiently relevant recommendations were found.

Cover for Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Abstract

Research in mechanistic interpretability seeks to explain behaviors of machine learning models in terms of their internal components. However, most previous work either focuses on simple behaviors in small models, or describes complicated behaviors in larger models with broad strokes. In this work, we bridge this gap by presenting an explanation for how GPT-2 small performs a natural language task called indirect object identification (IOI). Our explanation encompasses 26 attention heads grouped into 7 main classes, which we discovered using a combination of interpretability approaches relying on causal interventions. To our knowledge, this investigation is the largest end-to-end attempt at reverse-engineering a natural behavior "in the wild" in a language model. We evaluate the reliability of our explanation using three quantitative criteria--faithfulness, completeness and minimality. Though these criteria support our explanation, they also point to remaining gaps in our understanding. Our work provides evidence that a mechanistic understanding of large ML models is feasible, opening opportunities to scale our understanding to both larger models and more complex tasks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Circuits and Knockouts
  • 3 Discovering the Circuit
  • 3.1 Which heads directly affect the output? (Name Mover Heads)
  • 3.2 Which heads affect the Name Mover Heads’ attention? (S-Inhibition Heads)
  • 3.3 Which heads affect S-Inhibition values?
  • 3.4 Did we miss anything? The Story of the Backup Name Movers Heads
  • 4 Experimental validation
  • 4.1 Completeness
  • 4.2 Minimality
  • 4.3 Comparison with a baseline circuit
  • 4.4 Designing adversarial examples
  • 5 Discussion
  • References
  • 6 Appendix
  • A Disentangling token and positional signal in the output of S-Inhibition Heads
  • B Path patching
  • C Direct effect on S-Inhibition Heads’ keys
  • D Identification of Previous Token Heads
  • E IOI Templates
  • F Backup Name Mover Heads
  • G GPT-2 small full architecture
  • H Validation of the induction mechanism on sequences of random tokens
  • I Validation of Duplicate Token Heads
  • J Role of MLPs in the task
  • K Minimality sets
  • L Template for adversarial examples
  • M Greedy Algorithm

Knowls

  1. Knowl 1 — A 26-head circuit implements indirect-object identification in GPT-2 small

    model/method

    The indirect-object identification (IOI) task asks a model to complete a sentence with the indirect object (IO), rather than the subject (S) of the main clause. For example, in “When Mary and John went to the store, John gave a drink to”, the target is Mary. The task can be solved by identifying the names in the sentence, removing the duplicated name, and outputting the name that remains.

    The authors identify a circuit of 26 attention heads in 12-layer GPT-2 small, organized into seven classes: Previous Token Heads (2.2, 4.11); Duplicate Token Heads (0.1, 3.0, and fuzzy head 0.10); Induction Heads (5.5, 6.9, and fuzzy heads 5.8, 5.9); S-Inhibition Heads (7.3, 7.9, 8.6, 8.10); Name Mover Heads (9.6, 9.9, 10.0); Negative Name Mover Heads (10.7, 11.10); and Backup Name Mover Heads (9.0, 9.7, 10.1, 10.2, 10.6, 10.10, 11.2, 11.9). The circuit routes information about repeated names and their positions to the final-token position, where Name Mover Heads copy the remaining name toward the output. It accounts for the bulk, but not all, of GPT-2 small’s IOI performance.

    The evaluation sentences were generated from 15 templates with random single-token names, places, and objects. On 100,000 IOI examples, the paper reports a mean IO-versus-S logit difference of 3.56, an average IO probability of 49%, and an IO logit greater than the S logit on 99.3% of examples.

  2. Knowl 2 — Path patching measures a head’s direct causal effect along selected paths

    model/method

    Path patching estimates the direct effect of a sender attention head on selected receiver components while excluding routes mediated by other attention heads. The method uses an original input xorigx_{\mathrm{orig}} from the IOI distribution pIOIp_{\mathrm{IOI}} and a corresponding input xnewx_{\mathrm{new}} from pABCp_{\mathrm{ABC}}, where the names are replaced by three unrelated names but the template is preserved. A receiver set RR can contain attention-head queries, keys, or values, or the final residual-stream state.

    The procedure is: (1) record activations for both inputs; (2) run a modified pass on xorigx_{\mathrm{orig}} in which the sender head is replaced with its activation from xnewx_{\mathrm{new}} and other attention heads are frozen to their original-input activations; (3) recompute intervening MLPs and receiver components, and save the receivers’ resulting activations; and (4) run a pass on xorigx_{\mathrm{orig}} with those saved receiver activations substituted in, recomputing the remaining computation to obtain logits. Freezing intermediate attention heads isolates the sender-to-receiver paths that do not pass through another attention head. The reported path effects are averages over more than 200 original/new input pairs, measured by the change in IO-versus-S logit difference.

  3. Knowl 3 — Name Mover Heads copy attended names, while Negative Name Movers oppose them

    empirical result

    Path patching from attention heads at the final-token position to the logits identifies heads 9.6, 9.9, and 10.0 as positively contributing to the IO-versus-S logit difference. These Name Mover Heads attend strongly to the IO token (average attention probability 0.59) and copy the name they attend to: their copy scores exceed 95%, compared with below 20% for an average attention head. Here, copy score is the proportion of tested inputs for which a head’s output, simulated through its OV matrix, places the input name token among the top five logits. Across 500 examples, attention probability on a name correlates with the head output’s projection in that name’s unembedding direction at ρ>0.81\rho > 0.81.

    Heads 10.7 and 11.10 have the opposite direct effect on the logit difference. They attend to names but write against the direction of the attended name; their negative copy score is 98%, compared with 12% for an average head. The paper labels them Negative Name Mover Heads and notes that their broader function is not established.

  4. Knowl 4 — S-Inhibition Heads steer Name Movers away from the repeated subject

    empirical result

    Four S-Inhibition Heads—7.3, 7.9, 8.6, and 8.10—act at the final-token position and influence the queries of the Name Mover Heads. They attend preferentially to the second occurrence of the subject, S2; their average attention from the final position to S2 is 0.51. Patching their effect on Name Mover queries with activations from pABCp_{\mathrm{ABC}} shifts Name Mover attention away from the IO and toward the first subject occurrence, S1. This supports the interpretation that S-Inhibition Heads help the Name Movers select the IO over the repeated subject.

    Their outputs carry two distinguishable signals: a token signal identifying the subject name, and a positional signal identifying the S1 position. Counterfactual patching experiments indicate that both signals contribute, with the positional signal having the larger effect. If SposS_{\mathrm{pos}} is 11 or −1-1 for an original or inverted positional signal, and StokS_{\mathrm{tok}} is 11, 00, or −1-1 for an original, uncorrelated, or inverted token signal, respectively, the patched logit difference is approximated by 2.31Spos+0.99Stok2.31S_{\mathrm{pos}} + 0.99S_{\mathrm{tok}}. The approximation has a mean error of 7% relative to the baseline logit difference. Inverting both signals reverses the logit difference, making the model favor S over IO.

  5. Knowl 5 — Duplicate-token and induction mechanisms supply positional information to S-Inhibition Heads

    empirical result

    Duplicate Token Heads 0.1 and 3.0 attend from S2 to the previous occurrence of the duplicated token at S1. The authors interpret their output as carrying information about the earlier token’s position. Two Induction Heads, 5.5 and 6.9, attend from S2 to S1+1, the token immediately following S1; they use information supplied at S1+1 by Previous Token Heads 2.2 and 4.11. In this IOI circuit, the induction heads’ output serves as a positional signal for S-Inhibition Heads, rather than simply performing next-token copying. Heads 5.8 and 5.9 are also included as fuzzy induction heads in the circuit.

    Evidence for these roles comes from activation-patching and repeated random-token tests. Patching the outputs of the Duplicate Token and Induction Heads from examples with altered S1 positions causes a logit-difference drop at least 88% as large as patching the S-Inhibition Heads; changing the names while keeping the relevant positions fixed causes less than 8% of that drop. On repeated random-token sequences, heads 0.1 and 3.0 rank among the three strongest duplicate-attention heads; 5.5 and 6.9 rank among the five strongest induction-attention heads and show copying behavior; and 2.2 and 4.11 have the highest previous-token attention scores. The proposed pointer-like positional mechanism is supported by these interventions, but the paper does not fully explain how the heads implement it.

  6. Knowl 6 — Backup Name Mover Heads compensate when the main Name Movers are ablated

    empirical result

    When the three regular Name Mover Heads (9.6, 9.9, and 10.0) are knocked out together, the IOI logit difference falls by only 5%. Path patching after this knockout identifies eight Backup Name Mover Heads: 9.0, 9.7, 10.1, 10.2, 10.6, 10.10, 11.2, and 11.9. The heads were selected as strong post-knockout contributors, with an effect-size threshold above 2%.

    Their behavior is heterogeneous: four resemble regular Name Movers; two attend equally to IO and S and copy both; one attends more to S1 and writes toward S; and one attends to S2 and writes negatively. The compensation is conditional: these heads do not normally perform the same output role, but become important after the regular Name Movers are removed. The authors hypothesize that training with dropout may have encouraged this robustness, but do not establish its cause.

  7. Knowl 7 — Faithfulness, completeness, and minimality test different properties of a circuit

    definition

    Let MM be the full model, CC a circuit (a subgraph of model components), and XX an input drawn from the IOI distribution pIOIp_{\mathrm{IOI}}. A knockout removes nodes outside the specified graph; in the experiments, removed head-position nodes are mean-ablated using activations averaged over pABCp_{\mathrm{ABC}} samples of the same template. Let f(C(X);X)f(C(X);X) be the logit difference between the IO and S tokens when the circuit is run on XX, and define circuit performance as

    F(C)=EX∼pIOI[f(C(X);X)].F(C) = \mathbb{E}_{X\sim p_{\mathrm{IOI}}}[f(C(X);X)].

    Faithfulness measures whether a circuit performs like the full model, by comparing F(C)F(C) with F(M)F(M). Completeness asks whether this similarity persists after any subset KK of circuit nodes is removed, using the discrepancy ∣F(C∖K)−F(M∖K)∣|F(C\setminus K)-F(M\setminus K)|. Minimality asks whether each circuit node vv can be shown to matter: for each v∈Cv\in C, there should be some K⊆C∖{v}K\subseteq C\setminus\{v\} for which removing vv changes performance, measured by ∣F(C∖(K∪{v}))−F(C∖K)∣|F(C\setminus(K\cup\{v\}))-F(C\setminus K)|. Faithfulness concerns task performance, completeness concerns whether the model and circuit retain matching behavior under knockouts, and minimality concerns whether included nodes are behaviorally relevant.

  8. Knowl 8 — The proposed circuit is mostly faithful and minimally relevant, but completeness tests reveal gaps

    empirical result

    For the 26-head circuit, the difference from full-model performance is ∣F(M)−F(C)∣=0.46|F(M)-F(C)|=0.46, or 13% of the full-model logit difference F(M)=3.56F(M)=3.56; thus the circuit retains 87% of the full model’s measured IOI performance. Randomly sampled knockout sets and whole-class knockouts yielded small completeness discrepancies. However, a greedy search for knockout sets found discrepancies as large as 3.09, equivalent to 87% of the original logit difference. The resulting sets typically mixed heads from several classes and were not readily interpretable, so the search demonstrates gaps without identifying a simple missing component.

    The authors found a knockout set demonstrating a nontrivial minimality effect for every included head; each reported effect was at least 1% of the original logit difference, though individual head contributions varied substantially. A simpler baseline circuit—excluding Backup and Negative Name Mover Heads—also had high task performance, but was more readily shown to be incomplete using random and whole-class knockouts. These results show why matching the model’s unablated score alone does not establish that a circuit fully captures its computation.

  9. Knowl 9 — Duplicating the indirect object sharply degrades IOI performance

    empirical result

    The authors used the circuit’s emphasis on duplicate detection to construct sentences containing an extra occurrence of the indirect object (IO), inserted in a natural middle sentence. On this adversarial distribution, the mean IO-versus-S logit difference is 1.23, the IO probability is 0.36, and the model assigns S a higher logit than IO on 23.4% of examples. For comparison, the IOI baseline distribution has logit difference 3.55, IO probability 0.49, and an S-higher-than-IO rate of 0.7%. A control distribution with an extra occurrence of S instead has logit difference 3.64, IO probability 0.59, and an S-higher-than-IO rate of 0.4%.

    The contrast is consistent with the circuit-based expectation that additional IO duplication can disrupt the model’s selection of the non-subject name, while merely adding another occurrence of S does not produce the same failure. The authors caution that this attack could likely have been found without their circuit analysis and that they do not fully understand the circuit’s behavior on the adversarial sentences.

  10. Knowl 10 — The circuit analysis leaves MLP roles and several mechanisms unresolved

    limitation

    The proposed circuit focuses on attention heads and does not include MLPs, layer norms, or embedding matrices. Individual knockout experiments suggest that MLP0 has a large effect on IOI performance, while later MLPs do not have large effects when removed individually; knocking out all MLPs after the first nevertheless reduces the logit difference to −1.1-1.1. The MLP contribution is therefore not captured by the attention-head circuit.

    The authors also do not fully explain how S-Inhibition Heads encode positional information, how Duplicate Token Heads implement duplicate detection at the parameter level, or how induction heads are repurposed to provide positional signals in this task. The compensation mechanism of Backup Name Movers remains unexplained, and the paper’s analysis is limited to GPT-2 small; scaling the reverse-engineering method to larger models is left open.

Coverage note — The paper’s detailed template listings, head-by-head plots, greedy-search pseudocode, and auxiliary repeated-token analyses are not separate knowls because their reconstructive findings and relevant quantitative evidence are incorporated into the circuit, validation, and limitation knowls.

References

  1. 1.Boaz Barak, Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang. Hidden progress in deep learning: Sgd learns parities near the computational limit. arXiv preprint arXiv:2207.08799, 2022.
  2. 2.Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda B. Viegas, and Martin Wattenberg. An interpretability illusion for BERT. CoRR, abs/2104.07143, 2021. URL https://arxiv.org/abs/2104.07143.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  4. 4.Guendalina Caldarini, Sardar Jaf, and Kenneth McGarry. A literature survey of recent advances in chatbots. Information, 13(1):41, 2022.
  5. 5.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021. https://transformer-circuits.pub/2021/framework/index.html.
  6. 6.Atticus Geiger, Hanson Lu, Thomas F Icard, and Christopher Potts. Causal abstractions of neural networks. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, 2021. URL https://openreview.net/forum?id=RmuXDtjDhG.
  7. 7.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. arXiv preprint arXiv:2012.14913, 2020.
  8. 8.Dan Hendrycks and Mantas Mazeika. X-risk analysis for ai research. arXiv, abs/2206.05862, 2022.
  9. 9.Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas. Natural language descriptions of deep visual features. In International Conference on Learning Representations, 2021.
  10. 10.Sarthak Jain and Byron C. Wallace. Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 3543–3556, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1357. URL https://aclanthology.org/N19-1357.
  11. 11.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in gpt. arXiv preprint arXiv:2202.05262, 2022.
  12. 12.Paul Michel, Omer Levy, and Graham Neubig. Are sixteen heads really better than one? Advances in neural information processing systems, 32, 2019.
  13. 13.Jesse Mu and Jacob Andreas. Compositional explanations of neurons. Advances in Neural Information Processing Systems, 33:17153–17163, 2020.
  14. 14.Neel Nanda and Tom Lieberum. A mechanistic interpretability analysis of grokking, 2022. URL https://www.alignmentforum.org/posts/N6WM6hs7RQMKDhYjB/a-mechanistic-interpretability-analysis-of-grokking.
  15. 15.Chris Olah. Mechanistic interpretability, variables, and the importance of interpretable bases. https://www.transformer-circuits.pub/2022/mech-interp-essay, 2022. Accessed: 2022-15-09.
  16. 16.Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: An introduction to circuits. Distill, 2020. doi: 10.23915/distill.00024.001. https://distill.pub/2020/circuits/zoom-in.
  17. 17.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. In-context learning and induction heads. arXiv preprint arXiv:2209.11895, 2022.
  18. 18.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  19. 19.Tilman Rauker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. Toward transparent ai: A survey on interpreting the inner structures of deep neural networks, 2022. URL https://arxiv.org/abs/2207.13243.
  20. 20.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pp. 5998–6008, 2017.
  21. 21.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. Investigating gender bias in language models using causal mediation analysis. Advances in Neural Information Processing Systems, 33:12388–12401, 2020.
  22. 22.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. ArXiv, abs/2206.07682, 2022.
  23. 23.Angela Zhang, Lei Xing, James Zou, and Joseph C Wu. Shifting machine learning for healthcare from development to deployment and from models to data. Nature Biomedical Engineering, pp. 1–16, 2022.

Citation

MLA
Wang, K., et al. “Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small”. arXiv, 2022, http://arxiv.org/abs/2211.00593v1.
APA
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., & Steinhardt, J. (2022). Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small. arXiv. http://arxiv.org/abs/2211.00593v1
Chicago
Wang, K., A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt. 2022. “Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small”. arXiv. http://arxiv.org/abs/2211.00593v1.
Harvard
Wang, K. et al. (2022) “Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.00593v1.
Vancouver
1. Wang K, Variengien A, Conmy A, Shlegeris B, Steinhardt J (2022) Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small. arXiv

BibTeX

@article{wang2022interpretability,
  title = {Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small},
  author = {Wang, Kevin and Variengien, Alexandre and Conmy, Arthur and Shlegeris, Buck and Steinhardt, Jacob},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.00593v1},
  eprint = {2211.00593}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/