Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Fred ZhangNeel Nanda

article2024ICLR345 citations

Demonstrates how arbitrary choices in corruption methods and evaluation metrics lead to conflicting activation patching results, establishing concrete best practices for reliable circuit discovery and localization in language models.

Listen

As large language models see wider deployment in critical systems, understanding their inner workings has become essential for debugging errors, preventing unwanted behavior, and ensuring safety. A foundational tool for this work is activation patching—an intervention technique designed to pinpoint which internal model components causally drive specific outputs. Despite its widespread use across the AI research community, there is no standardized protocol for how to corrupt baseline inputs or measure intervention effects. The article systematically evaluates how these implementation choices affect interpretability results, demonstrating that common practices can yield conflicting conclusions about how models operate.

The researchers conducted systematic experiments across multiple models (including GPT-2 variants and the 6-billion parameter GPT-J) and several benchmark tasks, such as factual recall, indirect object identification, Python code completion, and multi-digit arithmetic. They compared two primary methods for corrupting prompts—adding random Gaussian noise versus symmetric token replacement using natural counterfactual words—and evaluated different scoring metrics, notably output probability versus logit difference (the difference in raw model output scores between correct and incorrect answers). They also analyzed sliding window patching, which intervenes across multiple adjacent layers simultaneously.

The findings show that methodological choices substantially alter interpretability outcomes. First, Gaussian noise corruption frequently pushes the model outside its normal data distribution, disrupting internal mechanisms and producing noisy or contradictory findings; for instance, in factual recall, Gaussian noise created an apparent peak in early-to-middle layers that was two to five times higher than when using clean token replacement. Second, the probability metric fails to identify "negative" components that actively harm performance whenever corruption drives the baseline probability close to zero, whereas logit difference reliably captures both positive and negative components. Third, sliding window patching produces maximum peak effects at least 20 percent higher than summing individual layer interventions, artificially amplifying weak single-layer signals due to non-linear multi-layer interactions. Finally, the choice of which specific token positions to corrupt strongly dictates which internal circuits are detected.

These results carry direct implications for AI governance, model editing, and safety engineering. When practitioners rely on flawed interpretability hyperparameters, they risk editing the wrong model weights or drawing false conclusions about model reliability and compliance. To establish robust best practices, the article recommends using symmetric token replacement over random noise to maintain valid data distributions, adopting logit difference rather than raw probabilities as the primary evaluation metric, testing single-layer interventions before multi-layer windows, and varying the target tokens corrupted to ensure complete circuit discovery.

The findings are supported by consistent results across several diverse model architectures and reasoning tasks. However, users should note that the analysis focused on decoder-only language models up to 6 billion parameters and examined overriding corrupted activations with clean ones, rather than the reverse direction. Decision-makers should treat localization findings derived purely from Gaussian noise or multi-layer window patching with caution until validated against standardized, in-distribution methods.

arXiv: 2309.16042
Cover for Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Abstract

Mechanistic interpretability seeks to understand the internal mechanisms of machine learning models, where localization -- identifying the important model components -- is a key step. Activation patching, also known as causal tracing or interchange intervention, is a standard technique for this task (Vig et al., 2020), but the literature contains many variants with little consensus on the choice of hyperparameters or methodology. In this work, we systematically examine the impact of methodological details in activation patching, including evaluation metrics and corruption methods. In several settings of localization and circuit discovery in language models, we find that varying these hyperparameters could lead to disparate interpretability results. Backed by empirical observations, we give conceptual arguments for why certain metrics or methods may be preferred. Finally, we provide recommendations for the best practices of activation patching going forwards.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Activation patching
  • 2.2 Problem settings
  • 3 Corruption methods
  • 3.1 Results on corruption methods
  • 3.2 Evidence for OOD behavior in Gaussian noise corruption
  • 4 Evaluation Metrics
  • 4.1 Localizing factual recall with logit difference
  • 4.2 Circuit discovery with probability
  • 5 Sliding window patching
  • 6 Discussion and recommendations
  • 7 Related work
  • 8 Conclusion
  • References
  • A Review of Transformer Architecture
  • B Details on Experimental Settings
  • C Results on arithmetic reasoning in GPT-J
  • D Results on Python docstring circuit
  • E Results on the greater-than circuit in GPT-2 small
  • F Which tokens to corrupt matters
  • G Further details on factual association
  • G.1 Plots on MLP patching at the last subject token in GPT-2 XL
  • G.2 Plots on MLP patching at all token positions in GPT-2 XL
  • G.3 Plots on sliding window patching in GPT2-XL
  • G.4 Plots on activation patching of MLP layers on GPT-2 large
  • G.5 Plots on activation patching of MLP layers in GPT-J
  • H Further details on IOI circuit discovery
  • H.1 Detailed plots on activation patching
  • H.2 Details on detections
  • H.3 Detailed plots on fully random corruption
  • I Dataset Samples

Knowls

  1. Knowl 1 — Activation Patching Procedure and Metric Formulations

    model/method

    Activation patching (also known as causal tracing or interchange intervention) localizes the causal role of internal model components across three forward passes:

    1. Clean run: The model processes a clean prompt XcleanX_{\text{clean}} with target answer rr, and the intermediate activations of candidate components (e.g., attention head outputs or MLP layers) are cached.
    2. Corrupted run: The model processes a corrupted prompt XcorruptX_{\text{corrupt}} (with answer r′r' under token replacement), recording corrupted baseline outputs.
    3. Patched run: The model processes XcorruptX_{\text{corrupt}} while replacing the activation of a specific component at chosen token positions with its cached value from the clean run.

    Let cl\text{cl}, ∗*, and pt\text{pt} denote the clean, corrupted, and patched runs, respectively. The patching effect is evaluated using one of three standard metrics:

    • Probability: Difference in target token probability: ΔP(r)=Ppt(r)−P∗(r)\Delta P(r) = P_{\text{pt}}(r) - P_{*}(r)

    • Normalized Logit Difference: Difference between target logit and counterfactual logit, normalized to [0,1][0, 1] where 11 denotes full restoration: LD(r,r′)=Logit(r)−Logit(r′)\text{LD}(r, r') = \text{Logit}(r) - \text{Logit}(r') EffectLD=LDpt(r,r′)−LD∗(r,r′)LDcl(r,r′)−LD∗(r,r′)\text{Effect}_{\text{LD}} = \frac{\text{LD}_{\text{pt}}(r, r') - \text{LD}_{*}(r, r')}{\text{LD}_{\text{cl}}(r, r') - \text{LD}_{*}(r, r')}

    • Kullback-Leibler (KL) Divergence: Reduction in divergence from the clean predictive distribution: ΔDKL=DKL(Pcl∥P∗)−DKL(Pcl∥Ppt)\Delta D_{\text{KL}} = D_{\text{KL}}(P_{\text{cl}} \parallel P_{*}) - D_{\text{KL}}(P_{\text{cl}} \parallel P_{\text{pt}})

  2. Knowl 2 — Out-of-Distribution Mechanism Disruption Under Gaussian Noise Corruption

    empirical result

    Gaussian Noising (GN) corruption—adding Gaussian noise N(0,ν)\mathcal{N}(0, \nu) with ν=3σemb\nu = 3\sigma_{\text{emb}} to token embeddings—drives transformer hidden states out-of-distribution (OOD) and disrupts downstream circuit mechanisms in GPT-2 small on Indirect Object Identification (IOI):

    1. Disruption of Attention Routing: In clean IOI runs, Name Mover (NM) heads at the final token assign an average attention probability of 0.580.58 to the indirect object (IO) token. Under Symmetric Token Replacement (STR), which swaps the second subject (S2) with IO, the attention pattern cleanly switches to attend to the subject (S1S1). Under GN corruption, NM attention degrades into an unnatural split between IO (0.260.26) and S1S1 (0.210.21).
    2. Failure of Upstream Value Restoration: Patching the clean value activations of S-Inhibition heads into the corrupted run restores the model's IO logit difference to 1.041.04 under STR, but recovers only 0.490.49 under GN because corrupted signals continue to distort NM behavior along unpatched paths.
    3. Spurious Negative Head Detections: Under GN, Duplicate Token head 0.100.10 is detected as negatively contributing to model performance under Probability and KL divergence metrics (and negatively under Logit Difference). Under STR, head 0.100.10 is correctly identified as contributing positively.
  3. Knowl 3 — Failure Mode of Probability Metric for Negative Circuit Components

    empirical result

    Evaluating activation patching with the probability metric ΔP(r)=Ppt(r)−P∗(r)\Delta P(r) = P_{\text{pt}}(r) - P_{*}(r) systematically fails to detect negative model components (components whose normal operation reduces the probability of the correct output) whenever prompt corruption drives the corrupted baseline probability P∗(r)P_{*}(r) close to zero.

    Because probability is bounded below by 00, the maximum possible negative patching effect is strictly lower-bounded by −P∗(r)-P_{*}(r):

    ΔP(r)≥−P∗(r)\Delta P(r) \ge -P_{*}(r)

    In GPT-2 small on Indirect Object Identification (IOI):

    • Under STR corruption, P∗(IO)=0.03P_{*}(\text{IO}) = 0.03. For a detection threshold of 22 standard deviations below the mean (threshold −0.027-0.027 with mean 0.0030.003 and SD=0.015\text{SD} = 0.015), Negative Name Mover head 11.1011.10 achieves an effect of −0.022-0.022 and is missed.
    • Under random 3-name replacement (pABCp_{\text{ABC}} corruption), P∗(IO)≈5×10−4P_{*}(\text{IO}) \approx 5 \times 10^{-4}. As a result, probability patching fails to detect both Negative Name Mover heads (10.710.7 and 11.1011.10).

    In contrast, normalized logit difference has no bounded floor at −P∗(r)-P_*(r) and successfully detects both negative heads across all corruption regimes.

  4. Knowl 4 — Disparity in MLP Factual Recall Localization Between GN and STR

    empirical result

    In GPT-2 XL and GPT-2 large on factual recall (evaluated on the PAIREDFACTS dataset containing 145 paired counterfactual prompts), localizing MLP layer computation at the subject's final token produces contrasting results depending on the corruption method:

    • Under Gaussian Noising (GN), a pronounced localization peak appears in early-middle MLP layers (centered around layer 16 in GPT-2 XL), reproducing prior causal tracing findings.
    • Under Symmetric Token Replacement (STR), which replaces subject entities with semantically similar in-distribution entities of equal sequence length, the early-middle layer peak is substantially weaker or absent.

    Across multiple sliding window sizes (33, 55, and 1010), the peak patching effect under GN is 2×2\times to 5×5\times larger than under STR for both probability and logit difference metrics, showing that early-middle MLP factual storage localization in causal tracing is heavily influenced by the choice of noise corruption.

  5. Knowl 5 — Sliding Window Patching Amplifies Apparent Localization Relative to Single-Layer Patching

    empirical result

    Sliding window patching—restoring the activations of kk adjacent MLP layers simultaneously—produces substantially larger localization peaks than summing the effects of patching each layer individually across that same window.

    In GPT-2 XL on factual recall using Gaussian Noising (GN) and probability evaluation, sliding window patching yields peak values that exceed the sum of individual layer patchings by:

    • 1.40×1.40\times for window size k=3k = 3
    • 1.75×1.75\times for window size k=5k = 5
    • 1.59×1.59\times for window size k=10k = 10

    Across combinations of corruption methods, metrics, and window sizes, sliding window patching generates at least 20%20\% more peak effect than linear summation. Patching individual MLP layers in isolation reveals minimal or no middle-layer peak (e.g., at layer 15 in GPT-2 XL). The sharp concentration in multi-layer patching is an emergent consequence of joint non-linear restoration across layers—such as simultaneously suppressing corrupted information flow across the full window—rather than localization within a single layer.

  6. Knowl 6 — Discrepancy in Token-Position Attribution Between Probability and Logit Difference

    empirical result

    In GPT-2 XL and GPT-J on factual association, the choice of evaluation metric alters the relative importance assigned to different token positions. The probability metric concentrates significantly more causal attribution onto the last subject token than normalized logit difference does.

    Measuring the ratio of the summed layer patching effects at the last subject token to the summed layer patching effects across middle subject tokens (using a sliding window of size 5 in GPT-2 XL):

    • Under STR corruption: Probability yields a ratio of 4.33×4.33\times, whereas Logit Difference yields 1.22×1.22\times.
    • Under GN corruption: Probability yields a ratio of 1.74×1.74\times, whereas Logit Difference yields 0.77×0.77\times.

    Probability highlights the last subject token as the primary computational hub, whereas logit difference reveals that preceding subject tokens contribute comparable causal influence to output generation.

  7. Knowl 7 — IOI Circuit Attention Head Detections Across Activation Patching Configurations

    data/table

    In GPT-2 small on the Indirect Object Identification (IOI) task, attention head outputs were patched individually while corrupting the second subject token (S2S2). Detections were defined as heads whose patching effect fell at least 22 standard deviations from the mean effect across all heads.

    Corruption Metric NM DT SI Negative NM Induction
    STR Probability 1/3 0/2 3/4 1/2 1/2
    GN Probability 0/3 1/2 2/4 2/2 1/2
    STR Logit difference 1/3 0/2 3/4 2/2 1/2
    GN Logit difference 1/3 1/2 3/4 2/2 1/2
    STR KL divergence 1/3 0/2 3/4 2/2 1/2
    GN KL divergence 0/3 0/2 2/4 2/2 1/2

    Across head classes:

    • Name Mover (NM) (heads 9.6, 9.9, 10.0): S2 corruption misses at least 2 out of 3 NM heads across all settings.
    • Duplicate Token (DT) (heads 0.1, 0.10): Under GN, head 0.10 is incorrectly flagged as a negative head under Probability and KL divergence.
    • S-Inhibition (SI) (heads 7.9, 8.6, 8.10, 9.9): STR consistently recovers 3 out of 4 SI heads across metrics, while GN recovers only 2 out of 4 under Probability and KL.
    • Negative NM (heads 10.7, 11.10): STR Probability misses head 11.10 due to probability bounding.
  8. Knowl 8 — Circuit Discovery Sensitivity to the Choice of Corrupted Prompt Tokens

    empirical result

    In GPT-2 small on Indirect Object Identification (IOI), the specific tokens selected for corruption determine which sub-circuits activation patching uncovers:

    1. Corrupting S2: Replacing S2 with IO (in STR) or adding noise to S2 (in GN) isolates duplicate detection and name suppression mechanisms, successfully identifying S-Inhibition heads (7.9,8.6,8.107.9, 8.6, 8.10) but failing to detect Name Mover heads 9.69.6 and 10.010.0.
    2. Corrupting S1 and IO: Replacing both S1S1 and IO with random names (in STR) or adding noise to their embeddings (in GN) shifts the intervention to the name extraction and writing pipeline. This configuration discovers all three Name Mover heads (9.6,9.9,10.09.6, 9.9, 10.0) and both Negative Name Mover heads (10.7,11.1010.7, 11.10), but completely misses the S-Inhibition heads.

    Corrupting different tokens forces activation patching to trace distinct computational pathways in the model's graph.

  9. Knowl 9 — Artifacts of Relative Difference Metrics in Arithmetic Reasoning

    empirical result

    In GPT-J (6B) on 2-shot arithmetic prompts (X1+Y1=Z1,X2+Y2=Z2,X3+Y3=X_1 + Y_1 = Z_1, X_2 + Y_2 = Z_2, X_3 + Y_3 =), patching single MLP layers at the last token position using the relative change metric:

    Effect=12[Ppt(r)−P∗(r)P∗(r)+P∗(r′)−Ppt(r′)Ppt(r′)]\text{Effect} = \frac{1}{2} \left[ \frac{P_{\text{pt}}(r) - P_{*}(r)}{P_{*}(r)} + \frac{P_{*}(r') - P_{\text{pt}}(r')}{P_{\text{pt}}(r')} \right]

    produces an artificially exaggerated localization peak under Symmetric Token Replacement (STR) where values reach into the thousands. This occurs because STR replaces integer operands with random integers, making the baseline probability of the clean answer in the corrupted run P∗(r)P_*(r) negligible. The tiny denominator P∗(r)P_*(r) inflates small absolute probability shifts into large multipliers.

    When evaluated with standard probability or logit difference:

    • For addition and subtraction, STR still yields up to 4×4\times sharper concentration than Gaussian Noising (GN).
    • For multiplication, STR and GN produce nearly identical localization curves.
  10. Knowl 10 — Methodological Recommendations for Activation Patching in Language Models

    model/method

    Based on empirical comparisons across factual recall, arithmetic, indirect object identification, docstring completion, and greater-than tasks, activation patching should follow four methodological practices:

    1. Corruption Method: Prefer Symmetric Token Replacement (STR) over Gaussian Noising (GN). STR generates counterfactual prompts that remain within the training distribution, preventing out-of-distribution internal breakdowns and spurious head detections. GN should only be used when token length alignment or lack of valid semantic substitutes precludes STR.
    2. Evaluation Metric: Prefer Normalized Logit Difference (LD(r,r′)=Logit(r)−Logit(r′)\text{LD}(r, r') = \text{Logit}(r) - \text{Logit}(r')) over output Probability. Logit difference successfully identifies negative components and controls for modules that non-specifically boost broad token classes (e.g., all names).
    3. Layer Granularity: Evaluate single-layer interventions before applying multi-layer sliding window patching. Multi-layer patching captures joint non-linear interactions across layers that inflate peak magnitudes.
    4. Multi-Token Corruption: Systematically vary which prompt tokens are corrupted when tasks contain multiple informative tokens, as different corrupted tokens reveal complementary sub-circuits.

Coverage note — Specific per-head localization heatmaps for Python docstring completion and the greater-than task were omitted as separate knowls because their core findings (GN noisiness vs STR localization and sensitivity to metrics) are fully covered within the broader knowls.

References

  1. 1.Hritik Bansal, Karthik Gopalakrishnan, Saket Dingliwal, Sravan Bodapati, Katrin Kirchhoff, and Dan Roth. Rethinking the role of scale for in-context learning: An interpretability-based case study at 66 billion scale. In Annual Meeting of the Association for Computational Linguistics (ACL), 2023.
  2. 2.Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang. Hidden progress in deep learning: SGD learns parities near the computational limit. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  3. 3.Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html, 2023.
  4. 4.Davis Brown, Nikhil Vyas, and Yamini Bansal. On privileged and convergent bases in neural network representations. arXiv preprint arXiv:2307.12941, 2023.
  5. 5.Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim. Thread: Circuits. Distill, 5(3):e24, 2020.
  6. 6.Stephen Casper, Tilman Rauker, Anson Ho, and Dylan Hadfield-Menell. Toward transparent AI: A survey on interpreting the inner structures of deep neural networks. In IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2022.
  7. 7.Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldwosky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas. Causal scrubbing, a method for rigorously testing interpretability hypotheses. AI Alignment Forum, 2022. https://www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing.
  8. 8.Bilal Chughtai, Lawrence Chan, and Neel Nanda. A toy model of universality: Reverse engineering how networks learn group operations. In International Conference on Machine Learning (ICML), 2023.
  9. 9.Arthur Conmy, Augustine N Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  10. 10.Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023.
  11. 11.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. Knowledge neurons in pretrained transformers. In Annual Meeting of the Association for Computational Linguistics (ACL), 2022.
  12. 12.Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. Analyzing transformers in embedding space. In Annual Meeting of the Association for Computational Linguistics (ACL), 2023.
  13. 13.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021. https://transformer-circuits.pub/2021/framework/index.html.
  14. 14.Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M Shieber, Tal Linzen, and Yonatan Belinkov. Causal analysis of syntactic agreement mechanisms in neural language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), 2021.
  15. 15.Atticus Geiger, Kyle Richardson, and Christopher Potts. Neural natural language inference models partially embed theories of lexical entailment and negation. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, 2020.
  16. 16.Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  17. 17.Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts. Inducing causal structure for interpretable neural networks. In International Conference on Machine Learning (ICML), 2022.
  18. 18.Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D Goodman. Finding alignments between interpretable causal variables and distributed neural representations. arXiv preprint arXiv:2303.02536, 2023.
  19. 19.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021.
  20. 20.Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022.
  21. 21.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767, 2023.
  22. 22.Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora. Localizing model behavior with path patching. arXiv preprint arXiv:2304.05969, 2023.
  23. 23.Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. Finding neurons in a haystack: Case studies with sparse probing. arXiv preprint arXiv:2305.01610, 2023.
  24. 24.Michael Hanna, Ollie Liu, and Alexandre Variengien. How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  25. 25.Peter Hase, Harry Xie, and Mohit Bansal. The out-of-distribution problem in explainability and search methods for feature importance explanations. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  26. 26.Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. Does localization inform editing? Surprising differences in causality-based localization vs. knowledge editing in language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  27. 27.Stefan Heimersheim and Jett Janiak. A circuit for Python docstrings in a 4-layer attention-only transformer. https://www.alignmentforum.org/posts/u6KXXmKFbXfWzoAXn/a-circuit-for-python-docstrings-in-a-4-layer-attention-only, 2023.
  28. 28.Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas. Natural language descriptions of deep visual features. In International Conference on Learning Representations (ICLR), 2021.
  29. 29.Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. A benchmark for interpretability methods in deep neural networks. In Advances in neural information processing systems (NeurIPS), 2019.
  30. 30.Dominik Janzing, Lenon Minorics, and Patrick Blöbaum. Feature relevance quantification in explainable ai: A causal problem. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2020.
  31. 31.Shahar Katz and Yonatan Belinkov. Interpreting transformer’s attention dynamic memory and visualizing the semantic information flow of GPT. arXiv preprint arXiv:2305.13417, 2023.
  32. 32.Michael A Lepori, Ellie Pavlick, and Thomas Serre. NeuroSurgeon: A toolkit for subnetwork analysis. arXiv preprint arXiv:2309.00244, 2023.
  33. 33.Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. In International Conference on Learning Representations (ICLR), 2023a.
  34. 34.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. In Advances in Neural Information Processing Systems (NeurIPS), 2023b.
  35. 35.Yuchen Li, Yuanzhi Li, and Andrej Risteski. How do transformers learn topic structure: Towards a mechanistic understanding. In International Conference on Machine Learning (ICML), 2023c.
  36. 36.Tom Lieberum, Matthew Rahtz, János Kramár, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik. Does circuit analysis interpretability scale? Evidence from multiple choice capabilities in Chinchilla. arXiv preprint arXiv:2307.09458, 2023.
  37. 37.Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. The hydra effect: Emergent self-repair in language model computations. arXiv preprint arXiv:2307.15771, 2023.
  38. 38.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  39. 39.Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. Language models implement simple word2vec-style vector arithmetic. arXiv preprint arXiv:2305.16130, 2023.
  40. 40.Jesse Mu and Jacob Andreas. Compositional explanations of neurons. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  41. 41.Neel Nanda and Joseph Bloom. TransformerLens. https://github.com/neelnanda-io/TransformerLens, 2022.
  42. 42.Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. In International Conference on Learning Representations (ICLR), 2023a.
  43. 43.Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023b.
  44. 44.Chris Olah. Mechanistic interpretability, variables, and the importance of interpretable bases. https://transformer-circuits.pub/2022/mech-interp-essay/index.html, 2022.
  45. 45.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. In-context learning and induction heads. Transformer Circuits Thread, 2022. https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html.
  46. 46.Judea Pearl. Direct and indirect effects. In Conference on Uncertainty and Artificial Intelligence (UAI), 2001.
  47. 47.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 2019.
  48. 48.Adam Scherlis, Kshitij Sachan, Adam S Jermyn, Joe Benton, and Buck Shlegeris. Polysemanticity and capacity in neural networks. arXiv preprint arXiv:2210.01892, 2022.
  49. 49.Paul Soulos, R Thomas McCoy, Tal Linzen, and Paul Smolensky. Discovering the compositional structure of vector representations with role learning networks. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, 2020.
  50. 50.Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan. Understanding arithmetic reasoning in language models using causal mediation analysis. arXiv preprint arXiv:2305.15054, 2023.
  51. 51.Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar. Explaining grokking through circuit efficiency. arXiv preprint arXiv:2309.02390, 2023.
  52. 52.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  53. 53.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  54. 54.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  55. 55.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In International Conference on Learning Representations (ICLR), 2023.
  56. 56.Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski. (Un)interpretability of transformers: a case study with Dyck grammars. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  57. 57.Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah D Goodman. Interpretability at scale: Identifying causal mechanisms in alpaca. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  58. 58.Zeyu Yun, Yubei Chen, Bruno Olshausen, and Yann LeCun. Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors. In Proceedings of Deep Learning Inside Out (DeeLIO): The 2nd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, 2021.
  59. 59.Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas. The clock and the pizza: Two stories in mechanistic explanation of neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2023.

Citation

MLA
Zhang, F., and N. Nanda. “Towards Best Practices of Activation Patching in Language Models: Metrics and Methods”. arXiv, 2023, http://arxiv.org/abs/2309.16042v2.
APA
Zhang, F., & Nanda, N. (2023). Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. arXiv. http://arxiv.org/abs/2309.16042v2
Chicago
Zhang, F., and N. Nanda. 2023. “Towards Best Practices of Activation Patching in Language Models: Metrics and Methods”. arXiv. http://arxiv.org/abs/2309.16042v2.
Harvard
Zhang, F. and Nanda, N. (2023) “Towards Best Practices of Activation Patching in Language Models: Metrics and Methods”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.16042v2.
Vancouver
1. Zhang F, Nanda N (2023) Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. arXiv

BibTeX

@article{zhang2023towards,
  title = {Towards Best Practices of Activation Patching in Language Models: Metrics and Methods},
  author = {Zhang, Fred and Nanda, Neel},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.16042v2},
  eprint = {2309.16042}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors