Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Fred ZhangNeel Nanda
Demonstrates how arbitrary choices in corruption methods and evaluation metrics lead to conflicting activation patching results, establishing concrete best practices for reliable circuit discovery and localization in language models.
As large language models see wider deployment in critical systems, understanding their inner workings has become essential for debugging errors, preventing unwanted behavior, and ensuring safety. A foundational tool for this work is activation patching—an intervention technique designed to pinpoint which internal model components causally drive specific outputs. Despite its widespread use across the AI research community, there is no standardized protocol for how to corrupt baseline inputs or measure intervention effects. The article systematically evaluates how these implementation choices affect interpretability results, demonstrating that common practices can yield conflicting conclusions about how models operate.
The researchers conducted systematic experiments across multiple models (including GPT-2 variants and the 6-billion parameter GPT-J) and several benchmark tasks, such as factual recall, indirect object identification, Python code completion, and multi-digit arithmetic. They compared two primary methods for corrupting prompts—adding random Gaussian noise versus symmetric token replacement using natural counterfactual words—and evaluated different scoring metrics, notably output probability versus logit difference (the difference in raw model output scores between correct and incorrect answers). They also analyzed sliding window patching, which intervenes across multiple adjacent layers simultaneously.
The findings show that methodological choices substantially alter interpretability outcomes. First, Gaussian noise corruption frequently pushes the model outside its normal data distribution, disrupting internal mechanisms and producing noisy or contradictory findings; for instance, in factual recall, Gaussian noise created an apparent peak in early-to-middle layers that was two to five times higher than when using clean token replacement. Second, the probability metric fails to identify "negative" components that actively harm performance whenever corruption drives the baseline probability close to zero, whereas logit difference reliably captures both positive and negative components. Third, sliding window patching produces maximum peak effects at least 20 percent higher than summing individual layer interventions, artificially amplifying weak single-layer signals due to non-linear multi-layer interactions. Finally, the choice of which specific token positions to corrupt strongly dictates which internal circuits are detected.
These results carry direct implications for AI governance, model editing, and safety engineering. When practitioners rely on flawed interpretability hyperparameters, they risk editing the wrong model weights or drawing false conclusions about model reliability and compliance. To establish robust best practices, the article recommends using symmetric token replacement over random noise to maintain valid data distributions, adopting logit difference rather than raw probabilities as the primary evaluation metric, testing single-layer interventions before multi-layer windows, and varying the target tokens corrupted to ensure complete circuit discovery.
The findings are supported by consistent results across several diverse model architectures and reasoning tasks. However, users should note that the analysis focused on decoder-only language models up to 6 billion parameters and examined overriding corrupted activations with clean ones, rather than the reverse direction. Decision-makers should treat localization findings derived purely from Gaussian noise or multi-layer window patching with caution until validated against standardized, in-distribution methods.
- Paper: Towards Automated Circuit Discovery for Mechanistic Interpretability, Arthur Conmy et al. (2023). Introduces automated edge-level circuit discovery via activation patching in transformers, establishing the fundamental interpretability framework and patching paradigms that the source paper systematically audits and refines.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). Develops top-down representation reading and causal steering interventions in language models, providing crucial conceptual grounding for interpreting internal activations.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). Applies causal mediation and activation patching on task-specific language model representations, illustrating the standard localization methodologies whose metric and corruption choices are analyzed in the source.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). Demonstrates linear probing and internal activation intervention across attention heads, providing key context on how causal interventions are used to identify and manipulate functional circuits.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). Synthesizes broader theoretical pitfalls and non-identifiability challenges across mechanistic interpretability methods, contextualizing the empirical fragility and metric sensitivity uncovered in the source paper.
- Paper: RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations, Jing Huang et al. (2024). Introduces a standardized benchmark using counterfactual interchange interventions to quantitatively evaluate representation disentanglement, building directly upon causal intervention best practices.
- Paper: How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models, Dun Li Chan et al. (2026). Applies multi-level internal tracking and activation patching to study how various input perturbations propagate through language model hidden states and attention heads.
- Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). Employs mechanistic interpretability and activation analysis to examine whether fine-tuning rewires underlying circuits or merely adds superficial output transformations.
- Paper: What Makes and Breaks Safety Fine-tuning? A Mechanistic Study, Samyak Jain et al. (2024). Uses mechanistic representation analyses to isolate localized internal weight changes and explain how safety alignment alters internal feature processing.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). Leverages internal layer representations and attention traces to construct lightweight self-evaluation mechanisms that predict model failure modes.
