Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Kevin WangAlexandre VariengienArthur ConmyBuck ShlegerisJacob Steinhardt
Reverse-engineers the 26-attention-head circuit that GPT-2 small uses for indirect object identification, establishing quantitative criteria to validate mechanistic interpretability in transformer language models.
Modern transformer language models are deployed across critical commercial and public systems, yet their internal decision-making processes remain largely opaque black boxes. This opacity introduces operational, safety, and compliance risks because users cannot reliably anticipate or audit out-of-distribution behaviors and systematic failures. The article aims to demonstrate that mechanistic interpretability can reverse-engineer how a model performs a specific, natural language task by isolating the human-understandable algorithmic "circuit" embedded within its parameters.
To achieve this, the authors investigated GPT-2 small on an indirect object identification task, which requires predicting the indirect object in sentences with duplicated subjects (such as predicting "Mary" in "When Mary and John went to the store, John gave a drink to"). The analysis used causal intervention techniques, primarily "path patching" and mean-activation ablations, across 100,000 synthetic template samples and repeated token sequences. By systematically tracing pathways backward from the output logits, the authors evaluated how individual attention heads read, process, and write information across residual streams.
Key findings reveal that a sparse subnetwork of 26 attention heads—just 1.1% of the model's total head-position pairs—performs the core task across seven functional roles. Primary "Name Mover Heads" copy the target name to the output, while "S-Inhibition Heads" suppress attention to duplicated subjects using both positional and token signals. Furthermore, the analysis uncovered unexpected internal dynamics: "Backup Name Mover Heads" automatically compensate and restore performance when primary heads are knocked out (resulting in only a 5% drop in logit difference), while "Negative Name Mover Heads" systematically suppress confidence, likely to hedge against high loss on incorrect predictions. In quantitative evaluations, the isolated circuit achieved 87% of the full model's task performance.
These findings prove that complex linguistic capabilities in language models decompose into modular, understandable algorithms rather than uninterpretable statistical noise. However, the presence of redundant backup pathways and hedging heads highlights that standard ablation tests can be misleading. Using these mechanistic insights, the authors designed targeted adversarial examples: introducing an extra duplicate of the indirect object caused the model to fail and choose the incorrect subject 23.4% of the time, compared to an error rate of just 0.7% on standard inputs.
For practitioners and researchers, the authors recommend incorporating structured circuit-evaluation criteria—specifically faithfulness, completeness, and minimality—when validating internal model explanations. Organizations building safety and assurance audits should leverage circuit discovery to craft targeted adversarial tests rather than relying solely on surface-level evaluations. Next steps require extending these causal techniques to larger, modern models and analyzing the role of multilayer perceptrons and layer normalization, which were excluded from this head-level analysis.
No sufficiently relevant recommendations were found.
- Paper: Towards Automated Circuit Discovery for Mechanistic Interpretability, Arthur Conmy et al. (2023). Building on manually mapped circuits such as the IOI circuit, this paper automates edge-level circuit discovery and tests recovery against published ground truths.
- Paper: Towards Best Practices of Activation Patching in Language Models: Metrics and Methods, Fred Zhang et al. (2024). It turns the source’s causal patching approach into a methodological evaluation, showing how corruption choices and scoring metrics can change circuit conclusions.
- Paper: Information Flow Routes: Automatically Interpreting Language Models at Scale, Javier Ferrando et al. (2024). It extends circuit tracing to faster, automated attribution-based routes and evaluates them on the same indirect-object-identification benchmark.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). It critiques mechanistic circuit explanations through the problem of redundant pathways, directly challenging how findings like backup heads should be interpreted.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). It applies causal interventions to IOI while moving the analysis from attention-head circuits toward interpretable features learned by sparse autoencoders.
