MIB: A Mechanistic Interpretability Benchmark

Aaron MuellerAtticus GeigerSarah WiegreffeDana AradIván ArcuschinAdam BelfkiYik Siu ChanJaden Fiotto-KaufmanTal HaklayMichael Hanna

article2025ICML55 citations

Establishes a standardized benchmark to rigorously evaluate mechanistic interpretability methods across circuit and causal variable localization, revealing critical performance differences among popular techniques like sparse autoencoders and attribution patching.

Listen

Mechanistic interpretability aims to reverse-engineer neural language models by mapping their internal representations to causal mechanisms and human-understandable concepts. Despite rapid growth in proposed interpretability techniques, the field lacks standardized evaluation practices. New methods are frequently assessed using customized metrics across narrow, ad-hoc tasks, making it difficult to determine whether newer techniques offer genuine advancements over existing baselines. Establishing standard evaluation benchmarks is essential for verifying whether interpretability tools can reliably assist in auditing, safety evaluations, and fine-grained control of production models.

The article introduces the Mechanistic Interpretability Benchmark (MIB) to establish a rigorous, standardized framework for evaluating interpretability methods. The benchmark evaluates these techniques across two distinct tracks: circuit localization, which tests how effectively methods locate sparse subgraphs of model components that drive specific behaviors, and causal variable localization, which assesses how well methods isolate features within internal activations and align them with high-level conceptual variables. The authors evaluated four open-weight language models of varying scales—ranging from GPT-2 Small to Llama-3.1 8B—across four diverse tasks: Indirect Object Identification, multi-digit Arithmetic, synthetic Multiple-Choice Question Answering, and the grade-school science AI2 Reasoning Challenge, utilizing standardized counterfactual input datasets and public leaderboards with private evaluation splits.

Key findings show significant differences in performance among leading interpretability techniques. For circuit localization, edge-level attribution patching using integrated gradients over inputs (EAP-IG-inputs) combined with counterfactual ablations consistently yielded the highest faithfulness across models and tasks, balancing localization accuracy and computational cost. Conversely, node-level methods performed poorly because they fail to achieve necessary circuit sparsity, and exact activation patching suffered from prohibitive computational costs without providing superior accuracy. For causal variable localization, the supervised Distributed Alignment Search (DAS) consistently outperformed unsupervised featurization techniques. Furthermore, popular unsupervised Sparse Autoencoders (SAEs) failed to identify features that outperformed standard internal vector dimensions (neurons) when aligning to task-relevant concepts, challenging current assumptions regarding SAE superiority in concept extraction.

These findings provide clarity on the actual capabilities of mechanistic interpretability tools, indicating that supervised methods and gradient-based attribution offer the most reliable pathways for model inspection. The relative underperformance of Sparse Autoencoders highlights a critical gap between unsupervised representation extraction and causally validated concept mapping, cautioning organizations against relying solely on SAE-based pipelines for auditing safety-critical behaviors without explicit causal validation.

Practitioners should prioritize integrated gradient edge-attribution methods when isolating model subgraphs and utilize supervised subspace methods like DAS for targeted concept probing. Future research should focus on refining non-linear and unsupervised feature extraction techniques to improve causal alignment, addressing computational bottlenecks in large model evaluations, and developing methods to jointly discover features and circuit structures. Confidence in the relative rankings of evaluated baselines is high across the standardized tasks, though readers should exercise caution when extrapolating these findings to unstudied tasks, complex multi-step reasoning, or models that exhibit low baseline task performance.

No sufficiently relevant recommendations were found.

Cover for MIB: A Mechanistic Interpretability Benchmark

Abstract

How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components - and connections between them - most important for performing a task (e.g., attribution patching or information flow routes). The causal variable localization track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAEs) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAE features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field.

Table of Contents

  • 1 Introduction
  • 2 Materials
  • 2.1 Tasks
  • 2.2 Counterfactual Inputs
  • 2.3 Models
  • 2.4 Leaderboard
  • 3 Circuit Localization Track
  • 3.1 Circuit Metrics
  • 3.2 Circuit Localization Baselines
  • 3.3 Results
  • 4 Causal Variable Localization Track
  • 4.1 Causal Abstraction
  • 4.2 Causal Variable Localization Baselines
  • 4.3 MCQA and ARC (Easy)
  • 4.4 Two-Digit Addition
  • 4.5 RAVEL
  • 4.6 Indirect Object Identification.
  • 4.7 General Discussion
  • 5 Related Work
  • 6 Conclusions
  • References
  • A Table of Notation
  • B Limitations
  • C Further Details on Materials
  • C.1 Indirect Object Identification (IOI)
  • C.2 Arithmetic
  • C.3 Multiple-choice question answering (MCQA)
  • C.4 AI2 Reasoning Challenge (ARC)
  • C.5 Model Performance
  • C.6 Leaderboard
  • D Details on Circuit Localization Track
  • D.1 Methods
  • D.2 Baselines
  • D.3 Further Circuit Localization Results
  • E Details on InterpBench Model Training
  • F Details on Causal Variable Localization Track
  • F.1 Causal Abstraction Analysis
  • F.2 Aligning Unsupervised Features to Causal Variables
  • F.3 Distributed Alignment Search
  • F.4 Hyperparameters
  • F.5 High-level Causal Models and Experimental Details for Each Task
  • F.5.1 Multiple-choice Question Answering
  • F.5.2 Arithmetic
  • F.5.3 RAVEL
  • F.5.4 Indirect Object Indentification

Knowls

  1. Knowl 1 — MIB standardizes evaluation across circuit and causal-variable localization

    experimental setup

    MIB evaluates mechanistic-interpretability methods in two tracks: circuit localization, which ranks model components and connections by their task relevance, and causal-variable localization, which tests whether features of hidden activations implement variables in a task-level causal model. The four shared tasks are indirect object identification (IOI), two-operand arithmetic, synthetic multiple-choice question answering (MCQA), and the Easy and Challenge subsets of the AI2 Reasoning Challenge (ARC); the causal-variable track also evaluates RAVEL city-attribute queries. Four open-weight language models are tested: GPT-2 Small (117M), Qwen-2.5 0.5B, Gemma-2 2B, and Llama-3.1 8B. An additional controlled IOI transformer with a known circuit supports ground-truth edge evaluation. The released data include training, validation, public-test, and private-test splits, with fixed instance-to-counterfactual mappings so methods are evaluated on the same interventions. For example, the split counts are 10,000/10,000/1,000/1,000 for IOI; 34,400/4,920/1,000/1,000 for arithmetic addition; 17,400/2,484/1,000/1,000 for subtraction; 110/50/50/50 for MCQA; and 2,251/570/1,188/1,188 for ARC Easy, where each sequence is train/validation/public test/private test. The private test supports held-out comparisons; IOI private examples also use names and direct objects absent from the public data.

  2. Knowl 2 — Circuit faithfulness is evaluated by two complementary size-integrated metrics

    equation

    For a circuit CC in a neural network NN, define its faithfulness as f(C,N;m)=m(C)−m(∅)m(N)−m(∅)f(C,N;m)=\frac{m(C)-m(\emptyset)}{m(N)-m(\emptyset)}. Here, mm is the task logit difference between the answer correct for the original input and the answer correct for its paired counterfactual; m(C)m(C) is evaluated with the circuit retained and other components ablated, m(N)m(N) is the full-model value, and m(∅)m(\emptyset) is the value when all components are ablated. The integrated circuit performance ratio (CPR) rewards circuits that recover positive task performance, whereas circuit-model distance (CMD) measures deviation from the full model, including deviations caused by harmful components. For a circuit-localization method that produces a circuit CkC_k at edge fraction k∈[0,1]k\in[0,1], the idealized metrics are

    CPR=∫01f(Ck,N;m) dk,CMD=∫01∣1−f(Ck,N;m)∣ dk.\mathrm{CPR}=\int_0^1 f(C_k,N;m)\,dk,\qquad \mathrm{CMD}=\int_0^1 |1-f(C_k,N;m)|\,dk.

    Higher CPR and lower CMD are preferred. The two metrics capture different goals: finding components that improve task performance versus finding components whose combined behavior matches the full model.

  3. Knowl 3 — Circuit size uses neuron-weighted edges and is evaluated over a range of sparsities

    model/method

    MIB measures circuit size by a weighted edge count so that circuits defined at different component granularities can be compared. Including a node is treated as including all its outgoing edges; including only some neurons of a source node gives its outgoing edges a proportional cost. Specifically,

    ∣C∣=∑(u,v)∈C∣Nu∩NC∣∣Nu∣,|C|=\sum_{(u,v)\in C}\frac{|N_u\cap N_C|}{|N_u|},

    where CC is a set of computation-graph edges, uu and vv are source and destination nodes, NuN_u is the set of neurons in source node uu, and NCN_C is the set of neurons included in the circuit. The count is normalized by the number of possible edges to obtain a circuit-size proportion. To approximate CPR and CMD, MIB evaluates circuits whose size is at most each of k∈{0.001,0.002,0.005,0.01,0.02,0.05,0.1,0.2,0.5,1}k\in\{0.001,0.002,0.005,0.01,0.02,0.05,0.1,0.2,0.5,1\}, computes faithfulness at each size, and applies the trapezoidal rule to the resulting curve. A separate controlled IOI transformer, trained to implement a known circuit, permits AUROC over edges because its ground-truth edge labels are known; that model has 6 layers, 4 heads per layer, dmodel=64d_{\mathrm{model}}=64, and dhead=16d_{\mathrm{head}}=16.

  4. Knowl 4 — Circuit-localization baselines span causal attribution, pruning, and non-causal routes

    model/method

    MIB compares exact edge activation patching (EActP), attribution patching (AP), integrated-gradient AP variants, information-flow routes (IFR), and a learned pruning-mask baseline. EActP estimates each edge’s indirect effect by intervening on that edge, at a cost of O(n)O(n) forward passes for nn possible edges. Edge AP (EAP) approximates an edge effect by multiplying the source activation change between original and counterfactual inputs by the gradient of the task logit metric with respect to the destination input; it requires O(1)O(1) forward passes. Integrated-gradient AP averages such gradients along an interpolation path: the input-interpolation variant (EAP-IG-inputs) interpolates input embeddings, while the activation-interpolation variant (EAP-IG-activations) interpolates target activations separately by layer. The benchmark uses Z=5Z=5 interpolation steps; the stated costs are O(Z)O(Z) forward passes for input interpolation and O(ZL)O(ZL) for activation interpolation in a model with LL layers. Node-level versions of attribution patching are also tested. IFR assigns connection scores from similarity between a source node’s output and a destination node’s input, without a counterfactual causal intervention. Uniform Gradient Sampling (UGS) learns edge-mask probabilities to preserve model behavior while encouraging sparsity; its sparsity-loss weight is 10−310^{-3}, selected from {10−2,10−3,…,10−7}\{10^{-2},10^{-3},\ldots,10^{-7}\} using validation data. Circuit scores are converted to circuits using top-ranked edges, or greedy path construction where needed; high scores are used for CPR and absolute score magnitudes for CMD.

  5. Knowl 5 — Attribution with input integrated gradients is the strongest circuit-localization baseline overall

    empirical result

    Across public and private test sets, EAP-IG-inputs generally achieves the strongest average CPR and CMD results among the tested circuit-localization methods. Its public-test CMD is 0.01 on Llama-3.1 IOI and 0.00 on Llama-3.1 arithmetic; corresponding CPR values are 2.08 and 0.99. It is not best in every individual setting: exact EActP leads on GPT-2 IOI, but not on Qwen-2.5 IOI or the known-circuit InterpBench model. On InterpBench, EAP-IG-activations has the highest reported AUROC, 0.81, compared with 0.71 for EAP-IG-inputs. Counterfactual-ablation circuits generally outperform circuits found with mean or learned optimal ablations. Node-level circuits are often weak, consistent with their larger edge cost at a given circuit size; IFR usually beats random circuits but tends to trail attribution methods. EAP-IG-activations and UGS remain competitive on CMD in some settings, although UGS is less competitive on CPR, which is consistent with its objective of preserving model behavior rather than selecting only positively impactful components. The broad pattern persists on the private test set.

  6. Knowl 6 — Causal-variable localization tests aligned interchange interventions

    definition

    A causal-variable submission specifies a task-level causal model HH, a hidden activation vector h∈Rdh\in\mathbb{R}^d from a language model NN, an invertible featurizer FF that maps hh into a feature space, and a selected set of feature indices ΠX\Pi_X aligned with a causal variable XX of HH. For a base input bb and paired counterfactual input cc, the high-level interchange intervention runs HH on bb while setting XX to the value it takes on cc. The aligned low-level intervention runs NN on bb while setting features ΠX\Pi_X to their values on cc. Interchange intervention accuracy (IIA) is the fraction of tested input pairs for which the two interventions produce the same output:

    IIA(X,ΠX,H,D)=1∣D∣∑(b,c)∈D1 ⁣[HX←X(c)(b)=NΠX←ΠX(c)(b)],\mathrm{IIA}(X,\Pi_X,H,D)=\frac{1}{|D|}\sum_{(b,c)\in D}\mathbf{1}\!\left[H_{X\leftarrow X(c)}(b)=N_{\Pi_X\leftarrow\Pi_X(c)}(b)\right],

    where DD is the set of base-counterfactual pairs and 1[⋅]\mathbf{1}[\cdot] equals 1 when the outputs agree and 0 otherwise. The metric is used for MCQA, ARC, arithmetic, and RAVEL. IOI instead uses mean-squared error between the causal model’s predicted logit difference and the language model’s logit difference. Evaluations exclude examples where the model predicts incorrectly on the base input or any counterfactual used.

  7. Knowl 7 — Causal-variable baselines distinguish unsupervised features from supervised alignment

    model/method

    MIB compares five ways to obtain features and align them with task-level variables. The Full Vector baseline uses the identity featurizer and intervenes on the entire hidden vector. PCA and sparse autoencoders (SAEs) create unsupervised feature spaces; GemmaScope and LlamaScope SAEs are used for the corresponding models. Differential Binary Masking (DBM) supplies supervision only for feature selection: it learns a binary mask over standard dimensions, PCA components, or SAE features to maximize causal-variable faithfulness on training data. SAE interventions retain the base activation’s reconstruction error so that the intervention remains compatible with the reconstruction-based featurizer. Distributed Alignment Search (DAS) is supervised featurization: it learns orthogonal directions in activation space directly from causal-model interchange interventions, with no separate feature-selection stage. MIB searches candidate token locations and layers for non-IOI tasks, evaluates alignment quality across layers, and reports both layer averages and the best layer. DAS dimensionalities are 16 for the MCQA/ARC ordering variable and arithmetic carry variable, 32 for IOI subject-token and subject-position variables, half the residual-stream dimension for MCQA/ARC answer-token variables, and one eighth of the residual-stream dimension for RAVEL.

  8. Knowl 8 — Task-level causal hypotheses specify what each localized variable should do

    model/method

    For MCQA and ARC, the hypothesized computation separates determining the correct option’s position (XOrderX_{\mathrm{Order}}) from retrieving the answer token at that position (OAnswerO_{\mathrm{Answer}}). Counterfactuals that change option order or labels distinguish these variables: intervening on the position should select the base prompt’s token at the counterfactual position, whereas intervening on the answer token should transfer the counterfactual answer token. For two-digit addition, the causal model parses ones and tens digits, computes the ones output and a carry variable XCarryX_{\mathrm{Carry}}, and uses the carry with the tens digits to compute higher output digits. The benchmark targets XCarryX_{\mathrm{Carry}} and includes both random counterfactuals and pairs designed to change the carry while holding parts of the input and output fixed. For RAVEL, a prompt identifies a city and queried attribute; separate country, continent, and language variables feed an output mechanism that selects the queried attribute. For IOI, the high-level model extracts the subject token (STokS_{\mathrm{Tok}}) and its position (SPosS_{\mathrm{Pos}}), then predicts the subject-versus-indirect-object logit difference from whether token identity or position has been inverted. On the curated data, the fitted logit-difference model is 0.048+0.768 TokenSignal+2.005 PositionSignal0.048+0.768\,\mathrm{TokenSignal}+2.005\,\mathrm{PositionSignal}, where each signal is +1+1 when unchanged and −1-1 when inverted. IOI localization targets these two variables in the four previously identified S-inhibition attention heads.

  9. Knowl 9 — DAS usually outperforms unsupervised features, while carry localization remains weak

    empirical result

    DAS has the strongest layer-averaged causal-variable results in most tested settings, and DBM on ordinary hidden dimensions is a stronger unsupervised baseline than DBM on PCA or SAE features in general. In MCQA, DAS reaches mean IIA of 95% for the answer-token variable and 77% for answer position on Gemma-2; on Llama-3.1 the corresponding means are 94% and 77%. For ARC Easy, DAS means are 88% and 76% for those variables on Gemma-2, and 88% and 74% on Llama-3.1. On RAVEL, DAS means for continent, country, and language are 75%, 57%, and 62% on Gemma-2 and 75%, 58%, and 63% on Llama-3.1. For IOI, DAS gives the lowest reported mean-squared errors: 2.20 for subject position and 2.08 for subject token, versus 2.45 and 2.82 for Full Vector. Two-digit carry localization is much weaker: mean IIA is 31% on Gemma-2 and 54% on Llama-3.1 with DAS, compared with 29% and 35% for Full Vector. The authors note that the apparent carry signal in Llama-3.1 can reflect correlated tens-digit output; performance fails with random counterfactuals. SAE features generally do not improve on standard hidden dimensions with DBM, and are often worse. For example, MCQA DAS achieves 95%/77% mean IIA on Gemma-2, while DBM-on-SAE achieves 73%/51%; the Full Vector baseline can have perfect best-layer performance for an answer-token variable but substantially lower layer-average scores, indicating that whole-vector interventions often fail to isolate a variable consistently.

  10. Knowl 10 — Benchmark conclusions are limited by causal-model assumptions and task coverage

    limitation

    MIB separates circuit discovery from feature construction for cleaner comparisons, although the two problems can interact: a feature decomposition could change which circuits are discoverable, and feature choice may depend on the downstream task. The causal-variable evaluations assume that the proposed high-level causal variables exist in the language model’s computation; the task-level causal model may be an imperfect account of the model even when its intervention scores are high. Circuit faithfulness is not a complete recovery measure on ordinary models: without ground-truth components, MIB cannot provide a general automated completeness score, so edge AUROC is available only for the controlled known-circuit model. Causal-variable evaluations also filter out model failures, and therefore do not measure whether the causal model explains incorrect behavior. Finally, the benchmark covers language models rather than other modalities, and its baseline hyperparameter searches were not exhaustive.

Coverage note — Detailed prompt templates and counterfactual examples, leaderboard submission mechanics, and per-method hyperparameter and runtime constraints beyond the main baseline comparisons are omitted because they support benchmark use but do not add load-bearing findings to the core contribution.

References

  1. 1.Abraham, E. D., D’Oosterlinck, K., Feder, A., Gat, Y. O., Geiger, A., Potts, C., Reichart, R., and Wu, Z. CE-Bab: Estimating the causal effects of real-world concepts on NLP model behavior. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=3AbigH4s-ml.
  2. 2.Amini, A., Pimentel, T., Meister, C., and Cotterell, R. Naturalistic causal probing for morpho-syntax. Transactions of the Association for Computational Linguistics, 11:384–403, 2023. doi: 10.1162/tacl_a_00554. URL https://aclanthology.org/2023.tacl-1.23/.
  3. 3.Arora, A., Jurafsky, D., and Potts, C. CausalGym: Benchmarking causal interpretability methods on linguistic tasks. In Ku, L., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 14638–14663. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL-LONG.785. URL https://doi.org/10.18653/v1/2024.acl-long.785.
  4. 4.Atanasova, P., Camburu, O.-M., Lioma, C., Lukasiewicz, T., Simonsen, J. G., and Augenstein, I. Faithfulness tests for natural language explanations. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 283–294, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-short.25. URL https://aclanthology.org/2023.acl-short.25/.
  5. 5.Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C. Towards monosemanticity: Decomposing language models with dictionary learning. In Transformer Circuits Thread, 2023. URL https://transformer-circuits.pub/2023/monosemantic-features/index.html.
  6. 6.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  7. 7.Cao, N. D., Schlichtkrull, M. S., Aziz, W., and Titov, I. How do decisions emerge across layers in neural models? Interpretation with differentiable masking. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 3243–3255, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.262. URL https://aclanthology.org/2020.emnlp-main.262/.
  8. 8.Cao, N. D., Schmid, L., Hupkes, D., and Titov, I. Sparse interventions in language models with differentiable masking. In Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, 2022. URL https://doi.org/10.18653/v1/2022.blackboxnlp-1.2.
  9. 9.Chan, L., Garriga-Alonso, A., Goldwosky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N. Causal scrubbing, a method for rigorously testing interpretability hypotheses. AI Alignment Forum, 2022. https://www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing.
  10. 10.Chaudhary, M. and Geiger, A. Evaluating open-source sparse autoencoders on disentangling factual knowledge in GPT-2 small. CoRR, abs/2409.04478, 2024. URL https://arxiv.org/abs/2409.04478.
  11. 11.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? Try ARC, the AI2 Reasoning Challenge. arXiv preprint arXiv:1803.05457, 2018. URL https://arxiv.org/abs/1803.05457.
  12. 12.Cohen, R., Biran, E., Yoran, O., Globerson, A., and Geva, M. Evaluating the ripple effects of knowledge editing in language models. Transactions of the Association for Computational Linguistics, 12:283–298, 2024. doi: 10.1162/tacl_a_00644. URL https://aclanthology.org/2024.tacl-1.16/.
  13. 13.Conmy, A., Mavor-Parker, A., Lynch, A., Heimersheim, S., and Garriga-Alonso, A. Towards automated circuit discovery for mechanistic interpretability. Advances in Neural Information Processing Systems, 36:16318–16352, 2023.
  14. 14.Csordas, R., van Steenkiste, S., and Schmidhuber, J. Are neural nets modular? Inspecting functional modularity through differentiable weight masks. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=7uVcpu-gMD.
  15. 15.Csordas, R., Potts, C., Manning, C. D., and Geiger, A. Recurrent neural networks learn to store and generate sequences using non-linear representations. In The 7th BlackboxNLP Workshop, 2024. URL https://openreview.net/forum?id=NUQeYgg8x4.
  16. 16.Dai, Q., Heinzerling, B., and Inui, K. Representational analysis of binding in language models. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 17468–17493, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.967. URL https://aclanthology.org/2024.emnlp-main.967/.
  17. 17.Davies, X., Nadeau, M., Prakash, N., Shaham, T. R., and Bau, D. Discovering variable binding circuitry with desiderata. CoRR, abs/2307.03637, 2023. doi: 10.48550/ARXIV.2307.03637. URL https://doi.org/10.48550/arXiv.2307.03637.
  18. 18.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The Llama 3 herd of models. CoRR, abs/2407.21783, 2024. URL https://arxiv.org/abs/2407.21783.
  19. 19.Ferrando, J. and Voita, E. Information flow routes: Automatically interpreting language models at scale. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 17432–17445, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.965. URL https://aclanthology.org/2024.emnlp-main.965/.
  20. 20.Ferrando, J., Sarti, G., Bisazza, A., and Costa-jussa, M. R. A primer on the inner workings of transformer-based language models. CoRR, abs/2405.00208, 2024. doi: 10.48550/ARXIV.2405.00208. URL https://doi.org/10.48550/arXiv.2405.00208.
  21. 21.Finlayson, M., Mueller, A., Gehrmann, S., Shieber, S., Linzen, T., and Belinkov, Y. Causal analysis of syntactic agreement mechanisms in neural language models. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 1828–1843, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.144. URL https://aclanthology.org/2021.acl-long.144/.
  22. 22.Geiger, A., Richardson, K., and Potts, C. Neural natural language inference models partially embed theories of lexical entailment and negation. In Alishahi, A., Belinkov, Y., Chrupała, G., Hupkes, D., Pinter, Y., and Sajjad, H. (eds.), Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pp. 163–173, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.blackboxnlp-1.16. URL https://aclanthology.org/2020.blackboxnlp-1.16/.
  23. 23.Geiger, A., Lu, H., Icard, T., and Potts, C. Causal abstractions of neural networks. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 9574–9586, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/4f5c422f4d49a5a807eda27434231040-Abstract.html.
  24. 24.Geiger, A., Ibeling, D., Zur, A., Chaudhary, M., Chauhan, S., Huang, J., Arora, A., Wu, Z., Goodman, N., Potts, C., and Icard, T. Causal abstraction: A theoretical foundation for mechanistic interpretability. CoRR, abs/2301.04709, 2024a. URL https://arxiv.org/abs/2301.04709.
  25. 25.Geiger, A., Wu, Z., Potts, C., Icard, T., and Goodman, N. D. Finding alignments between interpretable causal variables and distributed neural representations. In Locatello, F. and Didelez, V. (eds.), Causal Learning and Reasoning, 1-3 April 2024, Los Angeles, California, USA, volume 236 of Proceedings of Machine Learning Research, pp. 160–187. PMLR, 2024b. URL https://proceedings.mlr.press/v236/geiger24a.html.
  26. 26.Gupta, R., Arcuschin, I., Kwa, T., and Garriga-Alonso, A. InterpBench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. URL https://openreview.net/forum?id=R9gR9MPuD5.
  27. 27.Hanna, M., Pezzelle, S., and Belinkov, Y. Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms. In ICML 2024 Workshop on Mechanistic Interpretability, 2024. URL https://openreview.net/forum?id=grXgesr5dT.
  28. 28.He, Z., Shu, W., Ge, X., Chen, L., Wang, J., Zhou, Y., Liu, F., Guo, Q., Huang, X., Wu, Z., Jiang, Y., and Qiu, X. Llama Scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders. CoRR, abs/2410.20526, 2024. doi: 10.48550/ARXIV.2410.20526. URL https://doi.org/10.48550/arXiv.2410.20526.
  29. 29.Hewitt, J., Geirhos, R., and Kim, B. We can’t understand AI using our existing vocabulary. CoRR, abs/2502.07586, 2025. URL https://arxiv.org/abs/2502.07586.
  30. 30.Huang, J., Wu, Z., Potts, C., Geva, M., and Geiger, A. RAVEL: Evaluating interpretability methods on disentangling language model representations. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8669–8687, Bangkok, Thailand, August 2024a. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.470. URL https://aclanthology.org/2024.acl-long.470/.
  31. 31.Huang, Y., Hu, S., Han, X., Liu, Z., and Sun, M. Unified view of grokking, double descent and emergent abilities: A comprehensive study on algorithm task. In First Conference on Language Modeling, 2024b. URL https://openreview.net/forum?id=cG1EbmWiSs.
  32. 32.Huben, R., Cunningham, H., Smith, L. R., Ewart, A., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=F76bwRSLeK.
  33. 33.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7b. CoRR, abs/2310.06825, 2023. URL https://arxiv.org/abs/2310.06825.
  34. 34.Karpas, E., Abend, O., Belinkov, Y., Lenz, B., Lieber, O., Ratner, N., Shoham, Y., Bata, H., Levine, Y., Leyton-Brown, K., et al. MRKL systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning. arXiv preprint arXiv:2205.00445, 2022.
  35. 35.Karvonen, A., Rager, C., Lin, J., Tigges, C., Bloom, J., Chanin, D., Lau, Y.-T., Farrell, E., Conmy, A., McDougall, C., Ayonrinde, K., Wearden, M., Marks, S., and Nanda, N. SAEBench: A comprehensive benchmark for sparse autoencoders, 2025. URL https://www.neuronpedia.org/sae-bench.
  36. 36.Li, M. and Janson, L. Optimal ablation for interpretability. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=opt72TYzwZ.
  37. 37.Li, R. and Gao, Y. Anchored answers: Unravelling positional bias in GPT-2’s multiple-choice questions. arXiv preprint arXiv:2405.03205, 2024. URL https://arxiv.org/abs/2405.03205.
  38. 38.Lieberum, T., Rahtz, M., Kramar, J., Irving, G., Shah, R., and Mikulik, V. Does circuit analysis interpretability scale? Evidence from multiple choice capabilities in chinchilla. arXiv preprint arXiv:2307.09458, 2023a. URL https://arxiv.org/abs/2307.09458.
  39. 39.Lieberum, T., Rahtz, M., Kramar, J., Nanda, N., Irving, G., Shah, R., and Mikulik, V. Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla. CoRR, abs/2307.09458, 2023b. URL https://arxiv.org/abs/2307.09458.
  40. 40.Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramar, J., Dragan, A., Shah, R., and Nanda, N. Gemma Scope: Open sparse autoencoders everywhere all at once on Gemma 2. In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H. (eds.), Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pp. 278–300, Miami, Florida, US, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.blackboxnlp-1.19. URL https://aclanthology.org/2024.blackboxnlp-1.19/.
  41. 41.Liu, Z., Michaud, E. J., and Tegmark, M. Omnigrok: Grokking beyond algorithmic data. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=zDiHoIWa0q1.
  42. 42.Marks, S. and Tegmark, M. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=aajyHYjjsk.
  43. 43.Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A. Sparse feature circuits: Discovering and editing interpretable causal graphs in language models. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=I4e82CIDxv.
  44. 44.Meng, K., Bau, D., Andonian, A. J., and Belinkov, Y. Locating and editing factual associations in GPT. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=-h6WAS6eE4.
  45. 45.Merullo, J., Eickhoff, C., and Pavlick, E. Circuit component reuse across tasks in transformer language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=fpoAYV6Wsk.
  46. 46.Miller, J., Chughtai, B., and Saunders, W. Transformer circuit evaluation metrics are not robust. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=zSf8PJyQb2.
  47. 47.Mills, E., Su, S., Russell, S., and Emmons, S. Almanacs: A simulatability benchmark for language model explainability. CoRR, abs/2312.12747, 2023.
  48. 48.Mueller, A. Missed causes and ambiguous effects: Counterfactuals pose challenges for interpreting neural networks. In ICML 2024 Workshop on Mechanistic Interpretability, 2024. URL https://openreview.net/forum?id=pJs3ZiKBM5.
  49. 49.Mueller, A., Brinkmann, J., Li, M., Marks, S., Pal, K., Prakash, N., Rager, C., Sankaranarayanan, A., Sharma, A. S., Sun, J., Todd, E., Bau, D., and Belinkov, Y. The quest for the right mediator: A history, survey, and theoretical grounding of causal interpretability. CoRR, abs/2408.01416, 2024. URL https://arxiv.org/abs/2408.01416.
  50. 50.Nanda, N. Attribution Patching: Activation Patching At Industrial Scale, 2023. URL https://www.neelnanda.io/mechanistic-interpretability/attribution-patching.
  51. 51.Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J. Progress measures for grokking via mechanistic interpretability. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=9XFSbDPmdW.
  52. 52.Nikankin, Y., Reusch, A., Mueller, A., and Belinkov, Y. Arithmetic without algorithms: Language models solve math with a bag of heuristics. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=O9YTt26r2P.
  53. 53.Norlund, T., Hagstrom, L., and Johansson, R. Transferring knowledge from vision to language: How to achieve it and how to measure it? In Bastings, J., Belinkov, Y., Dupoux, E., Giulianelli, M., Hupkes, D., Pinter, Y., and Sajjad, H. (eds.), Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pp. 149–162, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.blackboxnlp-1.10. URL https://aclanthology.org/2021.blackboxnlp-1.10/.
  54. 54.Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. Zoom in: An introduction to circuits. Distill, 2020. doi: 10.23915/distill.00024.001. https://distill.pub/2020/circuits/zoom-in.
  55. 55.OpenAI. ChatGPT. https://openai.com/chatgpt, 2022.
  56. 56.Paik, C., Aroca-Ouellette, S., Roncone, A., and Kann, K. The World of an Octopus: How Reporting Bias Influences a Language Model‘s Perception of Color. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 823–835, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.63. URL https://aclanthology.org/2021.emnlp-main.63/.
  57. 57.Pearl, J. Direct and indirect effects. In Proceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence, UAI’01, pp. 411–420, San Francisco, CA, USA, 2001. Morgan Kaufmann Publishers Inc. ISBN 1558608001.
  58. 58.Prakash, N., Shaham, T. R., Haklay, T., Belinkov, Y., and Bau, D. Fine-tuning enhances existing mechanisms: A case study on entity tracking. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=8sKcAWOf2D.
  59. 59.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. Blog post, 2019. URL https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
  60. 60.Rauker, T., Ho, A., Casper, S., and Hadfield-Menell, D. Toward transparent AI: A survey on interpreting the inner structures of deep neural networks. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning, SaTML 2023, Raleigh, NC, USA, February 8-10, 2023, pp. 464–483. IEEE, 2023. doi: 10.1109/SATML54575.2023.00039. URL https://doi.org/10.1109/SaTML54575.2023.00039.
  61. 61.Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Rame, A., et al. Gemma 2: Improving open language models at a practical size. arXiv:2408.00118, 2024. URL https://arxiv.org/abs/2408.00118.
  62. 62.Saphra, N. and Wiegreffe, S. Mechanistic? In The 7th BlackboxNLP Workshop, 2024. URL https://openreview.net/forum?id=schAf4BPtD.
  63. 63.Schwettmann, S., Shaham, T. R., Materzynska, J., Chowdhury, N., Li, S., Andreas, J., Bau, D., and Torralba, A. FIND: A function description benchmark for evaluating interpretability methods. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers.nips.cc/paper_files/paper/2023/hash/ef0164c1112f56246224af540857348f-Abstract-Datasets_and_Benchmarks.html.
  64. 64.Sharkey, L., Chughtai, B., Batson, J., Lindsey, J., Wu, J., Bushnaq, L., Goldowsky-Dill, N., Heimersheim, S., Ortega, A., Bloom, J., Biderman, S., Garriga-Alonso, A., Conmy, A., Nanda, N., Rumbelow, J., Wattenberg, M., Schoots, N., Miller, J., Michaud, E. J., Casper, S., Tegmark, M., Saunders, W., Bau, D., Todd, E., Geiger, A., Geva, M., Hoogland, J., Murfet, D., and McGrath, T. Open problems in mechanistic interpretability. CoRR, abs/2501.16496, 2025. URL https://arxiv.org/abs/2501.16496.
  65. 65.Shi, C., Beltran-Velez, N., Nazaret, A., Zheng, C., Garriga-Alonso, A., Jesson, A., Makar, M., and Blei, D. Hypothesis testing the circuit hypothesis in LLMs. In ICML 2024 Workshop on Mechanistic Interpretability, 2024. URL https://openreview.net/forum?id=ibSNv9cldu.
  66. 66.Smolensky, P. Neural and conceptual interpretation of PDP models. In McClelland, J. L., Rumelhart, D. E., and the PDP Research Group (eds.), Parallel Distributed Processing: Explorations in the Microstructure of Cognition: Psychological and Biological Models, volume 2, pp. 390–431. MIT Press, 1986.
  67. 67.Stolfo, A., Belinkov, Y., and Sachan, M. A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7035–7052, 2023.
  68. 68.Sundararajan, M., Taly, A., and Yan, Q. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, pp. 3319–3328. JMLR.org, 2017.
  69. 69.Syed, A., Rager, C., and Conmy, A. Attribution patching outperforms automated circuit discovery. In Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., and Chen, H. (eds.), Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pp. 407–416, Miami, Florida, US, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.blackboxnlp-1.25. URL https://aclanthology.org/2024.blackboxnlp-1.25/.
  70. 70.Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N. Linear representations of sentiment in large language models. CoRR, abs/2310.15154, 2023. URL https://arxiv.org/abs/2310.15154.
  71. 71.Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. Investigating gender bias in language models using causal mediation analysis. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 12388–12401. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/92650b2e92217715fe312e6fa7b90d82-Paper.pdf.
  72. 72.Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/forum?id=NpsVSN6o4ul.
  73. 73.Wiegreffe, S., Tafjord, O., Belinkov, Y., Hajishirzi, H., and Sabharwal, A. Answer, assemble, ace: Understanding how LMs answer multiple choice questions. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=6NNA0MxhCH.
  74. 74.Wu, Z., Geiger, A., Icard, T., Potts, C., and Goodman, N. D. Interpretability at scale: Identifying causal mechanisms in alpaca. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers.nips.cc/paper_files/paper/2023/hash/f6a8b109d4d4fd64c75e94aaf85d9697-Abstract-Conference.html.
  75. 75.Wu, Z., Geiger, A., Arora, A., Huang, J., Wang, Z., Goodman, N., Manning, C. D., and Potts, C. pyvene: A library for understanding and improving pytorch models via interventions. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: System Demonstrations), pp. 158–165, 2024. URL https://aclanthology.org/2024.naacl-demo.16/.
  76. 76.Wu, Z., Arora, A., Geiger, A., Wang, Z., Huang, J., Jurafsky, D., Manning, C. D., and Potts, C. AxBench: Steering LLMs? even simple baselines outperform sparse autoencoders. CoRR, abs/2501.17148, 2025. URL https://arxiv.org/abs/2501.17148.
  77. 77.Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al. Qwen2.5 technical report. arXiv:2412.15115, 2024. URL https://arxiv.org/abs/2412.15115.
  78. 78.Zhang, F. and Nanda, N. Towards best practices of activation patching in language models: Metrics and methods. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Hf17y6u9BC.
  79. 79.Zhang, W., Wan, C., Zhang, Y., ming Cheung, Y., Tian, X., Shen, X., and Ye, J. Interpreting and improving large language models in arithmetic calculation. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=CfOtiepP8s.
  80. 80.Zhong, Z., Wu, Z., Manning, C., Potts, C., and Chen, D. MQuAKE: Assessing knowledge editing in language models via multi-hop questions. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 15686–15702, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.971. URL https://aclanthology.org/2023.emnlp-main.971/.

Citation

MLA
Mueller, A., et al. “MIB: A Mechanistic Interpretability Benchmark”. arXiv, 2025, http://arxiv.org/abs/2504.13151v2.
APA
Mueller, A., Geiger, A., Wiegreffe, S., Arad, D., Arcuschin, I., Belfki, A., Chan, Y. S., Fiotto-Kaufman, J., Haklay, T., Hanna, M., Huang, J., Gupta, R., Nikankin, Y., Orgad, H., Prakash, N., Reusch, A., Sankaranarayanan, A., Shao, S., Stolfo, A., … Belinkov, Y. (2025). MIB: A Mechanistic Interpretability Benchmark. arXiv. http://arxiv.org/abs/2504.13151v2
Chicago
Mueller, A., A. Geiger, S. Wiegreffe, et al. 2025. “MIB: A Mechanistic Interpretability Benchmark”. arXiv. http://arxiv.org/abs/2504.13151v2.
Harvard
Mueller, A. et al. (2025) “MIB: A Mechanistic Interpretability Benchmark”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2504.13151v2.
Vancouver
1. Mueller A, Geiger A, Wiegreffe S, et al (2025) MIB: A Mechanistic Interpretability Benchmark. arXiv

BibTeX

@article{mueller2025mib,
  title = {MIB: A Mechanistic Interpretability Benchmark},
  author = {Mueller, Aaron and Geiger, Atticus and Wiegreffe, Sarah and Arad, Dana and Arcuschin, Iván and Belfki, Adam and Chan, Yik Siu and Fiotto-Kaufman, Jaden and Haklay, Tal and Hanna, Michael and Huang, Jing and Gupta, Rohan and Nikankin, Yaniv and Orgad, Hadas and Prakash, Nikhil and Reusch, Anja and Sankaranarayanan, Aruna and Shao, Shun and Stolfo, Alessandro and Tutek, Martin and Zur, Amir and Bau, David and Belinkov, Yonatan},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2504.13151v2},
  eprint = {2504.13151}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/