MIB: A Mechanistic Interpretability Benchmark
Aaron MuellerAtticus GeigerSarah WiegreffeDana AradIván ArcuschinAdam BelfkiYik Siu ChanJaden Fiotto-KaufmanTal HaklayMichael Hanna
Establishes a standardized benchmark to rigorously evaluate mechanistic interpretability methods across circuit and causal variable localization, revealing critical performance differences among popular techniques like sparse autoencoders and attribution patching.
Mechanistic interpretability aims to reverse-engineer neural language models by mapping their internal representations to causal mechanisms and human-understandable concepts. Despite rapid growth in proposed interpretability techniques, the field lacks standardized evaluation practices. New methods are frequently assessed using customized metrics across narrow, ad-hoc tasks, making it difficult to determine whether newer techniques offer genuine advancements over existing baselines. Establishing standard evaluation benchmarks is essential for verifying whether interpretability tools can reliably assist in auditing, safety evaluations, and fine-grained control of production models.
The article introduces the Mechanistic Interpretability Benchmark (MIB) to establish a rigorous, standardized framework for evaluating interpretability methods. The benchmark evaluates these techniques across two distinct tracks: circuit localization, which tests how effectively methods locate sparse subgraphs of model components that drive specific behaviors, and causal variable localization, which assesses how well methods isolate features within internal activations and align them with high-level conceptual variables. The authors evaluated four open-weight language models of varying scales—ranging from GPT-2 Small to Llama-3.1 8B—across four diverse tasks: Indirect Object Identification, multi-digit Arithmetic, synthetic Multiple-Choice Question Answering, and the grade-school science AI2 Reasoning Challenge, utilizing standardized counterfactual input datasets and public leaderboards with private evaluation splits.
Key findings show significant differences in performance among leading interpretability techniques. For circuit localization, edge-level attribution patching using integrated gradients over inputs (EAP-IG-inputs) combined with counterfactual ablations consistently yielded the highest faithfulness across models and tasks, balancing localization accuracy and computational cost. Conversely, node-level methods performed poorly because they fail to achieve necessary circuit sparsity, and exact activation patching suffered from prohibitive computational costs without providing superior accuracy. For causal variable localization, the supervised Distributed Alignment Search (DAS) consistently outperformed unsupervised featurization techniques. Furthermore, popular unsupervised Sparse Autoencoders (SAEs) failed to identify features that outperformed standard internal vector dimensions (neurons) when aligning to task-relevant concepts, challenging current assumptions regarding SAE superiority in concept extraction.
These findings provide clarity on the actual capabilities of mechanistic interpretability tools, indicating that supervised methods and gradient-based attribution offer the most reliable pathways for model inspection. The relative underperformance of Sparse Autoencoders highlights a critical gap between unsupervised representation extraction and causally validated concept mapping, cautioning organizations against relying solely on SAE-based pipelines for auditing safety-critical behaviors without explicit causal validation.
Practitioners should prioritize integrated gradient edge-attribution methods when isolating model subgraphs and utilize supervised subspace methods like DAS for targeted concept probing. Future research should focus on refining non-linear and unsupervised feature extraction techniques to improve causal alignment, addressing computational bottlenecks in large model evaluations, and developing methods to jointly discover features and circuit structures. Confidence in the relative rankings of evaluated baselines is high across the standardized tasks, though readers should exercise caution when extrapolating these findings to unstudied tasks, complex multi-step reasoning, or models that exhibit low baseline task performance.
- Paper: Towards Automated Circuit Discovery for Mechanistic Interpretability, Arthur Conmy et al. (2023). Read ACDC first to understand automated edge-level circuit discovery and the activation-patching baselines that MIB evaluates against.
- Paper: Towards Best Practices of Activation Patching in Language Models: Metrics and Methods, Fred Zhang et al. (2024). Its analysis of counterfactual corruptions and patching metrics clarifies the intervention choices underlying MIB’s circuit-localization evaluations.
- Paper: Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small, Kevin Wang et al. (2022). This foundational IOI circuit study supplies the concrete sparse-circuit example and task that MIB uses to assess localization methods.
- Paper: RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations, Jing Huang et al. (2024). RAVEL establishes counterfactual evaluation of concept localization, preparing readers for MIB’s causal-variable track and its comparison of supervised and unsupervised methods.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). Integrated Gradients is a core attribution method behind MIB’s strongest edge-localization approach, so its axioms and path-based computation make those results easier to interpret.
No sufficiently relevant recommendations were found.
