CausalGym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman AroraDan JurafskyChristopher Potts
Introduces CausalGym, a benchmark that adapts syntactic evaluation tasks to evaluate causal interpretability methods and track how neural language models acquire complex linguistic mechanisms during training.
Language models are increasingly deployed in high-stakes environments and studied as computational models of language, yet the internal mechanisms governing their predictions remain opaque. While traditional linguistic evaluations measure only input–output behaviors, emerging interpretability tools aim to locate model-internal features. However, these interpretability methods have largely been tested on narrow, ad hoc datasets without rigorous standards for causal efficacy.
The article addresses this gap by introducing CausalGym, a multi-task benchmark designed to evaluate how effectively various interpretability techniques identify internal representations that causally govern linguistic behavior. The primary objective is to benchmark several feature-finding methods across model scales and uncover how complex linguistic mechanisms develop during training.
To construct the benchmark, the researchers adapted 28 test suites from SyntaxGym and created one novel task, yielding 29 distinct linguistic tasks covering syntactic agreement, licensing, garden path effects, and long-distance dependencies. They evaluated seven interpretability methods—including distributed alignment search (DAS), linear probing, difference-in-means, linear discriminant analysis (LDA), principal component analysis (PCA), k-means clustering, and a random vector baseline. Using one-dimensional distributed interchange interventions, the team manipulated internal representations at specific layers and token positions across the Pythia model family (ranging from 14 million to 6.9 billion parameters). They measured causal effect using the log odds-ratio of target predictions and introduced control tasks mapping inputs to arbitrary tokens to evaluate method selectivity.
The investigation produced four key findings. First, DAS achieved the highest raw causal efficacy across all model sizes, posting an overall log odds-ratio of 9.95 in the 6.9-billion-parameter model, compared to 3.42 for linear probing and 2.91 for difference-in-means. Second, when adjusting for expressivity using arbitrary control tasks, probing showed superior selectivity at larger scales (achieving a selectivity score of 3.79 on the 6.9-billion model compared to 2.48 for DAS), revealing that much of DAS’s raw advantage stems from its capacity to fit arbitrary input–output patterns. Third, unsupervised methods (PCA and k-means) and LDA showed weak causal efficacy, with LDA barely surpassing the random baseline (0.27 vs. 0.01 at 6.9B parameters). Fourth, detailed case studies on negative polarity item licensing and filler–gap dependencies in a 1-billion-parameter model revealed that internal linguistic mechanisms emerge in discrete, abrupt stages early in training rather than developing gradually, routing information through intermediate token positions across layers.
These findings provide direct evidence that internal linguistic features are organized linearly but require careful causal verification. The results show that high classification accuracy from a probe does not guarantee that a model actually uses that information downstream, while optimization-heavy methods like DAS risk overstating meaningful alignment due to sheer expressivity. For practitioners and decision-makers, this highlights that establishing model interpretability requires rigorous control baselines rather than relying on raw intervention strength.
The article recommends that researchers adopt standardized interventional benchmarks like CausalGym to validate interpretability methods and investigate internal learning dynamics. Organizations deploying language models should maintain cautious oversight, as demonstrating that a model possesses interpretable internal mechanisms does not inherently make it safe or suitable for autonomous decision-making in sensitive domains.
The findings carry high confidence within the evaluated scope, though the conclusions are bound by specific limitations: the benchmark focuses exclusively on English syntax, evaluates only one model family trained on a single data order, and restricts interventions to one-dimensional linear subspaces. Future work should validate these techniques on non-English languages, diverse model architectures, multi-dimensional subspaces, and non-linguistic tasks.
- Paper: Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small, Kevin Wang et al. (2022). Its causal circuit analysis of a language-model task shows how interventions can test whether internal components actually drive linguistic behavior, a foundation for CausalGym’s efficacy evaluations.
- Paper: Towards Best Practices of Activation Patching in Language Models: Metrics and Methods, Fred Zhang et al. (2024). Its systematic study of activation-patching choices and effect metrics prepares readers to understand why CausalGym standardizes interventions and measures causal impact.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). Its layer-by-layer probing of linguistic tasks provides the behavioral and representational groundwork that CausalGym advances by testing whether identified information causally governs predictions.
- Paper: MIB: A Mechanistic Interpretability Benchmark, Aaron Mueller et al. (2025). It broadens CausalGym’s benchmark agenda beyond linguistic tasks, standardizing causal-variable and circuit-localization evaluations across models and diverse behaviors.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). It develops a broader statistical-causal critique of interpretability methods, extending CausalGym’s warning that apparent explanatory success can be unstable or misleading.
