CausalGym: Benchmarking causal interpretability methods on linguistic tasks

Aryaman AroraDan JurafskyChristopher Potts

article2024ACL42 citationsOutstanding Paper Award, SAC Award: Interpretability and Analysis of Models for NLP

Introduces CausalGym, a benchmark that adapts syntactic evaluation tasks to evaluate causal interpretability methods and track how neural language models acquire complex linguistic mechanisms during training.

Listen

Language models are increasingly deployed in high-stakes environments and studied as computational models of language, yet the internal mechanisms governing their predictions remain opaque. While traditional linguistic evaluations measure only input–output behaviors, emerging interpretability tools aim to locate model-internal features. However, these interpretability methods have largely been tested on narrow, ad hoc datasets without rigorous standards for causal efficacy.

The article addresses this gap by introducing CausalGym, a multi-task benchmark designed to evaluate how effectively various interpretability techniques identify internal representations that causally govern linguistic behavior. The primary objective is to benchmark several feature-finding methods across model scales and uncover how complex linguistic mechanisms develop during training.

To construct the benchmark, the researchers adapted 28 test suites from SyntaxGym and created one novel task, yielding 29 distinct linguistic tasks covering syntactic agreement, licensing, garden path effects, and long-distance dependencies. They evaluated seven interpretability methods—including distributed alignment search (DAS), linear probing, difference-in-means, linear discriminant analysis (LDA), principal component analysis (PCA), k-means clustering, and a random vector baseline. Using one-dimensional distributed interchange interventions, the team manipulated internal representations at specific layers and token positions across the Pythia model family (ranging from 14 million to 6.9 billion parameters). They measured causal effect using the log odds-ratio of target predictions and introduced control tasks mapping inputs to arbitrary tokens to evaluate method selectivity.

The investigation produced four key findings. First, DAS achieved the highest raw causal efficacy across all model sizes, posting an overall log odds-ratio of 9.95 in the 6.9-billion-parameter model, compared to 3.42 for linear probing and 2.91 for difference-in-means. Second, when adjusting for expressivity using arbitrary control tasks, probing showed superior selectivity at larger scales (achieving a selectivity score of 3.79 on the 6.9-billion model compared to 2.48 for DAS), revealing that much of DAS’s raw advantage stems from its capacity to fit arbitrary input–output patterns. Third, unsupervised methods (PCA and k-means) and LDA showed weak causal efficacy, with LDA barely surpassing the random baseline (0.27 vs. 0.01 at 6.9B parameters). Fourth, detailed case studies on negative polarity item licensing and filler–gap dependencies in a 1-billion-parameter model revealed that internal linguistic mechanisms emerge in discrete, abrupt stages early in training rather than developing gradually, routing information through intermediate token positions across layers.

These findings provide direct evidence that internal linguistic features are organized linearly but require careful causal verification. The results show that high classification accuracy from a probe does not guarantee that a model actually uses that information downstream, while optimization-heavy methods like DAS risk overstating meaningful alignment due to sheer expressivity. For practitioners and decision-makers, this highlights that establishing model interpretability requires rigorous control baselines rather than relying on raw intervention strength.

The article recommends that researchers adopt standardized interventional benchmarks like CausalGym to validate interpretability methods and investigate internal learning dynamics. Organizations deploying language models should maintain cautious oversight, as demonstrating that a model possesses interpretable internal mechanisms does not inherently make it safe or suitable for autonomous decision-making in sensitive domains.

The findings carry high confidence within the evaluated scope, though the conclusions are bound by specific limitations: the benchmark focuses exclusively on English syntax, evaluates only one model family trained on a single data order, and restricts interventions to one-dimensional linear subspaces. Future work should validate these techniques on non-English languages, diverse model architectures, multi-dimensional subspaces, and non-linguistic tasks.

  • Paper: MIB: A Mechanistic Interpretability Benchmark, Aaron Mueller et al. (2025). It broadens CausalGym’s benchmark agenda beyond linguistic tasks, standardizing causal-variable and circuit-localization evaluations across models and diverse behaviors.
  • Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). It develops a broader statistical-causal critique of interpretability methods, extending CausalGym’s warning that apparent explanatory success can be unstable or misleading.
Cover for CausalGym: Benchmarking causal interpretability methods on linguistic tasks

Abstract

Language models (LMs) have proven to be powerful tools for psycholinguistic research, but most prior work has focused on purely behavioural measures (e.g., surprisal comparisons). At the same time, research in model interpretability has begun to illuminate the abstract causal mechanisms shaping LM behavior. To help bring these strands of research closer together, we introduce CausalGym. We adapt and expand the Syntax-Gym suite of tasks to benchmark the ability of interpretability methods to causally affect model behaviour. To illustrate how CausalGym can be used, we study the pythia models (14M–6.9B) and assess the causal efficacy of a wide range of interpretability methods, including linear probing and distributed alignment search (DAS). We find that DAS outperforms the other methods, and so we use it to study the learning trajectory of two difficult linguistic phenomena in pythia-1b: negative polarity item licensing and filler–gap dependencies. Our analysis shows that the mechanism implementing both of these tasks is learned in discrete stages, not gradually.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Benchmark
  • 3.1 Premise
  • 3.2 Templatising SyntaxGym
  • 3.3 Tasks
  • 3.4 Evaluation
  • 4 Methods
  • 4.1 Preliminaries
  • 4.2 Definitions
  • 5 Experiments
  • 5.1 Measuring causal efficacy
  • 5.2 Controlling for expressivity
  • 5.3 Results
  • 6 Case studies
  • 6.1 Training dynamics
  • 7 Conclusion
  • References
  • A Tasks
  • B Training and evaluation details
  • C Hyperparameter tuning
  • D Data and licensing
  • E Detailed odds-ratio results
  • E.1 Per-layer
  • E.2 Per-task
  • F Odds-ratio plots for all methods on selected tasks

Knowls

  1. Knowl 1 — CausalGym converts syntactic contrasts into scalable counterfactual pairs

    model/method

    CausalGym adapts SyntaxGym-style syntactic tests into grammatical minimal pairs for causal intervention experiments. Each pair varies a binary linguistic feature and its expected next-token label while keeping other template regions fixed. To generate pairs, a task template specifies feature types and the options available for the feature-bearing input region and label; the generator samples two different types, samples matching input and label options for each, and uses the same sampled options in all non-label regions. CausalGym contains 29 tasks: 4 agreement, 7 licensing, 6 garden-path, 4 gross-syntactic-state, and 8 long-distance-dependency tasks. Twenty-eight were templatised from SyntaxGym and agr_gender is novel. The authors excluded tasks without suitable grammatical paired sentences, including two center-embedding tasks, and merged six gendered reflexive-licensing suites into three non-gendered tasks.

  2. Knowl 2 — Interventions use one-dimensional distributed interchange

    equation

    CausalGym benchmarks one-dimensional distributed interchange interventions (1D DII). For a model component ff that outputs an nn-dimensional activation, a base input bb, a source input ss, and a feature direction a∈Rna\in\mathbb{R}^n, the intervention replaces the base activation with the base activation modified along aa according to the source activation's projection onto that direction:

    fa∗(b,s)=f(b)+(f(s)⊤a−f(b)⊤a)a.f_a^*(b,s)=f(b)+\big(f(s)^\top a-f(b)^\top a\big)a.

    Here, f(b)f(b) and f(s)f(s) are the component activations for the base and source inputs, respectively. In the experiments, ff is the residual-stream state after a Transformer layer, evaluated at the last token of a template region. If the paired inputs tokenize to different lengths, representations are aligned by template region before intervention. The method tests a single linear feature direction rather than replacing the entire activation.

  3. Knowl 3 — Causal efficacy is measured by changes in next-token log odds

    equation

    For an evaluation example e=⟨b,s,yb,ys⟩e=\langle b,s,y_b,y_s\rangle, let bb and ss be the base and source inputs, and let yby_b and ysy_s be their ground-truth next-token labels. Let pp be the original language model and pf←fa∗p_{f\leftarrow f_a^*} the model after the intervention at component ff. Causal effect is the change in the log odds of the source label relative to the base label:

    Odds⁡(p,pf←fa∗,e)=log⁡(p(yb∣b)p(ys∣b)⋅pf←fa∗(ys∣b,s)pf←fa∗(yb∣b,s)).\operatorname{Odds}(p,p_{f\leftarrow f_a^*},e)=\log\left(\frac{p(y_b\mid b)}{p(y_s\mid b)}\cdot\frac{p_{f\leftarrow f_a^*}(y_s\mid b,s)}{p_{f\leftarrow f_a^*}(y_b\mid b,s)}\right).

    For an evaluation set EE, the average effect is AvgOdds⁡=∣E∣−1∑e∈EOdds⁡(p,pf←fa∗,e)\operatorname{AvgOdds}=|E|^{-1}\sum_{e\in E}\operatorname{Odds}(p,p_{f\leftarrow f_a^*},e). A value of zero indicates no change in these odds, and larger values indicate a stronger causal effect in the direction of the source label. To summarize a feature-finding method mm across a model, the paper takes the maximum average odds over template regions r∈Rr\in R separately for each layer l∈Ll\in L, then averages those maxima across layers: OverallOdds⁡(p,m,E)=∣L∣−1∑l∈Lmax⁡r∈RAvgOdds⁡(p,pf(l,r)←fam∗,E)\operatorname{OverallOdds}(p,m,E)=|L|^{-1}\sum_{l\in L}\max_{r\in R}\operatorname{AvgOdds}(p,p_{f^{(l,r)}\leftarrow f^*_{a_m}},E). Here f(l,r)f^{(l,r)} is the residual-stream component at layer ll and region rr, and ama_m is the direction found by method mm.

  4. Knowl 4 — The benchmark compares causal and representation-based feature directions

    model/method

    CausalGym compares seven ways of selecting the 1D DII direction. Distributed alignment search (DAS) optimizes a direction on paired training examples to increase the intervened model's probability of the source next-token label, using cross-entropy loss while keeping the language-model weights fixed. A linear probe is trained to classify the base input's label, and its weight vector supplies the intervention direction. Difference-in-means uses the difference between mean activations for the two base-label classes; linear discriminant analysis uses the covariance-adjusted difference between class means. The unsupervised methods are the first principal component of centered activations and the difference between the centroids of a two-means clustering; a random direction is the baseline. Thus, DAS directly optimizes intervention outcomes, whereas the probe, mean-difference, LDA, and unsupervised directions are obtained without that causal training objective.

  5. Knowl 5 — Experiments evaluate nine Pythia scales across layers and regions

    experimental setup

    The main benchmark comparison uses nine Pythia checkpoints from 14 million to 6.9 billion parameters. The family was trained on the same data in the same order, allowing scale comparisons under a shared training sequence. For each of the 29 tasks, the authors train directions on 400 examples and evaluate on 100 non-overlapping examples; training examples include both directions of each sampled pair, and evaluation examples are resampled if they overlap with training. Each method is tested at every residual-stream layer and template region, with the region maximum then aggregated across layers. DAS is trained for one epoch with batch size 4, giving 100 updates, Adam, learning rate 5×10−35\times10^{-3}, and a linear schedule with 10% warmup followed by decay. Probe results for models larger than Pythia-70M use the better of two tested regularization settings.

  6. Knowl 6 — DAS has the strongest average causal efficacy, but not uniformly the best selectivity

    empirical result

    Across the CausalGym tasks, DAS produces the largest average overall odds-ratio at each of the nine reported Pythia scales. Its score rises from 3.94 on Pythia-14M to a peak of 10.74 on Pythia-1B, and is 9.95 on Pythia-6.9B. At Pythia-1B, the corresponding averages are 3.66 for probing, 3.17 for difference-in-means, 2.07 for PCA, 2.13 for k-means, 0.29 for LDA, and 0.03 for random directions. Probing and difference-in-means are generally the next strongest methods, while PCA and k-means are substantially weaker and LDA is close to the random baseline. Mean task accuracy increases from 0.62 at 14M to 0.89 at 6.9B. These results show that optimizing directly for intervention outcomes finds highly efficacious directions, but do not by themselves establish that those directions capture only the model's task-relevant mechanism.

  7. Knowl 7 — Arbitrary-mapping controls reveal an expressivity confound

    empirical result

    To test whether a direction's causal effect is specific to the linguistic behavior, CausalGym creates control tasks by replacing the two next-token labels with arbitrary tokens (for example, _dog and _give) while preserving which inputs belong to each label class. Selectivity is the original-task odds-ratio minus the control-task odds-ratio at each intervention site, aggregated using the same layer-and-region procedure as overall odds. DAS can also produce substantial effects on these arbitrary mappings, consistent with its ability to fit output behavior rather than only a pre-existing linguistic feature. On Pythia-1B, DAS has overall odds-ratio 10.74 but selectivity 3.34; probing has lower overall odds-ratio 3.66 but higher selectivity 4.24. Thus, controlling for arbitrary-mapping performance reduces DAS's apparent advantage and shows why causal efficacy alone is not a sufficient comparison of feature-finding methods.

  8. Knowl 8 — Final Pythia-1B mechanisms distribute linguistic information across positions

    empirical result

    At the final Pythia-1B checkpoint, layer-by-region intervention analyses of negative-polarity-item licensing and filler–gap dependencies show causal effects at multiple positions rather than at a single input or output location. In the npi_any_subj-relc task, the feature associated with sentence-level negation is causally effective at the complementizer that in early layers, at the auxiliary in middle layers, and at the main verb in later layers, before the model predicts any. The filler_gap_subj analysis likewise indicates a multi-step mechanism for tracking a distant extraction dependency. These results support the paper's conclusion that both behaviors involve movement of information through intermediate positions.

  9. Knowl 9 — The two Pythia-1B mechanisms emerge in discrete training stages

    empirical result

    Checkpoint analyses of Pythia-1B show staged rather than gradual development of the causal mechanisms for negative-polarity-item licensing and filler–gap tracking. For npi_any_subj-relc, an effect at the NPI and main verb is evident by training step 1,000; by step 2,000 the auxiliary becomes important in middle layers and the NPI effect shifts toward early layers; by step 3,000 an additional intermediate location appears. For filler_gap_subj, an initial mechanism is evident by step 2,000 at the filler position, the first determiner, and the final token; the main verb joins the mechanism after step 10,000. In both tasks the model first shows direct information transfer from the alternating input feature toward the output, with intermediate steps added later. Task accuracy also rises across the checkpoints, from 0.45 to 0.97 for NPI licensing and from 0.49 to 0.89 for filler–gap tracking.

  10. Knowl 10 — CausalGym's scope is limited to English and one-dimensional linear interventions

    limitation

    The benchmark contains English linguistic tasks only, and its results come from Pythia models trained on the same data in a fixed order; other languages, model families, or training data could produce different mechanisms and results. CausalGym also benchmarks only one-dimensional linear subspaces, leaving multidimensional linear and nonlinear interventions untested. Its task set is a selected range of linguistic behaviors rather than a coverage of non-linguistic model behaviors.

Coverage note — The exhaustive per-task score tables and hyperparameter-search grids are omitted because they are detailed diagnostics rather than additional general findings; their main comparative conclusions and the final experimental settings are included.

References

  1. 1.Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017. Fine-grained analysis of sentence embeddings using auxiliary prediction tasks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France.
  2. 2.Afra Amini, Tiago Pimentel, Clara Meister, and Ryan Cotterell. 2023. Naturalistic causal probing for morpho-syntax. Transactions of the Association for Computational Linguistics, 11:384–403.
  3. 3.Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219.
  4. 4.Nora Belrose. 2023. Diff-in-means concept editing is worst-case optimal. EleutherAI Blog.
  5. 5.Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023. LEACE: Perfect linear concept erasure in closed form. arXiv:2306.03819.
  6. 6.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, ICML 2023, volume 202 of Proceedings of Machine Learning Research, pages 2397–2430, Honolulu, Hawaii, USA. PMLR.
  7. 7.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc.
  8. 8.Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas. 2022. Causal scrubbing: A method for rigorously testing interpretability hypotheses. In Alignment Forum.
  9. 9.Angelica Chen, Ravid Schwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, and Naomi Saphra. 2023. Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs. arXiv:2309.07311.
  10. 10.Abhijith Chintam, Rahel Beloch, Willem Zuidema, Michael Hanna, and Oskar van der Wal. 2023. Identifying and adapting transformer-components responsible for gender bias in an English language model. In Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 379–394, Singapore. Association for Computational Linguistics.
  11. 11.Aaron Defazio, Francis Bach, and Simon Lacoste-Julien. 2014. SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives. In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc.
  12. 12.Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy models of superposition. Transformer Circuits Thread.
  13. 13.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread.
  14. 14.Allyson Ettinger, Ahmed Elgohary, and Philip Resnik. 2016. Probing for semantic evidence of composition by means of simple classification tasks. In Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP, pages 134–139, Berlin, Germany. Association for Computational Linguistics.
  15. 15.Richard Futrell, Ethan Wilcox, Takashi Morita, Peng Qian, Miguel Ballesteros, and Roger Levy. 2019. Neural language models as psycholinguistic subjects: Representations of syntactic state. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 32–42, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy. 2020. SyntaxGym: An online platform for targeted evaluation of language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 70–76, Online. Association for Computational Linguistics.
  17. 17.Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021. Causal abstractions of neural networks. In Advances in Neural Information Processing Systems, volume 34, pages 9574–9586. Curran Associates, Inc.
  18. 18.Atticus Geiger, Chris Potts, and Thomas Icard. 2023a. Causal abstraction for faithful model interpretation. arxiv:2301.04709.
  19. 19.Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, and Christopher Potts. 2022. Inducing causal structure for interpretable neural networks. In International Conference on Machine Learning, ICML 2022, volume 162 of Proceedings of Machine Learning Research, pages 7324–7338, Baltimore, Maryland, USA. PMLR.
  20. 20.Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman. 2023b. Finding alignments between interpretable causal variables and distributed neural representations. arXiv:2303.02536.
  21. 21.Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora. 2023. Localizing model behavior with path patching. arXiv:2304.05969.
  22. 22.Adam Goodkind and Klinton Bicknell. 2018. Predictive power of word surprisal for reading times is a linear function of language model quality. In Proceedings of the 8th Workshop on Cognitive Modeling and Computational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah. Association for Computational Linguistics.
  23. 23.Clément Guerner, Anej Svete, Tianyu Liu, Alexander Warstadt, and Ryan Cotterell. 2023. A geometric notion of causal probing. arXiv:2307.15054.
  24. 24.Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018. Colorless green recurrent networks dream hierarchically. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1195–1205, New Orleans, Louisiana. Association for Computational Linguistics.
  25. 25.Michael Hanna, Yonatan Belinkov, and Sandro Pezzelle. 2023. When language models fall in love: Animacy processing in transformer language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12120–12135, Singapore. Association for Computational Linguistics.
  26. 26.Sophie Hao and Tal Linzen. 2023. Verb conjugation in transformers is determined by linear encodings of subject number. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 4531–4539, Singapore. Association for Computational Linguistics.
  27. 27.John Hewitt and Percy Liang. 2019. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2733–2743, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy. 2020. A systematic assessment of syntactic generalization in neural language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1725–1744, Online. Association for Computational Linguistics.
  29. 29.Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova, and Roger Levy. 2024. Language models align with human judgments on key grammatical constructions. arXiv:2402.01676.
  30. 30.Julie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, and Christopher Potts. 2024. Mission: Impossible language models. arXiv:2401.06416.
  31. 31.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA.
  32. 32.Karim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau, and Ryan Cotterell. 2022. Probing for the usage of grammatical number. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8818–8831, Dublin, Ireland. Association for Computational Linguistics.
  33. 33.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023. Inference-time intervention: Eliciting truthful answers from a language model. In Advances in Neural Information Processing Systems, volume 36.
  34. 34.Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. Assessing the ability of LSTMs to learn syntax-sensitive dependencies. Transactions of the Association for Computational Linguistics, 4:521–535.
  35. 35.Samuel Marks and Max Tegmark. 2023. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv:2310.06824.
  36. 36.Rebecca Marvin and Tal Linzen. 2018. Targeted syntactic evaluation of language models. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1192–1202, Brussels, Belgium. Association for Computational Linguistics.
  37. 37.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems, volume 35, pages 17359–17372. Curran Associates, Inc.
  38. 38.Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 746–751, Atlanta, Georgia. Association for Computational Linguistics.
  39. 39.Neel Nanda, Andrew Lee, and Martin Wattenberg. 2023. Emergent linear representations in world models of self-supervised sequence models. In Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 16–30, Singapore. Association for Computational Linguistics.
  40. 40.Chris Olah. 2022. Mechanistic interpretability, variables, and the importance of interpretable bases. Transformer Circuits Thread.
  41. 41.Kiho Park, Yo Joong Choe, and Victor Veitch. 2023. The linear representation hypothesis and the geometry of large language models. arXiv:2311.03658.
  42. 42.Judea Pearl. 2009. Causality: Models, Reasoning, and Inference, 2nd edition. Cambridge University Press.
  43. 43.Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake VanderPlas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Edouard Duchesnay. 2011. Scikit-learn: Machine learning in Python. The Journal of Machine Learning Research, 12:2825–2830.
  44. 44.Cory Shain, Clara Meister, Tiago Pimentel, Ryan Cotterell, and Roger Levy. 2024. Large-scale evidence for logarithmic effects of word predictability on reading time. Proceedings of the National Academy of Sciences. To appear.
  45. 45.Nathaniel J. Smith and Roger Levy. 2013. The effect of word predictability on reading time is logarithmic. Cognition, 128(3):302–319.
  46. 46.Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023. Linear representations of sentiment in large language models. arXiv:2310.15154.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, W. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30, pages 5998–6008. Curran Associates, Inc.
  48. 48.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems, volume 33, pages 12388–12401. Curran Associates, Inc.
  49. 49.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda.
  50. 50.Alex Warstadt and Samuel R. Bowman. 2022. What artificial neural networks can tell us about human language acquisition. Algebraic Structures in Natural Language, pages 17–60.
  51. 51.Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020. BLiMP: The benchmark of linguistic minimal pairs for English. Transactions of the Association for Computational Linguistics, 8:377–392.
  52. 52.Ethan Wilcox, Clara Meister, Ryan Cotterell, and Tiago Pimentel. 2023a. Language model quality correlates with psychometric predictive power in multiple languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7503–7511, Singapore. Association for Computational Linguistics.
  53. 53.Ethan Gotlieb Wilcox, Richard Futrell, and Roger Levy. 2023b. Using computational models to test syntactic learnability. Linguistic Inquiry, pages 1–44.
  54. 54.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  55. 55.Zhengxuan Wu, Atticus Geiger, Aryaman Arora, Jing Huang, Noah D. Goodman, Christopher D. Manning, and Christopher Potts. 2024. pyvene: A library for understanding and improving PyTorch models via interventions. Under review.
  56. 56.Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah D. Goodman. 2023. Interpretability at scale: Identifying causal mechanisms in Alpaca. In Advances in Neural Information Processing Systems, volume 36.
  57. 57.Takateru Yamakoshi, James McClelland, Adele Goldberg, and Robert Hawkins. 2023. Causal interventions expose implicit situation models for commonsense language understanding. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13265–13293, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Arora, A., et al. “CausalGym: Benchmarking Causal Interpretability Methods on Linguistic Tasks”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 14638–63, https://doi.org/10.18653/v1/2024.acl-long.785.
APA
Arora, A., Jurafsky, D., & Potts, C. (2024). CausalGym: Benchmarking causal interpretability methods on linguistic tasks. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14638–14663. https://doi.org/10.18653/v1/2024.acl-long.785
Chicago
Arora, A., D. Jurafsky, and C. Potts. 2024. “CausalGym: Benchmarking Causal Interpretability Methods on Linguistic Tasks”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14638–63. https://doi.org/10.18653/v1/2024.acl-long.785.
Harvard
Arora, A., Jurafsky, D. and Potts, C. (2024) “CausalGym: Benchmarking causal interpretability methods on linguistic tasks”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14638–14663. Available at: https://doi.org/10.18653/v1/2024.acl-long.785.
Vancouver
1. Arora A, Jurafsky D, Potts C (2024) CausalGym: Benchmarking causal interpretability methods on linguistic tasks. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14638–14663

BibTeX

@inproceedings{arora-etal-2024-causalgym,
    title = "{C}ausal{G}ym: Benchmarking causal interpretability methods on linguistic tasks",
    author = "Arora, Aryaman  and
      Jurafsky, Dan  and
      Potts, Christopher",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.785/",
    doi = "10.18653/v1/2024.acl-long.785",
    pages = "14638--14663"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/