RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations

Jing HuangZhengxuan WuChristopher PottsMor GevaAtticus Geiger

article2024ACL83 citations

Presents a diagnostic benchmark and a multi-task alignment method to quantitatively evaluate and improve how interpretability techniques disentangle polysemantic, distributed representations across language model activations.

Listen

Modern artificial intelligence relies heavily on large language models, yet understanding how these complex systems internally represent distinct pieces of knowledge remains a critical challenge. Individual neural units within these models are polysemantic, meaning a single component participates in encoding multiple unrelated concepts simultaneously. While numerous interpretability techniques have been developed to map these complex internal states into distinct, understandable features, the field has lacked standardized, quantitative frameworks to measure how effectively these techniques actually isolate specific concepts without unintentionally altering others.

To address this gap, the article introduces the RAVEL diagnostic benchmark, designed to evaluate and compare interpretability techniques on their ability to localize and disentangle specific attributes of entities represented inside language models. Using counterfactual interchange interventions—which alter model states during text processing to observe causal behavioral changes—the benchmark measures whether a technique successfully causes a targeted concept to change while isolating and preserving unrelated attributes.

Across extensive experiments evaluating multiple interpretability families on the Llama2-7B language model, the article establishes four primary findings. First, methods utilizing counterfactual supervision significantly outperform unsupervised approaches; specifically, unsupervised techniques like Principal Component Analysis and sparse autoencoders struggled with disentanglement (scoring roughly 39% to 49%), whereas counterfactually supervised methods performed substantially better. Second, the article introduces Multi-task Distributed Alignment Search, which achieves the state-of-the-art disentanglement score on the benchmark (reaching 60.1% on unseen entities and 65.6% on unseen prompt contexts) by incorporating isolation criteria directly into the training objective. Third, certain real-world attribute pairs, such as country and language or latitude and longitude, remain consistently difficult for any method to separate due to fundamental entanglements in how models organize knowledge. Fourth, internal representations become progressively more disentangled in deeper layers of the model, with concept isolation peaking around intermediate and later layers.

These findings provide crucial guidance for technical leaders and teams deploying language models in high-stakes environments. They demonstrate that understanding and controlling model behavior requires analyzing distributed representations rather than assuming individual neurons hold discrete concepts. Relying on unsupervised feature extractors or simple probing risks mischaracterizing model mechanisms, which could lead to flawed model audits or ineffective safety interventions. Instead, utilizing multi-task causal alignment offers a more reliable, faithful approach to isolating and steering specific model behaviors.

Organizations seeking to interpret or edit language models should adopt multi-task causal intervention methods when isolating internal concepts and run evaluations across multiple prompt contexts to verify generalizability. For future development, technical teams should instantiate the RAVEL benchmark on other emerging model architectures and expand intervention analyses beyond entity-level tokens to explore multi-token and multi-layer dynamics across broader network components.

  • Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). This work formalizes the linear representation hypothesis using causal counterfactual concept pairs and metric inner products, providing theoretical justification for the causal alignment behaviors observed in RAVEL.
  • Paper: On the Origins of Linear Representations in Large Language Models, Yibo Jiang et al. (2024). This paper presents a theoretical framework explaining why language models naturally learn linear and orthogonal concept representations during training, providing mathematical foundations for the empirical disentanglement dynamics analyzed in RAVEL.
  • Paper: Scaling and evaluating sparse autoencoders, Leo Gao et al. (2025). This study scales sparse autoencoders using TopK activations to improve feature quality and disentanglement, addressing the key architectural limitations of unsupervised dictionary learning highlighted by RAVEL.
  • Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). This treatise provides a meta-level statistical and causal critique of unidentifiability and overdetermination across mechanistic interpretability techniques, extending RAVEL's warnings about the risks of unprincipled interpretability methods.
Cover for RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations

Abstract

Individual neurons participate in the representation of multiple high-level concepts. To what extent can different interpretability methods successfully disentangle these roles? To help address this question, we introduce RAVEL (Resolving Attribute–Value Entanglements in Language Models), a dataset that enables tightly controlled, quantitative comparisons between a variety of existing interpretability methods. We use the resulting conceptual framework to define the new method of Multi-task Distributed Alignment Search (MDAS), which allows us to find distributed representations satisfying multiple causal criteria. With Llama2-7B as the target language model, MDAS achieves state-of-the-art results on RAVEL, demonstrating the importance of going beyond neuron-level analyses to identify features distributed across activations. We release our benchmark at https://github.com/explanare/ravel.

Table of Contents

  • 1 Introduction
  • 2 The RAVEL Dataset
  • 2.1 Data Generation
  • 2.2 Interpretability Evaluation
  • 3 Interpretability Methods
  • 3.1 PCA
  • 3.2 Sparse Autoencoder
  • 3.3 Relaxed Linear Adversarial Probe
  • 3.4 Differential Binary Masking
  • 3.5 Distributed Alignment Search
  • 3.6 Multi-task DBM and DAS
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Results
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Dataset Details
  • A.1 Details of Entities and Attributes
  • A.2 The RAVEL Llama2-7B Instance
  • B Method Details
  • B.1 PCA
  • B.2 Sparse Autoencoder
  • B.3 RLAP
  • B.4 DBM-based and DAS-based Methods
  • B.5 Computational Cost
  • C Results
  • C.1 Breakdown of Benchmark Results
  • C.2 Additional Attribute Disentanglement Results

Knowls

  1. Knowl 1 — The RAVEL Benchmark for Attribute Disentanglement in Language Models

    experimental setup

    RAVEL (Resolving Attribute–Value Entanglements in Language Models) is a diagnostic interpretability benchmark designed to evaluate how well interpretability methods isolate and disentangle specific attributes of entities represented in language models.

    RAVEL contains 5 entity types covering factual, linguistic, and commonsense knowledge:

    • City: Attributes include Country, Language, Latitude, Longitude, Timezone, Continent (3,552 entities, 150 prompt templates).
    • Nobel Laureate: Attributes include Award Year, Birth Year, Country of Birth, Field, Gender (928 entities, 100 prompt templates).
    • Verb: Attributes include Definition, Past Tense, Pronunciation, Singular (986 entities, 60 prompt templates).
    • Physical Object: Attributes include Biological Category, Color, Size, Texture (563 entities, 60 prompt templates).
    • Occupation: Attributes include Duty, Gender Bias, Industry, Work Location (799 entities, 50 prompt templates).

    The dataset uses two types of prompt templates for each entity EE and attribute AA with value AEA_E:

    1. Attribute prompts PEAP_E^A: Queries that instruct the model to produce AEA_E (e.g., Paris is in the continent of).
    2. Entity prompts WEW_E: Sentences containing EE sampled from Wikipedia that do not query any attribute in the attribute set A\mathcal{A} (e.g., Tokyo is a large city.).

    The full prompt collection is D={x:x∈PEA∪WE,E∈E,A∈A}\mathcal{D} = \{x : x \in P_E^A \cup W_E, E \in \mathcal{E}, A \in \mathcal{A}\}.

    RAVEL evaluates generalization via two distinct 50%/25%/25% train/dev/test data splits:

    • Entity Split: Entities are partitioned randomly across splits while keeping prompt templates identical across splits.
    • Context Split: Prompt templates are partitioned randomly across splits while keeping the entity set identical across splits.

    When evaluating a target model MM, RAVEL filters for entities and prompt templates on which MM achieves high prediction accuracy (e.g., for Llama2-7B, 2,800 entities across types with average accuracies of 94.3%–97.1%).

  2. Knowl 2 — Causal Metrics for Attribute Disentanglement: Cause, Iso, and Disentangle Scores

    definition

    Let MM be a language model, τ(M(x))\tau(M(x)) be the token predicted by MM on input xx, NN be a set of neural components (e.g., residual stream activations at layer LL), F\mathcal{F} be a featurizer mapping NN into a feature space, and FAF_A be a feature representation corresponding to target attribute A∈AA \in \mathcal{A}.

    The interchange intervention II(M,FA,x,x′)\text{II}(M, F_A, x, x') replaces the feature FAF_A in M(x)M(x) with the value it takes under base/source input x′x':

    II(M,FA,x,x′)≜τ(MFA←GetFeature(M(x′),FA)(x))\text{II}(M, F_A, x, x') \triangleq \tau\left(M_{F_A \leftarrow \text{GetFeature}(M(x'), F_A)}(x)\right)

    Three evaluation metrics quantify disentanglement over a prompt distribution D\mathcal{D}:

    1. Cause Score: Measures whether intervening on FAF_A using an input x′∈PE′A′∪WE′x' \in P_{E'}^{A'} \cup W_{E'} successfully changes the predicted attribute value from AEA_E to AE′A_{E'} on target prompt x∈PEAx \in P_E^A:

    Cause(A,FA,M,D)≜Ex∈PEA,x′∈PE′A′∪WE′[I(II(M,FA,x,x′)=AE′)]\text{Cause}(A, F_A, M, \mathcal{D}) \triangleq \mathbb{E}_{x \in P_E^A, x' \in P_{E'}^{A'} \cup W_{E'}} \left[ \mathbb{I}\left(\text{II}(M, F_A, x, x') = A_{E'}\right) \right]

    1. Isolation (Iso) Score: Measures whether intervening on FAF_A leaves non-target attributes A∗∈A∖{A}A^* \in \mathcal{A} \setminus \{A\} unaffected when evaluated on prompt x∗∈PEA∗x^* \in P_E^{A^*}:

    Iso(A,FA,M,D)≜1∣A∖{A}∣∑A∗∈A∖{A}Ex∗∈PEA∗,x′∈PE′A′∪WE′[I(II(M,FA,x∗,x′)=AE∗)]\text{Iso}(A, F_A, M, \mathcal{D}) \triangleq \frac{1}{|\mathcal{A} \setminus \{A\}|} \sum_{A^* \in \mathcal{A} \setminus \{A\}} \mathbb{E}_{x^* \in P_E^{A^*}, x' \in P_{E'}^{A'} \cup W_{E'}} \left[ \mathbb{I}\left(\text{II}(M, F_A, x^*, x') = A^*_E\right) \right]

    1. Disentangle Score: The arithmetic mean of the Cause and Iso scores:

    Disentangle(A,FA,M,D)≜12(Cause(A,FA,M,D)+Iso(A,FA,M,D))\text{Disentangle}(A, F_A, M, \mathcal{D}) \triangleq \frac{1}{2} \left( \text{Cause}(A, F_A, M, \mathcal{D}) + \text{Iso}(A, F_A, M, \mathcal{D}) \right)

    The overall benchmark score for an entity type is the average Disentangle score across all its attributes.

  3. Knowl 3 — Multi-Task Distributed Alignment Search (MDAS)

    model/method

    Multi-task Distributed Alignment Search (MDAS) is an interpretability method that learns a linear subspace FAF_A of neural representations NN to isolate an attribute AA while preserving all other attributes A∗∈A∖{A}A^* \in \mathcal{A} \setminus \{A\}.

    Whereas standard Distributed Alignment Search (DAS) only optimizes the causal intervention loss LCause\mathcal{L}_{\text{Cause}}, MDAS jointly trains on a counterfactual cause objective and a multi-task isolation objective:

    LCause(A,FA,M)=CE(II(M,FA,x,x′),AE′)\mathcal{L}_{\text{Cause}}(A, F_A, M) = \text{CE}\left(\text{II}(M, F_A, x, x'), A_{E'}\right)

    LIso(A∗,FA,M)=CE(II(M(x),FA,x′),AE∗)\mathcal{L}_{\text{Iso}}(A^*, F_A, M) = \text{CE}\left(\text{II}(M(x), F_A, x'), A^*_E\right)

    LDisentangle(A,FA,M)=LCause(A,FA,M)+1∣A∖{A}∣∑A∗∈A∖{A}LIso(A∗,FA,M)\mathcal{L}_{\text{Disentangle}}(A, F_A, M) = \mathcal{L}_{\text{Cause}}(A, F_A, M) + \frac{1}{|\mathcal{A} \setminus \{A\}|} \sum_{A^* \in \mathcal{A} \setminus \{A\}} \mathcal{L}_{\text{Iso}}(A^*, F_A, M)

    where CE\text{CE} denotes cross-entropy loss, x∈PEAx \in P_E^A, x′∈PE′A′∪WE′x' \in P_{E'}^{A'} \cup W_{E'}, and x∗∈PEA∗x^* \in P_E^{A^*}. Through this multi-task loss, MDAS directly penalizes the feature FAF_A if an intervention on FAF_A perturbs the model's outputs for non-target attributes.

  4. Knowl 4 — Low-Rank Linear Subspace Interchange Interventions in DAS and MDAS

    equation

    To avoid instantiating and learning an expensive full n×nn \times n orthogonal rotation matrix QQ over hidden dimension nn, Distributed Alignment Search (DAS) and Multi-task DAS (MDAS) parameterize the target feature subspace FAF_A using an orthonormal projection matrix W∈Rk×nW \in \mathbb{R}^{k \times n} containing k≪nk \ll n orthonormal row vectors (WW⊤=IkW W^\top = I_k).

    The interchange intervention on neural activation vector GetVals(M(x),N)∈Rn\text{GetVals}(M(x), N) \in \mathbb{R}^n with source activation GetVals(M(x′),N)∈Rn\text{GetVals}(M(x'), N) \in \mathbb{R}^n is computed via orthogonal projection:

    II(M,FA,x,x′)=(In−W⊤W)GetVals(M(x),N)+W⊤WGetVals(M(x′),N)\text{II}(M, F_A, x, x') = \left(I_n - W^\top W\right) \text{GetVals}(M(x), N) + W^\top W \text{GetVals}(M(x'), N)

    Here, W⊤W∈Rn×nW^\top W \in \mathbb{R}^{n \times n} is the rank-kk projection operator onto the feature subspace FAF_A, and In−W⊤WI_n - W^\top W projects onto the orthogonal complement (null space) of FAF_A.

  5. Knowl 5 — Multi-Task Differential Binary Masking (MDBM)

    model/method

    Multi-task Differential Binary Masking (MDBM) selects a sparse subset of individual neurons to represent attribute AA while minimizing interference on other attributes.

    MDBM parameterizes the neuron mask using continuous learnable vector m∈Rnm \in \mathbb{R}^n with temperature TT annealed during training from 10−210^{-2} to 10−710^{-7}. The intervened activation vector is:

    n=(1−σ(m/T))∘GetVals(M(x),N)+σ(m/T)∘GetVals(M(x′),N)n = \left(1 - \sigma(m/T)\right) \circ \text{GetVals}(M(x), N) + \sigma(m/T) \circ \text{GetVals}(M(x'), N)

    where σ\sigma is the sigmoid function and ∘\circ denotes element-wise multiplication.

    While standard DBM minimizes LCause=CE(τ(MN←n(x)),AE′)+λ∥m∥1\mathcal{L}_{\text{Cause}} = \text{CE}(\tau(M_{N \leftarrow n}(x)), A_{E'}) + \lambda \|m\|_1, MDBM optimizes the multi-task objective:

    LDisentangle(A,FA,M)=LCause(A,FA,M)+1∣A∖{A}∣∑A∗∈A∖{A}LIso(A∗,FA,M)\mathcal{L}_{\text{Disentangle}}(A, F_A, M) = \mathcal{L}_{\text{Cause}}(A, F_A, M) + \frac{1}{|\mathcal{A} \setminus \{A\}|} \sum_{A^* \in \mathcal{A} \setminus \{A\}} \mathcal{L}_{\text{Iso}}(A^*, F_A, M)

    The multi-task isolation objective naturally encourages sparsity by penalizing unnecessary neurons, eliminating the need for an explicit L1L_1 penalty (optimal λ=0\lambda = 0 for MDBM).

  6. Knowl 6 — Comparative Evaluation of Interpretability Methods on RAVEL

    data/table

    Disentanglement scores on RAVEL for Llama2-7B across 5 entity types under the Entity split (unseen entities) and Context split (unseen prompt templates):

    Method Supervision Entity (%) Context (%)
    Full Rep. None 40.5 39.5
    PCA None 39.5 39.1
    Sparse Autoencoder (SAE) None 48.6 46.8
    RLAP Attribute 48.8 50.9
    DBM Counterfactual 52.2 49.8
    DAS Counterfactual 56.5 57.3
    MDBM Counterfactual 53.7 53.9
    MDAS Counterfactual 60.1 65.6

    These results establish that:

    1. MDAS achieves state-of-the-art disentanglement on both splits (60.1% Entity, 65.6% Context).
    2. Methods utilizing counterfactual supervision with interventions (DBM, DAS, MDBM, MDAS) outperform methods based on linear attribute probing (RLAP) and unsupervised featurizers (PCA, SAE).
    3. Adding multi-task isolation objectives improves overall disentanglement performance: MDBM improves upon DBM by +1.5% (Entity) / +4.1% (Context), and MDAS improves upon DAS by +3.6% (Entity) / +8.3% (Context).
  7. Knowl 7 — Dimensionality Trade-off Between Causal Efficacy and Concept Isolation

    empirical result

    Across all evaluated interpretability methods on RAVEL (PCA, Sparse Autoencoder, RLAP, DBM, MDBM, DAS, MDAS), the size of the identified feature subspace FAF_A governs an inverse relationship between the Cause score and the Iso score:

    • Intervening on a larger fraction of dimensions (e.g., ≈50%\approx 50\% of the activation dimension) increases the Cause score (often exceeding 0.70–0.80) but decreases the Iso score (dropping below 0.40–0.50), because larger subspaces inevitably perturb other entity attributes.
    • Intervening on very few dimensions (e.g., ≈1%\approx 1\%) produces high Iso scores but very low Cause scores (often <0.20< 0.20), failing to reliably change the model's behavior to the target counterfactual.
    • MDAS achieves the highest Disentangle score while intervening on a compact subspace consisting of only 4%4\% of the residual stream dimension (k=128k = 128 out of n=4096n = 4096), balancing Cause and Iso significantly better than neuron-aligned masks (DBM/MDBM).
  8. Knowl 8 — Layerwise Dynamics of Attribute Disentanglement in Transformer Residual Streams

    empirical result

    In Llama2-7B, attribute representations across the 32 Transformer layers show progressive disentanglement:

    • In early layers (layers 1–7), features identified at the last entity token fail to generalize to unseen entities, resulting in near-zero Cause scores.
    • Intermediate representations (around layer 8) begin achieving high Cause scores (>0.60> 0.60), but exhibit low Isolation (Iso ≈0.50\approx 0.50), indicating that multiple attributes are still tightly coupled in the same subspace.
    • As representations pass through deeper layers, attribute disentanglement increases: Iso score rises from 0.500.50 at layer 8 to 0.800.80 at layer 16 for city attributes, reaching the maximum Disentangle score at layer 16.
    • For other entity types (Nobel laureates, verbs, physical objects, occupations), the highest Disentangle scores occur around layer 7.
  9. Knowl 9 — Asymmetric Attribute Entanglement and Ripple Effects in Language Models

    empirical result

    Supervised interpretability methods on RAVEL reveal that certain attribute pairs within an entity type are intrinsically entangled in Llama2-7B's internal representations, regardless of the method used (DAS, MDAS, RLAP, DBM, MDBM):

    • Inseparable pairs: Interventions targeting the country attribute also change language (and vice versa), and interventions targeting latitude strongly affect longitude. Intervening on features for these attributes produces unavoidable ripple effects on each other.
    • Asymmetrically entangled attributes: In standard DAS, intervening on any city attribute feature alters continent and timezone predictions >60%>60\% of the time. MDAS successfully isolates continent and timezone from other attributes, reducing off-diagonal Cause scores from 0.50–0.850.50\text{--}0.85 down to 0.01–0.240.01\text{--}0.24.
    • Separable pairs: Certain attributes such as language and continent can be almost completely disentangled from each other.
  10. Knowl 10 — Architectural and Positional Limitations of the RAVEL Baseline Evaluation

    limitation

    The experimental baselines on RAVEL have two primary limitations:

    1. Model Scope: Experiments are evaluated on Llama2-7B, a decoder-only Transformer. Other architectures or training regimes may organize internal representations differently, requiring new benchmark instantiations.
    2. Intervention Site Scope: The intervention search space is restricted exclusively to the residual stream activation at the last token position of the entity (tEt_E). Entity attribute representations in autoregressive models can be distributed across other token positions in the prompt sequence or across multiple layers.

Coverage note — No substantial contributed material was omitted. Baseline implementation details (e.g., linear probe classification accuracies, scikit-learn regularization hyperparameters, and full per-attribute breakdown tables for all 5 domains) are summarized in the main benchmark, method, and empirical result knowls.

References

  1. 1.Eldar David Abraham, Karel D’Oosterlinck, Amir Feder, Yair Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu. 2022. CEBaB: Estimating the causal effects of real-world concepts on NLP model behavior. In Advances in Neural Information Processing Systems (NeurIPS).
  2. 2.Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018. Linear algebraic structure of word senses, with applications to polysemy. In Transactions of the Association of Computational Linguistics (TACL).
  3. 3.Pepa Atanasova, Oana-Maria Camburu, Christina Lioma, Thomas Lukasiewicz, Jakob Grue Simonsen, and Isabelle Augenstein. 2023. Faithfulness tests for natural language explanations. In Association for Computational Linguistics (ACL).
  4. 4.Sander Beckers, Frederick Eberhardt, and Joseph Y. Halpern. 2020. Approximate causal abstractions. In Uncertainty in Artificial Intelligence Conference (UAI).
  5. 5.Sander Beckers and Joseph Y. Halpern. 2019. Abstracting causal models. In Conference on Artificial Intelligence (AAAI).
  6. 6.Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023. LEACE: Perfect linear concept erasure in closed form. In Advances in Neural Information Processing Systems (NeurIPS).
  7. 7.Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda B. Viégas, and Martin Wattenberg. 2021. An interpretability illusion for BERT. In arXiv preprint arXiv:2104.07143.
  8. 8.Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. 2023. Towards monosemanticity: Decomposing language models with dictionary learning. In Transformer Circuits Thread.
  9. 9.Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim. 2020. Thread: Circuits. In Distill.
  10. 10.Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, and Ivan Titov. 2020. How do decisions emerge across layers in neural models? Interpretation with differentiable masking. In Empirical Methods in Natural Language Processing (EMNLP).
  11. 11.Nicola De Cao, Leon Schmid, Dieuwke Hupkes, and Ivan Titov. 2022. Sparse interventions in language models with differentiable masking. In Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP.
  12. 12.Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas. 2022. Causal scrubbing: a method for rigorously testing interpretability hypotheses. In Alignment Forum Blog post.
  13. 13.Pattarawat Chormai, Jan Herrmann, Klaus-Robert Müller, and Grégoire Montavon. 2022. Disentangled explanations of neural network predictions by finding relevant subspaces. In arXiv preprint arXiv:2212.14855.
  14. 14.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? An analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP.
  15. 15.Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2024. Evaluating the ripple effects of knowledge editing in language models. In Transactions of the Association of Computational Linguistics (TACL).
  16. 16.Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023. Towards automated circuit discovery for mechanistic interpretability. In Advances in Neural Information Processing Systems (NeurIPS).
  17. 17.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Association for Computational Linguistics (ACL).
  18. 18.Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber. 2021. Are neural nets modular? Inspecting functional modularity through differentiable weight masks. In International Conference on Learning Representations (ICLR).
  19. 19.Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2024. Sparse autoencoders find highly interpretable features in language models. In International Conference on Learning Representations (ICLR).
  20. 20.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Association for Computational Linguistics (ACL).
  21. 21.Xander Davies, Max Nadeau, Nikhil Prakash, Tamar Rott Shaham, and David Bau. 2023. Discovering variable binding circuitry with desiderata. In arXiv preprint arXiv:2307.03637.
  22. 22.Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021. Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals. In Transactions of the Association of Computational Linguistics (TACL).
  23. 23.Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy models of superposition. In arXiv preprint arXiv:2209.10652.
  24. 24.Jiahai Feng and Jacob Steinhardt. 2024. How do language models bind entities in context? In International Conference on Learning Representations (ICLR).
  25. 25.Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov. 2021. Causal analysis of syntactic agreement mechanisms in neural language models. In Association for Computational Linguistics and International Joint Conference on Natural Language Processing (ACL-IJCNLP).
  26. 26.Atticus Geiger, Hanson Lu, Thomas F Icard, and Christopher Potts. 2021. Causal abstractions of neural networks. In Advances in Neural Information Processing Systems (NeurIPS).
  27. 27.Atticus Geiger, Christopher Potts, and Thomas Icard. 2023a. Causal abstraction for faithful model interpretation. Ms., Stanford University.
  28. 28.Atticus Geiger, Kyle Richardson, and Chris Potts. 2020. Neural natural language inference models partially embed theories of lexical entailment and negation. In Proceedings of the 2020 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP.
  29. 29.Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman. 2023b. Finding alignments between interpretable causal variables and distributed neural representations. In Causal Learning and Reasoning (CLeaR).
  30. 30.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. Dissecting recall of factual associations in auto-regressive language models. In Empirical Methods in Natural Language Processing (EMNLP).
  31. 31.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Empirical Methods in Natural Language Processing (EMNLP).
  32. 32.Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024. Patchscopes: A unifying framework for inspecting hidden representations of language models. In arXiv preprint arXiv:2401.06102.
  33. 33.Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora. 2023. Localizing model behavior with path patching. In arXiv preprint arXiv:2304.05969.
  34. 34.Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023. Finding neurons in a haystack: Case studies with sparse probing. In Transactions on Machine Learning Research (TMLR).
  35. 35.Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023. How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model. In Advances in Neural Information Processing Systems (NeurIPS).
  36. 36.Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models. In Advances in Neural Information Processing Systems (NeurIPS).
  37. 37.Roee Hendel, Mor Geva, and Amir Globerson. 2023. In-context learning creates task vectors. In Empirical Methods in Natural Language Processing (EMNLP).
  38. 38.Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. 2024. Linearity of relation decoding in transformer language models. In International Conference on Learning Representations (ICLR).
  39. 39.John Hewitt, Kawin Ethayarajh, Percy Liang, and Christopher Manning. 2021. Conditional probing: measuring usable information beyond a baseline. In Empirical Methods in Natural Language Processing (EMNLP).
  40. 40.Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu, and Christopher Potts. 2023. Rigorously assessing natural language explanations of neurons. In Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP.
  41. 41.Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018. Visualisation and “diagnostic classifiers” reveal how recurrent and recursive neural networks process hierarchical structure. In Journal of Artificial Intelligence Research (JAIR).
  42. 42.Belinda Z. Li, Maxwell I. Nye, and Jacob Andreas. 2021. Implicit representations of meaning in neural language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 1813–1827. Association for Computational Linguistics.
  43. 43.Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik. 2023. Does circuit analysis interpretability scale? Evidence from multiple choice capabilities in chinchilla. In arXiv preprint arXiv:2307.09458.
  44. 44.Francesco Locatello, Stefan Bauer, Mario Lucic, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. 2018. Challenging common assumptions in the unsupervised learning of disentangled representations. CoRR, abs/1811.12359.
  45. 45.Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020. Emergent linguistic structure in artificial neural networks trained by self-supervision. In Proceedings of the National Academy of Sciences (PNAS).
  46. 46.Samuel Marks and Max Tegmark. 2023. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In arXiv preprint arXiv:2310.06824.
  47. 47.J. L. McClelland, D. E. Rumelhart, and PDP Research Group, editors. 1986. Parallel Distributed Processing. Volume 2: Psychological and Biological Models. MIT Press, Cambridge, MA.
  48. 48.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems (NeurIPS).
  49. 49.Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2023. A mechanism for solving relational tasks in transformer language models. In arXiv preprint arXiv:2305.16130.
  50. 50.Edmund Mills, Shiye Su, Stuart Russell, and Scott Emmons. 2023. Almanacs: A simulatability benchmark for language model explainability. In arXiv preprint arXiv:2312.12747.
  51. 51.Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020. Zoom in: An introduction to circuits. In Distill.
  52. 52.Kiho Park, Yo Joong Choe, and Victor Veitch. 2023. The linear representation hypothesis and the geometry of large language models. In arXiv preprint arXiv:2311.03658.
  53. 53.Judea Pearl. 2001. Direct and indirect effects. In Proceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence, UAI’01, pages 411–420, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
  54. 54.Judea Pearl. 2009. Causality. Cambridge University Press.
  55. 55.Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018. Dissecting contextual word embeddings: Architecture and representation. In Empirical Methods in Natural Language Processing (EMNLP).
  56. 56.Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020. Information-theoretic probing for linguistic structure. In Association for Computational Linguistics (ACL).
  57. 57.Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020. Null it out: Guarding protected attributes by iterative nullspace projection. In Association for Computational Linguistics (ACL).
  58. 58.Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell. 2022. Linear adversarial concept erasure. In International Conference on Machine Learning (ICML).
  59. 59.D. E. Rumelhart, J. L. McClelland, and PDP Research Group, editors. 1986. Parallel Distributed Processing. Volume 1: Foundations. MIT Press, Cambridge, MA.
  60. 60.Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio. 2021. Toward causal representation learning. Proc. IEEE, 109(5):612–634.
  61. 61.Sarah Schwettmann, Tamar Rott Shaham, Joanna Materzynska, Neil Chowdhury, Shuang Li, Jacob Andreas, David Bau, and Antonio Torralba. 2023. Find: A function description benchmark for evaluating interpretability methods. In Advances in Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks.
  62. 62.Paul Smolensky. 1988. On the proper treatment of connectionism. Behavioral and Brain Sciences, 11(1):1–23.
  63. 63.Peter Spirtes, Clark Glymour, and Richard Scheines. 2000. Causation, Prediction, and Search. MIT Press.
  64. 64.Alex Tamkin, Mohammad Taufeeque, and Noah D. Goodman. 2023. Codebook features: Sparse and discrete interpretability for neural networks. In arXiv preprint arXiv:2310.17230.
  65. 65.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT rediscovers the classical NLP pipeline. In Association for Computational Linguistics (ACL).
  66. 66.Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023. Linear representations of sentiment in large language models. In arXiv preprint arXiv:2310.15154.
  67. 67.Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. 2024. Function vectors in large language models. In International Conference on Learning Representations (ICLR).
  68. 68.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. In arXiv preprint arXiv:2307.09288.
  69. 69.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS).
  70. 70.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems (NeurIPS).
  71. 71.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In International Conference on Learning Representations (ICLR).
  72. 72.Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah D. Goodman. 2023. Interpretability at scale: Identifying causal mechanisms in alpaca. In Advances in Neural Information Processing Systems (NeurIPS).
  73. 73.Zexuan Zhong, Zhengxuan Wu, Christopher Manning, Christopher Potts, and Danqi Chen. 2023. MQuAKE: Assessing knowledge editing in language models via multi-hop questions. In Empirical Methods in Natural Language Processing (EMNLP).

Citation

MLA
Huang, J., et al. “RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 8669–87, https://doi.org/10.18653/v1/2024.acl-long.470.
APA
Huang, J., Wu, Z., Potts, C., Geva, M., & Geiger, A. (2024). RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8669–8687. https://doi.org/10.18653/v1/2024.acl-long.470
Chicago
Huang, J., Z. Wu, C. Potts, M. Geva, and A. Geiger. 2024. “RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8669–87. https://doi.org/10.18653/v1/2024.acl-long.470.
Harvard
Huang, J. et al. (2024) “RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8669–8687. Available at: https://doi.org/10.18653/v1/2024.acl-long.470.
Vancouver
1. Huang J, Wu Z, Potts C, Geva M, Geiger A (2024) RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8669–8687

BibTeX

@inproceedings{huang-etal-2024-ravel,
    title = "{RAVEL}: Evaluating Interpretability Methods on Disentangling Language Model Representations",
    author = "Huang, Jing  and
      Wu, Zhengxuan  and
      Potts, Christopher  and
      Geva, Mor  and
      Geiger, Atticus",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.470/",
    doi = "10.18653/v1/2024.acl-long.470",
    pages = "8669--8687"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/