RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
Jing HuangZhengxuan WuChristopher PottsMor GevaAtticus Geiger
Presents a diagnostic benchmark and a multi-task alignment method to quantitatively evaluate and improve how interpretability techniques disentangle polysemantic, distributed representations across language model activations.
Modern artificial intelligence relies heavily on large language models, yet understanding how these complex systems internally represent distinct pieces of knowledge remains a critical challenge. Individual neural units within these models are polysemantic, meaning a single component participates in encoding multiple unrelated concepts simultaneously. While numerous interpretability techniques have been developed to map these complex internal states into distinct, understandable features, the field has lacked standardized, quantitative frameworks to measure how effectively these techniques actually isolate specific concepts without unintentionally altering others.
To address this gap, the article introduces the RAVEL diagnostic benchmark, designed to evaluate and compare interpretability techniques on their ability to localize and disentangle specific attributes of entities represented inside language models. Using counterfactual interchange interventions—which alter model states during text processing to observe causal behavioral changes—the benchmark measures whether a technique successfully causes a targeted concept to change while isolating and preserving unrelated attributes.
Across extensive experiments evaluating multiple interpretability families on the Llama2-7B language model, the article establishes four primary findings. First, methods utilizing counterfactual supervision significantly outperform unsupervised approaches; specifically, unsupervised techniques like Principal Component Analysis and sparse autoencoders struggled with disentanglement (scoring roughly 39% to 49%), whereas counterfactually supervised methods performed substantially better. Second, the article introduces Multi-task Distributed Alignment Search, which achieves the state-of-the-art disentanglement score on the benchmark (reaching 60.1% on unseen entities and 65.6% on unseen prompt contexts) by incorporating isolation criteria directly into the training objective. Third, certain real-world attribute pairs, such as country and language or latitude and longitude, remain consistently difficult for any method to separate due to fundamental entanglements in how models organize knowledge. Fourth, internal representations become progressively more disentangled in deeper layers of the model, with concept isolation peaking around intermediate and later layers.
These findings provide crucial guidance for technical leaders and teams deploying language models in high-stakes environments. They demonstrate that understanding and controlling model behavior requires analyzing distributed representations rather than assuming individual neurons hold discrete concepts. Relying on unsupervised feature extractors or simple probing risks mischaracterizing model mechanisms, which could lead to flawed model audits or ineffective safety interventions. Instead, utilizing multi-task causal alignment offers a more reliable, faithful approach to isolating and steering specific model behaviors.
Organizations seeking to interpret or edit language models should adopt multi-task causal intervention methods when isolating internal concepts and run evaluations across multiple prompt contexts to verify generalizability. For future development, technical teams should instantiate the RAVEL benchmark on other emerging model architectures and expand intervention analyses beyond entity-level tokens to explore multi-token and multi-layer dynamics across broader network components.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). This paper establishes the foundational use of sparse autoencoders to resolve polysemanticity and extract monosemantic feature dictionaries in language models, which RAVEL directly benchmarks and evaluates for concept disentanglement.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). This work introduces representation engineering and linear directional reading methods like PCA for model concepts, providing the core representation-level interpretability baselines evaluated by the RAVEL benchmark.
- Paper: Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV), Been Kim et al. (2018). This foundational paper introduces concept activation vectors and directional derivatives for quantitative concept probing, establishing the linear concept representation paradigms tested in RAVEL.
- Paper: A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance Estimation, Thomas Fel et al. (2023). This work formalizes unsupervised concept extraction as dictionary learning, defining the extraction and importance frameworks that RAVEL assesses using causal interventions.
- Paper: Towards Automated Circuit Discovery for Mechanistic Interpretability, Arthur Conmy et al. (2023). This paper formalizes causal intervention and activation patching for discovering model mechanisms, which underpins the counterfactual interchange intervention methodology utilized throughout RAVEL.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). This paper introduces unsupervised contrastive probing of internal representations, establishing the unsupervised representation probing setting that RAVEL shows struggles with attribute disentanglement.
- Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). This work formalizes the linear representation hypothesis using causal counterfactual concept pairs and metric inner products, providing theoretical justification for the causal alignment behaviors observed in RAVEL.
- Paper: On the Origins of Linear Representations in Large Language Models, Yibo Jiang et al. (2024). This paper presents a theoretical framework explaining why language models naturally learn linear and orthogonal concept representations during training, providing mathematical foundations for the empirical disentanglement dynamics analyzed in RAVEL.
- Paper: Scaling and evaluating sparse autoencoders, Leo Gao et al. (2025). This study scales sparse autoencoders using TopK activations to improve feature quality and disentanglement, addressing the key architectural limitations of unsupervised dictionary learning highlighted by RAVEL.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). This treatise provides a meta-level statistical and causal critique of unidentifiability and overdetermination across mechanistic interpretability techniques, extending RAVEL's warnings about the risks of unprincipled interpretability methods.
