SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Adam KarvonenCan RagerJohnny LinCurt TiggesJoseph Isaac BloomDavid ChaninYeu-Tong LauEoin FarrellCallum McDougallKola Ayonrinde

article2025ICML136 citations

Introduces a standardized evaluation suite of eight metrics and over 200 trained models to assess sparse autoencoder architectures on feature disentanglement, interpretability, and practical applications beyond traditional proxy loss.

Listen

As large language models become central to critical software systems, understanding their internal decision-making processes is essential for safety, compliance, and risk management. Researchers frequently employ Sparse Autoencoders (SAEs)—specialized neural network components that break down complex model internal states into distinct, understandable concepts. However, progress in the field has historically relied on simple unsupervised proxy metrics, primarily the trade-off between how accurately an SAE reconstructs the original model data and how sparsely it does so. This approach leaves major uncertainties about whether architectural improvements actually deliver practical interpretability and clean concept separation.

The article addresses this gap by introducing SAEBench, a standardized and extensible evaluation framework designed to assess SAE performance across diverse, practically relevant criteria. The primary objective is to evaluate over 200 trained SAEs across seven leading architectures and multiple model scales to determine how standard proxy metrics compare against real-world downstream tasks, including concept detection, interpretability, and practical applications like knowledge unlearning.

To conduct this evaluation, the researchers trained a comprehensive suite of SAEs sweeping across different dictionary capacities (4,000 to 65,000 features, and up to 1 million in existing public models) and activation sparsity levels on the Gemma-2-2B and Pythia-160M language models. The benchmark evaluates each system across eight standardized metrics categorized into four core dimensions: concept detection, automated interpretability judged by a language model, activation reconstruction fidelity, and feature disentanglement. The suite also introduces novel diagnostic techniques to measure how cleanly independent concepts can be isolated and ablated without causing unintended side effects.

The findings demonstrate that traditional proxy metrics do not reliably predict practical performance. First, Matryoshka SAEs—a hierarchical design—substantially outperform other architectures on concept detection and feature disentanglement, beating alternative approaches on five out of eight metrics and outperforming standard architectures by margins of 30% to 40% on specific concept removal tasks, despite showing lower reconstruction accuracy on traditional curves. Second, standard ReLU-based autoencoders perform the worst on five out of eight evaluations, confirming they are largely obsolete compared to modern alternatives. Third, increasing the dictionary size improves basic reconstruction and per-feature interpretability across all models, but it degrades feature disentanglement in non-hierarchical architectures due to excessive concept fragmentation; Matryoshka was the only architecture whose disentanglement capability improved with scale. Finally, the analysis shows that optimal sparsity is highly task-dependent, though a moderate sparsity range of 50 to 150 active features offers the best overall compromise across capabilities.

These results carry significant implications for the deployment and oversight of AI interpretability tools. Optimizing exclusively for reconstruction fidelity creates a false sense of progress while obscuring critical failure modes like feature absorption, where concepts become entangled or hidden. For organizations investing in AI safety, auditing, or targeted knowledge removal (such as erasing proprietary or hazardous data), selecting the appropriate architecture and evaluating across multidimensional metrics is critical to avoid wasted compute and unreliable safety guarantees.

The article recommends that practitioners adopt multi-metric evaluation suites rather than single-score proxies when developing or selecting SAEs. Development teams should test systems across a range of sparsity levels (specifically 20 to 200 active features) with directly comparable baselines, and consider hierarchical architectures like Matryoshka for tasks demanding precise concept isolation. For future work, the evaluation suite should be expanded to larger model architectures, additional network layers, and non-text modalities such as vision and biology models.

Confidence in these findings is high for medium-sized language models and the specific datasets evaluated, as results remained consistent across multiple seeds and scale sweeps. However, readers should note key limitations: supervised metrics currently depend on a limited set of ground-truth concepts (such as profession or syntax), unlearning evaluations are constrained by the baseline capabilities of the underlying model, and quantitative scoring cannot fully replace human-led qualitative analysis during deep safety investigations.

adamkarvonen/SAEBenchKarvonen et al (2025).pdf
Cover for SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Abstract

Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most prior work evaluates progress using unsupervised proxy metrics with unclear practical relevance. We introduce SAEBench, a comprehensive evaluation suite that measures SAE performance across eight diverse metrics, spanning interpretability, feature disentanglement and practical applications like unlearning. To enable systematic comparison, we open-source a suite of over 200 SAEs across seven recently proposed SAE architectures and training algorithms. Our evaluation reveals that gains on proxy metrics do not reliably translate to better practical performance. For instance, while Matryoshka SAEs slightly underperform on existing proxy metrics, they substantially outperform other architectures on feature disentanglement metrics; moreover, this advantage grows with SAE scale. By providing a standardized framework for measuring progress in SAE development, SAEBench enables researchers to study scaling trends and make nuanced comparisons between different SAE architectures and training methodologies. Our interactive interface enables researchers to flexibly visualize relationships between metrics across hundreds of open-source SAEs at neuropedia.org/sae-bench. Code and models available at: github.com/adamkarvonen/SAEBench

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 2.1. SAEs for Interpretability
  • 2.2. SAE Evaluations
  • 3. SAEBench: A Comprehensive Benchmark
  • 3.1. Metrics
  • 3.2. Existing Metrics
  • 3.2.1. TRADITIONAL SPARSITY-FIDELITY TRADEOFF
  • 3.2.2. AUTOMATED INTERPRETABILITY
  • 3.2.3. k-SPARSE PROBING
  • 3.2.4. RAVEL
  • 3.3. Adapted Metrics
  • 3.3.1. FEATURE ABSORPTION
  • 3.3.2. UNLEARNING CAPABILITY
  • 3.4. Novel Metrics
  • 3.5. Practitioner's Guide
  • 4. Results
  • 4.1. Comparing SAE Architectures
  • 4.2. Dictionary Size Scaling Dynamics
  • 4.3. Task-Dependent Optimal Sparsity
  • 4.4. Monitoring Training
  • 4.5. Model Scale Effects
  • 4.6. Unexpected Findings and Limitations
  • 5. Limitations
  • 6. Conclusion
  • Impact Statement
  • Acknowledgements
  • Author Contributions
  • References
  • A. Computational Requirements
  • B. SAE Training Details
  • C. Extended Related Work
  • C.1. SAE Benchmarks
  • D. Further Evaluation Details
  • PCA Baseline Implementation
  • Core Evaluation Metrics
  • LLMScoring / Automated Interpretability
  • Sparse Probing
  • RAVEL
  • Feature Absorption
  • Unlearning
  • Spurious Correlation Removal (SCR)
  • Targeted Probe Perturbation (TPP)
  • E. Baseline Comparison: SAEs on Randomly Initialized vs. Fully Trained Models
  • F. Additional Dictionary Scaling Analysis
  • G. Training Dynamics
  • H. Intervention Set Size Analysis
  • I. Gemma-Scope Evaluation Results
  • J. Further SAE Bench Evaluation Results

Knowls

  1. Knowl 1 — SAEBench evaluates four distinct dimensions of SAE quality

    model/method

    SAEBench is a benchmark for sparse autoencoders (SAEs) that treats quality as a multidimensional property rather than a single reconstruction–sparsity trade-off. It groups eight evaluations into four capabilities: concept detection (sparse probing and feature absorption), interpretability (automated interpretability), reconstruction (Loss Recovered), and feature disentanglement (unlearning, RAVEL, Spurious Correlation Removal, and Targeted Probe Perturbation). The evaluations include both existing measures and adapted or newly developed tests, and are intended to expose trade-offs that one proxy metric can miss.

  2. Knowl 2 — The benchmark compares a broad SAE suite under a shared training protocol

    experimental setup

    The authors evaluated more than 200 SAEs spanning seven architectures: ReLU, Matryoshka BatchTopK, TopK, BatchTopK, Gated, JumpReLU, and P-Annealing. Their main trained suite covered widths of 4k, 16k, and 65k latents, targeting L0L_0 sparsities of 20, 40, 80, 160, 320, and 640. The SAEs were trained on Gemma-2-2B residual-stream activations at layer 12 and Pythia-160M activations at layer 8. Training used 500M tokens from The Pile, a batch size of 2,048, context length 1,024, and learning rate 3×10−43\times10^{-4}; corresponding runs used the same data and data order. A separate evaluation covered Gemma-Scope SAEs with widths from 16k to 1M on Gemma-2-2B and Gemma-2-9B.

  3. Knowl 3 — Matryoshka BatchTopK leads on several concept and disentanglement tests

    empirical result

    On the 65k-width Gemma-2-2B suite, Matryoshka BatchTopK performed best on five of the eight benchmark metrics in the typical L0=40L_0=40–200200 range. Its strongest results included feature absorption, RAVEL, sparse probing, Spurious Correlation Removal (SCR), and Targeted Probe Perturbation (TPP), even though it ranked below TopK and BatchTopK on the sparsity–fidelity frontier. Its advantage was less consistent above L0=200L_0=200. The original ReLU SAE was outperformed by other architectures on five of eight metrics and had the weakest Loss Recovered results, but at L0>200L_0>200 it performed best on single-latent sparse probing and remained comparable in the unlearning evaluation. In the low-L0L_0 sparsity–fidelity regime, BatchTopK ranked ahead of TopK, JumpReLU, Gated, Matryoshka, P-Annealing, and ReLU; these differences narrowed at moderate L0L_0. Thus, rankings on the reconstruction–sparsity proxy did not reliably predict rankings on other tasks.

  4. Knowl 4 — Increasing dictionary width improves reconstruction but often harms concept isolation

    empirical result

    Across the evaluated architectures, scaling width from 4k to 16k to 65k latents generally improved Loss Recovered and automated-interpretability scores. For most non-hierarchical architectures, however, larger dictionaries worsened feature absorption and reduced SCR performance at a fixed intervention budget. Matryoshka was the exception: it showed only minor degradation in absorption and generally improved SCR with width, making it the only evaluated architecture that improved on feature-disentanglement measures with scale. The SCR degradation for other architectures persisted when intervention sizes were varied, so it was not explained solely by ablating a smaller fraction of a wider dictionary. The authors hypothesize that feature splitting or absorption contributes to the decline, but do not establish this as its cause. Width scaling used a fixed number of training steps and tokens, so compute increased with dictionary size.

  5. Knowl 5 — Feature absorption measures when auxiliary latents compensate for an incomplete feature

    model/method

    The feature-absorption evaluation tests whether an SAE represents a concept partly through latents other than its principal feature latents. The authors use first-letter classification as the target task: they train logistic-regression probes on residual-stream activations for tokens containing English letters, then identify principal SAE latents for each letter using sparse probing. Additional latents are counted as feature splits when adding them increases probe F1 by more than 0.030.03.

    On held-out tokens that the residual-stream probe classifies correctly, absorption is flagged when the principal latents account for less of the probe-direction signal than the residual stream does, while other SAE latents with positive projection onto the probe direction account for at least the configured minimum share of that signal. The score on a flagged token is the portion of the feature’s SAE representation supplied by those other, absorbing latents rather than the principal latents; unflagged tokens receive a score of zero. The aggregate absorption score averages over correctly classified test tokens, and SAEBench reports its complement, 1−absorption score1-\text{absorption score}, so higher values indicate less absorption. This procedure is designed to detect partial absorption and compensation shared across multiple latents, not just complete suppression of a principal latent.

  6. Knowl 6 — SCR tests whether ablating spurious-signal latents improves a biased classifier

    model/method

    Spurious Correlation Removal (SCR) adapts the SHIFT procedure to measure whether an SAE separates a classifier’s intended signal from a correlated, unwanted signal. The evaluation creates biased binary datasets—for example, profession and gender in Bias in Bios, or product category and sentiment in Amazon Reviews—and trains a linear classifier on the biased examples. It identifies SAE latents associated with the spurious signal using absolute attribution from a probe trained to predict that signal, then zero-ablates those latents in the classifier’s SAE representation. The modified classifier is evaluated on a balanced dataset, where improved accuracy on the intended label indicates that removing the spurious signal helped.

    The normalized SCR score is

    SSHIFT=Aabl−AbaseAoracle−Abase,S_{\mathrm{SHIFT}}=\frac{A_{\mathrm{abl}}-A_{\mathrm{base}}}{A_{\mathrm{oracle}}-A_{\mathrm{base}}},

    where AablA_{\mathrm{abl}} is accuracy after latent ablation, AbaseA_{\mathrm{base}} is the biased classifier’s accuracy before ablation, and AoracleA_{\mathrm{oracle}} is the accuracy of a probe trained directly on the intended label. The main analysis uses 20-latent interventions, with sweeps over other intervention sizes.

  7. Knowl 7 — TPP measures whether class-specific latent ablations selectively damage matching probes

    model/method

    Targeted Probe Perturbation (TPP) extends the zero-ablation logic of SCR to multiclass datasets without requiring correlated labels. For a dataset with mm classes, the method trains a linear probe for each class jj and selects a set of SAE latents LiL_i for each class ii using the largest signed importance scores for that class. It evaluates every class probe both normally and after ablating each class-specific latent set. Let AjA_j be the accuracy of the unablated probe for class jj, and let Ai,jA_{i,j} be its accuracy after ablating LiL_i. The score is

    STPP=mean⁡i=j(Ai,j−Aj)−mean⁡i≠j(Ai,j−Aj).S_{\mathrm{TPP}}=\operatorname{mean}_{i=j}(A_{i,j}-A_j)-\operatorname{mean}_{i\ne j}(A_{i,j}-A_j).

    A high score indicates that ablating latents selected for one class disproportionately reduces the matching class probe’s accuracy while leaving other class probes relatively unaffected. The evaluation uses 4,000 training and 1,000 test sequences per task, truncated to 128 tokens; the main comparisons use 20-latent ablations.

  8. Knowl 8 — RAVEL evaluates attribute changes while checking for unintended effects

    model/method

    The RAVEL evaluation measures whether SAE-latent interventions can change a selected attribute of an entity while preserving other attributes. For example, an intervention might change a model’s predicted country for a city without changing its predicted language. SAEBench evaluates cities with country, continent, and language attributes, and Nobel Prize winners with country of birth, field, and gender attributes. It selects entities and templates with high base-model prediction accuracy, then uses a Multitask Differentiable Binary Mask (MDBM) to select latents while jointly optimizing the Cause and Isolation metrics. The reported disentanglement score averages those two metrics: Cause measures whether the requested attribute changes, and Isolation measures whether other attributes remain intact. Interventions transfer latent values at the entity’s final token; the evaluation does not add the SAE reconstruction-error term back during RAVEL interventions.

  9. Knowl 9 — The unlearning test rewards domain removal with retained general performance

    model/method

    SAEBench evaluates SAE-assisted unlearning by targeting biology knowledge in WMDP-bio while checking unrelated capabilities on selected MMLU categories. It identifies candidate latents by comparing activation frequency on a forget set (WMDP-bio corpus) and a retain set (WikiText). Latents exceeding the retain-set sparsity threshold are excluded; among the rest, the most frequent forget-set latents are selected and clamped to a negative value whenever they activate. The tested settings use retain thresholds of 0.0010.001 or 0.010.01, 10 or 20 latents, and negative-clamp magnitudes of 25, 50, 100, or 200. Evaluation considers only questions the unmodified model answers correctly across all option permutations. For each SAE, the reported score is the lowest WMDP-bio accuracy among settings that preserve MMLU accuracy above 0.990.99; lower target-domain accuracy under that constraint indicates more effective selective unlearning.

  10. Knowl 10 — Loss Recovered normalizes the effect of SAE reconstruction on model loss

    equation

    Loss Recovered measures how well replacing a model activation with its SAE reconstruction preserves next-token prediction performance, relative to zero-ablating that activation. It is defined as

    Loss Recovered=H∗−H0Horig−H0,\text{Loss Recovered}=\frac{H^*-H_0}{H_{\mathrm{orig}}-H_0},

    where HorigH_{\mathrm{orig}} is the model’s cross-entropy loss with its original activation, H∗H^* is the loss after substituting the SAE reconstruction for that activation, and H0H_0 is the loss after zero-ablating the activation. The losses are evaluated on the same prediction task and data. A value closer to 1 means the reconstruction recovers more of the loss increase caused by zero ablation.

  11. Knowl 11 — Automated interpretability scores whether an LLM can predict feature activations from a description

    model/method

    The automated-interpretability evaluation uses an LLM to generate a natural-language description of an SAE latent from sequences that activate it, then tests whether another LLM judgment can use that description to identify activating sequences. In the reported setup, GPT-4o-mini generates explanations from high-activation and activation-weighted web-text examples. For each latent, the shuffled test set contains 10 randomly sampled sequences, two maximally activating sequences, and two importance-weighted sequences. A judge predicts which sequences activate the latent; prediction accuracy is the interpretability score. The evaluation samples 1,000 non-dead latents and uses 2 million activation-dataset tokens with 128-token context.

  12. Knowl 12 — Sparse probing tests concept detection using a small selected set of SAE latents

    model/method

    Sparse probing asks whether a small number of SAE latents can support classification of a labeled concept, even though the SAE was not trained with that label. For each one-versus-all task, the method encodes the examples, mean-pools latent activations across non-padding tokens, and ranks latents by the difference in mean activation between positive and negative examples. It trains a logistic-regression probe on the top kk latents and measures held-out accuracy; SAEBench evaluates k∈{1,2,5}k\in\{1,2,5\} and emphasizes k=1k=1. The suite covers 35 binary tasks from Bias in Bios, Amazon Reviews, Europarl, GitHub, and AG News, including profession, product, sentiment, language, programming-language, and news-topic classification. Each task uses 4,000 training and 1,000 test examples, with inputs truncated to 128 tokens.

  13. Knowl 13 — Moderate sparsity is a practical compromise, but the best level depends on the task

    empirical result

    No single L0L_0 level performed best across SAEBench tasks. Lower L0L_0 (sparser codes) often improved automated interpretability; higher L0L_0 generally improved reconstruction, RAVEL, and TPP, and reduced feature absorption. Moderate sparsity sometimes balanced these trends and sometimes performed best on sparse probing or SCR. The authors recommend L0L_0 values around 50–150 as a reasonable cross-task compromise, not as a universal optimum. They also observed that RAVEL and especially TPP could favor L0>400L_0>400, outside the usual 20–200 range, while the related SCR measure did not show the same preference.

  14. Knowl 14 — Many evaluation metrics approach most of their performance by 50 million training tokens

    empirical result

    Training checkpoints of 16k-width TopK and ReLU SAEs on Gemma-2-2B were evaluated at 0, 5M, 15M, 50M, 150M, and 500M training tokens. Across many metrics, the largest gains occurred by 50M tokens, followed by slower improvement through later checkpoints. The results do not imply that training beyond 50M tokens is unnecessary: the authors note that small numerical changes may correspond to important qualitative differences and that larger dictionaries may benefit from greater training budgets.

  15. Knowl 15 — Metric behavior depends on the model, layer, labels, and limits of quantitative evaluation

    limitation

    SAEBench’s supervised measures cover only concepts with available reliable labels, a small subset of the concepts represented in language models; this limited coverage can make some results noisy. Results on Gemma-2-2B and Pythia-160M also differed: reconstruction trends were similar, but supervised metrics such as SCR, TPP, sparse probing, and absorption behaved differently, and the Matryoshka SCR advantage seen on Gemma-2-2B did not appear on Pythia-160M. The authors suggest that these measures depend on the model’s ability to represent the tested concepts, so performance on a small model may constrain the evaluation itself. In Gemma-Scope evaluations, reconstruction and automated interpretability generally improved with width, whereas absorption, SCR, and TPP often degraded; unlearning also varied by layer and was near zero on the final evaluated layers. The benchmark does not directly measure all qualitative aspects of interpretability, and its metrics have different scales and noise levels, so the authors do not combine them into a single score.

Coverage note — The exploratory comparison of SAEs trained on randomly initialized versus fully trained models and the per-evaluation runtime accounting were omitted because they are ancillary to the benchmark design, core methods, and main SAE comparison findings.

References

  1. 1.Anthropic Interpretability Team. Training sparse autoencoders. https://transformer-circuits.pub/2024/april-update/index.html#training-saes, 2024a. [Accessed January 20, 2025].
  2. 2.Anthropic Interpretability Team. Circuits updates — august 2024. Transformer Circuits Thread, 2024b. URL https://transformer-circuits.pub/2024/august-update/index.html.
  3. 3.Ayonrinde, K. Adaptive sparse allocation with mutual choice & feature choice sparse autoencoders, 2024. URL https://arxiv.org/abs/2411.02124.
  4. 4.Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O. Pythia: A suite for analyzing large language models across training and scaling, 2023. URL https://arxiv.org/abs/2304.01373.
  5. 5.Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. https://transformer-circuits.pub/2023/monosemantic-features/index.html.
  6. 6.Bussmann, B., Leask, P., and Nanda, N. Batchtopk: A simple improvement for topk-saes, 2024a. URL https://www.alignmentforum.org/posts/Nkx6yWZNbAsfvic98/batchtopk-a-simple-improvement-for-topk-saes.
  7. 7.Bussmann, B., Leask, P., and Nanda, N. Learning multi-level features with matryoshka saes, December 19 2024b. URL https://www.alignmentforum.org/posts/rKM9b6B2LqwSB5ToN/learning-multi-level-features-with-matryoshka-saes. Alignment Forum.
  8. 8.Bussmann, B., Pearce, M., Leask, P., Bloom, J. I., Sharkey, L., and Nanda, N. Showing sae latents are not atomic using meta-saes, 2024c. URL https://www.alignmentforum.org/posts/TMAmHh4DdMr4nCSr5/showing-sae-latents-are-not-atomic-using-meta-saes.
  9. 9.Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J. A is for absorption: Studying feature splitting and absorption in sparse autoencoders, 2024a. URL https://arxiv.org/abs/2409.14507.
  10. 10.Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J. A is for absorption: Studying feature splitting and absorption in sparse autoencoders, 2024b. URL https://arxiv.org/abs/2409.14507.
  11. 11.Chaudhary, M. and Geiger, A. Evaluating open-source sparse autoencoders on disentangling factual knowledge in gpt-2 small, 2024. URL https://arxiv.org/abs/2409.04478.
  12. 12.Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models, 2023. URL https://arxiv.org/abs/2309.08600.
  13. 13.Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M. Not all language model features are linear, 2024. URL https://arxiv.org/abs/2405.14860.
  14. 14.Farrell, E., Lau, Y.-T., and Conmy, A. Applying sparse autoencoders to unlearn knowledge in language models, 2024. URL https://arxiv.org/abs/2410.19278.
  15. 15.Gao, L., Dupre la Tour, T., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J. Scaling and evaluating sparse autoencoders. arXiv preprint arXiv:2406.04093, 2024. URL https://arxiv.org/abs/2406.04093.
  16. 16.Gemma Team, Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Rame, A., Ferret, J., Liu, P., Tafti, P., Friesen, A., Casbon, M., Ramos, S., Kumar, R., Lan, C. L., Jerome, S., Tsitsulin, A., Vieillard, N., Stanczyk, P., Girgin, S., Momchev, N., Hoffman, M., Thakoor, S., Grill, J.-B., Neyshabur, B., Bachem, O., Walton, A., Severyn, A., Parrish, A., Ahmad, A., Hutchison, A., Abdagic, A., Carl, A., Shen, A., Brock, A., Coenen, A., Laforge, A., Paterson, A., Bastian, B., Piot, B., Wu, B., Royal, B., Chen, C., Kumar, C., Perry, C., Welty, C., Choquette-Choo, C. A., Sinopalnikov, D., Weinberger, D., Vijaykumar, D., Rogozinska, D., Herbison, D., Bandy, E., Wang, E., Noland, E., Moreira, E., Senter, E., Eltyshev, E., Visin, F., Rasskin, G., Wei, G., Cameron, G., Martins, G., Hashemi, H., Klimczak-Plucinska, H., Batra, H., Dhand, H., Nardini, I., Mein, J., Zhou, J., Svensson, J., Stanway, J., Chan, J., Zhou, J. P., Carrasqueira, J., Iljazi, J., Becker, J., Fernandez, J., van Amersfoort, J., Gordon, J., Lipschultz, J., Newlan, J., yeong Ji, J., Mohamed, K., Badola, K., Black, K., Millican, K., McDonell, K., Nguyen, K., Sodh ia, K., Greene, K., Sjoesund, L. L., Usui, L., Sifre, L., Heuermann, L., Lago, L., McNealus, L., Soares, L. B., Kilpatrick, L., Dixon, L., Martins, L., Reid, M., Singh, M., Iverson, M., Gorner, M., Velloso, M., Wirth, M., Davidow, M., Miller, M., Rahtz, M., Watson, M., Risdal, M., Kazemi, M., Moynihan, M., Zhang, M., Kahng, M., Park, M., Rahman, M., Khatwani, M., Dao, N., Bardoliwalla, N., Devanathan, N., Dumai, N., Chauhan, N., Wahltinez, O., Botarda, P., Barnes, P., Barham, P., Michel, P., Jin, P., Georgiev, P., Culliton, P., Kuppala, P., Comanescu, R., Merhej, R., Jana, R., Rokni, R. A., Agarwal, R., Mullins, R., Saadat, S., Carthy, S. M., Cogan, S., Perrin, S., Arnold, S. M. R., Krause, S., Dai, S., Garg, S., Sheth, S., Ronstrom, S., Chan, S., Jordan, T., Yu, T., Eccles, T., Hennigan, T., Kocisky, T., Doshi, T., Jain, V., Yadav, V., Meshram, V., Dharmadhikari, V., Barkley, W., Wei, W., Ye, W., Han, W., Kwon, W., Xu, X., Shen, Z., Gong, Z., Wei, Z., Cotruta, V., Kirk, P., Rao, A., Giang, M., Peran, L., Warkentin, T., Collins, E., Barral, J., Ghahramani, Z., Hadsell, R., Sculley, D., Banks, J., Dragan, A., Petrov, S., Vinyals, O., Dean, J., Hassabis, D., Kavukcuoglu, K., Farabet, C., Buchatskaya, E., Borgeaud, S., Fiedel, N., Joulin, A., Kenealy, K., Dadashi, R., and Andreev, A. Gemma 2: Improving open language models at a practical size, 2024. URL https://arxiv.org/abs/2408.00118.
  17. 17.Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D. Finding neurons in a haystack: Case studies with sparse probing, 2023. URL https://arxiv.org/abs/2305.01610.
  18. 18.Heap, T., Lawson, T., Farnik, L., and Aitchison, L. Sparse autoencoders can interpret randomly initialized transformers, 2025. URL https://arxiv.org/abs/2501.17727.
  19. 19.Huang, J., Wu, Z., Potts, C., Geva, M., and Geiger, A. Ravel: Evaluating interpretability methods on disentangling language model representations, 2024. URL https://arxiv.org/abs/2402.17700.
  20. 20.Karvonen, A., Wright, B., Rager, C., Angell, R., Brinkmann, J., Smith, L., Verdun, C. M., Bau, D., and Marks, S. Measuring progress in dictionary learning for language model interpretability with board game models, 2024. URL https://arxiv.org/abs/2408.00113.
  21. 21.Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramar, J., Dragan, A., Shah, R., and Nanda, N. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024. URL https://arxiv.org/abs/2408.05147.
  22. 22.Makelov, A., Lange, G., and Nanda, N. Towards principled evaluations of sparse autoencoders for interpretability and control, 2024. URL https://arxiv.org/abs/2405.08366.
  23. 23.Marks, L., Paren, A., Krueger, D., and Barez, F. Enhancing neural network interpretability with feature-aligned sparse autoencoders, 2024a. URL https://arxiv.org/abs/2411.01220.
  24. 24.Marks, S., Karvonen, A., and Mueller, A. dictionary learning. https://github.com/saprmarks/dictionary_learning, 2024b.
  25. 25.Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A. Sparse feature circuits: Discovering and editing interpretable causal graphs in language models, 2024c. URL https://arxiv.org/abs/2403.19647.
  26. 26.Mudide, A., Engels, J., Michaud, E. J., Tegmark, M., and Schroeder de Witt, C. Efficient dictionary learning with switch sparse autoencoders. arXiv preprint arXiv:2410.08201, 2024. URL https://arxiv.org/abs/2410.08201.
  27. 27.Nanda, N. Open Source Replication & Commentary on Anthropic’s Dictionary Learning Paper, Oct 2023. URL https://www.alignmentforum.org/posts/aPTgTKC45dWvL9XBF/open-source-replication-and-commentary-on-anthropic-s.
  28. 28.Paulo, G., Mallen, A., Juang, C., and Belrose, N. Automatically interpreting millions of features in large language models, 2024. URL https://arxiv.org/abs/2410.13928.
  29. 29.Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramar, J., Shah, R., and Nanda, N. Improving dictionary learning with gated sparse autoencoders. arXiv preprint arXiv:2404.16014, 2024a. URL https://arxiv.org/abs/2404.16014.
  30. 30.Rajamanoharan, S., Lieberum, T., Sonnerat, N., Conmy, A., Varma, V., Kramar, J., and Nanda, N. Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders. arXiv preprint arXiv:2407.14435, 2024b. URL https://arxiv.org/abs/2407.14435.
  31. 31.Taggart, G. M. Prolu: A nonlinearity for sparse autoencoders, 2024. https://www.alignmentforum.org/posts/HEpufTdakGTTKgoYF/prolu-a-nonlinearity-for-sparse-autoencoders.
  32. 32.Venhoff, C., Calinescu, A., Torr, P., and de Witt, C. S. Sage: Scalable ground truth evaluations for large sparse autoencoders, 2024. URL https://arxiv.org/abs/2410.07456.
  33. 33.Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022. URL https://arxiv.org/abs/2211.00593.

Citation

MLA
Karvonen, A., et al. “SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability”. arXiv, 2025, https://doi.org/10.48550/arxiv.2503.09532.
APA
Karvonen, A., Rager, C., Lin, J., Tigges, C., Bloom, J., Chanin, D., Lau, Y.-T., Farrell, E., McDougall, C., Ayonrinde, K., Till, D., Wearden, M., Conmy, A., Marks, S., & Nanda, N. (2025). SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. arXiv. https://doi.org/10.48550/arxiv.2503.09532
Chicago
Karvonen, A., C. Rager, J. Lin, et al. 2025. “SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2503.09532.
Harvard
Karvonen, A. et al. (2025) “SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability”. arXiv. Available at: https://doi.org/10.48550/arxiv.2503.09532.
Vancouver
1. Karvonen A, Rager C, Lin J, et al (2025) SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. https://doi.org/10.48550/arxiv.2503.09532

BibTeX

@misc{https://doi.org/10.48550/arxiv.2503.09532,
  doi = {10.48550/ARXIV.2503.09532},
  url = {https://arxiv.org/abs/2503.09532},
  author = {Karvonen, Adam and Rager, Can and Lin, Johnny and Tigges, Curt and Bloom, Joseph and Chanin, David and Lau, Yeu-Tong and Farrell, Eoin and McDougall, Callum and Ayonrinde, Kola and Till, Demian and Wearden, Matthew and Conmy, Arthur and Marks, Samuel and Nanda, Neel},
  keywords = {Machine Learning (cs.LG), Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/