SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
Adam KarvonenCan RagerJohnny LinCurt TiggesJoseph Isaac BloomDavid ChaninYeu-Tong LauEoin FarrellCallum McDougallKola Ayonrinde
Introduces a standardized evaluation suite of eight metrics and over 200 trained models to assess sparse autoencoder architectures on feature disentanglement, interpretability, and practical applications beyond traditional proxy loss.
As large language models become central to critical software systems, understanding their internal decision-making processes is essential for safety, compliance, and risk management. Researchers frequently employ Sparse Autoencoders (SAEs)—specialized neural network components that break down complex model internal states into distinct, understandable concepts. However, progress in the field has historically relied on simple unsupervised proxy metrics, primarily the trade-off between how accurately an SAE reconstructs the original model data and how sparsely it does so. This approach leaves major uncertainties about whether architectural improvements actually deliver practical interpretability and clean concept separation.
The article addresses this gap by introducing SAEBench, a standardized and extensible evaluation framework designed to assess SAE performance across diverse, practically relevant criteria. The primary objective is to evaluate over 200 trained SAEs across seven leading architectures and multiple model scales to determine how standard proxy metrics compare against real-world downstream tasks, including concept detection, interpretability, and practical applications like knowledge unlearning.
To conduct this evaluation, the researchers trained a comprehensive suite of SAEs sweeping across different dictionary capacities (4,000 to 65,000 features, and up to 1 million in existing public models) and activation sparsity levels on the Gemma-2-2B and Pythia-160M language models. The benchmark evaluates each system across eight standardized metrics categorized into four core dimensions: concept detection, automated interpretability judged by a language model, activation reconstruction fidelity, and feature disentanglement. The suite also introduces novel diagnostic techniques to measure how cleanly independent concepts can be isolated and ablated without causing unintended side effects.
The findings demonstrate that traditional proxy metrics do not reliably predict practical performance. First, Matryoshka SAEs—a hierarchical design—substantially outperform other architectures on concept detection and feature disentanglement, beating alternative approaches on five out of eight metrics and outperforming standard architectures by margins of 30% to 40% on specific concept removal tasks, despite showing lower reconstruction accuracy on traditional curves. Second, standard ReLU-based autoencoders perform the worst on five out of eight evaluations, confirming they are largely obsolete compared to modern alternatives. Third, increasing the dictionary size improves basic reconstruction and per-feature interpretability across all models, but it degrades feature disentanglement in non-hierarchical architectures due to excessive concept fragmentation; Matryoshka was the only architecture whose disentanglement capability improved with scale. Finally, the analysis shows that optimal sparsity is highly task-dependent, though a moderate sparsity range of 50 to 150 active features offers the best overall compromise across capabilities.
These results carry significant implications for the deployment and oversight of AI interpretability tools. Optimizing exclusively for reconstruction fidelity creates a false sense of progress while obscuring critical failure modes like feature absorption, where concepts become entangled or hidden. For organizations investing in AI safety, auditing, or targeted knowledge removal (such as erasing proprietary or hazardous data), selecting the appropriate architecture and evaluating across multidimensional metrics is critical to avoid wasted compute and unreliable safety guarantees.
The article recommends that practitioners adopt multi-metric evaluation suites rather than single-score proxies when developing or selecting SAEs. Development teams should test systems across a range of sparsity levels (specifically 20 to 200 active features) with directly comparable baselines, and consider hierarchical architectures like Matryoshka for tasks demanding precise concept isolation. For future work, the evaluation suite should be expanded to larger model architectures, additional network layers, and non-text modalities such as vision and biology models.
Confidence in these findings is high for medium-sized language models and the specific datasets evaluated, as results remained consistent across multiple seeds and scale sweeps. However, readers should note key limitations: supervised metrics currently depend on a limited set of ground-truth concepts (such as profession or syntax), unlearning evaluations are constrained by the baseline capabilities of the underlying model, and quantitative scoring cannot fully replace human-led qualitative analysis during deep safety investigations.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). Read this foundational demonstration of SAE feature extraction and evaluation first to understand the methods and interpretability claims that SAEBench tests across architectures.
- Paper: Scaling and evaluating sparse autoencoders, Leo Gao et al. (2025). Its large-scale SAE training methods and measures of feature usefulness provide key context for SAEBench’s architecture, sparsity, and evaluation comparisons.
- Paper: RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations, Jing Huang et al. (2024). RAVEL’s tests of concept disentanglement establish an earlier evaluation problem that helps frame SAEBench’s diagnostics for feature separation and unintended effects.
- Paper: Insights into a radiology-specialised multimodal large language model with sparse autoencoders, Kenza Bouzid et al. (2025). This radiology-model study applies SAEs to clinical concepts and model steering, showing how benchmarked feature interpretability translates into a high-stakes multimodal setting.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). This SAE-based steering method carries feature analysis into knowledge-selection interventions, extending SAEBench’s interest in practical downstream use.
