Are Sparse Autoencoders Useful? A Case Study in Sparse Probing

Subhash KantamneniJoshua EngelsSenthooran RajamanoharanMax TegmarkNeel Nanda

article2025ICML73 citations

Demonstrates across 113 classification tasks that sparse autoencoder probes fail to outperform standard linear baselines in challenging regimes such as data scarcity and distribution shift, critically questioning the practical downstream utility of sparse dictionary learning in language models.

Listen

Mechanistic interpretability aims to understand how large language models represent and process information, often to enhance AI safety, model steering, and reliability. Sparse autoencoders (SAEs) have emerged as a prominent technique for decomposing complex internal model representations into interpretable, human-understandable concepts. However, validating whether SAEs extract true representations remains difficult because there is no ground-truth benchmark for internal model representations. The article systematically evaluates whether SAEs provide a practical, measurable advantage over conventional methods in the real-world downstream task of activation probing—training classifiers to predict specific concepts from internal model states.

The investigation benchmarked SAE-based probes against standard classification baselines across 113 diverse binary classification datasets using open-weight models, primarily Gemma-2-9B and Llama-3.1-8B. The evaluation examined standard operating conditions alongside four challenging operational environments where the inductive bias of interpretable features was hypothesized to provide an advantage: data scarcity, extreme class imbalance, severe label corruption, and out-of-distribution covariate shifts. To ensure a robust assessment, the authors used a portfolio selection methodology called the "Quiver of Arrows," simulating a practitioner who selects the best-performing method using validation data and evaluates the marginal test-set improvement from adding SAEs to existing tools.

The findings show that SAE probes fail to deliver a consistent performance benefit over standard baselines. In standard conditions, adding SAEs slightly reduced average test accuracy. In challenging regimes—including limited training data, unbalanced classes, noisy labels, and distribution shifts—SAEs did not meaningfully outperform simple baselines like logistic regression. Furthermore, while SAEs initially appeared uniquely capable of uncovering dataset errors, identifying spurious correlations, and pooling across multiple tokens, subsequent experiments demonstrated that standard baseline classifiers achieved equivalent insights when properly designed. Comparing eight successive SAE architectural variants released over recent years also revealed only minor, statistically insignificant performance gains.

These results carry significant implications for AI engineering, research investment, and governance. Deploying current SAE techniques as downstream classifiers or audit mechanisms adds computational overhead without improving diagnostic or predictive accuracy. The findings caution decision-makers against assuming that interpretable representations automatically enhance task performance or out-of-distribution robustness. They also underscore that interpretability techniques must be benchmarked against well-tuned baselines rather than naive comparisons to avoid misleading conclusions.

For practical applications, teams deploying linear probes should rely on established, lightweight methods such as regularized logistic regression and attention-based pooling rather than SAE latents. Organizations funding or conducting interpretability research should prioritize developing rigorous benchmarks and testing on controlled environments where underlying model features are known. While the article's conclusions are strongly supported across multiple models and extensive datasets, the authors note that probing is a proxy task; future research should continue testing whether other downstream applications, such as model steering or direct feature ablation, realize distinct benefits from SAE architectures.

Kantamneni et al (2025).pdf
Cover for Are Sparse Autoencoders Useful? A Case Study in Sparse Probing

Abstract

Sparse autoencoders (SAEs) are a popular method for interpreting concepts represented in large language model (LLM) activations. However, there is a lack of evidence regarding the validity of their interpretations due to the lack of a ground truth for the concepts used by an LLM, and a growing number of works have presented problems with current SAEs. One alternative source of evidence would be demonstrating that SAEs improve performance on downstream tasks beyond existing baselines. We test this by applying SAEs to the real-world task of LLM activation probing in four regimes: data scarcity, class imbalance, label noise, and covariate shift. Due to the difficulty of detecting concepts in these challenging settings, we hypothesize that SAEs’ basis of interpretable, concept-level latents should provide a useful inductive bias. However, although SAEs occasionally perform better than baselines on individual datasets, we are unable to ensemble SAEs and baselines to consistently improve over just baseline methods. Additionally, although SAEs initially appear promising for identifying spurious correlations, detecting poor dataset quality, and training multi-token probes, we are able to achieve similar results with simple non-SAE baselines as well. Though we cannot discount SAEs’ utility on other tasks, our findings highlight the shortcomings of current SAEs and the need to rigorously evaluate interpretability methods on downstream tasks with strong baselines.

Table of Contents

  • 1. Introduction
  • 2. Methodology
  • 2.1. Classification Datasets
  • 2.2. Probing Strategy
  • 2.3. Probing Methods
  • 2.4. Experimental Setup: Quiver of Arrows
  • 3. Comparing Probing Techniques in Different Regimes
  • 3.1. Standard Conditions
  • 3.2. Data Scarcity, Class Imbalance, and Label Noise
  • 3.3. Covariate Shift
  • 4. Interpretability
  • 4.1. Probe Interpretability: Pruning and Latent Generalization
  • 4.2. Latent Interpretability
  • 4.3. Detecting Dataset Quality Issues
  • 4.3.1. GLUE COLA
  • 4.3.2. AI VS HUMAN
  • 5. Assessing Improvements in SAE Architectures
  • 6. Conclusion
  • Acknowledgments
  • Impact Statement
  • References
  • A. Related Work
  • A.1. Probing
  • A.2. Challenges in Neural Network Interpretability
  • A.3. SAE and SAE Applications
  • B. Disentangled Representations
  • C. Classification Datasets
  • D. Probing Setup
  • D.1. AUC
  • D.2. Probing Method Validation Details
  • D.3. Probing Method Hyperparameter Details
  • D.4. Additional Discussion of Quiver of Arrows
  • E. Normal Conditions
  • E.1. Cross Layer-wise Comparisons
  • E.2. Choosing which SAEs to Test
  • E.3. Plotting Dataset Performance vs. K
  • F. Additional Results for Various Regimes
  • F.1. Data Scarcity
  • F.2. Class Imbalance
  • F.3. Label Corruption
  • G. Intuition for SAE Probe Utility Across Regimes
  • H. Evaluating the Validity of the Quiver of Arrows
  • I. Reproducing Core Results on Llama-3.1-8b
  • J. Interpretability
  • J.1. Pruning
  • J.2. Latent Investigation
  • J.3. GLUE CoLA
  • J.4. AI Made
  • K. Multiple Tokens Results
  • L. SAE Architectural Improvements
  • M. Assessing SAE Probe Binarization

Knowls

  1. Knowl 1 — Sparse SAE probing selects latents by class-conditional activation difference

    model/method

    For each binary task, the authors extract the language model’s last-token activation for every prompt and encode it with a pretrained sparse autoencoder (SAE). For each latent, they calculate the absolute difference between its mean activation on positive training examples and its mean activation on negative training examples, then select the kk latents with the largest differences. If zj,iz_{j,i} is latent ii’s activation for training example jj, and T1T_1 and T0T_0 are the index sets for positive and negative examples, respectively, the selected set is

    I=arg top k⁡i∣1∣T1∣∑j∈T1zj,i−1∣T0∣∑j∈T0zj,i∣.I = \operatorname{arg\,top\,k}_{i}\left|\frac{1}{|T_1|}\sum_{j\in T_1}z_{j,i}-\frac{1}{|T_0|}\sum_{j\in T_0}z_{j,i}\right|.

    A logistic-regression classifier with L1L_1 regularization is then trained on the selected latent activations to predict the binary target. Here kk is the number of selected latents and ∣T1∣|T_1|, ∣T0∣|T_0| are the numbers of training examples in each class. For Gemma experiments, the authors normalize each latent by its average activation when active before calculating selection scores; they report that this change improves selection for small kk but does not materially change the overall comparisons.

  2. Knowl 2 — The quiver-of-arrows evaluation tests whether SAEs add value to an existing toolkit

    definition

    The quiver-of-arrows procedure compares a collection of probe methods with and without SAE probes, modeling a practitioner who can choose among available methods. For each method, hyperparameters are selected using validation AUC; the method with the highest validation AUC is chosen, and its held-out test AUC is reported. The SAE contribution is assessed by comparing the test performance of the baseline collection against the expanded collection containing the SAE probes. This validation-based selection avoids choosing methods using test-set results and therefore tests whether adding SAE probes improves the practitioner’s available options, rather than whether a hand-picked SAE can win on a particular test set.

  3. Knowl 3 — The evaluation spans 113 tasks, two language models, and multiple strong probe baselines

    experimental setup

    The study evaluates 113 binary classification tasks using hidden activations from Gemma-2-9B, with core results replicated on Llama-3.1-8B. Prompts range from 5 to 1,024 tokens, and probes use the final-token activation. The Gemma experiments use JumpReLU SAEs from Gemma Scope; the Llama experiments use TopK SAEs from Llama Scope. The five non-SAE baselines are logistic regression, PCA followed by logistic regression, kk-nearest neighbors, XGBoost, and a multilayer perceptron; SAE probes use L1L_1-regularized logistic regression. Probe quality is measured by test-set area under the ROC curve (AUC). The challenging evaluation regimes are limited training data, altered positive-class frequency, randomly corrupted training labels, and covariate shift in test prompts.

  4. Knowl 4 — Across core probing regimes, adding SAE probes does not consistently improve test performance

    empirical result

    Under the quiver-of-arrows evaluation on Gemma-2-9B, mean test AUC for baseline-only, SAE-only, and combined baseline-plus-SAE method collections was, respectively: 0.940, 0.930, and 0.939 under standard conditions; 0.819, 0.806, and 0.812 under data scarcity; and 0.921, 0.906, and 0.916 under class imbalance. Thus, in these comparisons, the combined toolkit did not exceed the baseline-only toolkit, and SAE-only methods performed worse on average. In standard conditions, an SAE method was selected for 14 of 113 tasks, yet adding SAE probes slightly reduced aggregate test performance. For label-noise experiments, corrupted validation labels made validation-based method selection unreliable, so the authors compared logistic regression directly with a width-16,000, k=128k=128 SAE probe; this comparison likewise showed no meaningful average SAE advantage. The authors also report no statistically significant advantage from newer SAE architectures: tests of eight architectures showed a slight positive trend with release date, but variation across datasets was larger than the trend. Higher SAE L0L_0 and larger kk generally improved SAE probe performance, while SAE width had only a small effect.

  5. Knowl 5 — SAE probes generalize worse than logistic regression under the tested covariate shifts

    empirical result

    The authors evaluated covariate shift on eight datasets using 300 shifted test examples per dataset. The shifts included two GLUE-X extreme-task sets, translated prompts for three tasks, and syntactic or character changes to name-related prompts for three tasks. Probes were trained under standard conditions and compared using logistic regression and an SAE probe built from a width-131,000, L0=114L_0=114 SAE. Logistic regression outperformed the SAE probe across the OOD evaluation overall. This result provides no evidence that the sparse, interpretable SAE basis made probes more robust to the tested changes in prompt distribution.

  6. Knowl 6 — Pruning SAE latents modestly helps some OOD probes, but feature non-transfer explains larger failures

    empirical result

    To investigate OOD failures, the authors selected the eight most class-discriminative latents for three tasks and ranked them using descriptions generated by automated latent interpretation and relevance assessments. Restricting probes to higher-ranked latents improved OOD AUC from the eight-latent probe to the one-latent probe by 0.024 on the living-room task and 0.052 on GLUE QNLI; human ranking produced similar conclusions. These gains were modest relative to the ID-to-OOD performance drops, suggesting that removing spurious latents was not the main remedy. On the living-room task, latent 122774 had ID test AUC 0.99 and OOD test AUC 0.64. It was strongly active for the English phrase “living room” but did not activate for its French translation, illustrating how a highly predictive latent can fail to represent the target robustly across languages.

  7. Knowl 7 — SAE and baseline classifiers both expose likely CoLA label errors

    empirical result

    On the CoLA grammatical-acceptability task, SAE latent 369585 had test AUC 0.76 and appeared to activate on ungrammatical sentences, including some examples labeled grammatical. A logistic-regression probe identified similar suspicious examples, showing that this diagnostic was not unique to SAE features. To estimate the extent of label problems, three language models independently judged 1,000 random CoLA examples; the majority judgment disagreed with the dataset label on 22% of examples. When classifiers trained on original CoLA labels were evaluated on the subset where the majority judgment disagreed with those labels, the single-latent classifier outperformed both a dense SAE probe and logistic regression. The 22% figure is disagreement with the language-model ensemble, rather than an independently verified error rate.

  8. Knowl 8 — An SAE feature for AI-generated text tracks punctuation, a correlation also found by logistic regression

    empirical result

    For a dataset distinguishing human-written text from ChatGPT-3.5 text, SAE latent 105150 achieved AUC 0.82 but appeared to respond primarily to punctuation rather than text authorship or semantic content. The dataset’s final-token patterns supplied a plausible spurious cue: AI-generated examples more often ended in a period, while human examples more often ended in a space. A logistic-regression classifier trained for the same task and evaluated on more than 2.8 million tokens from the Pile also activated most strongly on punctuation tokens. Thus, SAE interpretation exposed a dataset artifact, but the authors reproduced the same diagnostic with a non-SAE baseline.

  9. Knowl 9 — A stronger token-pooling baseline reduces the apparent advantage of multi-token SAE probes

    empirical result

    In experiments on 60 randomly selected datasets, the authors compared SAE probes that max-pool each latent across prompt tokens with activation-based probes. Against a last-token activation baseline, a max-pooled SAE probe beat the baseline by more than 0.005 AUC on 19.6% of tasks, compared with a 2.2% win rate for last-token SAE probes. However, max-pooling activation dimensions is a weak baseline, so the authors also tested an attention-pooled activation probe using learned query and value vectors to score and aggregate token activations. When the evaluation selected between pooled and last-token strategies for both SAE and activation probes, the SAE win rate fell to 8.7%. Binarizing pooled SAE activations at threshold 1 did not improve the comparison: the corresponding win rates were 13.0% against the last-token baseline and 6.5% in the strategy-selection comparison.

  10. Knowl 10 — Probe performance is a proxy for SAE utility, and the tested SAE-probe choices limit the conclusion

    limitation

    The study evaluates SAE utility through probing performance, which the authors regard as more task-relevant than reconstruction error but still only a proxy: better or worse probe performance does not establish whether SAE latents recover the model’s true concepts. The authors did not test every possible SAE-probe configuration; in particular, their main SAE probes use logistic regression and select at most 128 latents, and the principal comparisons use last-token activations. They therefore do not rule out gains from other probe methods, larger latent sets, or other downstream tasks. Their findings support a conclusion about the tested SAE methods and probing setups, not a wholesale judgment about the SAE approach.

Coverage note — The full 113-task catalog, exhaustive per-task scores, and detailed hyperparameter grids are omitted because they are supporting diagnostics rather than distinct main findings.

References

  1. 1.Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B. Sanity checks for saliency maps. Advances in neural information processing systems, 31, 2018.
  2. 2.AI, T. and Ishii, D. Spam Text Message Classification — kaggle.com. https://www.kaggle.com/datasets/team-ai/spam-text-message-classification. [Accessed 21-01-2025].
  3. 3.Alain, G. Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644, 2016.
  4. 4.AllenAI. allenai/basic arithmetic · Datasets at Hugging Face — huggingface.co. https://huggingface.co/datasets/allenai/basic_arithmetic. [Accessed 21-01-2025].
  5. 5.Anthropic. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku, 2024. URL https://www.anthropic.com/news/3-5-models-and-computer-use. Accessed: 2025-01-30.
  6. 6.Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. Linear algebraic structure of word senses, with applications to polysemy. Transactions of the Association for Computational Linguistics, 6:483–495, 2018.
  7. 7.Ben Zhou, Daniel Khashabi, Q. N. and Roth, D. “going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding. In EMNLP, 2019.
  8. 8.Bengio, Y., Courville, A., and Vincent, P. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8): 1798–1828, 2013.
  9. 9.Bereska, L. and Gavves, E. Mechanistic interpretability for ai safety–a review. arXiv preprint arXiv:2404.14082, 2024.
  10. 10.Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html, 2023.
  11. 11.Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020.
  12. 12.Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viegas, F., and Wattenberg, M. An interpretability illusion for bert. arXiv preprint arXiv:2104.07143, 2021.
  13. 13.Braun, D., Taylor, J., Goldowsky-Dill, N., and Sharkey, L. Identifying functionally important features with end-to-end sparse dictionary learning, 2024. URL https://arxiv.org/abs/2405.12241.
  14. 14.Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. https://transformer-circuits.pub/2023/monosemantic-features/index.html.
  15. 15.Bricken, T., Marcus, J., Mishra-Sharma, S., Tong, M., Perez, E., Sharma, M., Rivoire, K., and Henighan, T. Using dictionary learning features as classifiers, October 2024a. URL https://transformer-circuits.pub/2024/features-as-classifiers/index.html. Accessed: 2025-01-23.
  16. 16.Bricken, T., Marcus, J., Mishra-Sharma, S., Tong, M., Perez, E., Sharma, M., Rivoire, K., and Henighan, T. Using dictionary learning features as classifiers, October 2024b. URL https://transformer-circuits.pub/2024/features-as-classifiers/index.html. Transformer Circuits.
  17. 17.Bussmann, B., Leask, P., and Nanda, N. Batchtopk sparse autoencoders, 2024a.
  18. 18.Bussmann, B., Leask, P., and Nanda, N. Learning multi-level features with matryoshka saes. https://www.lesswrong.com/posts/rKM9b6B2LqwSB5ToN/learning-multi-level-features-with-matryoshka-saes, 2024b. Accessed: 2025-01-23.
  19. 19.Chalnev, S., Siu, M., and Conmy, A. Improving steering vectors by targeting sparse autoencoder features. arXiv preprint arXiv:2411.02193, 2024.
  20. 20.Chanin, D., Wilken-Smith, J., Dulka, T., Bhatnagar, H., and Bloom, J. A is for absorption: Studying feature splitting and absorption in sparse autoencoders. arXiv preprint arXiv:2409.14507, 2024.
  21. 21.Chaudhary, M. and Geiger, A. Evaluating open-source sparse autoencoders on disentangling factual knowledge in gpt-2 small. arXiv preprint arXiv:2409.04478, 2024.
  22. 22.Chen, T. and Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pp. 785–794, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 9781450342322. doi: 10.1145/2939672.2939785. URL https://doi.org/10.1145/2939672.2939785.
  23. 23.clmentbisaillon. fake-and-real-news-dataset — kaggle.com. https://www.kaggle.com/datasets/clmentbisaillon/fake-and-real-news-dataset. [Accessed 21-01-2025].
  24. 24.Conerly, T., Templeton, A., Bricken, T., Marcus, J., and Henighan, T. Circuits Updates - April 2024 — transformer-circuits.pub. https://transformer-circuits.pub/2024/april-update/index.html#training-saes, 2024. [Accessed 15-02-2025].
  25. 25.Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023.
  26. 26.Davidson, T., Warmsley, D., Macy, M., and Weber, I. Automated hate speech detection and the problem of offensive language. In Proceedings of the 11th International AAAI Conference on Web and Social Media, ICWSM ’17, pp. 512–515, 2017.
  27. 27.Dublish, T. Text classification documentation — kaggle.com. https://www.kaggle.com/datasets/tanishqdublish/text-classification-documentation. [Accessed 21-01-2025].
  28. 28.Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175, March 2021. ISSN 2307-387X. doi: 10.1162/tacl_a_00359. URL http://dx.doi.org/10.1162/tacl_a_00359.
  29. 29.Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M. Not all language model features are linear, 2024a. URL https://arxiv.org/abs/2405.14860.
  30. 30.Engels, J., Riggs, L., and Tegmark, M. Decomposing the dark matter of sparse autoencoders. arXiv preprint arXiv:2410.14670, 2024b.
  31. 31.Falgunipatel19. Medical Text Dataset -Cancer Doc Classification — kaggle.com. https://www.kaggle.com/datasets/falgunipatel19/biomedical-text-publication-classification. [Accessed 21-01-2025].
  32. 32.Farrell, E., Lau, Y.-T., and Conmy, A. Applying sparse autoencoders to unlearn knowledge in language models, 2024.
  33. 33.Gaggar, R., Bhagchandani, A., and Oza, H. Machine-generated text detection using deep learning, 2023.
  34. 34.Gallifant, J., Chen, S., Sasse, K., Aerts, H., Hartvigsen, T., and Bitterman, D. S. Sparse autoencoder features for classifications and transferability, 2025.
  35. 35.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  36. 36.Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J. Scaling and evaluating sparse autoencoders, 2024. URL https://arxiv.org/abs/2406.04093.
  37. 37.Gerami, S. AI Vs Human Text — kaggle.com. https://www.kaggle.com/datasets/shanegerami/ai-vs-human-text/data. [Accessed 21-01-2025].
  38. 38.Goh, A. IT Service Ticket Classification Dataset — kaggle.com. https://www.kaggle.com/datasets/adisongoh/it-service-ticket-classification-dataset. [Accessed 21-01-2025].
  39. 39.Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  40. 40.Gupta, P. Emotion Detection from Text — kaggle.com. https://www.kaggle.com/datasets/pashupatigupta/emotion-detection-from-text. [Accessed 21-01-2025].
  41. 41.Gurnee, W. and Tegmark, M. Language models represent space and time. arXiv preprint arXiv:2310.02207, 2023.
  42. 42.Gurnee, W. and Tegmark, M. Language models represent space and time. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=jE8xbmvFin.
  43. 43.Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D. Finding neurons in a haystack: Case studies with sparse probing. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=JYs1R9IMJr.
  44. 44.He, Z., Shu, W., Ge, X., Chen, L., Wang, J., Zhou, Y., Liu, F., Guo, Q., Huang, X., Wu, Z., Jiang, Y.-G., and Qiu, X. Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024.
  45. 45.Heap, T., Lawson, T., Farnik, L., and Aitchison, L. Sparse autoencoders can interpret randomly initialized transformers, 2025. URL https://arxiv.org/abs/2501.17727.
  46. 46.Heinzerling, B. and Inui, K. Monotonic representation of numeric properties in language models. arXiv preprint arXiv:2403.10381, 2024.
  47. 47.Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., and Steinhardt, J. Aligning ai with shared human values. Proceedings of the International Conference on Learning Representations (ICLR), 2021.
  48. 48.Higgins, I., Amos, D., Pfau, D., Racaniere, S., Matthey, L., Rezende, D., and Lerchner, A. Towards a definition of disentangled representations. arXiv preprint arXiv:1812.02230, 2018.
  49. 49.Huang, J., Wu, Z., Potts, C., Geva, M., and Geiger, A. Ravel: Evaluating interpretability methods on disentangling language model representations. arXiv preprint arXiv:2402.17700, 2024.
  50. 50.Johannes Welbl, Nelson F. Liu, M. G. Crowdsourcing multiple choice science questions. 2017.
  51. 51.Joshi, S., Dittadi, A., Lachapelle, S., and Sridhar, D. Identifiable steering via sparse autoencoding of multi-concept shifts. arXiv preprint arXiv:2502.12179, 2025.
  52. 52.Karvonen, A., Wright, B., Rager, C., Angell, R., Brinkmann, J., Smith, L., Verdun, C. M., Bau, D., and Marks, S. Measuring progress in dictionary learning for language model interpretability with board game models, 2024. URL https://arxiv. org/abs/2408.00113, pp. 16.
  53. 53.Karvonen, A., Pai, D., Wang, M., and Keigwin, B. Sieve: Saes beat baselines on a real-world task (a code generation case study). Tilde Research Blog, 12 2024a. URL https://www.tilderesearch.com/blog/sieve. Blog post.
  54. 54.Karvonen, A., Rager, C., Lin, J., Tigges, C., Bloom, J., Chanin, D., Lau, Y.-T., Farrell, E., Conmy, A., McDougall, C., Ayonrinde, K., Wearden, M., Marks, S., and Nanda, N. Saebench: A comprehensive benchmark for sparse autoencoders, December 2024b. URL https://www.neuronpedia.org/sae-bench/info. Accessed: 2025-01-20.
  55. 55.Karvonen, A., Wright, B., Rager, C., Angell, R., Brinkmann, J., Smith, L., Verdun, C. M., Bau, D., and Marks, S. Measuring progress in dictionary learning for language model interpretability with board game models. arXiv preprint arXiv:2408.00113, 2024c.
  56. 56.Kasaraneni, C. K. Medical Text — kaggle.com. https://www.kaggle.com/datasets/chaitanyakck/medical-text. [Accessed 21-01-2025].
  57. 57.Kotari, R. rkotari/clickbait · Datasets at Hugging Face — huggingface.co. https://huggingface.co/datasets/rkotari/clickbait. [Accessed 21-01-2025].
  58. 58.Lachapelle, S., Deleu, T., Mahajan, D., Mitliagkas, I., Bengio, Y., Lacoste-Julien, S., and Bertrand, Q. Synergies between disentanglement and sparsity: Generalization and identifiability in multi-task learning. In International Conference on Machine Learning, pp. 18171–18206. PMLR, 2023.
  59. 59.Li, Y., Michaud, E. J., Baek, D. D., Engels, J., Sun, X., and Tegmark, M. The geometry of concepts: Sparse autoencoder feature structure. arXiv preprint arXiv:2410.19750, 2024.
  60. 60.Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kramar, J., Dragan, A., Shah, R., and Nanda, N. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024.
  61. 61.Lin, J. Neuronpedia: Interactive reference and tooling for analyzing neural networks, 2023. URL https://www.neuronpedia.org. Software available from neuronpedia.org.
  62. 62.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-long.229/.
  63. 63.Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023.
  64. 64.MacDiarmid, M., Maxwell, T., Schiefer, N., Mu, J., Kaplan, J., Duvenaud, D., Bowman, S., Tamkin, A., Perez, E., Sharma, M., Denison, C., and Hubinger, E. Simple probes can catch sleeper agents, 2024. URL https://www.anthropic.com/news/probes-catch-sleeper-agents.
  65. 65.Marks, S. and Tegmark, M. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824, 2023.
  66. 66.Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A. Sparse feature circuits: Discovering and editing interpretable causal graphs in language models. arXiv preprint arXiv:2403.19647, 2024.
  67. 67.Mudide, A., Engels, J., Michaud, E. J., Tegmark, M., and de Witt, C. S. Efficient dictionary learning with switch sparse autoencoders, 2024. URL https://arxiv.org/abs/2410.08201.
  68. 68.Mur, M., Bandettini, P. A., and Kriegeskorte, N. Revealing representational content with pattern-information fmri—an introductory guide. Social cognitive and affective neuroscience, 4(1):101–109, 2009.
  69. 69.Nanda, N., Lee, A., and Wattenberg, M. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023.
  70. 70.Olmo, J., Wilson, J., Forsey, M., Hepner, B., Howe, T. V., and Wingate, D. Features that make a difference: Leveraging gradients for improved dictionary learning, 2024. URL https://arxiv.org/abs/2411.10397.
  71. 71.OpenAI. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/. [Accessed 27-01-2025].
  72. 72.OpenAI. Introducing openai o1. https://openai.com/o1/, 2024. [Accessed 29-01-2025].
  73. 73.Pang, B. and Lee, L. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the ACL, 2005.
  74. 74.Paulo, G., Mallen, A., Juang, C., and Belrose, N. Automatically interpreting millions of features in large language models. arXiv preprint arXiv:2410.13928, 2024.
  75. 75.Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramar, J., Shah, R., and Nanda, N. Improving dictionary learning with gated sparse autoencoders. arXiv preprint arXiv:2404.16014, 2024a.
  76. 76.Rajamanoharan, S., Lieberum, T., Sonnerat, N., Conmy, A., Varma, V., Kramar, J., and Nanda, N. Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders, 2024b. URL https://arxiv.org/abs/2407.14435.
  77. 77.Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., et al. Gemma 2: Improving open language models at a practical size, 2024.
  78. 78.Rogers, A., Kovaleva, O., Downey, M., and Rumshisky, A. Getting closer to AI complete question answering: A set of prerequisite real tasks. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pp. 8722–8731. AAAI Press, 2020. URL https://aaai.org/ojs/index.php/AAAI/article/view/6398.
  79. 79.Sharkey, L., Chughtai, B., Batson, J., Lindsey, J., Wu, J., Bushnaq, L., Goldowsky-Dill, N., Heimersheim, S., Ortega, A., Bloom, J., Biderman, S., Garriga-Alonso, A., Conmy, A., Nanda, N., Rumbelow, J., Wattenberg, M., Schoots, N., Miller, J., Michaud, E. J., Casper, S., Tegmark, M., Saunders, W., Bau, D., Todd, E., Geiger, A., Geva, M., Hoogland, J., Murfet, D., and McGrath, T. Open problems in mechanistic interpretability, 2025. URL https://arxiv.org/abs/2501.16496.
  80. 80.Smith, L. R. and Brinkmann, J. Interpreting preference models w/ sparse autoencoders, July 2024. URL https://www.alignmentforum.org/posts/5XmxmszdjzBQzqpmz/interpreting-\protect\penalty\z@preference-models-\protect\penalty\z@w-sparse-autoencoders. Accessed: 2025-01-23.
  81. 81.Stathead. Stathead: Your all-access ticket to the Sports Reference database. — Stathead.com — stathead.com. https://stathead.com/. [Accessed 21-01-2025].
  82. 82.Suter, R., Miladinovic, D., Scholkopf, B., and Bauer, S. Robustly disentangled causal mechanisms: Validating deep representations for interventional robustness. In International Conference on Machine Learning, pp. 6056–6065. PMLR, 2019.
  83. 83.Tafjord, O., Gardner, M., Lin, K., and Clark, P. ”quartz: An open-domain dataset of qualitative relationship questions”. ”2019”.
  84. 84.Talmor, A., Yoran, O., Bras, R. L., Bhagavatula, C., Goldberg, Y., Choi, Y., and Berant, J. Commonsenseqa 2.0: Exposing the limits of ai through gamification. arXiv preprint arXiv:2201.05320, 2022.
  85. 85.Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. Transformer Circuits Thread, 2024. URL https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html.
  86. 86.Van Steenkiste, S., Locatello, F., Schmidhuber, J., and Bachem, O. Are disentangled representations helpful for abstract visual reasoning? Advances in neural information processing systems, 32, 2019.
  87. 87.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJ4km2R5t7.
  88. 88.Warstadt, A., Singh, A., and Bowman, S. R. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641, 2019. doi: 10.1162/tacl_a_00290. URL https://aclanthology.org/Q19-1040/.
  89. 89.Wu, Z., Arora, A., Geiger, A., Wang, Z., Huang, J., Jurafsky, D., Manning, C. D., and Potts, C. Axbench: Steering llms? even simple baselines outperform sparse autoencoders, 2025. URL https://arxiv.org/abs/2501.17148.
  90. 90.Yun, Z., Chen, Y., Olshausen, B. A., and LeCun, Y. Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors. arXiv preprint arXiv:2103.15949, 2021.
  91. 91.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023.

Citation

MLA
Kantamneni, S., et al. “Are Sparse Autoencoders Useful? A Case Study in Sparse Probing”. arXiv, 2025, https://doi.org/10.48550/arxiv.2502.16681.
APA
Kantamneni, S., Engels, J., Rajamanoharan, S., Tegmark, M., & Nanda, N. (2025). Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. arXiv. https://doi.org/10.48550/arxiv.2502.16681
Chicago
Kantamneni, S., J. Engels, S. Rajamanoharan, M. Tegmark, and N. Nanda. 2025. “Are Sparse Autoencoders Useful? A Case Study in Sparse Probing”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2502.16681.
Harvard
Kantamneni, S. et al. (2025) “Are Sparse Autoencoders Useful? A Case Study in Sparse Probing”. arXiv. Available at: https://doi.org/10.48550/arxiv.2502.16681.
Vancouver
1. Kantamneni S, Engels J, Rajamanoharan S, Tegmark M, Nanda N (2025) Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. https://doi.org/10.48550/arxiv.2502.16681

BibTeX

@misc{https://doi.org/10.48550/arxiv.2502.16681,
  doi = {10.48550/ARXIV.2502.16681},
  url = {https://arxiv.org/abs/2502.16681},
  author = {Kantamneni, Subhash and Engels, Joshua and Rajamanoharan, Senthooran and Tegmark, Max and Nanda, Neel},
  keywords = {Machine Learning (cs.LG), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Are Sparse Autoencoders Useful? A Case Study in Sparse Probing},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/