Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering

Yu ZhaoAlessio DevotoGiwon HongXiaotang DuAryo Pradipta GemaHongru WangXuanli HeKam-Fai WongPasquale Minervini

article2025NAACL78 citations

Introduces SPARE, a training-free representation engineering method that leverages sparse auto-encoders to detect mid-layer conflict signals and steer whether large language models rely on parametric memory or contextual evidence during question answering.

Listen

Large language models store extensive factual knowledge in their internal parameters, but they are also frequently paired with external text retrieval systems to provide up-to-date context. When retrieved external text directly contradicts a model's internal memory—a problem known as a context-memory knowledge conflict—models typically default to following the provided context. This dynamic creates significant real-world risks, such as propagating external misinformation or overriding accurate internal knowledge.

The main objective of the article is to demonstrate how to detect these knowledge conflicts inside language models and introduce a practical, inference-time method to precisely steer whether a model relies on external context or its internal parametric memory.

To achieve this, the authors developed SPARE (Sparse Auto-Encoder-based Representation Engineering), a training-free intervention approach. Sparse auto-encoders are auxiliary neural networks that break down a model's dense, tangled internal representations into isolated, interpretable individual features. The researchers evaluated internal states across multiple models (Llama2-7B, Llama3-8B, and Gemma2-9B) using two benchmark question-answering datasets designed around knowledge conflicts (NQSwap and Macnoise). They identified which isolated features correlate with relying on context versus relying on memory, and then selectively added desired features and removed undesired features at middle model layers during the generation process.

The key findings demonstrate clear performance gains over standard techniques. First, internal probing revealed that knowledge conflicts produce strong, detectable signals concentrated specifically within the middle layers of the models. Second, steering models with SPARE outperformed existing activation-editing techniques by approximately 10 percentage points and outperformed contrastive decoding methods by roughly 15 percentage points in accuracy. Third, the method proved effective using an extremely small subset of features—manipulating less than 0.05% of available auto-encoder features in Gemma2-9B. Fourth, while previous methods struggled to enforce reliance on internal memory when misleading context was present, SPARE successfully steered models toward parametric memory (e.g., reaching 47.5% accuracy on Llama3-8B compared to 26.6% without control) while also achieving high context faithfulness when context was desired (reaching up to 92.2% on Macnoise). Ablation tests confirmed that both adding target features and subtracting conflicting features are essential to avoid severe performance degradation.

These findings indicate that organizations can regulate the behavior and trustworthiness of deployed language models at runtime without expensive model fine-tuning or prompt-based multi-turn verification that slows down responses. This provides a low-latency mechanism to mitigate misinformation risks and enforce organizational policies regarding when to trust external inputs over stored model memory. The discovery that middle layers contain the critical control points confirms earlier mechanistic findings while offering a more precise tool for direct model steering.

Organizations evaluating this approach should consider piloting sparse auto-encoder interventions in retrieval-augmented applications where misinformation defense is critical. Prior to widespread deployment, practitioners should verify feature calibration on representative domain data and establish external decision logic to determine whether context or internal memory is more likely to be accurate for a given query.

Certain limitations remain. The method depends on the availability or upfront training of sparse auto-encoders, which require non-trivial compute resources if not already publicly accessible. In addition, the empirical evaluations were focused on direct open-domain question-answering datasets; further investigation is needed to confirm generalizability to complex multi-step reasoning, long-form text generation, and nuanced, non-binary truth scenarios.

arXiv: 2410.15999

No sufficiently relevant recommendations were found.

Cover for Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering

Abstract

Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided in the context—this phenomenon, known as context-memory knowledge conflicts 1, can lead to undesirable model behaviour, such as reliance on outdated or incorrect information. Analysing the internal activations of LLMs, we find that they can internally register the signals of knowledge conflict at mid-layers. Such signals allow us to detect whether a knowledge conflict occurs and use inference-time intervention strategies to resolve it. In this work, we propose SPARE, a training-free representation engineering method that uses pre-trained sparse auto-encoders (SAEs) to control the knowledge selection behaviour of LLMs. SPARE identifies the functional features that control the knowledge selection behaviours and applies them to edit the internal activations of LLMs at inference time. Our experimental results show that SPARE can effectively control the usage of either knowledge source to resolve knowledge conflict in open-domain question-answering tasks, surpassing existing representation engineering methods (+10%) as well as contrastive decoding methods (+15%).

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Detection of Knowledge Conflicts
  • 4 Resolving Knowledge Conflicts by Representation Engineering
  • 4.1 Collecting Activations with Different Knowledge Selection Behaviours
  • 4.2 Identifying Functional SAE Activations
  • 4.3 Editing Activations to Steer Behaviours
  • 5 Experimental Results
  • 5.1 Settings
  • 5.2 Overall Performance Comparison
  • 5.3 Multi-Perspective Controlling Analysis
  • 6 Analysis and Discussion
  • 6.1 Analysing the Layer Choice
  • 6.2 Analysing the Residual Stream
  • 7 Related Works
  • 8 Conclusions
  • Limitations
  • Acknowledgments
  • References
  • A More Analysis of Knowledge Conflict Probing
  • A.1 Details of Probing Model
  • A.2 More Probing Results
  • B Sparse Auto-Encoders Details
  • C Implementation Details
  • C.1 Collecting Activations
  • C.2 Identifying Functional Activations
  • C.3 Impact of the Size of the Collected Activations
  • C.4 Development Set and Demonstrations
  • C.5 Implementation Details of Representation Engineering Baselines
  • C.6 Searching Hyperparameters
  • D Selected SAEs Activations of Gemma2-9B
  • E Distribution of Mutual Information
  • F Distribution Patterns of the Residual Stream Under Knowledge Conflict
  • F.1 Skewness of Residual Stream
  • F.2 L1 Norm and L2 Norm Pattern

Knowls

  1. Knowl 1 — Conflict-conditioned activation collection for SPARE

    model/method

    SPARE is a training-free inference-time method for steering whether a language model uses contextual knowledge or its parametric memory when the two conflict. For an open-domain question-answering instance, let QQ be the question, ECE_C be evidence whose answer CC conflicts with the model’s memorized answer MM, and let DEC={(Q,EC)}D_{E_C}=\{(Q,E_C)\} denote conflict inputs. SPARE runs the model on these inputs and partitions them into DCD_C, where the generated answer follows the context and equals CC, and DMD_M, where the generated answer follows parametric memory and equals MM.

    At the last input position—the position used to predict the first answer token—SPARE collects residual-stream hidden states hCj,hMj∈Rdh_C^j,h_M^j\in\mathbb{R}^d from NN examples in DCD_C and DMD_M. A pretrained sparse autoencoder maps each hidden state to a nonnegative latent vector z=fθ(h)∈Rnz=f_\theta(h)\in\mathbb{R}^n and decodes latent vectors with gϕ(z)=Wϕzg_\phi(z)=W_\phi z; the columns of WϕW_\phi are the learned SAE feature directions. The resulting latent vectors are averaged separately to obtain zˉC\bar z_C and zˉM\bar z_M, which summarize activation patterns associated with contextual and parametric knowledge selection.

    In the implementation, the averages can be confidence-weighted. For an example whose hidden state produces CC, the weight is proportional to log⁡P(C∣hCj)/(log⁡P(C∣hCj)+log⁡P(M∣hCj))\log P(C\mid h_C^j)/(\log P(C\mid h_C^j)+\log P(M\mid h_C^j)); the analogous weight for a parametric example uses P(M∣hMj)P(M\mid h_M^j) in the numerator. The normalized weights are used to form zˉC\bar z_C and zˉM\bar z_M. The collected activations are then filtered into functional features and used for inference-time editing.

  2. Knowl 2 — Mutual-information selection of functional SAE features

    model/method

    SPARE identifies a small set of SAE latents associated with knowledge selection rather than editing every latent dimension. For each SAE latent ZiZ_i, where i∈{1,…,n}i\in\{1,\ldots,n\}, it estimates the mutual information with the binary generated-answer label Y∈{C,M}Y\in\{C,M\}:

    I(Zi;Y)=∑zi∑y∈{C,M}P(zi,y)log⁡P(zi,y)P(zi)P(y).I(Z_i;Y)=\sum_{z_i}\sum_{y\in\{C,M\}}P(z_i,y)\log\frac{P(z_i,y)}{P(z_i)P(y)}.

    The latents are sorted in descending order of I(Zi;Y)I(Z_i;Y) separately for every edited layer. Given a user-selected information proportion K∈(0,1]K\in(0,1], SPARE selects the smallest kk satisfying

    ∑i=1kI(Zi;Y)∑j=1nI(Zj;Y)≥K.\frac{\sum_{i=1}^{k}I(Z_i;Y)}{\sum_{j=1}^{n}I(Z_j;Y)}\ge K.

    Let zˉC\bar z_C and zˉM\bar z_M be the contextual- and parametric-selection mean latent vectors collected from conflict examples. Among the selected latents, a feature is assigned to contextual selection when its expected activation difference satisfies EC[Zi]−EM[Zi]>0\mathbb{E}_C[Z_i]-\mathbb{E}_M[Z_i]>0; otherwise, it is assigned to parametric selection. SPARE constructs sparse functional vectors zC,zM∈Rnz_C,z_M\in\mathbb{R}^n by retaining zˉC,i\bar z_{C,i} only for contextual-associated latents and setting other coordinates to zero, while retaining zˉM,i\bar z_{M,i} only for parametric-associated latents and setting other coordinates to zero. Thus, zCz_C and zMz_M are orthogonal in their retained coordinates and represent the two opposing knowledge-selection behaviours.

  3. Knowl 3 — Inference-time SAE editing for knowledge selection

    model/method

    SPARE edits the residual-stream hidden state at the last input position using the current SAE activation and the two functional activation vectors. To steer a conflict input toward parametric knowledge, let h∈Rdh\in\mathbb{R}^d be the current hidden state and z=fθ(h)∈Rnz=f_\theta(h)\in\mathbb{R}^n its nonnegative SAE activation. SPARE removes only the currently present contextual-associated activation and adds only the missing parametric-associated activation:

    zi−=min⁡{zi,zC,i},zi+=max⁡{zM,i−zi,0},z_i^- = \min\{z_i,z_{C,i}\},\qquad z_i^+=\max\{z_{M,i}-z_i,0\},

    where zi−z_i^- is the amount removed and zi+z_i^+ is the amount added for latent coordinate ii. The edited hidden state is

    h′=h+α[−gϕ(z−)+gϕ(z+)],h'=h+\alpha\left[-g_\phi(z^-)+g_\phi(z^+)\right],

    where gϕg_\phi is the SAE decoder and α≥0\alpha\ge 0 controls intervention strength. The minimum and maximum operations prevent the edited latent from becoming negative or from exceeding the target functional activation. To steer toward contextual knowledge instead, SPARE swaps the roles of zCz_C and zMz_M in the removal and addition rules. The method edits the original hidden state directly rather than decoding a fully replaced latent vector, because full SAE reconstruction can lose information and because nonnegative SAE activations would prevent flexible scaling by α\alpha. For the reported ODQA experiments, editing is applied only at the final input position.

  4. Knowl 4 — Mid-layer residual-stream signals reveal knowledge conflicts

    empirical result

    The paper shows that context-memory knowledge conflicts can be detected from internal activations before answer generation. It compares non-conflict inputs DEM={(Q,EM)}D_{E_M}=\{(Q,E_M)\}, where evidence EME_M agrees with the model’s memorized answer, with conflict inputs DEC={(Q,EC)}D_{E_C}=\{(Q,E_C)\}, where evidence contradicts that memory. For each layer of Llama2-7B and Gemma2-9B, a logistic-regression probe is trained on held-out examples to distinguish the two groups using the residual-stream hidden state, MLP activation, or self-attention activation at the final input position.

    Across the models and activation types, probing AUROC rises from the early layers to the middle layers, indicating that the residual stream internally registers whether contextual and parametric knowledge conflict. Probe performance declines in later layers, particularly for MLP and self-attention activations, suggesting that these modules add little further conflict signal after the middle layers. This result motivates applying SPARE in middle layers, where the conflict signal and the functional knowledge-selection representation are most accessible.

  5. Knowl 5 — Experimental design for evaluating SPARE

    experimental setup

    SPARE is evaluated on two conflict-aware open-domain question-answering datasets, NQSwap and Macnoise, using Llama3-8B, Llama2-7B, and Gemma2-9B. Llama3-8B and Gemma2-9B use publicly available pretrained SAEs; the authors train SAEs for Llama2-7B. The Llama2-7B and Llama3-8B SAEs have n=131072n=131072 latent dimensions and hidden size d=4096d=4096; the Llama2-7B SAE is pretrained on 10B RedPajama tokens. The method is evaluated with greedy decoding and three in-context demonstrations that align answer format without indicating which knowledge source should be trusted.

    The evaluation reports exact-match accuracy for steering toward the contextual answer CC (EMCEM_C) and toward the parametric answer MM (EMMEM_M). Baselines include TaskVec, ActAdd, linear and squared-exponential SEA, DoLa, CAD, and in-context learning. SPARE is applied to layers 13–16 of Llama3-8B, layers 12–15 of Llama2-7B, and layers 23–25 plus 29–31 of Gemma2-9B. The selected hyperparameters are (K,α)=(0.07,2)(K,\alpha)=(0.07,2) for Llama3-8B, (0.06,2.2)(0.06,2.2) for Llama2-7B, (0.01,3)(0.01,3) when steering Gemma2-9B toward contextual knowledge, and (0.01,1.8)(0.01,1.8) when steering Gemma2-9B toward parametric knowledge. The experiments use five demonstration seeds and report means with deviations.

  6. Knowl 6 — SPARE outperforms competing knowledge-selection controls

    data/table

    The main comparison measures exact-match accuracy for producing the parametric answer (EMMEM_M) or contextual answer (EMCEM_C) under conflict evidence. Values are reported as mean ±\pm deviation over demonstration seeds. SPARE is generally the strongest inference-time method and is especially more effective than contrastive decoding when steering toward parametric memory.

    Could not parse LaTeX table

    SPARE obtains the best result in every reported model–dataset combination for both steering directions. Its EMMEM_M values range from 30.72 to 47.51 and its EMCEM_C values range from 69.32 to 92.24. The results support the paper’s claim that sparse feature editing controls both knowledge sources more reliably than dense representation engineering, contrastive decoding, or demonstration-based control.

  7. Knowl 7 — SPARE changes behaviour while preserving compatible model decisions

    empirical result

    The paper evaluates whether a controller can both switch the model’s original knowledge choice and avoid damaging examples where the original choice already matches the requested target. For conflict examples where an uncontrolled model generates contextual answer CC, EMC→MEM_{C\to M} is the accuracy of changing that answer to parametric answer MM; for examples where the uncontrolled model generates MM, EMM→CEM_{M\to C} is the accuracy of changing it to contextual answer CC. For preservation, EMM→MEM_{M\to M} measures retaining MM when parametric steering is requested on examples that already produce MM, and EMC→CEM_{C\to C} measures retaining CC when contextual steering is requested on examples that already produce CC.

    On NQSwap, the plotted results show SPARE in the upper-right region for switching behaviour, outperforming the representation-engineering and contrastive-decoding baselines in jointly changing both directions. All methods find it harder to switch toward parametric knowledge than toward contextual knowledge, consistent with the models’ tendency to follow presented context under conflict. In the preservation analysis, SPARE retains behaviour at a level close to in-context learning and introduces less unnecessary change than dense activation-editing baselines. CAD preserves contextual behaviour particularly well but loses substantially more accuracy when preserving parametric behaviour.

  8. Knowl 8 — Both removal and addition are necessary for SPARE

    empirical result

    An ablation on NQSwap tests three simplified versions of SPARE: input-independent editing that applies fixed functional vectors without computing input-dependent removal and addition amounts; remove-only editing that subtracts undesired-behaviour features; and add-only editing that adds desired-behaviour features without removing undesired features. The full SPARE method computes both z−z^- and z+z^+ from the current input activation.

    The input-independent variant performs close to the uncontrolled baseline and therefore fails to reliably steer knowledge usage. The remove-only variant obtains zero accuracy for both EMMEM_M and EMCEM_C, showing that removing the features associated with the original behaviour alone does not produce the desired alternative answer or preserve the original answer. The add-only variant performs worse than no control. These results indicate that effective steering requires input-dependent magnitude constraints together with simultaneous removal of the undesired functional features and addition of the desired functional features.

  9. Knowl 9 — Middle layers are the most effective intervention sites

    empirical result

    Layer-wise experiments apply SPARE separately to individual layers of Llama3-8B and Gemma2-9B and measure exact-match performance for steering toward contextual and parametric answers. The strongest control occurs at middle layers rather than at the earliest or latest layers. These are also the layers where the conflict probe achieves its highest discrimination between conflict and non-conflict inputs.

    The result supports the paper’s interpretation that middle-layer residual representations contain functional features governing knowledge selection. It also explains why the main experiments edit contiguous middle-layer ranges: layers 13–16 for Llama3-8B, layers 12–15 for Llama2-7B, and layers 23–25 plus 29–31 for Gemma2-9B. The authors report that this is the first use in their study of pretrained SAEs to extract and edit such a knowledge-selection feature.

  10. Knowl 10 — SPARE produces distinct residual-stream distribution patterns

    empirical result

    The paper analyzes Llama3-8B conflict inputs edited at layer 15. When SPARE steers toward parametric knowledge, the residual-stream conflict-probe score decreases immediately, meaning that the edited representations become more similar to representations from inputs whose context agrees with parametric memory. When SPARE steers toward contextual knowledge, the probe score increases, indicating a stronger remaining conflict signal and representations further from the non-conflict group.

    The residual stream also develops different distributional patterns after intervention. Using kurtosis as the reported skewness-related measure, contextual steering makes later residual representations substantially more skewed beginning around layer 19, whereas parametric steering makes them less skewed. In uncontrolled conflict examples, the group that freely selects contextual answers has a more skewed residual stream than the group that selects parametric answers from approximately layer 19 onward. Additional Hoyer- and Gini-based analyses support the hidden-state skewness difference, while L1- and L2-norm curves do not show a distinct difference between the two selection behaviours. The paper presents these distributional observations empirically and does not establish their mechanistic cause.

  11. Knowl 11 — Scope limitations of SAE-based knowledge steering

    limitation

    SPARE depends on a suitable pretrained sparse autoencoder for the target language model. When such an SAE is unavailable, training one can be expensive; the authors’ Llama2-7B SAE pretraining required about 300 80-GB A100 GPU-hours per hidden-state layer. The method’s generality is therefore limited by SAE availability and training cost.

    The evaluation covers only two open-domain question-answering datasets with context-memory conflicts. It remains unclear how well SPARE transfers to other conflict types, complex reasoning, multi-hop questions, or long-form generation. In addition, the current controller treats knowledge selection as a binary choice between contextual and parametric sources. Practical systems may need a graded or more elaborate trust decision, potentially involving a separate critic model.

Coverage note — Omitted the exact per-layer Gemma2-9B SAE index lists, detailed baseline hyperparameter-search procedures, supplementary probing metrics, and additional norm/skewness plots because they are implementation-level or supporting analyses rather than load-bearing contributions.

References

  1. 1.Zeyuan Allen-Zhu and Yuanzhi Li. 2023. Physics of language models: Part 1, context-free grammar. arXiv preprint arXiv:2305.13673.
  2. 2.Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  3. 3.Leonard Bereska and Efstratios Gavves. 2024. Mechanistic interpretability for AI safety - A review. CoRR, abs/2404.14082.
  4. 4.Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L. Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E. Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah. 2023. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread.
  5. 5.Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  6. 6.Sviatoslav Chalnev, Matthew Siu, and Arthur Conmy. 2024. Improving steering vectors by targeting sparse autoencoder features. ArXiv, abs/2411.02193.
  7. 7.Canyu Chen and Kai Shu. 2023a. Can llm-generated misinformation be detected? arXiv preprint arXiv:2309.13788.
  8. 8.Canyu Chen and Kai Shu. 2023b. Combating misinformation in the age of llms: Opportunities and challenges. AI Magazine.
  9. 9.Hung-Ting Chen, Michael J. Q. Zhang, and Eunsol Choi. 2022. Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. In EMNLP, pages 2292–2307. Association for Computational Linguistics.
  10. 10.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024. Dola: Decoding by contrasting layers improves factuality in large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net.
  11. 11.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single vector: Probing sentence embeddings for linguistic properties. arXiv preprint arXiv:1805.01070.
  12. 12.Alessio Devoto, Yu Zhao, Simone Scardapane, and Pasquale Minervini. 2024. A simple and effective l2 norm-based strategy for KV cache compression. CoRR, abs/2406.11430.
  13. 13.Yibing Du, Antoine Bosselut, and Christopher D. Manning. 2022. Synthetic disinformation attacks on automated fact verification systems. In AAAI, pages 10581–10589. AAAI Press.
  14. 14.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  15. 15.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1):12.
  16. 16.Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024. Scaling and evaluating sparse autoencoders. CoRR, abs/2406.04093.
  17. 17.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495.
  18. 18.Yoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger, and Mor Geva. 2025. Enhancing automated interpretability with output-centric feature descriptions. arXiv preprint arXiv:2501.08319.
  19. 19.Roee Hendel, Mor Geva, and Amir Globerson. 2023. In-context learning creates task vectors. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 9318–9333. Association for Computational Linguistics.
  20. 20.Giwon Hong, Jeonghwan Kim, Junmo Kang, Sung-Hyon Myaeng, and Joyce Jiyoung Whang. 2024. Why so gullible? enhancing the robustness of retrieval-augmented models against counterfactual noise. In Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024, pages 2474–2495. Association for Computational Linguistics.
  21. 21.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. CoRR, abs/2311.05232.
  22. 22.Robert Huben, Hoagy Cunningham, Logan Riggs, Aidan Ewart, and Lee Sharkey. 2024. Sparse autoencoders find highly interpretable features in language models. In ICLR. OpenReview.net.
  23. 23.Shadi Iskander, Kira Radinsky, and Yonatan Belinkov. 2023. Shielded representations: Protecting sensitive attributes through iterative gradient-based projection. In ACL (Findings), pages 5961–5977. Association for Computational Linguistics.
  24. 24.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  25. 25.Zhuoran Jin, Pengfei Cao, Hongbang Yuan, Yubo Chen, Jiexin Xu, Huaijun Li, Xiaojian Jiang, Kang Liu, and Jun Zhao. 2024. Cutting off the head ends the conflict: A mechanism for interpreting and mitigating knowledge conflicts in language models. arXiv preprint arXiv:2402.18154.
  26. 26.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6769–6781. Association for Computational Linguistics.
  27. 27.Michael Lan, Philip Torr, Austin Meek, Ashkan Khakzar, David Krueger, and Fazl Barez. 2024. Sparse autoencoders reveal universal feature spaces across large language models. arXiv preprint arXiv:2410.06981.
  28. 28.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  29. 29.Kenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023. Inference-time intervention: Eliciting truthful answers from a language model. In NeurIPS.
  30. 30.Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca D. Dragan, Rohin Shah, and Neel Nanda. 2024. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2. CoRR, abs/2408.05147.
  31. 31.Sheng Liu, Haotian Ye, Lei Xing, and James Y. Zou. 2024. In-context vectors: Making in context learning more effective and controllable through latent space steering. In ICML. OpenReview.net.
  32. 32.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7052–7063, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  33. 33.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 9802–9822. Association for Computational Linguistics.
  34. 34.Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024. Sparse feature circuits: Discovering and editing interpretable causal graphs in language models. arXiv preprint arXiv:2403.19647.
  35. 35.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  36. 36.Julian Minder, Kevin Du, Niklas Stoehr, Giovanni Monea, Chris Wendler, Robert West, and Ryan Cotterell. 2024. Controllable context sensitivity and the knob behind it. arXiv preprint arXiv:2411.07404.
  37. 37.Chris Olah. 2023. Distributed representations: Composition & superposition. Transformer Circuits Thread.
  38. 38.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022. In-context learning and induction heads. arXiv preprint arXiv:2209.11895.
  39. 39.Francesco Ortu, Zhijing Jin, Diego Doimo, Mrinmaya Sachan, Alberto Cazzaniga, and Bernhard Schölkopf. 2024. Competition of mechanisms: Tracing how language models handle facts and counterfactuals. arXiv preprint arXiv:2402.11655.
  40. 40.Jane Pan, Tianyu Gao, Howard Chen, and Danqi Chen. 2023a. What in-context learning "learns" in-context: Disentangling task recognition and task learning. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, pages 8298–8319. Association for Computational Linguistics.
  41. 41.Liangming Pan, Wenhu Chen, Min-Yen Kan, and William Yang Wang. 2023b. Attacking open-domain question answering by injecting misinformation. In IJCNLP (1), pages 525–539. Association for Computational Linguistics.
  42. 42.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 2463–2473. Association for Computational Linguistics.
  43. 43.Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo M. Ponti, and Shay B. Cohen. 2024. Spectral editing of activations for large language model alignment. CoRR, abs/2405.09719.
  44. 44.Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda. 2024. Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders. CoRR, abs/2407.14435.
  45. 45.Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020. Null it out: Guarding protected attributes by iterative nullspace projection. In ACL, pages 7237–7256. Association for Computational Linguistics.
  46. 46.Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2023. Steering llama 2 via contrastive activation addition. CoRR, abs/2312.06681.
  47. 47.Morgane Rivière, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, Anton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, Matt Hoffman, Shantanu Thakoor, Jean-Bastien Grill, Behnam Neyshabur, Olivier Bachem, Alanna Walton, Aliaksei Severyn, Alicia Parrish, Aliya Ahmad, Allen Hutchison, Alvin Abdagic, Amanda Carl, Amy Shen, Andy Brock, Andy Coenen, Anthony Laforge, Antonia Paterson, Ben Bastian, Bilal Piot, Bo Wu, Brandon Royal, Charlie Chen, Chintu Kumar, Chris Perry, Chris Welty, Christopher A. Choquette-Choo, Danila Sinopalnikov, David Weinberger, Dimple Vijaykumar, Dominika Rogozinska, Dustin Herbison, Elisa Bandy, Emma Wang, Eric Noland, Erica Moreira, Evan Senter, Evgenii Eltyshev, Francesco Visin, Gabriel Rasskin, Gary Wei, Glenn Cameron, Gus Martins, Hadi Hashemi, Hanna Klimczak-Plucinska, Harleen Batra, Harsh Dhand, Ivan Nardini, Jacinda Mein, Jack Zhou, James Svensson, Jeff Stanway, Jetha Chan, Jin Peng Zhou, Joana Carrasqueira, Joana Iljazi, Jocelyn Becker, Joe Fernandez, Joost van Amersfoort, Josh Gordon, Josh Lipschultz, Josh Newlan, Ju-yeong Ji, Kareem Mohamed, Kartikeya Badola, Kat Black, Katie Millican, Keelin McDonell, Kelvin Nguyen, Kiranbir Sodhia, Kish Greene, Lars Lowe Sjösund, Lauren Usui, Laurent Sifre, Lena Heuermann, Leticia Lago, and Lilly McNealus. 2024. Gemma 2: Improving open language models at a practical size. CoRR, abs/2408.00118.
  48. 48.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2024. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36.
  49. 49.Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024. Trusting your evidence: Hallucinate less with context-aware decoding. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Short Papers, NAACL 2024, Mexico City, Mexico, June 16-21, 2024, pages 783–791. Association for Computational Linguistics.
  50. 50.Zhaochen Su, Jun Zhang, Xiaoye Qu, Tong Zhu, Yanshu Li, Jiashuo Sun, Juntao Li, Min Zhang, and Yu Cheng. 2024. Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm. arXiv preprint arXiv:2408.12076.
  51. 51.Nishant Subramani, Nivedita Suresh, and Matthew E. Peters. 2022. Extracting latent steering vectors from pretrained language models. In ACL (Findings), pages 566–581. Association for Computational Linguistics.
  52. 52.Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024a. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. Transformer Circuits Thread.
  53. 53.Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024b. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. Transformer Circuits Thread.
  54. 54.Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. 2024. Function vectors in large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net.
  55. 55.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  56. 56.Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023a. Activation addition: Steering language models without optimization. CoRR, abs/2308.10248.
  57. 57.Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023b. Activation addition: Steering language models without optimization. CoRR, abs/2308.10248.
  58. 58.Yike Wang, Shangbin Feng, Heng Wang, Weijia Shi, Vidhisha Balachandran, Tianxing He, and Yulia Tsvetkov. 2023. Resolving knowledge conflicts in large language models. CoRR, abs/2310.00935.
  59. 59.Yuxiang Wu, Yu Zhao, Baotian Hu, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. 2022. An efficient memory-augmented transformer for knowledge-intensive NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 5184–5196. Association for Computational Linguistics.
  60. 60.Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024a. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net.
  61. 61.Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024b. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. In ICLR. OpenReview.net.
  62. 62.Rongwu Xu, Zehan Qi, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. 2024. Knowledge conflicts for llms: A survey. arXiv preprint arXiv:2403.08319.
  63. 63.Zeyu Yun, Yubei Chen, Bruno Olshausen, and Yann LeCun. 2021. Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors. In Proceedings of Deep Learning Inside Out (DeeLIO): The 2nd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 1–10, Online. Association for Computational Linguistics.
  64. 64.Michael J. Q. Zhang and Eunsol Choi. 2021. Situatedqa: Incorporating extra-linguistic contexts into QA. In EMNLP (1), pages 7371–7387. Association for Computational Linguistics.
  65. 65.Wanru Zhao, Vidit Khazanchi, Haodi Xing, Xuanli He, Qiongkai Xu, and Nicholas Donald Lane. 2024a. Attacks on third-party apis of large language models. In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models.
  66. 66.Yu Zhao, Xiaotang Du, Giwon Hong, Aryo Pradipta Gema, Alessio Devoto, Hongru Wang, Xuanli He, Kam-Fai Wong, and Pasquale Minervini. 2024b. Analysing the residual stream of language models under knowledge conflicts. CoRR, abs/2410.16090.
  67. 67.Yu Zhao, Yuanbin Qu, Konrad Staniszewski, Szymon Tworkowski, Wei Liu, Piotr Miłos, Yuxiang Wu, and Pasquale Minervini. 2024c. Analysing the impact of sequence composition on language model pre-training. arXiv preprint arXiv:2402.13991.
  68. 68.Zexuan Zhong, Ziqing Huang, Alexander Wettig, and Danqi Chen. 2023. Poisoning retrieval corpora by injecting adversarial passages. arXiv preprint arXiv:2310.19156.
  69. 69.Zeyuan Allen Zhu and Yuanzhi Li. 2023. Physics of language models: Part 3.1, knowledge storage and extraction. arXiv preprint arXiv:2309.14316.
  70. 70.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023a. Representation engineering: A top-down approach to AI transparency. CoRR, abs/2310.01405.
  71. 71.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023b. Representation engineering: A top-down approach to AI transparency. CoRR, abs/2310.01405.
  72. 72.Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024. Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models. arXiv preprint arXiv:2402.07867.

Citation

MLA
Zhao, Y., et al. “Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 5117–36, https://doi.org/10.18653/v1/2025.naacl-long.264.
APA
Zhao, Y., Devoto, A., Hong, G., Du, X., Gema, A. P., Wang, H., He, X., Wong, K.-F., & Minervini, P. (2025). Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5117–5136. https://doi.org/10.18653/v1/2025.naacl-long.264
Chicago
Zhao, Y., A. Devoto, G. Hong, et al. 2025. “Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5117–36. https://doi.org/10.18653/v1/2025.naacl-long.264.
Harvard
Zhao, Y. et al. (2025) “Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5117–5136. Available at: https://doi.org/10.18653/v1/2025.naacl-long.264.
Vancouver
1. Zhao Y, Devoto A, Hong G, Du X, Gema AP, Wang H, He X, Wong K-F, Minervini P (2025) Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5117–5136

BibTeX

@inproceedings{zhao-etal-2025-steering,
    title = "Steering Knowledge Selection Behaviours in {LLM}s via {SAE}-Based Representation Engineering",
    author = "Zhao, Yu  and
      Devoto, Alessio  and
      Hong, Giwon  and
      Du, Xiaotang  and
      Gema, Aryo Pradipta  and
      Wang, Hongru  and
      He, Xuanli  and
      Wong, Kam-Fai  and
      Minervini, Pasquale",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.264/",
    doi = "10.18653/v1/2025.naacl-long.264",
    pages = "5117--5136",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/