Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
Yu ZhaoAlessio DevotoGiwon HongXiaotang DuAryo Pradipta GemaHongru WangXuanli HeKam-Fai WongPasquale Minervini
Introduces SPARE, a training-free representation engineering method that leverages sparse auto-encoders to detect mid-layer conflict signals and steer whether large language models rely on parametric memory or contextual evidence during question answering.
Large language models store extensive factual knowledge in their internal parameters, but they are also frequently paired with external text retrieval systems to provide up-to-date context. When retrieved external text directly contradicts a model's internal memory—a problem known as a context-memory knowledge conflict—models typically default to following the provided context. This dynamic creates significant real-world risks, such as propagating external misinformation or overriding accurate internal knowledge.
The main objective of the article is to demonstrate how to detect these knowledge conflicts inside language models and introduce a practical, inference-time method to precisely steer whether a model relies on external context or its internal parametric memory.
To achieve this, the authors developed SPARE (Sparse Auto-Encoder-based Representation Engineering), a training-free intervention approach. Sparse auto-encoders are auxiliary neural networks that break down a model's dense, tangled internal representations into isolated, interpretable individual features. The researchers evaluated internal states across multiple models (Llama2-7B, Llama3-8B, and Gemma2-9B) using two benchmark question-answering datasets designed around knowledge conflicts (NQSwap and Macnoise). They identified which isolated features correlate with relying on context versus relying on memory, and then selectively added desired features and removed undesired features at middle model layers during the generation process.
The key findings demonstrate clear performance gains over standard techniques. First, internal probing revealed that knowledge conflicts produce strong, detectable signals concentrated specifically within the middle layers of the models. Second, steering models with SPARE outperformed existing activation-editing techniques by approximately 10 percentage points and outperformed contrastive decoding methods by roughly 15 percentage points in accuracy. Third, the method proved effective using an extremely small subset of features—manipulating less than 0.05% of available auto-encoder features in Gemma2-9B. Fourth, while previous methods struggled to enforce reliance on internal memory when misleading context was present, SPARE successfully steered models toward parametric memory (e.g., reaching 47.5% accuracy on Llama3-8B compared to 26.6% without control) while also achieving high context faithfulness when context was desired (reaching up to 92.2% on Macnoise). Ablation tests confirmed that both adding target features and subtracting conflicting features are essential to avoid severe performance degradation.
These findings indicate that organizations can regulate the behavior and trustworthiness of deployed language models at runtime without expensive model fine-tuning or prompt-based multi-turn verification that slows down responses. This provides a low-latency mechanism to mitigate misinformation risks and enforce organizational policies regarding when to trust external inputs over stored model memory. The discovery that middle layers contain the critical control points confirms earlier mechanistic findings while offering a more precise tool for direct model steering.
Organizations evaluating this approach should consider piloting sparse auto-encoder interventions in retrieval-augmented applications where misinformation defense is critical. Prior to widespread deployment, practitioners should verify feature calibration on representative domain data and establish external decision logic to determine whether context or internal memory is more likely to be accurate for a given query.
Certain limitations remain. The method depends on the availability or upfront training of sparse auto-encoders, which require non-trivial compute resources if not already publicly accessible. In addition, the empirical evaluations were focused on direct open-domain question-answering datasets; further investigation is needed to confirm generalizability to complex multi-step reasoning, long-form text generation, and nuanced, non-binary truth scenarios.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). This foundational work demonstrates that sparse autoencoders can decompose language model activations into interpretable, manipulable features, providing the exact mechanistic foundation that SPARE builds upon to steer knowledge selection.
- Paper: Scaling and evaluating sparse autoencoders, Leo Gao et al. (2025). This paper establishes scalable methods for training high-capacity sparse autoencoders to extract reliable internal features from language models, underlying SPARE's use of pre-trained SAE representations.
- Paper: Locating and Editing Factual Associations in GPT, Kevin Meng et al. (2022). It introduces the locate-and-edit framework and demonstrates that middle-layer activations in LLMs mediate factual knowledge associations, which motivates SPARE's targeted mid-layer activation interventions.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). It characterizes the behavioral divergence and memory boundaries between an LLM's internal parametric knowledge and external non-parametric retrieval context, defining the context-memory knowledge conflict setting addressed by SPARE.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). It provides a method to discover and probe latent factual knowledge directly from internal activations without fine-tuning, motivating SPARE's detection of internal conflict signals.
No sufficiently relevant recommendations were found.
