Built independently by an author, for readers. Read the story and support ChapterPal

keyword

SAE-based representation engineering

SAE-based representation engineering is a model-steering and interpretability framework that uses sparse autoencoders to analyze and manipulate the internal hidden activations of neural networks, particularly large language models. Rather than modifying model parameters through retraining or intervening directly on dense, polysemantic activation vectors, this approach projects internal representations into an expanded, sparse latent space where individual features correspond to interpretable concepts or functional behaviors. By identifying the specific latent features associated with targeted behaviors or cognitive states, practitioners can modify activations during inference to reliably steer model outputs, resolve internal information conflicts, and enhance behavioral alignment without altering the underlying model weights.

1 item

Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering

Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering

Yu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, Pasquale Minervini

OrganizationsMiniml.AISapienza University of RomeThe Chinese University of Hong KongUniversity College LondonUniversity of Edinburgh

Why you should read this

Introduces SPARE, a training-free representation engineering method that leverages sparse auto-encoders to detect mid-layer conflict signals and steer whether large language models rely on parametric memory or contextual evidence during question answering.

Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided in the context—this phenomenon, known as context-memory knowledge conflicts 1, can lead to undesirable model behaviour, such as reliance on outdated or incorrect information. Analysing the internal activations of LLMs, we find that they can internally register the signals of knowledge conflict at mid-layers. Such signals allow us to detect whether a knowledge conflict occurs and use inference-time intervention strategies to resolve it. In this work, we propose SPARE, a training-free representation engineering method that uses pre-trained sparse auto-encoders (SAEs) to control the knowledge selection behaviour of LLMs. SPARE identifies the functional features that control the knowledge selection behaviours and applies them to edit the internal activations of LLMs at inference time. Our experimental results show that SPARE can effectively control the usage of either knowledge source to resolve knowledge conflict in open-domain question-answering tasks, surpassing existing representation engineering methods (+10%) as well as contrastive decoding methods (+15%).

Added

2026-09-26