Word Embeddings Are Steers for Language Models
Chi HanJialiang XuManling LiYi FungChenkai SunNan JiangTarek F. AbdelzaherHeng Ji
Proposes LM-Steer, a lightweight approach that applies linear transformations to output word embeddings to achieve parameter-efficient, transferable, and continuous style control across language models without retraining internal weights.
Large language models often exhibit undesirable generation behaviors, including producing toxic language, social biases, or off-target styles inherited from their pre-training data. Controlling these behaviors typically requires computationally expensive model retraining, large-scale fine-tuning, or slow, complex external classifier guidance at runtime. The article investigates the theoretical and practical role of output word embeddings in model generation and introduces LM-Steer, an extremely lightweight method that steers generation style and safety by applying a simple linear transformation to output word embeddings.
The authors evaluated the approach across multiple open-source language model families ranging from 14 million to 7 billion parameters, including GPT-2, Pythia, GPT-J, and Llama-2. Training was conducted using standard benchmark datasets for detoxification and sentiment control. The primary assessments measured reduction in toxic output, sentiment polarity adherence, generation fluency (measured by text perplexity), generation diversity, and decoding speed relative to leading controlled-generation and fine-tuning baselines.
The findings establish that LM-Steer consistently achieves superior or competitive control compared to existing baselines while preserving text quality. On the language detoxification benchmark, LM-Steer reduced average maximum toxicity by more than 6 absolute percentage points relative to strong baselines while maintaining high fluency and diversity. The method proved exceptionally resource efficient: it requires training only about 0.2% of the original model parameters (less than one-tenth the parameters of low-rank adaptation, or LoRA) and achieves strong detoxification with as few as 30 training examples. Furthermore, LM-Steer allows continuous intensity adjustment, compositional multi-attribute control (such as simultaneously adjusting sentiment and toxicity), and direct transfer to other language models via an explicit mathematical transformation without retraining.
These results provide a low-cost, low-latency mechanism to enhance AI safety and content moderation in production environments. Because LM-Steer operates as an efficient output layer transformation, it incurs minimal computational overhead during decoding, making it practical for real-time deployment without altering core model parameters. Additionally, decomposing the learned steer matrix offers interpretability by highlighting specific vocabulary dimensions and text spans driving stylistic or toxic attributes.
Decision-makers should consider piloting LM-Steer for rapid, cost-effective safety alignment and style customization on self-hosted or open-source models, especially where training budgets or inference latencies are constrained. However, stakeholders should note key limitations: LM-Steer primarily operates at the lexical and wording level and is not designed for complex multi-step reasoning or deep syntactic structural changes. In addition, deployment is restricted to environments with direct access to model embedding weights, precluding its direct use with closed, black-box third-party model APIs.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Read this foundational toxicity benchmark first to understand the generation failures and detoxification approaches that motivate LM-Steer’s evaluation.
- Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). Its causal account of linear representations and output-vocabulary geometry supplies useful theoretical grounding for steering behavior through embedding transformations.
- Paper: Aligning Large Language Models with Representation Editing: A Control Perspective, Lingkai Kong et al. (2024). This later control-theoretic method shifts steering from LM-Steer’s output embeddings to intermediate representations, extending lightweight inference-time alignment.
- Paper: In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering, Sheng Liu et al. (2024). This later work extends controllable generation by extracting task directions from demonstrations and applying them to internal representations rather than output embeddings.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). This later method uses sparse autoencoder features to steer knowledge selection, extending representation-based control beyond LM-Steer’s lexical output-layer interventions.
