In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
Sheng LiuHaotian YeLei XingJames Y. Zou
Proposes replacing prompt demonstrations with in-context vectors that steer language model latent states directly, significantly cutting context length while enabling precise behavioral control and multi-task composition through vector arithmetic.
Large language models frequently rely on in-context learning, where example demonstrations are placed directly into the input prompt to steer model outputs. However, standard in-context learning consumes valuable prompt window capacity, remains difficult to steer precisely, and exhibits high sensitivity to prompt wording. Meanwhile, conventional fine-tuning requires updating model parameters, which is computationally expensive and risks overfitting. The article addresses these challenges by developing a scalable, inference-only steering technique that extracts task knowledge directly into internal representations.
The main objective of the article is to introduce and evaluate the In-Context Vector method, an alternative steering framework that extracts task information from demonstration examples into a single directional vector and applies it directly to the model's internal representations during generation. The article demonstrates how this technique performs across tasks such as language detoxification, text style transfer, role-playing, and model jailbreaking compared to standard prompting and parameter fine-tuning.
To evaluate this framework, the authors conducted experimental evaluations across several open-source large language models, including Falcon-7B, LLaMA-7B, LLaMA-13B, and Vicuna-7B. The approach extracts task vectors from a small set of paired or unpaired demonstration examples by calculating differences in internal token representations across network layers using principal component analysis or contrastive gradients. During inference on new queries, the prompts remain free of demonstration examples; instead, the vector is added directly across all network layers at each generated token, scaled by an adjustable steering strength parameter.
The key findings show significant improvements across multiple performance and efficiency dimensions. First, in language detoxification on the ParaDetox benchmark using only five examples, the vector method reduced toxic outputs by approximately 45% to 50% compared to baseline models, cutting toxicity from 79.84% to 34.77% on Falcon-7B and outperforming standard prompting (73.09%) and fine-tuning (52.78%) while maintaining high semantic similarity. Second, in style transfer, the method increased formality scores to 48.30% (compared to 32.96% for standard prompting) and sentiment positivity to 75.28% (compared to 63.42% for standard prompting). Third, in role-playing tasks, the vector approach achieved higher win rates under automated evaluations than both fine-tuning and prompting, with performance scaling positively with model size. Fourth, the vectors support direct arithmetic operations, allowing practitioners to combine distinct behavioral vectors (such as adding safety while subtracting politeness) or negate vectors to reverse behaviors without additional training. Finally, the study demonstrated security risks: applying adversarial vectors achieved up to a 99% jailbreak attack success rate within seconds, bypassing safety alignments.
These findings indicate that internal vector steering offers a high-performance, cost-effective alternative to prompt stuffing and parameter tuning. By removing demonstration text from input prompts, organizations can reduce input token processing costs and bypass context length limits. The method also enables fine-grained control via an adjustable scaling factor, letting developers systematically balance steering intensity against text fluency. However, the demonstrated jailbreak capability highlights a serious safety vulnerability, proving that aligned open-source models can be compromised at inference time without retraining.
Organizations deploying open-source language models should consider adopting internal vector steering for style alignment, detoxification, and multi-attribute behavioral tuning to reduce inference latency and fine-tuning overhead. In deployment, engineering teams must tune the steering magnitude parameter carefully, as excessive strength degrades output coherence and semantic retention. Furthermore, safety teams must establish robust safeguards to monitor and defend against adversarial activation manipulation.
The findings are constrained by several boundaries. The method requires direct access to internal model representations and layer activations, restricting its use to open-source or self-hosted models and making it inapplicable to closed, black-box commercial programming interfaces. While the experimental evidence across the evaluated benchmarks provides high confidence in the method's efficacy, further testing on broader enterprise applications and complex reasoning tasks is recommended before large-scale production adoption.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). It analyzes the essential components of in-context learning demonstrations, establishing foundational insights into how demonstration context influences language model representations.
- Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). It introduces a framework for meta-training language models to infer task semantics directly from context demonstrations, providing necessary background on in-context adaptation.
- Paper: Learning by Distilling Context, Charlie Snell et al. (2022). It formulates the problem of mitigating in-context prompt overhead by condensing demonstration information into compact representations.
- Paper: Meta-learning via Language Model In-context Tuning, Yanda Chen et al. (2022). It explores tuning models to learn from few-shot context examples, clarifying the mechanics of in-context task transfer.
- Paper: Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering, Zhiyong Wu et al. (2023). It investigates information compression and demonstration selection in in-context learning, offering theoretical context for compressing demonstrations into latent states.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). It details the internal latent space geometry and contextual representation dynamics of language models that enable vector-based steering.
- Paper: The Linear Representation Hypothesis and the Geometry of Large Language Models, Kiho Park et al. (2024). It formalizes the linear representation hypothesis and causal inner products, providing the geometric theory that justifies linear latent steering mechanisms like in-context vectors.
- Paper: Aligning Large Language Models with Representation Editing: A Control Perspective, Lingkai Kong et al. (2024). It advances latent steering by applying dynamic control theory to edit intermediate representations during generation without model fine-tuning.
- Paper: Stay on Topic with Classifier-Free Guidance, Guillaume Sanchez et al. (2024). It explores inference-time guidance in next-token probability distributions, offering an alternative activation steering technique for prompt adherence.
- Paper: A Survey on In-context Learning, Qingxiu Dong et al. (2024). It provides a comprehensive survey and taxonomy of in-context learning mechanisms, contextualizing activation steering and demonstration compression within the broader field.
- Paper: Sample-Efficient Learning from Agent Experience, Chenhui Gou et al. (2026). It extends the principle of distilling context and experience into compact model behaviors for interactive and agentic environments.
