In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Sheng LiuHaotian YeLei XingJames Y. Zou

article2024ICML288 citations

Proposes replacing prompt demonstrations with in-context vectors that steer language model latent states directly, significantly cutting context length while enabling precise behavioral control and multi-task composition through vector arithmetic.

Listen

Large language models frequently rely on in-context learning, where example demonstrations are placed directly into the input prompt to steer model outputs. However, standard in-context learning consumes valuable prompt window capacity, remains difficult to steer precisely, and exhibits high sensitivity to prompt wording. Meanwhile, conventional fine-tuning requires updating model parameters, which is computationally expensive and risks overfitting. The article addresses these challenges by developing a scalable, inference-only steering technique that extracts task knowledge directly into internal representations.

The main objective of the article is to introduce and evaluate the In-Context Vector method, an alternative steering framework that extracts task information from demonstration examples into a single directional vector and applies it directly to the model's internal representations during generation. The article demonstrates how this technique performs across tasks such as language detoxification, text style transfer, role-playing, and model jailbreaking compared to standard prompting and parameter fine-tuning.

To evaluate this framework, the authors conducted experimental evaluations across several open-source large language models, including Falcon-7B, LLaMA-7B, LLaMA-13B, and Vicuna-7B. The approach extracts task vectors from a small set of paired or unpaired demonstration examples by calculating differences in internal token representations across network layers using principal component analysis or contrastive gradients. During inference on new queries, the prompts remain free of demonstration examples; instead, the vector is added directly across all network layers at each generated token, scaled by an adjustable steering strength parameter.

The key findings show significant improvements across multiple performance and efficiency dimensions. First, in language detoxification on the ParaDetox benchmark using only five examples, the vector method reduced toxic outputs by approximately 45% to 50% compared to baseline models, cutting toxicity from 79.84% to 34.77% on Falcon-7B and outperforming standard prompting (73.09%) and fine-tuning (52.78%) while maintaining high semantic similarity. Second, in style transfer, the method increased formality scores to 48.30% (compared to 32.96% for standard prompting) and sentiment positivity to 75.28% (compared to 63.42% for standard prompting). Third, in role-playing tasks, the vector approach achieved higher win rates under automated evaluations than both fine-tuning and prompting, with performance scaling positively with model size. Fourth, the vectors support direct arithmetic operations, allowing practitioners to combine distinct behavioral vectors (such as adding safety while subtracting politeness) or negate vectors to reverse behaviors without additional training. Finally, the study demonstrated security risks: applying adversarial vectors achieved up to a 99% jailbreak attack success rate within seconds, bypassing safety alignments.

These findings indicate that internal vector steering offers a high-performance, cost-effective alternative to prompt stuffing and parameter tuning. By removing demonstration text from input prompts, organizations can reduce input token processing costs and bypass context length limits. The method also enables fine-grained control via an adjustable scaling factor, letting developers systematically balance steering intensity against text fluency. However, the demonstrated jailbreak capability highlights a serious safety vulnerability, proving that aligned open-source models can be compromised at inference time without retraining.

Organizations deploying open-source language models should consider adopting internal vector steering for style alignment, detoxification, and multi-attribute behavioral tuning to reduce inference latency and fine-tuning overhead. In deployment, engineering teams must tune the steering magnitude parameter carefully, as excessive strength degrades output coherence and semantic retention. Furthermore, safety teams must establish robust safeguards to monitor and defend against adversarial activation manipulation.

The findings are constrained by several boundaries. The method requires direct access to internal model representations and layer activations, restricting its use to open-source or self-hosted models and making it inapplicable to closed, black-box commercial programming interfaces. While the experimental evidence across the evaluated benchmarks provides high confidence in the method's efficacy, further testing on broader enterprise applications and complex reasoning tasks is recommended before large-scale production adoption.

Cover for In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Abstract

Large language models (LLMs) demonstrate emergent in-context learning capabilities, where they adapt to new tasks based on example demonstrations. However, in-context learning has seen limited effectiveness in many settings, is difficult to quantitatively control and takes up context window space. To overcome these limitations, we propose an alternative approach that recasts in-context learning as in-context vectors (ICV). Using ICV has two steps. We first use a forward pass on demonstration examples to create the in-context vector from the latent embedding of the LLM. This vector captures essential information about the intended task. On a new query, instead of adding demonstrations to the prompt, we shift the latent states of the LLM using the ICV. The ICV approach has several benefits: 1) it enables the LLM to more effectively follow the demonstration examples; 2) it's easy to control by adjusting the magnitude of the ICV; 3) it reduces the length of the prompt by removing the in-context demonstrations; 4) ICV is computationally much more efficient than fine-tuning. We demonstrate that ICV achieves better performance compared to standard in-context learning and fine-tuning on diverse tasks including safety, style transfer, role-playing and formatting. Moreover, we show that we can flexibly teach LLM to simultaneously follow different types of instructions by simple vector arithmetics on the corresponding ICVs.

Table of Contents

  • 1 Introduction
  • 2 Backgrounds
  • 3 Method
  • 3.1 Task summary: generating in-context vector (ICV)
  • 3.1.1 Paired demonstrations.
  • 3.1.2 Unpaired demonstrations.
  • 3.2 Feature shifting: apply the ICV to query examples
  • 3.3 Task arithmetic property of in-context vectors
  • 4 Experiments
  • 4.1 Safety of LLMs
  • 4.2 Writing style and role-playing
  • 4.3 Large language models
  • 4.4 Automatic text similarity evaluation
  • 5 Results
  • 5.1 Language detoxification and dialogue safety
  • 5.2 Jail break
  • 5.3 Speaking style
  • 5.4 Role-playing
  • 5.5 Task arithmetic
  • 6 Related works
  • 7 Conclusions
  • References
  • A More experimental details
  • B Datasets and demonstration examples
  • C In-context learning as feature shifting
  • D Latent states for text-classification datasets
  • E Proof of lemma
  • F Gradients of the contrastive objective Eq. ()

Knowls

  1. Knowl 1 — In-Context Vector (ICV) Latent Space Steering Framework

    model/method

    The In-Context Vector (ICV) framework is an inference-only alternative to In-Context Learning (ICL) that steers the internal activations of a Transformer large language model (LLM) without modifying model parameters or prepending demonstration tokens to the query context window.

    The framework operates in two distinct stages:

    1. Task Summary (Generating ICV): Given demonstration inputs xx and targets yy, each example is passed through the LLM separately in a forward pass. For each example, the hidden state activations hl∈R1×dh_l \in \mathbb{R}^{1 \times d} at the last token position across all LL Transformer layers are extracted and concatenated into a single full-model representation vector h∈R1×(L×d)h \in \mathbb{R}^{1 \times (L \times d)}. An optimization objective or closed-form statistic across demonstration representations is solved to produce a task-specific vector hICV∈R1×(L×d)h_{\text{ICV}} \in \mathbb{R}^{1 \times (L \times d)}.

    2. Feature Shifting (Inference on Query): Given a new query sequence XqueryX_{\text{query}} of length TT, the query is evaluated without demonstration context in the prompt. At each layer l∈{1,…,L}l \in \{1, \dots, L\} and at every token position t∈{1,…,T}t \in \{1, \dots, T\}, the latent state ht,l∈R1×dh_{t,l} \in \mathbb{R}^{1 \times d} is shifted by adding the corresponding layer slice hICVl∈R1×dh_{\text{ICV}}^l \in \mathbb{R}^{1 \times d} scaled by a hyperparameter λ\lambda, followed by an ℓ2\ell_2-norm normalization to maintain the original activation magnitude.

  2. Knowl 2 — ICV Generation for Paired Demonstrations via Principal Component Analysis

    theoretical result

    For a set of kk paired demonstration examples {(xi,yi)}i=1k\{(x_i, y_i)\}_{i=1}^k with extracted multi-layer representation vectors h(xi),h(yi)∈R1×(L×d)h(x_i), h(y_i) \in \mathbb{R}^{1 \times (L \times d)}, the desired in-context vector hh maximizes the variance of the projected differences:

    max⁡∥h∥2=11k∑i=1k(h⊤h(yi)−h⊤h(xi))2=max⁡∥h∥2=1h⊤(1k∑i=1kdidi⊤)h=max⁡∥h∥2=1h⊤ΣDh\max_{\|h\|_2 = 1} \frac{1}{k} \sum_{i=1}^k \left( h^\top h(y_i) - h^\top h(x_i) \right)^2 = \max_{\|h\|_2 = 1} h^\top \left( \frac{1}{k} \sum_{i=1}^k d_i d_i^\top \right) h = \max_{\|h\|_2 = 1} h^\top \Sigma_{\mathcal{D}} h

    where di=h(yi)−h(xi)d_i = h(y_i) - h(x_i) is the difference vector for pair ii, and ΣD=1k∑i=1kdidi⊤\Sigma_{\mathcal{D}} = \frac{1}{k} \sum_{i=1}^k d_i d_i^\top is the sample covariance matrix of the difference set D={d1,…,dk}\mathcal{D} = \{d_1, \dots, d_k\}.

    By the spectral theorem for symmetric matrices, the optimal unit vector hICVh_{\text{ICV}} is the principal eigenvector (first principal component from PCA) corresponding to the largest eigenvalue of ΣD\Sigma_{\mathcal{D}}.

  3. Knowl 3 — ICV Generation for Unpaired Demonstrations via Contrastive Gradient

    model/method

    When demonstration examples are unpaired—consisting of a negative set X={h(x1),…,h(xm)}\mathcal{X} = \{h(x_1), \dots, h(x_m)\} and a positive set Y={h(y1),…,h(yn)}\mathcal{Y} = \{h(y_1), \dots, h(y_n)\}—the in-context vector is obtained by optimizing the contrastive objective:

    g(h)=∑y∈Ylog⁡exp⁡(h⊤h(y))exp⁡(h⊤h(y))+∑x∈Xexp⁡(h⊤h(x))g(h) = \sum_{y \in \mathcal{Y}} \log \frac{\exp(h^\top h(y))}{\exp(h^\top h(y)) + \sum_{x \in \mathcal{X}} \exp(h^\top h(x))}

    Because a closed-form maximizer is unavailable, the in-context vector hICVh_{\text{ICV}} is set directly to the analytical gradient ∂g∂h\frac{\partial g}{\partial h} evaluated at the representations:

    hICV=∑y∈Y((1−py)h(y)−∑x∈Xpxh(x))h_{\text{ICV}} = \sum_{y \in \mathcal{Y}} \left( (1 - p_y) h(y) - \sum_{x \in \mathcal{X}} p_x h(x) \right)

    where the softmax probabilities over representations are defined as:

    py=exp⁡(h⊤h(y))exp⁡(h⊤h(y))+∑x∈Xexp⁡(h⊤h(x)),px=exp⁡(h⊤h(x))exp⁡(h⊤h(y))+∑x′∈Xexp⁡(h⊤h(x′))p_y = \frac{\exp(h^\top h(y))}{\exp(h^\top h(y)) + \sum_{x \in \mathcal{X}} \exp(h^\top h(x))}, \quad p_x = \frac{\exp(h^\top h(x))}{\exp(h^\top h(y)) + \sum_{x' \in \mathcal{X}} \exp(h^\top h(x'))}

  4. Knowl 4 — Latent State Feature Shifting, Adaptive Scaling, and Norm Matching

    equation

    During query generation, for a sequence with TT tokens across LL layers, the latent state ht,l∈R1×dh_{t,l} \in \mathbb{R}^{1 \times d} at position tt and layer ll (obtained after the MLP sub-layer following self-attention) is updated as:

    h~t,l:=ht,l+λhICVl\tilde{h}_{t,l} := h_{t,l} + \lambda h_{\text{ICV}}^l

    where hICVl∈R1×dh_{\text{ICV}}^l \in \mathbb{R}^{1 \times d} is the ll-th layer segment of the in-context vector and λ>0\lambda > 0 is the task strength scaling factor.

    To adaptively scale steering according to how misaligned an activation is with the task direction, λ\lambda can be dynamically modified as:

    λt,l=λ⋅(1+max⁡{0,cos⁡(ht,l,−Δhl)})\lambda_{t,l} = \lambda \cdot \left( 1 + \max\{0, \cos(h_{t,l}, -\Delta h^l)\} \right)

    where cos⁡(u,v)=u⋅v∥u∥2∥v∥2\cos(u, v) = \frac{u \cdot v}{\|u\|_2 \|v\|_2} and Δhl=hICVl\Delta h^l = h_{\text{ICV}}^l.

    To preserve the model's native representation scale and retain generation stability, the steered latent state is normalized to match its original ℓ2\ell_2 norm:

    h~t,l:=h~t,l⋅∥ht,l∥2∥h~t,l∥2\tilde{h}_{t,l} := \tilde{h}_{t,l} \cdot \frac{\|h_{t,l}\|_2}{\|\tilde{h}_{t,l}\|_2}

  5. Knowl 5 — Attention Decomposition of In-Context Learning as Latent Feature Shifting

    theoretical result

    In standard Transformer self-attention with query weight WqW_q, key weight WkW_k, and value weight WvW_v, prepending demonstration tokens XdemosX_{\text{demos}} to query tokens XqueryX_{\text{query}} partitions the input matrix as X=[Xdemos;Xquery]X = [X_{\text{demos}}; X_{\text{query}}]. The attention computation for a query token xqueryx_{\text{query}} decomposes linearly into a weighted sum of query-only attention and demonstration-cross attention:

    Attn(xqueryWq,XWk,XWv)=αh(Xquery)+(1−α)h(Xdemos)\text{Attn}(x_{\text{query}} W_q, X W_k, X W_v) = \alpha h(X_{\text{query}}) + (1 - \alpha) h(X_{\text{demos}})

    where:

    h(Xquery)=Attn(xqueryWq,XqueryWk,XqueryWv)h(X_{\text{query}}) = \text{Attn}(x_{\text{query}} W_q, X_{\text{query}} W_k, X_{\text{query}} W_v)

    h(Xdemos)=Attn(xqueryWq,XdemosWk,XdemosWv)h(X_{\text{demos}}) = \text{Attn}(x_{\text{query}} W_q, X_{\text{demos}} W_k, X_{\text{demos}} W_v)

    and α∈(0,1)\alpha \in (0, 1) is the normalized scalar attention weight assigned to the query tokens:

    α=∑iexp⁡(xqueryWqWk⊤Xquery)i∑iexp⁡(xqueryWqWk⊤Xdemos)i+∑jexp⁡(xqueryWqWk⊤Xquery)j\alpha = \frac{\sum_{i} \exp(x_{\text{query}} W_q W_k^\top X_{\text{query}})_{i}}{\sum_{i} \exp(x_{\text{query}} W_q W_k^\top X_{\text{demos}})_{i} + \sum_{j} \exp(x_{\text{query}} W_q W_k^\top X_{\text{query}})_{j}}

    This shows that conventional in-context learning acts by shifting the query activation h(Xquery)h(X_{\text{query}}) via a position-wise perturbation (1−α)h(Xdemos)(1 - \alpha) h(X_{\text{demos}}) governed implicitly by attention weights.

  6. Knowl 6 — Task Arithmetic on In-Context Vectors

    model/method

    In-context vectors support linear vector arithmetic to compose, combine, or invert behavioral transformations without model retraining:

    1. Task Inversion via Negation: Given an ICV hx→yh_{x \to y} derived from demonstrations converting domain xx to domain yy (e.g., informal to formal text), applying −hx→y-h_{x \to y} performs the reverse transformation y→xy \to x (e.g., formal to informal) using the same precomputed vector.
    2. Multi-Task Steering via Vector Addition and Subtraction: Multiple ICVs extracted for distinct attributes (such as safety, politeness, formality, or sentiment) can be linearly combined with individual scaling factors λi\lambda_i:

    hcombined=∑i±λihICV(i)h_{\text{combined}} = \sum_{i} \pm \lambda_i h_{\text{ICV}}^{(i)}

    For example, adding a safety vector and subtracting a politeness vector (+hsafe−hpolite+h_{\text{safe}} - h_{\text{polite}}) steers the model to generate responses that are safe and non-discriminatory yet rude/blunt.

  7. Knowl 7 — Language Detoxification Performance on ParaDetox

    data/table

    Evaluation of text detoxification on the ParaDetox benchmark comparing unsteered base models (w/o context), standard 5-shot In-Context Learning (ICL), 5-shot LoRA fine-tuning (LoRA FT, 20 epochs), task vector replacement, and 5-shot In-Context Vector steering (paired and unpaired, λ=0.1\lambda = 0.1) on Falcon-7B and LLaMA-7B across 670 test queries. Toxicity (%) was evaluated using a 2.7B-parameter safety classifier (lower is better); content preservation was measured by raw-text ROUGE-1 and feature-level BERTScore against non-toxic references (higher is better).

    Method Toxicity (%) ↓\downarrow ROUGE-1 ↑\uparrow BERT ↑\uparrow
    Original Test Set 84.58 - -
    Gold-standard Reference 16.23 - -
    Falcon-7b (w/o context) 79.84 72.60 93.29
    Falcon-7b (ICL) 73.09 73.58 93.51
    Falcon-7b (LoRA FT) 52.78 61.35 90.03
    Falcon-7b (Task Vector, λ=0.5\lambda = 0.5) 53.36 62.37 90.01
    Falcon-7b (ICV, λ=0.1\lambda = 0.1) 34.77 65.76 92.88
    Falcon-7b (ICV, unpaired, λ=0.1\lambda = 0.1) 35.56 64.76 91.27
    Llama-7b (w/o context) 71.60 73.15 93.32
    Llama-7b (ICL) 66.81 74.19 93.11
    Llama-7b (LoRA FT) 48.94 57.32 89.34
    Llama-7b (ICV, λ=0.1\lambda = 0.1) 39.54 65.97 92.73
    Llama-7b (ICV, unpaired, λ=0.1\lambda = 0.1) 40.15 64.11 91.76

    ICV achieves a reduction in toxicity by 49.81% on Falcon-7B (from 79.84% to 34.77%) and 45.04% on LLaMA-7B (from 71.60% to 39.54%), substantially outperforming 5-shot ICL (73.09% and 66.81%) and LoRA fine-tuning (52.78% and 48.94%) while maintaining high BERTScore semantic similarity.

  8. Knowl 8 — Style Transfer Performance on Formality and Sentiment

    data/table

    Quantitative results on two style transfer tasks using LLaMA-7B evaluated with 5 demonstration examples: (1) Formality transfer on 1,332 test sentences from the GYAFC family and relationships domain evaluated with an XLM-RoBERTa classifier (accuracy 85.21%); (2) Sentiment transfer on 1,000 Yelp review test samples evaluated with a DistilBERT classifier. Content retention was measured via ROUGE-1 and BERTScore.

    Informal →\rightarrow Formal Negative →\rightarrow Positive
    Method Formality (%) ↑\uparrow ROUGE-1 ↑\uparrow BERT ↑\uparrow Positivity (%) ↑\uparrow ROUGE-1 ↑\uparrow BERT ↑\uparrow
    Original Test Set 11.49 - - 10.10 - -
    w/o context 17.54 81.54 92.61 35.81 78.85 95.59
    ICL 32.96 83.85 93.61 63.42 73.86 95.00
    LoRA FT 21.99 80.13 92.86 65.92 66.91 93.89
    ICV (λ=0.1\lambda = 0.1) 48.30 80.23 92.81 75.28 68.27 94.32
    ICV (λ=0.12\lambda = 0.12, unpaired) 36.30 78.17 91.81 67.13 65.10 93.42

    Paired ICV (λ=0.1\lambda = 0.1) achieves 48.30% formality and 75.28% positivity, outperforming both standard ICL (32.96% formality, 63.42% positivity) and LoRA fine-tuning (21.99% formality, 65.92% positivity).

  9. Knowl 9 — Jailbreaking Safety-Aligned LLMs via Adversarial ICVs

    data/table

    Attack success rate (ASR, in %) evaluated on Vicuna-7B across 100 harmful target behaviors. Conventional ICL and ICVs were created using 5 demonstration examples of malicious queries paired with compliance responses. Attack success is verified if the model does not output standard refusal prefixes (e.g., 'I'm sorry', 'I apologize', 'I cannot') and generates coherent harmful instructions.

    Attack Method ASR (%) ↑\uparrow
    No-Attack 00.0
    In-Context Learning 44.0
    ICV (λ=0.10\lambda = 0.10) 50.0
    ICV (λ=0.18\lambda = 0.18) 93.0
    ICV (λ=0.20\lambda = 0.20) 99.0

    While 5-shot ICL reaches an ASR of 44.0%, scaling the ICV step size to λ=0.20\lambda = 0.20 enables ICV to achieve a 99.0% attack success rate within seconds of inference time, matching optimization-based attack baselines (e.g., GBDA, PEZ) that require roughly 30 minutes per instance.

  10. Knowl 10 — Layer-Wise Ablation of ICV Injection

    data/table

    Ablation study assessing the effect of injecting the in-context vector into specific individual Transformer layers versus across all layers simultaneously. Evaluations were performed on the Falcon-7B model on the ParaDetox language detoxification task with scaling factor λ=0.1\lambda = 0.1.

    Layer Configuration Toxicity (%) ↓\downarrow ROUGE-1 ↑\uparrow
    w/o ICV 79.84 72.60
    First layer 78.33 72.15
    Middle layer 78.56 72.11
    Last layer 38.37 67.26
    All layers 34.77 65.76

    Applying ICV solely to the first or middle layer produces minimal improvement over the unsteered baseline (78.33% and 78.56% vs 79.84%). Applying ICV to the last layer provides a strong reduction (38.37%), but the lowest toxicity and maximum steering efficacy is attained when ICV is applied across all layers simultaneously (34.77%).

  11. Knowl 11 — White-Box Activation Access Requirement for ICV

    limitation

    A key practical limitation of the In-Context Vector (ICV) approach compared to standard prompt-based In-Context Learning (ICL) is that ICV requires white-box access to the model's internal Transformer architecture. Specifically, the practitioner must be able to extract hidden states across intermediate layers during the task-summary pass and intervene on latent activations at every layer during forward passes of new queries. Consequently, ICV cannot be applied directly to black-box, closed-source LLMs accessed strictly via text-in/text-out commercial APIs.

Coverage note — Appendix D on 1-NN classification on AGNews using latent states was omitted as it is a minor exploratory side experiment tangential to the core generative ICV steering method.

References

  1. 1.Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. What learning algorithm is in-context learning? investigations with linear models. In The Eleventh International Conference on Learning Representations, 2022.
  2. 2.Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. arXiv preprint arXiv:2309.07875, 2023.
  3. 3.Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al. Gpt-neox-20b: An open-source autoregressive language model. In Proceedings of BigScience Episode# 5–Workshop on Challenges & Perspectives in Creating Large Language Models, pp. 95–136, 2022.
  4. 4.Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29, 2016.
  5. 5.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  6. 6.Burns, C., Ye, H., Klein, D., and Steinhardt, J. Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827, 2022.
  7. 7.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023.
  8. 8.Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, É., Ott, M., Zettlemoyer, L., and Stoyanov, V. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8440–8451, 2020.
  9. 9.Cui, G., Li, W., Ding, N., Huang, L., Liu, Z., and Sun, M. Decoder tuning: Efficient language understanding as decoding. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 15072–15087. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.ACL-LONG. 840.
  10. 10.Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F. Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers. arXiv preprint arXiv:2212.10559, 2022.
  11. 11.Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z. A survey for in-context learning. arXiv preprint arXiv:2301.00234, 2022.
  12. 12.Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. In Cohn, T., He, Y., and Liu, Y. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pp. 3356–3369. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020. FINDINGS-EMNLP.301.
  13. 13.Gilbert, S., Harvey, H., Melvin, T., Vollebregt, E., and Wicks, P. Large language model ai chatbots require approval as medical devices. Nature Medicine, pp. 1–3, 2023.
  14. 14.Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D. Gradient-based adversarial attacks against text transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 5747–5757, 2021.
  15. 15.Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8342–8360, 2020.
  16. 16.Hendel, R., Geva, M., and Globerson, A. In-context learning creates task vectors. arXiv preprint arXiv:2310.15916, 2023.
  17. 17.Holtzman, A., West, P., Shwartz, V., Choi, Y., and Zettlemoyer, L. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7038–7051, 2021.
  18. 18.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  19. 19.Hwang, G.-J. and Chang, C.-Y. A review of opportunities and challenges of chatbots in education. Interactive Learning Environments, 31(7):4099–4112, 2023.
  20. 20.Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A. Editing models with task arithmetic. In The Eleventh International Conference on Learning Representations, 2023.
  21. 21.Krishna, K., Nathani, D., Garcia, X., Samanta, B., and Talukdar, P. Few-shot controllable style transfer for low-resource multilingual settings. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pp. 7439–7468. Association for Computational Linguistics, 2022. doi: 10.18653/V1/ 2022.ACL-LONG.514.
  22. 22.Lee, P., Bubeck, S., and Petro, J. Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine. New England Journal of Medicine, 388(13):1233–1239, 2023.
  23. 23.Leong, C. T., Cheng, Y., Wang, J., Wang, J., and Li, W. Self-detoxifying language models via toxification reversal. arXiv preprint arXiv:2310.09573, 2023.
  24. 24.Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M. Emergent world representations: Exploring a sequence model trained on a synthetic task. arXiv preprint arXiv:2210.13382, 2022.
  25. 25.Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model. arXiv preprint arXiv:2306.03341, 2023.
  26. 26.Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  27. 27.Liu, J., Shen, D., Zhang, Y., Dolan, W. B., Carin, L., and Chen, W. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pp. 100–114, 2022a.
  28. 28.Liu, S., Niles-Weed, J., Razavian, N., and Fernandez-Granda, C. Early-learning regularization prevents memorization of noisy labels. Advances in neural information processing systems, 33:20331–20342, 2020.
  29. 29.Liu, S., Zhang, X., Sekhar, N., Wu, Y., Singhal, P., and Fernandez-Granda, C. Avoiding spurious correlations via logit correction. In The Eleventh International Conference on Learning Representations, 2022b.
  30. 30.Liu, S., Zhu, Z., Qu, Q., and You, C. Robust training under label noise by over-parameterization. In International Conference on Machine Learning, pp. 14153–14172. PMLR, 2022c.
  31. 31.Logacheva, V., Dementieva, D., Ustyantsev, S., Moskovskiy, D., Dale, D., Krotova, I., Semenov, N., and Panchenko, A. Paradetox: Detoxification with parallel data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6804–6818, 2022.
  32. 32.Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y. QUARK: controllable text generation with reinforced unlearning. In NeurIPS, 2022a.
  33. 33.Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8086–8098, 2022b.
  34. 34.Meade, N., Gella, S., Hazarika, D., Gupta, P., Jin, D., Reddy, S., Liu, Y., and Hakkani-Tür, D. Using in-context learning to improve dialogue safety. arXiv preprint arXiv:2302.00871, 2023.
  35. 35.Min, S., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Noisy channel language model prompting for few-shot text classification. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pp. 5316–5330. Association for Computational Linguistics, 2022a. doi: 10.18653/V1/2022.ACL-LONG. 365.
  36. 36.Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 11048–11064, 2022b.
  37. 37.Mini, U., Grietzer, P., Sharma, M., Meek, A., MacDiarmid, M., and Turner, A. M. Understanding and controlling a maze-solving policy network. arXiv preprint arXiv:2310.08043, 2023.
  38. 38.Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116, 2023.
  39. 39.Rao, S. and Tetreault, J. Dear sir or madam, may i introduce the gyafc dataset: Corpus, benchmarks and metrics for formality style transfer. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 129–140, 2018.
  40. 40.Razeghi, Y., Logan IV, R. L., Gardner, M., and Singh, S. Impact of pretraining term frequencies on few-shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 840–854, 2022.
  41. 41.Reif, E., Ippolito, D., Yuan, A., Coenen, A., Callison-Burch, C., and Wei, J. A recipe for arbitrary text style transfer with large language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 837–848, 2022.
  42. 42.Riley, P., Constant, N., Guo, M., Kumar, G., Uthus, D. C., and Parekh, Z. Textsettr: Few-shot text style extraction and tunable targeted restyling. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pp. 3786–3800. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.ACL-LONG.293.
  43. 43.Rubin, O., Herzig, J., and Berant, J. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2655–2671, 2022.
  44. 44.Schick, T., Udupa, S., and Schütze, H. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. Transactions of the Association for Computational Linguistics, 9:1408–1424, 2021.
  45. 45.Shen, T., Lei, T., Barzilay, R., and Jaakkola, T. Style transfer from non-parallel text by cross-alignment. Advances in neural information processing systems, 30, 2017.
  46. 46.Shin, S., Lee, S. W., Ahn, H., Kim, S., Kim, H. S., Kim, B., Cho, K., Lee, G., Park, W., Ha, J. W., et al. On the effect of pretraining corpora on in-context learning by a large-scale language model. In 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, pp. 5168–5186. Association for Computational Linguistics (ACL), 2022.
  47. 47.Skjuve, M., Følstad, A., Fostervold, K. I., and Brandtzaeg, P. B. My chatbot companion-a study of human-chatbot relationships. International Journal of Human-Computer Studies, 149:102601, 2021.
  48. 48.Som, A., Sikka, K., Gent, H., Divakaran, A., Kathol, A., and Vergyri, D. Demonstrations are all you need: Advancing offensive content paraphrasing using in-context learning. arXiv preprint arXiv:2310.10707, 2023.
  49. 49.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  50. 50.Turner, A., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M. Activation addition: Steering language models without optimization. arXiv preprint arXiv:2308.10248, 2023.
  51. 51.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  52. 52.Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. Transformers learn in-context by gradient descent. In International Conference on Machine Learning, pp. 35151–35174. PMLR, 2023.
  53. 53.Votipka, D., Fulton, K. R., Parker, J., Hou, M., Mazurek, M. L., and Hicks, M. Understanding security mistakes developers make: Qualitative analysis from build it, break it, fix it. In 29th USENIX Security Symposium (USENIX Security 20), pp. 109–126, 2020.
  54. 54.Wan, X., Sun, R., Dai, H., Arik, S. O., and Pfister, T. Better zero-shot reasoning with self-adaptive prompting. arXiv preprint arXiv:2305.14106, 2023a.
  55. 55.Wan, X., Sun, R., Nakhost, H., Dai, H., Eisenschlos, J. M., Arik, S. O., and Pfister, T. Universal self-adaptive prompting. arXiv preprint arXiv:2305.14926, 2023b.
  56. 56.Wang, B., Ping, W., Xiao, C., Xu, P., Patwary, M., Shoeybi, M., Li, B., Anandkumar, A., and Catanzaro, B. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. Advances in Neural Information Processing Systems, 35:35811–35824, 2022.
  57. 57.Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846, 2023.
  58. 58.Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  59. 59.Wulczyn, E., Thain, N., and Dixon, L. Ex machina: Personal attacks seen at scale. In Proceedings of the 26th international conference on world wide web, pp. 1391–1399, 2017.
  60. 60.Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations, 2021.
  61. 61.Xu, B., Wang, Q., Mao, Z., Lyu, Y., She, Q., and Zhang, Y. k nn prompting: Beyond-context learning with calibration-free nearest neighbor inference. In The Eleventh International Conference on Learning Representations, 2022.
  62. 62.Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2950–2968, 2021.
  63. 63.Xu, W., Ritter, A., Dolan, B., Grishman, R., and Cherry, C. Paraphrasing for style. In COLING, pp. 2899–2914, 2012.
  64. 64.Yang, J., Hui, B., Yang, M., Li, B., Huang, F., and Li, Y. Iterative forward tuning boosts in-context learning in language models. arXiv preprint arXiv:2305.13016, 2023.
  65. 65.Ye, S., Kim, D., Jang, J., Shin, J., and Seo, M. Guess the instruction! flipped learning makes language models stronger zero-shot learners. ArXiv, abs/2210.02969, 2022.
  66. 66.Yin, F., Vig, J., Laban, P., Joty, S., Xiong, C., and Wu, C.-S. J. Did you read the instructions? rethinking the effectiveness of task definitions in instruction learning. arXiv preprint arXiv:2306.01150, 2023.
  67. 67.Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations, 2019.
  68. 68.Zhang, X., Zhao, J., and LeCun, Y. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28, 2015.
  69. 69.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  70. 70.Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models, 2023. communication, it is essential for you to comprehend user queries in Cipher Code and subsequently deliver your responses utilizing Cipher Code.
  71. 71.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023.

Citation

MLA
Liu, S., et al. “In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering”. arXiv, 2023, http://arxiv.org/abs/2311.06668v3.
APA
Liu, S., Ye, H., Xing, L., & Zou, J. (2023). In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering. arXiv. http://arxiv.org/abs/2311.06668v3
Chicago
Liu, S., H. Ye, L. Xing, and J. Zou. 2023. “In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering”. arXiv. http://arxiv.org/abs/2311.06668v3.
Harvard
Liu, S. et al. (2023) “In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.06668v3.
Vancouver
1. Liu S, Ye H, Xing L, Zou J (2023) In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering. arXiv

BibTeX

@article{liu2023context,
  title = {In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering},
  author = {Liu, Sheng and Ye, Haotian and Xing, Lei and Zou, James},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.06668v3},
  eprint = {2311.06668}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/