Dual Operating Modes of In-Context Learning

Ziqian LinKangwook Lee

article2024ICML57 citations

Establishes a generalized probabilistic framework to mathematically explain how large language models transition between retrieving pretrained skills and learning new ones, providing theoretical foundations for puzzling empirical behaviors like the initial risk increase and the bounded efficacy of biased demonstrations.

Listen

Large language models demonstrate strong predictive capabilities when supplied with demonstration examples directly in the input prompt, a capability known as in-context learning. However, the fundamental mechanisms driving this behavior remain poorly understood, leading to counterintuitive model failures and unexpected performance shifts in deployment. The article establishes a rigorous mathematical framework to explain how in-context learning functions through two distinct operating modes: task retrieval, where the model locates and activates an existing skill acquired during pretraining, and task learning, where it acquires a genuinely novel skill directly from provided demonstration samples.

The article constructs a probabilistic generative model of pretraining data that represents latent task clusters through a Gaussian mixture distribution with task-dependent input distributions. By modeling the optimal next-token predictor as a Bayesian estimator that minimizes mean squared error, the authors obtain closed-form mathematical expressions for the transition from pretraining priors to test-time posteriors. They evaluate this analytical framework through synthetic mathematical simulations, neural network experiments with standard Transformer architectures, and empirical evaluations across leading production large language models, including GPT-4, Llama 2, Mistral 7B, and Mixtral 8x7B.

The analysis yields four central findings. First, the two operating modes correspond to distinct mathematical adjustments: task group re-weighting shifts the probability assigned to existing task clusters, which dominates when few demonstration examples are present (task retrieval), whereas task group shifting moves the internal function representations toward the demonstration task, dominating when many examples are present (task learning). Second, this duality explains the early ascent phenomenon, where prediction risk initially rises before falling as examples increase; a very small number of ambiguous demonstrations causes the model to rapidly retrieve an incorrect prior task, worsening error until sufficient examples force genuine task learning. Third, when prompts use biased or random labels, performance exhibits bounded efficacy: retrieval improves accuracy over the first few demonstrations, but performance degrades as more examples are added because the model shifts into task learning and memorizes incorrect or random labels. In controlled arithmetic tests with GPT-4, error rates under biased supervision fell from 75.0% at zero examples to 33.9% at two examples, but surged back to 85.1% at sixteen examples. Fourth, the experiments confirm that standard Transformer decoders closely approximate this theoretical Bayesian inference process across varied dimensions and noise levels.

These findings indicate that prompt engineering strategies relying on pseudolabels, unverified demonstrations, or minimal examples carry hidden operational risks. Practitioners cannot assume that adding more demonstration examples will monotonically improve model accuracy, nor that models are immune to label corruption over longer context windows. Furthermore, when evaluation benchmarks use short context windows, they fail to detect downstream failure modes caused by task learning degradation.

Organizations developing or deploying systems with in-context learning should audit prompt pipelines to avoid providing misleading or noisy demonstration samples in high-stakes workflows. When deploying zero-shot or pseudolabeling methods that use demonstration structure to retrieve skills, teams should strictly limit the number of demonstration examples to the retrieval regime (typically fewer than four to eight examples) to prevent the model from learning corrupt patterns. For complex tasks requiring genuine learning, developers should provide clean, high-coverage demonstrations with sufficient sample volume to bypass early retrieval misalignments.

The theoretical model assumes linear regression pretraining tasks with noiseless demonstration labels and unconstrained model capacity, whereas practical deployments involve complex, non-linear, categorical language generation. Nevertheless, the empirical replication of predicted early ascent and bounded efficacy phenomena across modern large language models provides strong confidence in the core operational dynamics identified by the article.

Cover for Dual Operating Modes of In-Context Learning

Abstract

In-context learning (ICL) exhibits dual operating modes: task learning, i.e. acquiring a new skill from in-context samples, and task retrieval, i.e., locating and activating a relevant pretrained skill. Recent theoretical work proposes various mathematical models to analyze ICL, but they cannot fully explain the duality. In this work, we analyze a generalized probabilistic model for pretraining data, obtaining a quantitative understanding of the two operating modes of ICL. Leveraging our analysis, we provide the first explanation of an unexplained phenomenon observed with real-world large language models (LLMs). Under some settings, the ICL risk initially increases and then decreases with more in-context examples. Our analysis offers a plausible explanation for this “early ascent” phenomenon: a limited number of in-context samples may lead to the retrieval of an incorrect skill, thereby increasing the risk, which will eventually diminish as task learning takes effect with more in-context samples. We also analyze ICL with biased labels, e.g., zero-shot ICL, where in-context examples are assigned random labels, and predict the bounded efficacy of such approaches. We corroborate our analysis and predictions with extensive experiments with Transformers and LLMs. The code is available at: https://github.com/UW-Madison-Lee-Lab/Dual_Operating_Modes_of_ICL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Pretraining and Data Generative Model
  • 3.1. Data Generative Model
  • 3.2. Bayes-Optimal Next-Token Predictor
  • 3.3. Gaussian/Linear Assumptions on Pretraining Data Generative Model
  • 4. Inference and Dual Operating Modes
  • 4.1. In-Context Task and In-Context Function
  • 4.2. Closed-Form Expression of Posterior
  • 4.3. Closed-Form Expression of ICL Prediction
  • 4.4. Prior Task Noises, CS, CR, and ICL Prediction
  • 4.5. Dual Operating Modes
  • 5. Early Ascent
  • 5.1. Fine-Grained Upper Bound
  • 5.2. The Effect of Dual Operating Modes on ICL Risk
  • 5.3. Early Ascent with Biased x Distribution
  • 6. Bounded Efficacy of Biased-Label ICL
  • 6.1. Upper Bound for ICL Risk with Biased Labels
  • 6.2. Bounded Efficacy of Biased-Label ICL in GPT-4
  • 6.3. Bounded Efficacy for Zero-Shot ICL
  • 7. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Notations
  • B. Prior Examples
  • B.1. Regular Polyhedrons
  • B.2. d-Dimensional Examples
  • B.3. Early Ascent Examples
  • C. Coarse Upper Bound for ICL Risk
  • D. Transformer Performance in Approximating Bayesian Inference
  • E. Additional Information for Bounded Efficacy in GPT-4
  • E.1. Experimental Setting
  • E.2. Additional Results
  • F. Bounded Efficacy in Zero-Shot ICL
  • G. The Derivation of Posterior
  • G.1. Prior to Posterior
  • G.2. Closed-Form Solution from Eq. 2 to Eq. 3
  • H. Detailed Analysis of Component Shifting and Re-Weighting
  • H.1. Analysis of Component Re-Weighting
  • H.3. Derivation Collection of Ψ_μ(α, β) and Ψ_w(α, β)
  • H.3.1. Derivation of Ψ_μ(α, β)
  • H.3.2. Derivation of Ψ_w(α, β)
  • H.4. Derivation Collection of μ̃_m and w̃_m
  • H.4.1. Derivation of μ̃_m
  • H.4.2. Derivation of w̃_m
  • I. Additional Experiments for Early Ascent
  • I.1. Early Ascent and Bounded Efficacy under Noisy Labels
  • I.2. Early Ascent under Non-Linear Regression and Discrete Token Prediction
  • I.2.1. Experiment Design for Non-Linear Regression
  • I.2.2. Experiment Design for Discrete Token Prediction
  • J. Mathematical Derivation for Early Ascent
  • J.1. A Specific Setting of Early Ascent
  • J.2. Theorem of Early Ascent
  • K. Proof Tools
  • K.1. Gaussian Tail Bound
  • K.2. Chi-Squared Tail Bound
  • K.3. Norm Tail Bound
  • K.4. Eigenvalue Concentration Bound
  • L. ICL to Learn the In-Context Function
  • L.1. Case When In-Context Input Variable Spans in Subspace
  • M. ICL with Biased Labels to Retrieve a Task
  • M.1. Non-Asymptotic Bound for the Bounded Efficacy Phenomenon
  • M.1.1. Bounded Efficacy - Part A
  • M.1.2. Bounded Efficacy - Part B
  • M.1.3. Bounded Efficacy - Part C
  • M.1.4. Bounded Efficacy - Summary
  • M.1.5. The Particular Interval
  • M.2. Asymptotic Bound
  • M.2.1. Asymptotic Bound - Part A′
  • M.2.2. Asymptotic Bound - Part B′
  • M.2.3. Asymptotic Bound - Part C′
  • M.2.4. Asymptotic Bound - Summary
  • N. Proof of Lemma 6.2
  • N.1. Proof of Lemma 6.2: Part A″
  • N.2. Proof of Lemma 6.2: Part B″
  • N.3. Proof of Lemma 6.2: Part C″
  • N.4. Proof of Lemma 6.2: Summary
  • N.5. The Particular Interval
  • O. Toy Example for Component Shifting and Component Re-Weighting
  • O.1. Toy Example: Pretraining Data Generative Model
  • O.2. Toy Example: Bayes-Optimal Next-Token Predictor
  • O.3. Toy Example: Gaussian Assumptions on Pretraining Data Generative Model
  • O.4. Toy Example: Posterior Analysis
  • O.5. Proof of Posterior Derivation in Toy Example

Knowls

  1. Knowl 1 — Clustered Gaussian-mixture model of pretraining tasks

    model/method

    The paper models pretraining sequences as samples from latent tasks grouped into MM Gaussian-mixture components. For a component mm, its center (μm,wm)∈Rd×Rd(\mu_m,w_m)\in\mathbb R^d\times\mathbb R^d has mixture probability πm\pi_m; task parameters satisfy μ∼N(μm,σμ2I)\mu\sim\mathcal N(\mu_m,\sigma_\mu^2 I) and w∼N(wm,σw2I)w\sim\mathcal N(w_m,\sigma_w^2 I). Given (μ,w)(\mu,w), inputs and labels are generated as x∼N(μ,σx2I)x\sim\mathcal N(\mu,\sigma_x^2 I) and y∣x,w∼N(w⊤x,σy2)y\mid x,w\sim\mathcal N(w^\top x,\sigma_y^2). Thus each task has its own input distribution as well as its own linear function. The components share these covariance and noise scales, their centers have unit norm, and their mixture weights are comparable: for some r>1r>1, 1/r≤πα/πβ≤r1/r\leq\pi_\alpha/\pi_\beta\leq r for all components α,β\alpha,\beta. A pretraining sequence first samples a task and then draws KK i.i.d. labeled pairs from it. At inference, kk labeled examples and a new unlabeled input are presented; the target in-context task has xi∼N(μ∗,τx2I)x_i\sim\mathcal N(\mu^*,\tau_x^2 I), noiseless labels yi=xi⊤w∗y_i=x_i^\top w^*, and ∥μ∗∥=∥w∗∥=1\|\mu^*\|=\|w^*\|=1.

  2. Knowl 2 — Closed-form Gaussian-mixture posterior

    equation

    Under the clustered linear-Gaussian pretraining model, observing kk labeled examples (xi,yi)i=1k(x_i,y_i)_{i=1}^k and the next unlabeled input xk+1x_{k+1} leaves the task posterior as an MM-component Gaussian mixture. For k≥1k\geq1, define δμ=σμ2/σx2\delta_\mu=\sigma_\mu^2/\sigma_x^2, δw=σw2/σy2\delta_w=\sigma_w^2/\sigma_y^2, μˉ=(k+1)−1∑i=1k+1xi\bar\mu=(k+1)^{-1}\sum_{i=1}^{k+1}x_i, S=k−1∑i=1kxixi⊤S=k^{-1}\sum_{i=1}^k x_ix_i^\top, and wˉ=k−1∑i=1kxiyi\bar w=k^{-1}\sum_{i=1}^k x_i y_i. Let Bμ=I+(k+1)δμIB_\mu=I+(k+1)\delta_\mu I and Bw=I+kδwSB_w=I+k\delta_w S. With ∥v∥B2=v⊤Bv\|v\|_B^2=v^\top Bv, component mm has posterior means and covariances

    μ~m=Bμ−1(μm+(k+1)δμμˉ),w~m=Bw−1(wm+kδwwˉ),\tilde\mu_m=B_\mu^{-1}(\mu_m+(k+1)\delta_\mu\bar\mu),\qquad \tilde w_m=B_w^{-1}(w_m+k\delta_w\bar w), Σ~μ=σμ2Bμ−1,Σ~w=σw2Bw−1.\tilde\Sigma_{\mu}=\sigma_\mu^2B_\mu^{-1},\qquad \tilde\Sigma_w=\sigma_w^2B_w^{-1}.

    Its unnormalized posterior mixture weight is πmcmμcmw\pi_m c_m^\mu c_m^w, where

    cmμ=exp⁡ ⁣[−∥μm∥2−∥μm+(k+1)δμμˉ∥Bμ−122σμ2],cmw=exp⁡ ⁣[−∥wm∥2−∥wm+kδwwˉ∥Bw−122σw2].c_m^\mu=\exp\!\left[-\frac{\|\mu_m\|^2-\|\mu_m+(k+1)\delta_\mu\bar\mu\|_{B_\mu^{-1}}^2}{2\sigma_\mu^2}\right],\qquad c_m^w=\exp\!\left[-\frac{\|w_m\|^2-\|w_m+k\delta_w\bar w\|_{B_w^{-1}}^2}{2\sigma_w^2}\right].

    Normalizing these weights gives π~m=πmcmμcmw/∑j=1Mπjcjμcjw\tilde\pi_m=\pi_m c_m^\mu c_m^w/\sum_{j=1}^M\pi_jc_j^\mu c_j^w. The posterior therefore updates both the locations of task-group centers and their probabilities; the unlabeled test input contributes information about the input-distribution parameter μ\mu.

  3. Knowl 3 — Bayes-optimal in-context prediction

    theoretical result

    For squared-loss pretraining on the task mixture, a sufficiently expressive next-token predictor trained to the population optimum predicts the conditional mean label. Given labeled in-context pairs (xi,yi)i=1k(x_i,y_i)_{i=1}^k and a test input xk+1x_{k+1}, its prediction is

    F∗(Sk⊕xk+1)=xk+1⊤w~,w~=∑m=1Mπ~mw~m.F^*(S_k\oplus x_{k+1})=x_{k+1}^\top\tilde w, \qquad \tilde w=\sum_{m=1}^M\tilde\pi_m\tilde w_m.

    Here Sk=[x1,y1,…,xk,yk]S_k=[x_1,y_1,\ldots,x_k,y_k], and π~m,w~m\tilde\pi_m,\tilde w_m are the posterior mixture weights and component means under the clustered Gaussian-mixture model. Thus the Bayes-optimal prediction is a posterior-weighted combination of the shifted component functions, rather than necessarily the prediction of a single task.

  4. Knowl 4 — Task retrieval and task learning as posterior updates

    definition

    The paper distinguishes two effects of in-context examples on the Gaussian-mixture task posterior. Component re-weighting changes the posterior probabilities π~m\tilde\pi_m of pretrained task groups; component shifting changes their posterior function centers w~m\tilde w_m (and input-distribution centers μ~m\tilde\mu_m). Task retrieval is the regime in which re-weighting is the more influential effect, so a few examples cause predictions to favor a pretrained skill. Task learning is the regime in which component shifting dominates: with many examples, the component function centers approach the in-context function w∗w^* and predictions depend less on the pretrained centers. The paper’s analysis and numerical examples show retrieval can dominate early when the pretraining task variances δμ,δw\delta_\mu,\delta_w are small, while the increasing evidence in many examples eventually makes shifting, and hence learning, effective.

  5. Knowl 5 — Fine-grained upper bound on in-context learning risk

    theoretical result

    Suppose in-context inputs are i.i.d. xi∼N(μ∗,τx2I)x_i\sim\mathcal N(\mu^*,\tau_x^2 I), labels are noiseless yi=xi⊤w∗y_i=x_i^\top w^*, and ∥μ∗∥=∥w∗∥=1\|\mu^*\|=\|w^*\|=1. For a population-optimal pretrained predictor, define the squared prediction risk Lk∗=(F∗(Sk⊕xk+1)−xk+1⊤w∗)2L_k^*=(F^*(S_k\oplus x_{k+1})-x_{k+1}^\top w^*)^2 and the positive-definite matrix A=(I+δw∑i=1kxixi⊤)−1A=(I+\delta_w\sum_{i=1}^k x_ix_i^\top)^{-1}. If wmw_m is the function center of prior component mm and λ1(A)\lambda_1(A) is the largest eigenvalue of AA, then the paper gives the bound

    E[Lk∗]<∑m=1M∥wm−w∗∥2 E ⁣[π~m∥xk+1∥2λ1(A)2].\mathbb E[L_k^*] <\sum_{m=1}^M\|w_m-w^*\|^2\,\mathbb E\!\left[\tilde\pi_m\|x_{k+1}\|^2\lambda_1(A)^2\right].

    The bound separates the contribution of each pretrained function’s distance from the target from its posterior probability and the amount of component shifting encoded by AA. Consequently, both re-weighting and shifting can affect risk; ignoring re-weighting yields a coarser bound. The result assumes the Gaussian in-context input model and noiseless in-context labels.

  6. Knowl 6 — Early ascent from retrieving a misleading task

    theoretical result

    For the same Gaussian in-context input model, define the most misleading prior component by

    α∈arg⁡min⁡m{∥μm−μ∗∥22σx2+∣(wm−w∗)⊤μ∗∣2+dτx2∥wm−w∗∥22σy2}.\alpha\in\arg\min_m\left\{ \frac{\|\mu_m-\mu^*\|^2}{2\sigma_x^2} +\frac{|(w_m-w^*)^\top\mu^*|^2+d\tau_x^2\|w_m-w^*\|^2}{2\sigma_y^2} \right\}.

    If the zero-example prediction has lower expected squared error on a fresh input x1x_1 than predicting with component wαw_\alpha, that is,

    E[(F∗(x1)−x1⊤w∗)2]<E[(x1⊤(wα−w∗))2],\mathbb E\big[(F^*(x_1)-x_1^\top w^*)^2\big] <\mathbb E\big[(x_1^\top(w_\alpha-w^*))^2\big],

    then for sufficiently small δμ\delta_\mu and δw\delta_w there is at least one k≥1k\geq1 for which the expected ICL risk exceeds the zero-example risk. This is an early ascent in risk: a small number of examples can re-weight the posterior toward a misleading pretrained function before component shifting has learned the target. In the paper’s numerical construction, the in-context input distribution is closer to a misleading component than the component whose function is the target; retrieval initially increases risk, and risk later falls as the predicted function moves toward the target. The plotted examples show retrieval dominating through about k=32k=32 examples.

  7. Knowl 7 — Biased-label and zero-shot ICL have bounded efficacy

    theoretical result

    The paper analyzes ICL whose labels encode a biased function w∗w^* even though the desired retrieved pretrained function is wα≠w∗w_\alpha\ne w^*. Assume component α\alpha is separated from every other prior component β\beta by positive margins: ∥μβ−μ∗∥2−∥μα−μ∗∥2≥dμ2\|\mu_\beta-\mu^*\|^2-\|\mu_\alpha-\mu^*\|^2\geq d_\mu^2, ∥wβ−w∗∥2−∥wα−w∗∥2≥dw2\|w_\beta-w^*\|^2-\|w_\alpha-w^*\|^2\geq d_w^2, and τx2∥wβ−w∗∥2−(1+τx2)∥wα−w∗∥2≥τx2uw2\tau_x^2\|w_\beta-w^*\|^2-(1+\tau_x^2)\|w_\alpha-w^*\|^2\geq\tau_x^2u_w^2. For sufficiently small δμ,δw\delta_\mu,\delta_w, on a specified nonempty range of example counts, the risk relative to retrieving wαw_\alpha, Lkα=(F∗(Sk⊕xk+1)−xk+1⊤wα)2L_k^\alpha=(F^*(S_k\oplus x_{k+1})-x_{k+1}^\top w_\alpha)^2, obeys

    E[Lkα]<C3exp⁡ ⁣[−k(dμ28σx2+uw2τx28σy2)]+48(1+dτx2)e−k/8+∥wα−w∗∥2(1+dτx2)min⁡ ⁣{1,4k2δw2(1+τx2)2},\mathbb E[L_k^\alpha] <C_3\exp\!\left[-k\left(\frac{d_\mu^2}{8\sigma_x^2}+\frac{u_w^2\tau_x^2}{8\sigma_y^2}\right)\right] +48(1+d\tau_x^2)e^{-\sqrt{k}/8} +\|w_\alpha-w^*\|^2(1+d\tau_x^2)\min\!\left\{1,4k^2\delta_w^2(1+\tau_x^2)^2\right\},

    where C3C_3 depends on the prior and in-context task. The first terms decay, whereas the last term grows and eventually limits efficacy: retrieval can help at first, but further examples cause the model to fit the biased labels. The paper also gives a zero-shot special case with two symmetric prior centers (μα,wα)=(−μβ,−wβ)(\mu_\alpha,w_\alpha)=(-\mu_\beta,-w_\beta) and uninformative labels yi=0y_i=0; its corresponding bound has decaying terms proportional to e−dμ2k/(8σx2)e^{-d_\mu^2k/(8\sigma_x^2)} and e−k/8e^{-\sqrt{k}/8}, plus an increasing term proportional to (1+dτx2)min⁡{1,k2δw2(1+τx2)2}(1+d\tau_x^2)\min\{1,k^2\delta_w^2(1+\tau_x^2)^2\}. This special case predicts bounded efficacy for zero-shot ICL under those idealized conditions.

  8. Knowl 8 — LLM experiments exhibit bounded efficacy with biased labels

    empirical result

    The paper reports bounded efficacy in GPT-4 on a synthetic arithmetic prompt. In-context examples use the biased rule ci=ai+bi+1c_i=a_i+b_i+1; the comparison skill is ordinary addition, a+ba+b. The reported error rates, in percent, for k=0,1,2,4,8,16k=0,1,2,4,8,16 examples are: ordinary “+” evaluation: 75.0,36.2,33.9,49.3,79.3,85.175.0, 36.2, 33.9, 49.3, 79.3, 85.1; biased “+” evaluation: 100.0,98.3,95.9,60.5,24.4,16.8100.0, 98.3, 95.9, 60.5, 24.4, 16.8. Thus, a few biased demonstrations first improve retrieval of ordinary addition, but as the prompt grows the addition error rises while the model increasingly learns the biased rule. The paper also evaluates classification with true versus random demonstration labels on Mistral 7B, Mixtral 8×7B, Llama 2 13B, Llama 2 70B, and GPT-4, using GLUE-MRPC, GLUE-RTE, TweetEval-Hate, SICK, and poem-sentiment. The random-label error curves turn upward at smaller example counts than the true-label curves; the true-label comparison indicates the effect is not simply worse performance on long sequences.

  9. Knowl 9 — Transformers approximate the model’s Bayesian predictor

    empirical result

    To test whether a trained Transformer can realize the Bayesian predictor implied by the clustered pretraining model, the paper trained a GPT-2-style decoder with 10 layers, 8 attention heads, and 1024-dimensional feedforward layers. Training used three epochs of 10,000 batches each, with 256 sequences per batch, AdamW weight decay 10−510^{-5}, and learning rate 10−510^{-5}. Experiments varied prior task noise, mixture-component count, and feature dimension. Across these settings, the Transformer’s prediction error approached that of numerical Bayesian inference. Approximation was harder with more feature dimensions and more mixture components. These experiments support the tractability of the posterior analysis as an account of behavior in a trained Transformer, rather than establishing that every LLM implements the same computation.

  10. Knowl 10 — Scope limitations of the analysis

    limitation

    The principal theoretical model assumes pretraining tasks are noisy linear regressions with Gaussian inputs and a Gaussian-mixture task prior, whereas real LLM pretraining includes complex nonlinear and categorical tasks. The main in-context analysis also assumes noiseless in-context labels. The paper explores selected nonlinear, discrete-token, and noisy-label examples experimentally, but these do not remove the stated gap between the tractable theory and general real-world pretraining or prompting.

Coverage note — Detailed posterior derivations, auxiliary numerical grids, and the scalar toy illustration were omitted because they support or visualize the main model, posterior, and dual-mode results rather than constituting separate contributions.

References

  1. 1.Ahn, K., Cheng, X., Daneshmand, H., and Sra, S. Transformers learn to implement preconditioned gradient descent for in-context learning. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  2. 2.Akyurek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. What learning algorithm is in-context learning? Investigations with linear models. In International Conference on Learning Representations (ICLR), 2023.
  3. 3.Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  4. 4.Barbieri, F., Camacho-Collados, J., Anke, L. E., and Neves, L. TweetEval: Unified benchmark and comparative evaluation for tweet classification. In Findings of the Association for Computational Linguistics: EMNLP, 2020.
  5. 5.Boucheron, S., Lugosi, G., and Massart, P. Concentration inequalities: A nonasymptotic theory of independence. Oxford University Press, 2013.
  6. 6.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  7. 7.Dagan, I., Glickman, O., and Magnini, B. The PASCAL recognising textual entailment challenge. In PASCAL Machine Learning Challenges Workshop (MLCW), 2005.
  8. 8.Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F. Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers. In Findings of the Association for Computational Linguistics (ACL), 2023.
  9. 9.Dolan, W. B. and Brockett, C. Automatically constructing a corpus of sentential paraphrases. In International Workshop on Paraphrasing (IWP@IJCNLP), 2005.
  10. 10.Garg, S., Tsipras, D., Liang, P. S., and Valiant, G. What can Transformers learn in-context? A case study of simple function classes. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  11. 11.Ghahramani, Z. and Jordan, M. Factorial hidden markov models. In Advances in Neural Information Processing Systems (NeurIPS), 1995.
  12. 12.Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D. Looped Transformers as programmable computers. In International Conference on Machine Learning (ICML), 2023.
  13. 13.Han, C., Wang, Z., Zhao, H., and Ji, H. In-context learning of large language models explained as kernel regression. arXiv preprint arXiv:2305.12766, 2023.
  14. 14.Jeon, H. J., Lee, J. D., Lei, Q., and Van Roy, B. An information-theoretic analysis of in-context learning. arXiv preprint arXiv:2401.15530, 2024.
  15. 15.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7B. arXiv preprint arXiv:2310.06825, 2023.
  16. 16.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  17. 17.Li, Y., Ildiz, M. E., Papailiopoulos, D., and Oymak, S. Transformers as algorithms: Generalization and stability in in-context learning. In International Conference on Machine Learning (ICML), 2023.
  18. 18.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019.
  19. 19.Lyu, X., Min, S., Beltagy, I., Zettlemoyer, L., and Hajishirzi, H. Z-ICL: Zero-shot in-context learning with pseudo-demonstrations. In Annual Meeting of the Association for Computational Linguistics (ACL), 2023.
  20. 20.Mahankali, A., Hashimoto, T. B., and Ma, T. One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention. In International Conference on Learning Representations (ICLR), 2024.
  21. 21.Marelli, M., Menini, S., Baroni, M., Bentivogli, L., Bernardi, R., and Zamparelli, R. A SICK cure for the evaluation of compositional distributional semantic models. In International Conference on Language Resources and Evaluation (LREC), 2014.
  22. 22.Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Rethinking the role of demonstrations: What makes in-context learning work? In Empirical Methods in Natural Language Processing (EMNLP), 2022.
  23. 23.OpenAI. GPT-4 technical report, 2023.
  24. 24.Pan, J., Gao, T., Chen, H., and Chen, D. What in-context learning “learns” in-context: Disentangling task recognition and task learning. In Findings of the Association for Computational Linguistics (ACL), 2023.
  25. 25.Rabiner, L. R. A tutorial on hidden markov models and selected applications in speech recognition. Proceedings of the IEEE, 1989.
  26. 26.Raventos, A., Paul, M., Chen, F., and Ganguli, S. The effects of pretraining task diversity on in-context learning of ridge regression. In ICLR Workshop on Mathematical and Empirical Understanding of Foundation Models (MEFoMo), 2023.
  27. 27.Razeghi, Y., IV, R. L. L., Gardner, M., and Singh, S. Impact of pretraining term frequencies on few-shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP, 2022.
  28. 28.Sheng, E. and Uthus, D. Investigating societal biases in a poetry composition system. In Workshop on Gender Bias in Natural Language Processing, 2020.
  29. 29.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  30. 30.Tsigler, A. and Bartlett, P. L. Benign overfitting in ridge regression. Journal of Machine Learning Research (JMLR), 2023.
  31. 31.Van Trees, H. L. Detection, estimation, and modulation theory, Part I: Detection, estimation, and linear modulation theory. John Wiley & Sons, 2004.
  32. 32.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), 2017.
  33. 33.Vershynin, R. High-dimensional probability: An introduction with applications in data science, volume 47. Cambridge university press, 2018.
  34. 34.von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. Transformers learn in-context by gradient descent. In International Conference on Machine Learning (ICML), 2023.
  35. 35.Wu, J., Zou, D., Chen, Z., Braverman, V., Gu, Q., and Bartlett, P. L. How many pretraining tasks are needed for in-context learning of linear regression? In International Conference on Learning Representations (ICLR), 2024.
  36. 36.Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. An explanation of in-context learning as implicit Bayesian inference. In International Conference on Learning Representations (ICLR), 2022.
  37. 37.Zhang, R., Frei, S., and Bartlett, P. L. Trained transformers learn linear models in-context. In Robustness of Few-shot and Zero-shot Learning in Large Foundation Models (R0-FoMo), 2023.

Citation

MLA
Lin, Z., and K. Lee. “Dual Operating Modes of In-Context Learning”. arXiv, 2024, http://arxiv.org/abs/2402.18819v2.
APA
Lin, Z., & Lee, K. (2024). Dual Operating Modes of In-Context Learning. arXiv. http://arxiv.org/abs/2402.18819v2
Chicago
Lin, Z., and K. Lee. 2024. “Dual Operating Modes of In-Context Learning”. arXiv. http://arxiv.org/abs/2402.18819v2.
Harvard
Lin, Z. and Lee, K. (2024) “Dual Operating Modes of In-Context Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.18819v2.
Vancouver
1. Lin Z, Lee K (2024) Dual Operating Modes of In-Context Learning. arXiv

BibTeX

@article{lin2024dual,
  title = {Dual Operating Modes of In-Context Learning},
  author = {Lin, Ziqian and Lee, Kangwook},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.18819v2},
  eprint = {2402.18819}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/