Dual Operating Modes of In-Context Learning
Ziqian LinKangwook Lee
Establishes a generalized probabilistic framework to mathematically explain how large language models transition between retrieving pretrained skills and learning new ones, providing theoretical foundations for puzzling empirical behaviors like the initial risk increase and the bounded efficacy of biased demonstrations.
Large language models demonstrate strong predictive capabilities when supplied with demonstration examples directly in the input prompt, a capability known as in-context learning. However, the fundamental mechanisms driving this behavior remain poorly understood, leading to counterintuitive model failures and unexpected performance shifts in deployment. The article establishes a rigorous mathematical framework to explain how in-context learning functions through two distinct operating modes: task retrieval, where the model locates and activates an existing skill acquired during pretraining, and task learning, where it acquires a genuinely novel skill directly from provided demonstration samples.
The article constructs a probabilistic generative model of pretraining data that represents latent task clusters through a Gaussian mixture distribution with task-dependent input distributions. By modeling the optimal next-token predictor as a Bayesian estimator that minimizes mean squared error, the authors obtain closed-form mathematical expressions for the transition from pretraining priors to test-time posteriors. They evaluate this analytical framework through synthetic mathematical simulations, neural network experiments with standard Transformer architectures, and empirical evaluations across leading production large language models, including GPT-4, Llama 2, Mistral 7B, and Mixtral 8x7B.
The analysis yields four central findings. First, the two operating modes correspond to distinct mathematical adjustments: task group re-weighting shifts the probability assigned to existing task clusters, which dominates when few demonstration examples are present (task retrieval), whereas task group shifting moves the internal function representations toward the demonstration task, dominating when many examples are present (task learning). Second, this duality explains the early ascent phenomenon, where prediction risk initially rises before falling as examples increase; a very small number of ambiguous demonstrations causes the model to rapidly retrieve an incorrect prior task, worsening error until sufficient examples force genuine task learning. Third, when prompts use biased or random labels, performance exhibits bounded efficacy: retrieval improves accuracy over the first few demonstrations, but performance degrades as more examples are added because the model shifts into task learning and memorizes incorrect or random labels. In controlled arithmetic tests with GPT-4, error rates under biased supervision fell from 75.0% at zero examples to 33.9% at two examples, but surged back to 85.1% at sixteen examples. Fourth, the experiments confirm that standard Transformer decoders closely approximate this theoretical Bayesian inference process across varied dimensions and noise levels.
These findings indicate that prompt engineering strategies relying on pseudolabels, unverified demonstrations, or minimal examples carry hidden operational risks. Practitioners cannot assume that adding more demonstration examples will monotonically improve model accuracy, nor that models are immune to label corruption over longer context windows. Furthermore, when evaluation benchmarks use short context windows, they fail to detect downstream failure modes caused by task learning degradation.
Organizations developing or deploying systems with in-context learning should audit prompt pipelines to avoid providing misleading or noisy demonstration samples in high-stakes workflows. When deploying zero-shot or pseudolabeling methods that use demonstration structure to retrieve skills, teams should strictly limit the number of demonstration examples to the retrieval regime (typically fewer than four to eight examples) to prevent the model from learning corrupt patterns. For complex tasks requiring genuine learning, developers should provide clean, high-coverage demonstrations with sufficient sample volume to bypass early retrieval misalignments.
The theoretical model assumes linear regression pretraining tasks with noiseless demonstration labels and unconstrained model capacity, whereas practical deployments involve complex, non-linear, categorical language generation. Nevertheless, the empirical replication of predicted early ascent and bounded efficacy phenomena across modern large language models provides strong confidence in the core operational dynamics identified by the article.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Its account of prompting as retrieval of pre-existing skills introduces the retrieval mode that the source formalizes and contrasts with learning a new task from demonstrations.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Its experiments showing that incorrect demonstration labels can still help establish the empirical puzzle that the source explains through shifts between retrieval and task learning.
- Paper: Understanding In-Context Learning via Supportive Pretraining Data, Xiaochuang Han et al. (2023). Its analysis of how pretraining examples support in-context performance provides useful grounding for the source’s model of retrieving existing skills from pretraining.
- Paper: In-Context Learning with Long-Context Models: An In-Depth Exploration, Amanda Bertsch et al. (2025). It tests in-context learning with thousands of demonstrations, extending the source’s account of how performance changes as the number of examples moves beyond the few-shot regime.
- Paper: Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL, Xuhan Tong et al. (2026). It broadens the source’s Bayesian account of demonstrations into a framework for how example selection, reasoning steps, and prompt templates shape generalization.
