Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL

Xuhan TongYuchen ZengJiawei Zhang

article2026arXiv0 citations

Establishes theoretical generalization bounds for in-context learning under mild assumptions, explaining how demonstration selection, Chain-of-Thought task decomposition, and prompt templates directly govern model performance on unseen tasks.

Listen

In-context learning allows large language models to adapt to downstream tasks using only input-output examples without updating internal model weights. Despite its widespread practical use, existing theoretical explanations often rely on unrealistic simplifications, such as linear attention or single-layer networks, and fail to explain how practical prompt design choices affect performance. The article addresses this gap by developing a unified theoretical framework under realistic, mild assumptions to evaluate how demonstration selection, Chain-of-Thought reasoning, demonstration counts, and prompt templates govern generalization on unseen tasks.

The authors analyze in-context learning mathematically by viewing adaptation as a path in representation space that connects test prompts to pretraining data, and they model multi-step Chain-of-Thought prompting as sequential task decomposition. They also frame demonstration-based prompting as Bayesian posterior inference over latent tasks. To validate these theoretical bounds, the authors conduct synthetic experiments using fine-tuned variants of Llama-3.2-3B and large pretrained Qwen models across multi-digit arithmetic and entity-retrieval benchmarks.

The article establishes several key findings. First, generalization error on unseen prompts is upper-bounded by three primary factors: the model's intrinsic baseline capability, the degree of task distribution shift, and an effective rate-of-change metric reflecting how stably the examples define the task. Informative examples that clearly identify a rule achieve substantially higher accuracy than ambiguous ones (for instance, yielding 56–60% accuracy on sports identification compared to 16–20% for ambiguous prompts). Second, Chain-of-Thought prompting succeeds only when a complex problem is broken into sub-tasks that align with operations the model already mastered during pretraining; decomposing a task into unfamiliar steps degrades performance below standard prompting without intermediate reasoning. Third, the model's sensitivity to prompt templates decays exponentially as more demonstrations are added, meaning that formatting differences and even consistently incorrect instructions become irrelevant once sufficient examples are provided. However, this stability completely breaks down when inconsistent, conflicting instructions are mixed across examples in the prompt.

These findings indicate that in-context learning functions primarily as a task retrieval and composition mechanism over capabilities acquired during pretraining. For organizations deploying language models, this means that investing effort into prompt formatting yields diminishing returns if many examples are present, whereas curating unambiguous demonstrations and aligning multi-step reasoning with familiar pretraining sub-tasks directly drives reliability and performance.

Practitioners should prioritize selecting clear, unambiguous demonstration pairs and ensure that Chain-of-Thought templates decompose workflows into simple sub-problems the model can reliably solve in isolation. Furthermore, engineering efforts should avoid mixing contradictory instructions across prompt examples. Because the empirical validation relies on synthetic arithmetic and structured identification benchmarks on selected open-weight models, future work should validate these theoretical boundaries across broader, more diverse real-world domain tasks before standardizing production pipelines.

No sufficiently relevant recommendations were found.

Cover for Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL

Abstract

In-Context Learning (ICL) enables pretrained LLMs to adapt to downstream tasks by conditioning on a small set of input-output demonstrations, without any parameter updates. Although there have been many theoretical efforts to explain how ICL works, most either rely on strong architectural or data assumptions, or fail to capture the impact of key practical factors such as demonstration selection, Chain-of-Thought (CoT) prompting, the number of demonstrations, and prompt templates. We address this gap by establishing a theoretical analysis of ICL under mild assumptions that links these design choices to generalization behavior. We derive an upper bound on the ICL test loss, showing that performance is governed by (i) the quality of selected demonstrations, quantified by Lipschitz constants of the ICL loss along paths connecting test prompts to pretraining samples, (ii) an intrinsic ICL capability of the pretrained model, and (iii) the degree of distribution shift. Within the same framework, we analyze CoT prompting as inducing a task decomposition and show that it is beneficial when demonstrations are well chosen at each substep and the resulting subtasks are easier to learn. Finally, we characterize how ICL performance sensitivity to prompt templates varies with the number of demonstrations. Together, our study shows that pretraining equips the model with the ability to generalize beyond observed tasks, while CoT enables the model to compose simpler subtasks into more complex ones, and demonstrations and instructions enable it to retrieve similar or complex tasks, including those that can be composed into more complex ones, jointly supporting generalization to unseen tasks. All theoretical insights are corroborated by experiments.

Table of Contents

  • 1 Introduction
  • 2 Lipschitz Characterization of Demonstration Effectiveness in ICL
  • 2.1 Background
  • 2.2 Lipschitz Generalization Bound
  • 3 CoT in ICL: A Lipschitz-Based Perspective
  • 3.1 Problem Setup
  • 3.2 ICL Generalization Bound in the Presence of CoT
  • 4 Synergistic Effects of Demonstration Count and Prompt Templates in ICL
  • 4.1 Problem Setup
  • 4.2 When Sensitivity to Prompt Templates Vanishes with Demonstrations
  • 4.3 When Sensitivity to Prompt Templates Does not Vanishes with Demonstrations
  • 5 Experiments
  • 5.1 Quantifying Demonstration Effectiveness under Prompt Shift
  • 5.2 Effect of Task Decomposition Quality in CoT
  • 5.3 Prompt Sensitivity vs. Demonstration Count
  • 6 Conclusion
  • References
  • A List of Common Notations
  • B Related Work
  • C Proof of
  • C.1 Preliminaries
  • C.2 Proof of
  • D Formal Justification of the Padding Reduction
  • D.1 Preliminaries
  • D.2 Stability of the Transformer Under Padding
  • E Proof of
  • E.1 Preliminaries
  • E.2 Proof of Theorem
  • E.3 Models can ignore irrelevant steps in CoT
  • F Proof of , Corollary and
  • F.1 Preliminaries
  • F.2 Proof of
  • F.3 Proof of Corollary
  • F.4 Proof of Corollary
  • G Proof of and Corollary
  • G.1 Preliminaries
  • G.2 Proof of
  • G.3 Proof of Corollary
  • H Experiments
  • H.1 Pretraining Dependence
  • H.2 Demonstration Choice
  • H.3 Prompt Sensitivity vs. Demonstration Count
  • H.4 CoT Effect

Knowls

  1. Knowl 1 — Lipschitz Generalization Bound for In-Context Learning on Unseen Prompts

    theoretical result

    Let V\mathcal{V} be a finite vocabulary, V∗\mathcal{V}^* the set of bounded-length token sequences embedded in a continuous representation space, and M\mathcal{M} a class of Transformer models. Let M∈MM \in \mathcal{M} be a pretrained model, M⋆:V∗→V∗M^\star: \mathcal{V}^* \to \mathcal{V}^* an oracle target model, and ℓICL(z):=∣M(z)−M⋆(z)∣\ell_{\text{ICL}}(z) := |M(z) - M^\star(z)| the in-context learning (ICL) loss for prompt zz. Let D0⊂V∗D_0 \subset \mathcal{V}^* denote the pretraining prompt domain, and consider an unseen test prompt z∈D∖D0z \in D \setminus D_0.

    Assume that for a reference pretraining prompt z′∈D0z' \in D_0, the straight-line path γ(t)=z′+t(z−z′)\gamma(t) = z' + t(z - z') for t∈[0,1]t \in [0, 1] lies in a compact set on which ℓICL\ell_{\text{ICL}} is Lipschitz continuous with constant Lz,z′L_{z,z'}, and that the pretraining coverage along the path Eγ:={t∈[0,1]:γ(t)∈D0}E_\gamma := \{t \in [0, 1] : \gamma(t) \in D_0\} has positive Lebesgue measure mes(Eγ)>0\text{mes}(E_\gamma) > 0. Let lγ:=Lz,z′∥z−z′∥l_\gamma := L_{z,z'} \|z - z'\| and define the amplification constant

    Aγ:=4mes([0,1])mes(Eγ)=4mes(Eγ)>1.A_\gamma := \frac{4 \text{mes}([0, 1])}{\text{mes}(E_\gamma)} = \frac{4}{\text{mes}(E_\gamma)} > 1.

    For any integer n≥1n \ge 1, the ICL loss on the unseen test prompt satisfies

    ℓICL(z)≤Aγnsup⁡z~∈D0ℓICL(z~)+(1+Aγn)lγ2n.\ell_{\text{ICL}}(z) \le A_\gamma^n \sup_{\tilde{z} \in D_0} \ell_{\text{ICL}}(\tilde{z}) + \left(1 + A_\gamma^n\right) \frac{l_\gamma}{2\sqrt{n}}.

    In particular, choosing n=⌈lγ2⌉n = \lceil l_\gamma^2 \rceil yields

    ℓICL(z)≤Aγ⌈lγ2⌉sup⁡z~∈D0ℓICL(z~)+1+Aγ⌈lγ2⌉2=O(κL2)sup⁡z′∈D0ℓICL(z′),\ell_{\text{ICL}}(z) \le A_\gamma^{\lceil l_\gamma^2 \rceil} \sup_{\tilde{z} \in D_0} \ell_{\text{ICL}}(\tilde{z}) + \frac{1 + A_\gamma^{\lceil l_\gamma^2 \rceil}}{2} = \mathcal{O}\left(\kappa^{L^2}\right) \sup_{z' \in D_0} \ell_{\text{ICL}}(z'),

    where κ>1\kappa > 1 quantifies the degree of prompt distribution shift and LL is the effective pathwise Lipschitz constant. This establishes that ICL generalization on unseen prompts is controlled by (i) the model's intrinsic ICL capability on pretraining prompts sup⁡z′∈D0ℓICL(z′)\sup_{z' \in D_0} \ell_{\text{ICL}}(z'), (ii) the prompt shift κ\kappa, and (iii) the effective Lipschitz constant LL, which measures how stably the demonstrations identify the underlying task.

  2. Knowl 2 — Error Propagation and Generalization Bound for Chain-of-Thought In-Context Learning

    theoretical result

    Consider a task decomposed into KK sequential reasoning subtasks. A prompt with kk intermediate reasoning steps generated for query xquery(0)x_{\text{query}}^{(0)} is denoted by z(k)=(x1(0),…,x1(K),…,xn(0),…,xn(K),xquery(0),…,xquery(k))z^{(k)} = (x_1^{(0)}, \dots, x_1^{(K)}, \dots, x_n^{(0)}, \dots, x_n^{(K)}, x_{\text{query}}^{(0)}, \dots, x_{\text{query}}^{(k)}), with z(0)z^{(0)} being the direct prompt. Let z∗(k)z^{*(k)} denote the rollout under the oracle model M⋆M_\star, where z(k+1)=f(z(k),M(z(k)))z^{(k+1)} = f(z^{(k)}, M(z^{(k)})) and z∗(k+1)=f(z∗(k),M⋆(z∗(k)))z^{*(k+1)} = f(z^{*(k)}, M_\star(z^{*(k)})) with z∗(0)=z(0)z^{*(0)} = z^{(0)}.

    Assume:

    1. Per-step error propagation: the rollout discrepancy Δk:=∥z(k)−z∗(k)∥\Delta_k := \|z^{(k)} - z^{*(k)}\| satisfies Δk+1≤βkℓICL(z(k))+αkΔk\Delta_{k+1} \le \beta_k \ell_{\text{ICL}}(z^{(k)}) + \alpha_k \Delta_k for constants αk,βk≥0\alpha_k, \beta_k \ge 0, with Δ0=0\Delta_0 = 0, where ℓICL(z):=∣M(z)−M⋆(z)∣\ell_{\text{ICL}}(z) := |M(z) - M_\star(z)|.
    2. Oracle smoothness: M⋆M_\star is Lk⋆L_k^\star-Lipschitz on the set of step-kk rollout states RkR_k.
    3. Single-step generalization: for each step mm, with pretraining prompt set D0(m)D_0^{(m)}, the step loss satisfies ℓICL(z(m))≤CmeL(m)2sup⁡z′∈D0(m)ℓICL(z′)\ell_{\text{ICL}}(z^{(m)}) \le C_m e^{L_{(m)}^2} \sup_{z' \in D_0^{(m)}} \ell_{\text{ICL}}(z') for constants Cm≥1C_m \ge 1 and effective Lipschitz parameters L(m)≥0L_{(m)} \ge 0.

    Then the total ICL test loss on the final output satisfies

    ℓICL(z(K−1))≤RK−1+LK−1⋆∑m=0K−2(∏j=m+1K−2αj)βmCmeL(m)2sup⁡z′∈D0(m)ℓICL(z′),\ell_{\text{ICL}}(z^{(K-1)}) \le R_{K-1} + L_{K-1}^\star \sum_{m=0}^{K-2} \left( \prod_{j=m+1}^{K-2} \alpha_j \right) \beta_m C_m e^{L_{(m)}^2} \sup_{z' \in D_0^{(m)}} \ell_{\text{ICL}}(z'),

    where RK−1:=∣M(z(K−1))−M⋆(z∗(K−1))∣R_{K-1} := |M(z^{(K-1)}) - M_\star(z^{*(K-1)})|. In simplified form,

    ℓICL(z)≤∑k=0K−1O(exp⁡(L(k)2))sup⁡z(k)∈D0(k)ℓICL(z(k)).\ell_{\text{ICL}}(z) \le \sum_{k=0}^{K-1} \mathcal{O}\left(\exp\left(L_{(k)}^2\right)\right) \sup_{z^{(k)} \in D_0^{(k)}} \ell_{\text{ICL}}(z^{(k)}).

    CoT improves generalization when the decomposition replaces a hard monolithic task (high LL) with simpler, well-learned subtasks that have substantially smaller effective Lipschitz constants L(k)L_{(k)}.

  3. Knowl 3 — Exponential Suppression of Prompt Template Sensitivity with Demonstration Count

    theoretical result

    Let T\mathcal{T} be a finite set of latent tasks with strictly positive prior π(t)>0\pi(t) > 0. Conditioned on task tt, input-output pairs are drawn i.i.d. as x∼Px∣t,y=ft(x)x \sim P_{x|t}, y = f_t(x), and instructions i∼Pi∣ti \sim P_{i|t}. Assume:

    1. Task Identifiability: for any t1≠t2t_1 \neq t_2, KL(Px,y∣t1∥Px,y∣t2)>0\text{KL}(P_{x,y|t_1} \parallel P_{x,y|t_2}) > 0.
    2. Bounded Instruction Sensitivity: ∥∇ilog⁡p(i∣t)∥2≤Li\|\nabla_i \log p(i \mid t)\|_2 \le L_i for all t∈Tt \in \mathcal{T} on a compact instruction domain.

    Consider prompt formats 1–5: (1) Single correct instruction, (2) Single incorrect instruction or pad, (3) Repeated correct instruction, (4) Repeated incorrect instruction, or (5) Varying correct instructions. Let zz contain NN demonstrations and t⋆t^\star be the true latent task.

    Then there exist constants α>0\alpha > 0 (where α=12min⁡t≠t⋆KL(Px,y∣t⋆∥Px,y∣t)\alpha = \frac{1}{2} \min_{t \neq t^\star} \text{KL}(P_{x,y|t^\star} \parallel P_{x,y|t})) and β>0\beta > 0 such that, with high probability, for all sufficiently large NN:

    ∣p(yquery∣z)−p(yquery∣xquery,t⋆)∣≤βexp⁡(−αN).|p(y_{\text{query}} \mid z) - p(y_{\text{query}} \mid x_{\text{query}}, t^\star)| \le \beta \exp(-\alpha N).

    Equivalently, for any two instructions i,i′i, i' in the prompt with all non-instruction tokens fixed,

    ∣p(yquery∣z(i))−p(yquery∣z(i′))∣≤2βexp⁡(−αN).|p(y_{\text{query}} \mid z(i)) - p(y_{\text{query}} \mid z(i'))| \le 2\beta \exp(-\alpha N).

    Furthermore, the gradient of the posterior predictive with respect to the instruction parameters satisfies

    ∥∇ip(yquery∣z)∥≤β′exp⁡(−α′N)\|\nabla_i p(y_{\text{query}} \mid z)\| \le \beta' \exp(-\alpha' N)

    for constants α′,β′>0\alpha', \beta' > 0. Thus, for formats 1–5, the influence of prompt templates vanishes exponentially as the number of demonstrations increases.

  4. Knowl 4 — Failure of Prompt Sensitivity Decay under Inconsistent Incorrect Instructions

    theoretical result

    Consider prompt format 6, where each of the NN demonstrations is assigned an independent, distinct, and incorrect instruction in∈Rdi_n \in \mathbb{R}^d, forming a concatenated instruction vector I=(i1,…,iN)∈RNdI = (i_1, \dots, i_N) \in \mathbb{R}^{Nd}. Let Di0=Br(Iinner)D_{i_0} = B_r(I^{\text{inner}}) be the inner instruction ball around Iinner=(i⋆,…,i⋆)I^{\text{inner}} = (i^\star, \dots, i^\star) of radius rr, and let Di=BR(Iinner)D_i = B_R(I^{\text{inner}}) be the extended instruction ball of radius R>rR > r.

    Let f(I):=p(yquery∣z(I))f(I) := p(y_{\text{query}} \mid z(I)) and f⋆:=p(yquery∣xquery,t⋆)f^\star := p(y_{\text{query}} \mid x_{\text{query}}, t^\star). Under task identifiability (Assumption 4.1) and bounded instruction sensitivity ∥∇inlog⁡p(in∣t)∥2≤G\|\nabla_{i_n} \log p(i_n \mid t)\|_2 \le G (Assumption 4.2), for any polynomial approximation degree n≥1n \ge 1 and all sufficiently large NN:

    ∥f−f⋆∥L∞(Di)≤exp⁡(C1kn)(β(N)e−αN+εn)+εn,\|f - f^\star\|_{L^\infty(D_i)} \le \exp(C_1 k n) \left( \beta(N) e^{-\alpha N} + \varepsilon_n \right) + \varepsilon_n,

    where k=Ndk = Nd, C1=log⁡(2R/r)C_1 = \log(2R/r), β(N)=O(e2GNr)\beta(N) = \mathcal{O}(e^{2G\sqrt{N}r}), and εn≤c0GRdNn\varepsilon_n \le c_0 G R \sqrt{d} \frac{N}{\sqrt{n}} for an absolute constant c0>0c_0 > 0.

    In consequence, bounding the approximation error εn=O(1)\varepsilon_n = \mathcal{O}(1) requires polynomial degree n=Ω(N2)n = \Omega(N^2), which drives the amplification prefactor exp⁡(C1kn)=exp⁡(Ω(dN3))\exp(C_1 kn) = \exp(\Omega(d N^3)) to diverge super-exponentially. Therefore, the uniform error satisfies

    ∥f−f⋆∥L∞(Di)≤O(exp⁡(Li2))(βe−αN+ϵ),\|f - f^\star\|_{L^\infty(D_i)} \le \mathcal{O}\left(\exp\left(L_i^2\right)\right) \left( \beta e^{-\alpha N} + \epsilon \right),

    and the gradient satisfies

    sup⁡I∈Di∥∇If(I)∥≤O(Li4exp⁡(Li2))(β′e−α′N+ϵ′)\sup_{I \in D_i} \|\nabla_I f(I)\| \le \mathcal{O}\left(L_i^4 \exp\left(L_i^2\right)\right) \left( \beta' e^{-\alpha' N} + \epsilon' \right)

    for strictly positive non-vanishing constants ϵ,ϵ′>0\epsilon, \epsilon' > 0. Sensitivity to prompt templates does not vanish as N→∞N \to \infty when instructions are inconsistent across demonstrations.

  5. Knowl 5 — Taxonomy of Instruction Formats in Latent Task In-Context Learning

    definition

    In a latent task model where inputs and outputs are generated according to x∼Px∣t,y=ft(x)x \sim P_{x|t}, y = f_t(x) with task t∈Tt \in \mathcal{T} and instruction likelihood p(i∣t)p(i \mid t) (where an instruction is correct for target task t⋆t^\star if p(i∣t⋆)>p(i∣t)p(i \mid t^\star) > p(i \mid t) for all t≠t⋆t \neq t^\star, and incorrect otherwise), an ICL prompt zz containing NN demonstrations and a query xqueryx_{\text{query}} is categorized into one of six formats:

    1. Single correct instruction (appears once at prompt start): z=(i,(x1,y1),…,(xN,yN),xquery)z = (i, (x_1, y_1), \dots, (x_N, y_N), x_{\text{query}})
    2. Single incorrect instruction (appears once or replaced by pad): z=(i~,(x1,y1),…,(xN,yN),xquery)z = (\tilde{i}, (x_1, y_1), \dots, (x_N, y_N), x_{\text{query}})
    3. Repeated correct instruction (shared across all examples and query): z=((i,x1,y1),…,(i,xN,yN),(i,xquery))z = ((i, x_1, y_1), \dots, (i, x_N, y_N), (i, x_{\text{query}}))
    4. Repeated incorrect instruction (shared across all examples and query): z=((i,x1,y1),…,(i,xN,yN),(i,xquery))z = ((i, x_1, y_1), \dots, (i, x_N, y_N), (i, x_{\text{query}}))
    5. Varying correct instructions (independent correct instructions per demonstration): z=((i1,x1,y1),…,(iN,xN,yN),(iquery,xquery))z = ((i_1, x_1, y_1), \dots, (i_N, x_N, y_N), (i_{\text{query}}, x_{\text{query}}))
    6. Varying incorrect instructions (inconsistent, incorrect instructions per demonstration): z=((i1,x1,y1),…,(iN,xN,yN),(iquery,xquery))z = ((i_1, x_1, y_1), \dots, (i_N, x_N, y_N), (i_{\text{query}}, x_{\text{query}}))
  6. Knowl 6 — Expected Instruction Stability Decomposition for Learned Predictors

    theoretical result

    Let μN\mu_N be the distribution over ICL prompts of formats 1–5 with exactly NN demonstrations. Let q⋆(z):=p(yquery∣z)q^\star(z) := p(y_{\text{query}} \mid z) denote the Bayes-optimal predictor under cross-entropy, and let qθ(z):=pθ(yquery∣z)q_\theta(z) := p_\theta(y_{\text{query}} \mid z) be a Transformer predictor learned by empirical cross-entropy minimization.

    Assume log-space realizability and positivity (Assumption F.5): there exists δ>0\delta > 0 such that q⋆(z)≥δq^\star(z) \ge \delta and qθ^(z)≥δq_{\hat{\theta}}(z) \ge \delta on the support of μN\mu_N, and the model class can approximate log⁡q⋆(z)\log q^\star(z) to arbitrary precision ϵ>0\epsilon > 0.

    Then, with high probability, for all sufficiently large NN and any instructions i,i′i, i':

    Ez[∣qθ(z(i))−qθ(z(i′))∣]≤2βe−αN⏟demonstrations effect+2Ez[∣qθ(z)−q⋆(z)∣]⏟pretraining effect,\mathbb{E}_z \left[ |q_\theta(z(i)) - q_\theta(z(i'))| \right] \le \underbrace{2\beta e^{-\alpha N}}_{\text{demonstrations effect}} + \underbrace{2\mathbb{E}_z \left[ |q_\theta(z) - q^\star(z)| \right]}_{\text{pretraining effect}},

    where α,β>0\alpha, \beta > 0 are the task convergence constants from Theorem 4.3. The first term represents task identification via demonstrations and decays exponentially with NN, while the second term reflects optimization and approximation error from pretraining.

  7. Knowl 7 — Stability of Transformer Attention Under Sequence Padding

    theoretical result

    Let yy and y′y' denote the softmax attention outputs at a fixed query position before and after inserting mm padding tokens ⟨pad⟩\langle\text{pad}\rangle with zero embedding into an original prompt zorigz_{\text{orig}}. Let xix_i be the attention logit assigned to token i∈Ii \in I, xpx_p the logit assigned to padding token p∈Pp \in P, zorig:=∑i∈Iexiz_{\text{orig}} := \sum_{i \in I} e^{x_i}, and zpad:=∑p∈Pexpz_{\text{pad}} := \sum_{p \in P} e^{x_p}. Let CV:=sup⁡i∥vi∥2C_V := \sup_{i} \|v_i\|_2 be a uniform bound on value vector norms, and let xpad,max⁡:=max⁡p∈Pxpx_{\text{pad},\max} := \max_{p \in P} x_p.

    The perturbation in the attention output is bounded by

    ∥y′−y∥2≤2CVzpadzorig+zpad≤2CVmexpad,max⁡zorig.\|y' - y\|_2 \le 2 C_V \frac{z_{\text{pad}}}{z_{\text{orig}} + z_{\text{pad}}} \le 2 C_V \frac{m e^{x_{\text{pad},\max}}}{z_{\text{orig}}}.

    Furthermore, if there exists an informative token j∈Ij \in I such that the attention gap Δ:=xj−xpad,max⁡>0\Delta := x_j - x_{\text{pad},\max} > 0, then for any target tolerance ε∈(0,2CV)\varepsilon \in (0, 2C_V), a padding length satisfying

    m≤ε2CV−εeΔm \le \frac{\varepsilon}{2C_V - \varepsilon} e^\Delta

    is sufficient to guarantee ∥y′−y∥2≤ε\|y' - y\|_2 \le \varepsilon. Hence, padding tokens act purely as normalization distractors whose effect is exponentially suppressed by the attention gap Δ\Delta.

  8. Knowl 8 — Empirical Effect of Demonstration Informativeness on Task Disambiguation

    data/table

    To evaluate how demonstration ambiguity affects ICL performance, experiments were conducted on two retrieval-style tasks using Qwen-30B-A3B and Qwen-235B-A22B: (1) Sport Identification (mapping a country to its sport label) and (2) People Identification (mapping a company to a canonical person name). Prompts compared an 'Identifying' pool (examples strongly specifying a unique latent rule) against an 'Ambiguous' pool (examples compatible with multiple latent rules), tested with k∈{2,4,8}k \in \{2, 4, 8\} demonstrations over 5 random seeds (100 test queries per seed).

    Task Demonstration Qwen-30B-A3B Qwen-235B-A22B
    Sport (k=2k=2) Ambiguous 0.122±0.0290.122 \pm 0.029 0.265±0.0730.265 \pm 0.073
    Sport (k=2k=2) Identifying 0.176±0.0620.176 \pm 0.062 0.421±0.0800.421 \pm 0.080
    Sport (k=4k=4) Ambiguous 0.151±0.0780.151 \pm 0.078 0.267±0.0840.267 \pm 0.084
    Sport (k=4k=4) Identifying 0.410±0.0590.410 \pm 0.059 0.542±0.0480.542 \pm 0.048
    Sport (k=8k=8) Ambiguous 0.164±0.0390.164 \pm 0.039 0.202±0.0370.202 \pm 0.037
    Sport (k=8k=8) Identifying 0.559±0.0740.559 \pm 0.074 0.602±0.0430.602 \pm 0.043
    People (k=2k=2) Ambiguous 0.735±0.0830.735 \pm 0.083 0.713±0.0740.713 \pm 0.074
    People (k=2k=2) Identifying 0.908±0.0590.908 \pm 0.059 0.940±0.0490.940 \pm 0.049
    People (k=4k=4) Ambiguous 0.775±0.0370.775 \pm 0.037 0.708±0.0960.708 \pm 0.096
    People (k=4k=4) Identifying 0.950±0.0580.950 \pm 0.058 0.931±0.0270.931 \pm 0.027
    People (k=8k=8) Ambiguous 0.820±0.0270.820 \pm 0.027 0.648±0.0080.648 \pm 0.008
    People (k=8k=8) Identifying 0.990±0.0010.990 \pm 0.001 0.845±0.0440.845 \pm 0.044

    Identifying demonstrations consistently outperform ambiguous demonstrations across all shot counts and models. The performance advantage widens as the number of demonstrations increases (e.g., reaching 0.5590.559 vs. 0.1640.164 on Sport and 0.9900.990 vs. 0.8200.820 on People at k=8k=8 for Qwen-30B-A3B), validating that informative demonstrations reduce effective Lipschitz constants and resolve latent task ambiguity.

  9. Knowl 9 — Empirical Superiority of In-Distribution Subtask Decomposition in Chain-of-Thought

    empirical result

    In a synthetic arithmetic experiment on 7-digit addition using Llama-3.2-3B fine-tuned via LoRA, Chain-of-Thought (CoT) prompting was tested under two subtask decomposition strategies across 0 to 5 demonstrations:

    1. In-Distribution Decomposition: decomposes 7-digit addition into sub-computations pre-trained to high proficiency (a 6-digit addition, a 1-digit addition, and multiplication by 10, trained on 500K samples).
    2. Out-of-Distribution Decomposition: decomposes 7-digit addition into ungrounded intermediate steps that still require 7-digit addition (e.g., 1234561=1000000+2345611234561 = 1000000 + 234561), preserving the original problem difficulty on subtasks the model cannot reliably solve in isolation.

    Empirical evaluation showed that in-distribution CoT substantially improved test accuracy over standard direct ICL, with performance increasing steadily with demonstration count. In contrast, out-of-distribution CoT performed worse than vanilla ICL without CoT and degraded further as the number of demonstrations increased. This confirms Theorem 3.1: CoT is beneficial only when subtasks align with low-error, low-Lipschitz routines in the pretraining distribution.

  10. Knowl 10 — Empirical Dynamics of Output Logit Sensitivity Across Prompt Formats and Demonstration Counts

    empirical result

    On a 3-digit addition task evaluated with Qwen3-235B-A22B and Qwen3-30B-A3B across 1 to 50 demonstrations, the discrete gradient magnitude of the output logits with respect to instruction tokens was measured under three prompt formats:

    1. Format 3 (Addition-only / correct instruction): the gradient of output logits with respect to instruction tokens remained near zero across all demonstration counts, demonstrating minimal prompt sensitivity.
    2. Format 5 (Multiplication-only / consistent incorrect instruction where '+' is replaced by '×' but labels are sums): the gradient was initially large at small demonstration counts (reflecting prior bias toward multiplication) but decayed smoothly toward zero as NN increased, as the demonstrations disambiguated the underlying addition task.
    3. Format 6 (Mixed-symbol / inconsistent incorrect instructions where operators '+', '×', '/', and '−' are randomly interleaved across demonstrations): gradient magnitudes remained large and irregular, exhibiting increased variance and even amplifying as NN increased.

    These results empirically corroborate the theoretical dichotomy between Theorem 4.3 (exponential decay of prompt sensitivity for consistent formats) and Theorem 4.7 (non-decaying sensitivity under inconsistent incorrect instructions).

  11. Knowl 11 — Interaction of Intrinsic ICL Capability and Distribution Shift in Multi-Digit Addition

    empirical result

    Three variants of Llama3.2-3B with differing intrinsic ICL capabilities were evaluated on multi-digit addition across 350 unseen test prompts with input lengths varying from 6 to 10 digits:

    1. Strong (Llama3.2-3B-Fine-Tuned on 20,000 separate 6-digit addition JSONL instances).
    2. Medium (Llama3.2-3B-Fine-Tuned on the same 20,000 instances concatenated into a single long prompt).
    3. Weak (Llama3.2-3B-Base without additional fine-tuning).

    Results demonstrated that on 6-digit addition (near the pretraining distribution), test accuracy strictly followed the ordering of intrinsic ICL capability (Strong > Medium > Weak). However, as digit length increased from 6 to 10 digits (increasing distribution shift κ\kappa), the performance gap between the three models narrowed, and all models converged to similarly low accuracy. This behavior corroborates the theoretical prediction of Theorem 2.3 that under large distribution shifts, the exponential Lipschitz factor O(κL2)\mathcal{O}(\kappa^{L^2}) dominates over the intrinsic capability term sup⁡z′∈D0ℓICL(z′)\sup_{z' \in D_0} \ell_{\text{ICL}}(z').

Coverage note — None was omitted; all key theoretical theorems (ICL Lipschitz bound, CoT bound, prompt sensitivity decay/non-decay, padding lemma) and empirical experiments (pretraining shift, demonstration informativeness, CoT decomposition quality, gradient dynamics, structured distractors) are fully covered.

References

  1. 1.Rishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet, Luis Rosias, Stephanie Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle. Many-shot in-context learning. Advances in Neural Information Processing Systems, 37:76930–76966, 2024.
  2. 2.Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra. Transformers learn to implement preconditioned gradient descent for in-context learning. Advances in Neural Information Processing Systems, 36, 2023.
  3. 3.Ekin Aky¨urek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. What learning algorithm is in-context learning? investigations with linear models. International Conference on Learning Representations, 2023.
  4. 4.Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  6. 6.Mark Chen. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  8. 8.Yingqian Cui, Pengfei He, Xianfeng Tang, Qi He, Chen Luo, Jiliang Tang, and Yue Xing. A theoretical understanding of chain-of-thought: Coherent reasoning and error-aware demonstration. In The 28th International Conference on Artificial Intelligence and Statistics, 2025.
  9. 9.Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers. ICLR 2023 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2023.
  10. 10.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021.
  11. 11.Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. Towards revealing the mystery behind chain of thought: A theoretical perspective. Advances in Neural Information Processing Systems, 2023.
  12. 12.Deqing Fu, Tian-Qi Chen, Robin Jia, and Vatsal Sharan. Transformers learn to achieve second-order convergence rates for in-context linear regression. Advances in Neural Information Processing Systems, 37:98675–98716, 2024.
  13. 13.Angeliki Giannou, Shashank Rajput, and Dimitris Papailiopoulos. The expressive power of tuning only the normalization layers. In Proceedings of Thirty Sixth Conference on Learning Theory, volume 195, pages 4130–4131, 2023.
  14. 14.Zixuan Gong, Xiaolin Hu, Huayi Tang, and Yong Liu. Towards auto-regressive next-token prediction: In-context learning emerges from generalization. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=gK1rl98VRp.
  15. 15.Jianliang He, Xintian Pan, Siyu Chen, and Zhuoran Yang. In-context linear regression demystified: Training dynamics and mechanistic interpretability of multi-head softmax attention. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=3TM3fxwTps.
  16. 16.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021a.
  17. 17.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021b.
  18. 18.Juno Kim and Taiji Suzuki. Transformers learn nonlinear features in context. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024. URL https://openreview.net/forum?id=WvvAu6XI1M.
  19. 19.Juno Kim, Tai Nakamaki, and Taiji Suzuki. Transformers are minimax optimal nonparametric in-context learners, 2024. URL https://arxiv.org/abs/2408.12186.
  20. 20.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners, 2023. URL https://arxiv.org/abs/2205.11916.
  21. 21.Hongbo Li, Lingjie Duan, and Yingbin Liang. Provable in-context learning of nonlinear regression with transformers, 2025. URL https://arxiv.org/abs/2507.20443.
  22. 22.Yingcong Li, Kartik Sreenivasan, Angeliki Giannou, Dimitris Papailiopoulos, and Samet Oymak. Dissecting chain-of-thought: Compositionality through in-context filtering and learning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=xEhKwsqxMa.
  23. 23.Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. Chain of thought empowers transformers to solve inherently serial problems. International Conference on Learning Representations, 2024.
  24. 24.Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B. Dolan, Lawrence Carin, and Weizhu Chen. What makes good in-context examples for gpt-3? Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, 2022.
  25. 25.G. G Lorentz. Approximation of functions. Chelsea Pub. Co., New York, N.Y., 1986.
  26. 26.Fei Lu and Yue Yu. Transformer learns the cross-task prior and regularization for in-context learning, 2025. URL https://arxiv.org/abs/2505.12138.
  27. 27.Arvind V. Mahankali, Tatsunori Hashimoto, and Tengyu Ma. One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention. In The Twelfth International Conference on Learning Representations, 2024.
  28. 28.William Merrill and Ashish Sabharwal. The expressive power of transformers with chain of thought. International Conference on Learning Representations, 2024.
  29. 29.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? Empirical Methods in Natural Language Processing, 2022.
  30. 30.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. Show your work: Scratchpads for intermediate computation with language models, 2021. URL https://arxiv.org/abs/2112.00114.
  31. 31.Kazusato Oko, Yujin Song, Taiji Suzuki, and Denny Wu. Pretrained transformer efficiently learns low-dimensional target functions in-context. Advances in Neural Information Processing Systems, 37:77316–77365, 2024.
  32. 32.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. In-context learning and induction heads, 2022. URL https://arxiv.org/abs/2209.11895.
  33. 33.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  34. 34.Madhur Panwar, Kabir Ahuja, and Navin Goyal. In-context learning through the bayesian prism. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=HX5ujdsSon.
  35. 35.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=Ti67584b98.
  36. 36.Evgeny J. Remez. Sur une propri´et´e extr´emale des polynˆomes de tchebychef. Communications de l’Institut des Sciences Math´ematiques et M´ecaniques de l’Universit´e de Kharkoff, 13:93–95, 1936.
  37. 37.Hongjin Su, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu. Selective annotation makes language models better few-shot learners. In The Eleventh International Conference on Learning Representations, 2023.
  38. 38.Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. Function vectors in large language models, 2024. URL https://arxiv.org/abs/2310.15213.
  39. 39.Rasul Tutunov, Antoine Grosnit, Juliusz Ziomek, Jun Wang, and Haitham Bou-Ammar. Why can large language models generate correct chain-of-thoughts?, 2024. URL https://arxiv.org/abs/2310.13571.
  40. 40.Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Jo˜ao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. Transformers learn in-context by gradient descent. International Conference on Machine Learning, 2023.
  41. 41.Xinyi Wang, Wanrong Zhu, and William Yang Wang. Large language models are implicitly topic models: Explaining and finding good demonstrations for in-context learning. arXiv preprint arXiv:2301.11916, 2023.
  42. 42.Albert Webson and Ellie Pavlick. Do prompt-based models really understand the meaning of their prompts? In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2300–2344, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.167. URL https://aclanthology.org/2022.naacl-main.167/.
  43. 43.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  44. 44.Don R Wilhelmsen. A markov inequality in several dimensions. Journal of Approximation Theory, 11(3):216–220, 1974. ISSN 0021-9045. doi: https://doi.org/10.1016/0021-9045(74)90012-4. URL https://www.sciencedirect.com/science/article/pii/0021904574900124.
  45. 45.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. An explanation of in-context learning as implicit bayesian inference. International Conference on Learning Representations, 2022.
  46. 46.Chenxiao Yang, Zhiyuan Li, and David Wipf. An in-context learning theoretic analysis of chain-of-thought. In ICML 2024 Workshop on In-Context Learning, 2024a.
  47. 47.Tong Yang, Yu Huang, Yingbin Liang, and Yuejie Chi. In-context learning with representations: Contextual generalization of trained transformers. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024b.
  48. 48.Seonghyeon Ye, Hyeonbin Hwang, Sohee Yang, Hyeongu Yun, Yireun Kim, and Minjoon Seo. Investigating the effectiveness of task-agnostic prefix prompt for instruction following. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19386–19394, 2024.
  49. 49.Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar. Are transformers universal approximators of sequence-to-sequence functions? International Conference on Learning Representations, 2020.
  50. 50.Yuchen Zeng and Kangwook Lee. The expressive power of low-rank adaptation. In The Twelfth International Conference on Learning Representations, 2024.
  51. 51.Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett. Trained transformers learn linear models in-context, 2023. URL https://arxiv.org/abs/2306.09927.
  52. 52.Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett. Trained transformers learn linear models in-context. Journal of Machine Learning Research, 25(49):1–55, 2024.
  53. 53.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. International Conference on Machine Learning, 139:12697–12706, 2021.

Citation

MLA
Tong, X., et al. “Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL”. arXiv, 2026, http://arxiv.org/abs/2603.19611v1.
APA
Tong, X., Zeng, Y., & Zhang, J. (2026). Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL. arXiv. http://arxiv.org/abs/2603.19611v1
Chicago
Tong, X., Y. Zeng, and J. Zhang. 2026. “Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL”. arXiv. http://arxiv.org/abs/2603.19611v1.
Harvard
Tong, X., Zeng, Y. and Zhang, J. (2026) “Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2603.19611v1.
Vancouver
1. Tong X, Zeng Y, Zhang J (2026) Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL. arXiv

BibTeX

@article{tong2026demonstrations,
  title = {Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL},
  author = {Tong, Xuhan and Zeng, Yuchen and Zhang, Jiawei},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2603.19611v1},
  eprint = {2603.19611}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/