An Explanation of In-context Learning as Implicit Bayesian Inference

Sang Michael XieAditi RaghunathanPercy LiangTengyu Ma

article2021ICLR1,161 citationsOutstanding Paper Honorable Mention

Explains in-context learning in language models as implicit Bayesian inference over latent concepts, proving how coherent pretraining distributions enable prompt-based task adaptation and reproducing key empirical phenomena on a synthetic dataset.

Listen

Large language models demonstrate a remarkable ability called in-context learning, where they execute downstream tasks simply by conditioning on a prompt containing several input-output examples. This learning occurs during inference without any explicit parameter updates or pretraining dedicated to task learning. Understanding how and why this capability emerges has remained a major challenge because modern models are trained on massive, unstructured web datasets and evaluate prompts that artificially concatenate independent examples, creating a significant mismatch with natural pretraining text. The article provides a mathematical and empirical explanation for this phenomenon by framing in-context learning as implicit Bayesian inference, where the model deduces a shared latent concept from the prompt to make its predictions.

The article establishes a theoretical framework using pretraining distributions generated from mixtures of Hidden Markov Models. When documents exhibit long-range coherence, a language model trained to predict the next token must implicitly infer a document-level concept across sequences of text. At test time, when the model is presented with a prompt composed of concatenated examples, it similarly identifies the shared latent concept to determine the appropriate output. To test and validate this theory in a controlled environment, the researchers introduced the Generative In-Context learning dataset, a synthetic benchmark comprising 1,000 documents with approximately 10 million tokens total across various vocabulary sizes. They evaluated both standard Transformer architectures ranging from 29 million to 115 million parameters and Long Short-Term Memory networks across varying prompt configurations.

The theoretical analysis proves that even though prompts create an artificial data distribution, the model asymptotically achieves optimal prediction error as the number of prompt examples increases, provided the latent concept is distinguishable. The expected error also decreases as the length of each example increases, showing that contextual information in the inputs aids concept inference beyond just the input-output mapping. Empirical evaluations on the synthetic dataset confirm that in-context accuracy consistently rises with more examples and longer example lengths. Crucially, ablation experiments show that in-context learning completely fails if pretraining data lacks latent concept structures or if prompts are generated from concepts absent during pretraining. The experiments also mirror real-world behaviors of large models: scaling up model size increases prompt accuracy from 60.2% to 84.7% even when pretraining loss remains constant, performance varies widely depending on the ordering of prompt examples, and zero-shot prompts occasionally outperform few-shot prompts when delimiter structures introduce interference.

These findings suggest that in-context learning does not require explicit meta-learning algorithms but arises naturally from pretraining on coherent, structured data. For teams designing language model applications, the results underscore that prompt engineering should focus on providing sufficiently rich context within examples and formatting prompts to minimize structural distribution shift relative to pretraining text. The evidence that larger models execute implicit Bayesian inference more effectively, even at equivalent training loss values, implies that capacity scaling benefits reasoning and concept extraction beyond simple memorization.

Organizations developing or deploying language models should structure pretraining datasets to maintain clear topic and concept coherence rather than relying solely on diverse but disorganized token transitions. Prompt designers should incorporate longer informative inputs and evaluate multiple example permutations to mitigate sensitivity to ordering. While the article's theoretical guarantees rely on bounded state transitions within Hidden Markov Models and discrete concept families, validation on larger benchmarks such as GPT-3 on the LAMBADA dataset supports the findings, showing that longer contextual examples improve accuracy by nearly 1%. Future work should investigate how models can extrapolate to entirely unseen concepts through the separation of syntax and semantics, as well as formalize how model architecture and parameter scaling interact with implicit inference.

No sufficiently relevant recommendations were found.

Cover for An Explanation of In-context Learning as Implicit Bayesian Inference

Abstract

Large language models (LMs) such as GPT-3 have the surprising ability to do in-context learning, where the model learns to do a downstream task simply by conditioning on a prompt consisting of input-output examples. The LM learns from these examples without being explicitly pretrained to learn. Thus, it is unclear what enables in-context learning. In this paper, we study how in-context learning can emerge when pretraining documents have long-range coherence. Here, the LM must infer a latent document-level concept to generate coherent next tokens during pretraining. At test time, in-context learning occurs when the LM also infers a shared latent concept between examples in a prompt. We prove when this occurs despite a distribution mismatch between prompts and pretraining data in a setting where the pretraining distribution is a mixture of HMMs. In contrast to messy large-scale datasets used to train LMs capable of in-context learning, we generate a small-scale synthetic dataset (GINC) where Transformers and LSTMs both exhibit in-context learning. Beyond the theory, experiments on GINC exhibit large-scale real-world phenomena including improved in-context performance with model scaling (despite the same pretraining loss), sensitivity to example order, and instances where zero-shot is better than few-shot in-context learning.

Table of Contents

  • 1 Introduction
  • 2 In-context learning setting
  • 2.1 Assumptions
  • 3 Theoretical analysis
  • 3.1 High-level approach
  • 3.2 Heuristic derivation
  • 3.3 Formal results
  • 3.3.1 In-context learning under distinguishability
  • 3.3.2 Non-distinguishable case
  • 4 Simulations
  • 5 Discussion and related work
  • 6 Conclusion
  • References
  • A Framework details
  • B Propositions for Theorem
  • C Convergence of the in-context predictor
  • D Proof of Theorem
  • E Non-distinguishable case
  • E.1 Proof of Theorem
  • E.2 Proof of Theorem
  • F Experimental details
  • F.1 GINC dataset
  • F.2 Transformer details
  • F.3 LSTM details
  • F.4 Varying the vocabulary size
  • F.5 Experiment on GPT-3

Knowls

  1. Knowl 1 — In-context learning as implicit Bayesian inference

    model/method

    The paper models pretraining documents as sequences generated under a latent concept θ\theta drawn from a prior. The concept governs token transitions, so predicting coherent continuations requires a language model (LM) to use evidence across a document to infer its concept. A prompt contains examples generated under a shared prompt concept θ∗\theta^*, even though concatenating independent examples can make the prompt atypical of a pretraining document. If the LM has learned the pretraining distribution pp exactly, its prediction for an output yy given training examples SnS_n and test input xtestx_{\mathrm{test}} is the posterior predictive distribution

    p(y∣Sn,xtest)=∫p(y∣Sn,xtest,θ) p(θ∣Sn,xtest) dθ.p(y\mid S_n,x_{\mathrm{test}})=\int p(y\mid S_n,x_{\mathrm{test}},\theta)\,p(\theta\mid S_n,x_{\mathrm{test}})\,d\theta.

    Here nn is the number of prompt examples, and the integral averages predictions over concepts according to their posterior probability. In-context learning can therefore arise without an explicit learning procedure: when the examples provide enough evidence for θ∗\theta^*, posterior averaging favors predictions under the shared prompt concept.

  2. Knowl 2 — HMM framework for prompts and prediction

    model/method

    A pretraining document is generated by drawing a concept θ\theta from a prior over a family Θ\Theta and sampling tokens from a hidden Markov model (HMM) whose transition probabilities depend on θ\theta. A prompt is generated differently: first draw a shared prompt concept θ∗∈Θ\theta^*\in\Theta, then independently generate nn examples and one test input under that concept. Each example consists of a length-kk token sequence Oi=[xi,yi]O_i=[x_i,y_i], where xi=Oi[1:k−1]x_i=O_i[1:k-1] is the input and yi=Oi[k]y_i=O_i[k] is its one-token target. A special delimiter token is inserted between training examples; the test input is xtest=xn+1x_{\mathrm{test}}=x_{n+1}. The prompt distribution is thus [x1,y1,odelim,…,xn,yn,odelim,xtest][x_1,y_1,o^{\mathrm{delim}},\ldots,x_n,y_n,o^{\mathrm{delim}},x_{\mathrm{test}}], with the examples independently started from a prompt-specific hidden-state distribution. The target is drawn from the prompt conditional pprompt(y∣xtest)p_{\mathrm{prompt}}(y\mid x_{\mathrm{test}}), while the LM predicts using the pretraining conditional p(y∣Sn,xtest)p(y\mid S_n,x_{\mathrm{test}}). The mismatch arises because concatenating independent examples creates transitions, including transitions around delimiters, that may be unlikely in a pretraining document.

  3. Knowl 3 — Distinguishability guarantees asymptotically optimal prediction

    theoretical result

    For a length-kk example, let ppromptjp^j_{\mathrm{prompt}} be the prompt-distribution conditional distribution of token O[j]O[j] given its preceding tokens, and let pθjp^j_\theta be the corresponding distribution under HMM concept θ\theta. Define

    KLj(θ∗∥θ)=EO[1:j−1]∼pprompt ⁣[KL(ppromptj∥pθj)].\mathrm{KL}_j(\theta^*\Vert\theta)=\mathbb{E}_{O[1:j-1]\sim p_{\mathrm{prompt}}}\!\left[\mathrm{KL}\left(p^j_{\mathrm{prompt}}\Vert p^j_\theta\right)\right].

    Let Δ\Delta be the probability margin between the most likely and second-most-likely outputs under pprompt(y∣xtest)p_{\mathrm{prompt}}(y\mid x_{\mathrm{test}}). The theorem assumes a well-specified prompt concept θ∗∈Θ\theta^*\in\Theta, positive and bounded HMM transition/start/emission probabilities, a delimiter that identifies delimiter hidden states, and a prompt start distribution within total variation distance Δ/4\Delta/4 of the HMM transition distribution from any delimiter state under θ∗\theta^*. Let c1,c2c_1,c_2 bound delimiter-state transition probabilities from below under θ∗\theta^* and from above under alternative concepts, let c3,c4c_3,c_4 bound the prior probabilities of delimiter states, and let c8c_8 lower-bound hidden-state start probabilities under θ∗\theta^*. Define the mismatch allowances ϵstart=log⁡(1/c8)\epsilon_{\mathrm{start}}=\log(1/c_8) and ϵdelim=2(log⁡c2−log⁡c1)+log⁡c4−log⁡c3\epsilon_{\mathrm{delim}}=2(\log c_2-\log c_1)+\log c_4-\log c_3. If, for every θ≠θ∗\theta\ne\theta^*,

    ∑j=1kKLj(θ∗∥θ)>ϵstart+ϵdelim,\sum_{j=1}^{k}\mathrm{KL}_j(\theta^*\Vert\theta)>\epsilon_{\mathrm{start}}+\epsilon_{\mathrm{delim}},

    then as the number of examples nn tends to infinity, the most likely prediction under the pretraining distribution converges to the most likely output under the prompt distribution: arg⁡max⁡yp(y∣Sn,xtest)→arg⁡max⁡ypprompt(y∣xtest)\arg\max_y p(y\mid S_n,x_{\mathrm{test}})\to\arg\max_y p_{\mathrm{prompt}}(y\mid x_{\mathrm{test}}). Consequently the predictor attains the minimum possible expected 0–1 error under the prompt distribution asymptotically. The condition says that evidence about the prompt concept in each example must outweigh the distribution mismatch introduced by prompt starts and delimiters.

  4. Knowl 4 — Excess-risk bound when concepts are not distinguishable

    theoretical result

    Let θ∈Rd\theta\in\mathbb{R}^d parameterize an HMM concept, with true prompt concept θ∗\theta^*. For each token position jj, let KLj(θ∗∥θ)\mathrm{KL}_j(\theta^*\Vert\theta) be the expected KL divergence between the prompt and concept-θ\theta next-token distributions, averaged over prompt histories. Define the non-distinguishable set BB as the concepts satisfying ∑j=1kKLj(θ∗∥θ)≤ϵstart+ϵdelim\sum_{j=1}^k\mathrm{KL}_j(\theta^*\Vert\theta)\leq\epsilon_{\mathrm{start}}+\epsilon_{\mathrm{delim}}, where the two ϵ\epsilon terms bound prompt-start and delimiter mismatch. Suppose the KL divergences have a second-order expansion around θ∗\theta^*, with Fisher information matrix Ij,θ∗I_{j,\theta^*} at position jj, and let

    γθ∗=max⁡jλmax⁡(Ij,θ∗)min⁡jλmin⁡(Ij,θ∗).\gamma_{\theta^*}=\frac{\max_j\lambda_{\max}(I_{j,\theta^*})}{\min_j\lambda_{\min}(I_{j,\theta^*})}.

    For k≥2k\geq2, as n→∞n\to\infty, the in-context predictor's expected 0–1 risk obeys

    lim⁡n→∞L0−1(fn)≤inf⁡fL0−1(f)+g−1 ⁣(O ⁣(γθ∗sup⁡θ∈B(ϵstart+ϵdelim)k−1)),\lim_{n\to\infty}L_{0-1}(f_n)\leq\inf_f L_{0-1}(f)+g^{-1}\!\left(O\!\left(\frac{\gamma_{\theta^*}\sup_{\theta\in B}(\epsilon_{\mathrm{start}}+\epsilon_{\mathrm{delim}})}{k-1}\right)\right),

    where fnf_n predicts the most likely output under the pretraining distribution conditioned on the prompt, and g(δ)=12[(1−δ)log⁡(1−δ)+(1+δ)log⁡(1+δ)]g(\delta)=\tfrac12[(1-\delta)\log(1-\delta)+(1+\delta)\log(1+\delta)] is the calibration function used to convert multiclass logistic-risk control to 0–1-risk control. This bound assumes the minimizers of the 0–1 and multiclass logistic risks coincide. Under the stated smoothness and conditioning assumptions, the excess-risk bound decreases approximately as 1/k1/k; better-conditioned Fisher information (smaller γθ∗\gamma_{\theta^*}) yields a tighter bound.

  5. Knowl 5 — Random-length test inputs retain an inverse-length risk bound

    theoretical result

    For each token position jj, let KLj(θ∗∥θ)\mathrm{KL}_j(\theta^*\Vert\theta) be the expected KL divergence between the prompt next-token distribution and the next-token distribution under HMM concept θ\theta, averaged over prompt histories. Let ϵstart\epsilon_{\mathrm{start}} and ϵdelim\epsilon_{\mathrm{delim}} denote the mismatch allowances for prompt starts and delimiter transitions, and define B={θ:∑j=1kKLj(θ∗∥θ)≤ϵstart+ϵdelim}B=\{\theta:\sum_{j=1}^{k}\mathrm{KL}_j(\theta^*\Vert\theta)\leq\epsilon_{\mathrm{start}}+\epsilon_{\mathrm{delim}}\}. If the test-input length is sampled uniformly from 2,3,…,k2,3,\ldots,k, then for k≥2k\geq2 and as the number of training examples nn tends to infinity, the in-context predictor satisfies

    lim⁡n→∞L0−1(fn)≤inf⁡fL0−1(f)+g−1 ⁣(O ⁣(sup⁡θ∈B(ϵstart+ϵdelim)k−1)),\lim_{n\to\infty}L_{0-1}(f_n)\leq\inf_f L_{0-1}(f)+g^{-1}\!\left(O\!\left(\frac{\sup_{\theta\in B}(\epsilon_{\mathrm{start}}+\epsilon_{\mathrm{delim}})}{k-1}\right)\right),

    where L0−1L_{0-1} is expected classification error under the prompt distribution and g(δ)=12[(1−δ)log⁡(1−δ)+(1+δ)log⁡(1+δ)]g(\delta)=\tfrac12[(1-\delta)\log(1-\delta)+(1+\delta)\log(1+\delta)]. The bound does not require the continuity assumption used for the fixed-length excess-risk result, but it assumes that the 0–1-risk and multiclass-logistic-risk minimizers coincide. The guarantee averages prediction error over positions from 2 through kk; it does not resolve the mismatch between fixed-length training examples and random-length test examples.

  6. Knowl 6 — GINC provides a controlled synthetic test of latent-concept learning

    experimental setup

    The Generative In-Context Learning dataset (GINC) is built from a uniform mixture of five HMM concepts. A hidden state consists of an entity index and a property index, each evolving as a Markov chain; the observed token is read deterministically from a shared entity-by-property memory matrix. The matrix has 10 entities and 10 properties, with the delimiter token occupying the first property column. Each concept changes the property transition matrix, while the entity transition matrix is shared across concepts and normally has the form 0.1T+0.9I0.1T+0.9I, encouraging entity persistence; TT is generated from a convex combination of 100 random permutation matrices. Mixture documents use vocabulary sizes 50, 100, or 150. The pretraining corpus has 1,000 documents of 10,240 tokens each (about 10 million tokens total), and validation uses 100 documents of 1,024 tokens each. Prompts use 0–64 examples, example lengths k∈{3,5,8,10}k\in\{3,5,8,10\}, and 2,500 prompts per setting. For each prompt, one of the five concepts is selected and shared across examples; the prompt start selects an entity uniformly and fixes a randomly selected starting property. The test target is the most likely output under the prompt distribution, rather than a sampled output, to remove intrinsic label noise. This prompt-start choice may violate the theory's total-variation assumption, although the authors report successful empirical learning.

  7. Knowl 7 — In-context accuracy improves with more and longer GINC examples

    empirical result

    Transformers and LSTMs pretrained on GINC both exhibit in-context learning: accuracy rises as the prompt contains more examples and as each example becomes longer. The comparison varies the number of examples from 0 to 64 and example length across k∈{3,5,8,10}k\in\{3,5,8,10\}; the effect appears for both model families, with the longer-example settings generally performing better. Results are averaged over five pretraining runs. This behavior matches the theory's prediction that more examples improve evidence accumulation about a shared concept and that additional tokens within each example can also carry concept information, beyond the input-output mapping alone.

  8. Knowl 8 — GINC ablations identify the mixture structure as necessary for learning

    empirical result

    Ablations with a four-layer Transformer show that exposure to varied token transitions alone does not produce GINC-style in-context learning. When pretraining uses only one concept, removing the mixture over concepts, in-context accuracy remains essentially flat as the prompt gains examples. When pretraining instead uses random transitions that expose the model to all possible token transitions, in-context learning also fails. A separate test prompts the same model with five random concepts outside the pretraining concept family; it fails to extrapolate to these unseen concepts. Together, these controlled comparisons support the role of a shared latent concept family, while showing that the observed learning does not automatically generalize to concepts outside that family.

  9. Knowl 9 — Model scale and architecture affect GINC accuracy beyond pretraining loss

    data/table

    The table reports pretraining losses and in-context accuracy for models evaluated with 64 examples of length k=10k=10. Each accuracy is reported with the paper's uncertainty interval; all values are reproduced as reported. Increasing Transformer depth generally improves in-context accuracy, including the vocabulary-50 comparison where 12- and 16-layer models both have validation loss 1.33 but accuracy rises from 81.2 to 84.7. LSTMs also outperform the Transformers listed here despite having fewer parameters. Results average five pretraining runs.

    Vocabulary Model Parameters Train loss Val. loss In-context accuracy
    50 Transformer (4 layer) 29M 1.49 1.50 60.2 ±\pm 5.7
    50 Transformer (12 layer) 85M 1.31 1.33 81.2 ±\pm 7.1
    50 Transformer (16 layer) 115M 1.31 1.33 84.7 ±\pm 3.4
    50 LSTM 28M 1.31 1.35 95.8 ±\pm 1.11
    100 Transformer (4 layer) 29M 1.58 1.59 67.4 ±\pm 4.7
    100 Transformer (12 layer) 85M 1.40 1.42 84.6 ±\pm 3.0
    100 Transformer (16 layer) 115M 1.41 1.43 88.7 ±\pm 1.6
    100 LSTM 28M 1.43 1.44 95.8 ±\pm 1.54
    150 Transformer (4 layer) 29M 1.44 1.45 92.8 ±\pm 1.9
    150 Transformer (12 layer) 85M 1.27 1.28 98.4 ±\pm 0.4
    150 Transformer (16 layer) 115M 1.27 1.28 98.1 ±\pm 0.5
    150 LSTM 28M 1.26 1.31 99.2 ±\pm 1.06
  10. Knowl 10 — Prompt example order can substantially change accuracy

    empirical result

    A four-layer Transformer trained on GINC with vocabulary size 50 shows marked sensitivity to the ordering of examples. For each of 10 sets of four training examples generated under one concept and prompt-start distribution, the experiment evaluates all 24 permutations. In-context accuracy differs by roughly 10–40 percentage points across permutations of the same example set. Thus, even when the examples and their labels are unchanged, their sequence in the prompt can materially affect the prediction.

  11. Knowl 11 — Zero-shot can outperform few-shot before recovering with more examples

    empirical result

    In a GINC setting with 12 concepts, vocabulary size 100, and transition-matrix temperature 0.01 (rather than the usual 0.1), a Transformer can perform worse with a small number of prompt examples than with no examples. Accuracy then recovers as more examples are added. The authors hypothesize that the prompt's concatenated-example structure initially distracts the model; the experiment establishes the non-monotonic pattern in this setting, not that explanation as a general mechanism.

  12. Knowl 12 — Longer GPT-3 examples help on a filtered LAMBADA task

    data/table

    GPT-3 was evaluated on a filtered LAMBADA test set containing examples 200–300 characters long. Prompts drew five examples from a short training pool (200–300 characters) or a long pool (500–600 characters); the two pools were equalized to 47 candidates. Longer prompt examples yielded 70.7% test accuracy versus 69.8% for short examples. To compare prompt length while controlling the number of distinct short examples, the study also used 10 short examples: duplicating five short examples gave 69.6%, while using 10 independent short examples gave 71.4%. The comparison indicates that longer examples can improve accuracy despite their length mismatch with the test set, and that simply repeating short examples does not provide the same benefit. The paper reports that five long examples close about 56% of the accuracy gap between five short examples and 10 independent short examples.

    Prompt examples Test accuracy (200–300 characters)
    5 short examples (200–300 characters) 69.8
    5 long examples (500–600 characters) 70.7
    10 short examples, duplicated 69.6
    10 short examples, independent 71.4
  13. Knowl 13 — Theory and experiments leave important generalization questions open

    limitation

    The formal analysis assumes that the pretrained LM fits its pretraining distribution exactly and studies prompts generated from a concept family represented in that distribution; it does not establish comparable guarantees for approximate model fitting, misspecification, or extrapolation to unseen concepts. GINC experiments show failure on prompts from random concepts outside the training family. The fixed-length theory also leaves open the mismatch between training examples of length kk and random-length test inputs. Finally, the experiments show that model size and architecture affect in-context accuracy, but the paper does not explain the mechanisms behind these effects.

Coverage note — Proof derivations and implementation-level optimizer and hardware settings are omitted because they support, but do not constitute, separate central contributions.

References

  1. 1.Leonard E. Baum and Ted Petrie. Statistical inference for probabilistic functions of finite state markov chains. The annals of mathematical statistics, 37(6):1554–1563, 1966.
  2. 2.D. Blei, Andrew Ng, and M. I. Jordan. Latent Dirichlet allocation. Journal of Machine Learning Research (JMLR), 3:993–1022, 2003.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  4. 4.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. Electra: Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations (ICLR), 2020.
  5. 5.A. P. Dempster, Laird N. M., and Rubin D. B. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B, 39(1):1–38, 1977.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Association for Computational Linguistics (ACL), pages 4171–4186, 2019.
  7. 7.Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. arXiv, 2021.
  8. 8.Zoubin Ghahramani and Michael Jordan. Factorial hidden Markov models. Machine Learning, 29:245–273, 1997.
  9. 9.Amit Gruber, Yair Weiss, and Michal Rosen-Zvi. Hidden topic Markov models. In Artificial Intelligence and Statistics (AISTATS), 2007.
  10. 10.M. Gunst and O. Shcherbakova. Asymptotic behavior of Bayes estimators for hidden Markov models with application to ion channels. Mathematical Methods of Statistics, 17, 2008.
  11. 11.Keith W. Hastings. Monte Carlo sampling methods using Markov chains and their applications. Biometrika, 57(1):97–109, 1970.
  12. 12.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997.
  13. 13.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations (ICLR), 2020.
  14. 14.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. Surface form competition: Why the highest probability answer isn’t always right, 2021.
  15. 15.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what language models know? In Association for Computational Linguistics (ACL), 2020.
  16. 16.Michael I. Jordan, Zoubin Ghahramani, Tommi S. Jaakkola, and Lawrence K. Saul. An introduction to variational methods for graphical models. Machine Learning, 37:183–233, 1999.
  17. 17.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Association for Computational Linguistics (ACL), 2017.
  18. 18.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  19. 19.B.J.K. Kleijn and A.W. van der Vaart. The Bernstein-von mises theorem under misspecification. Electronic Journal of Statistics, 6, 2012.
  20. 20.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  21. 21.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Association for Computational Linguistics (ACL), 2020.
  22. 22.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Association for Computational Linguistics (ACL), 2021.
  23. 23.Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham. Jurassic-1: Technical details and evaluation. Technical report, AI21 Labs, August 2021.
  24. 24.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  25. 25.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019.
  26. 26.Nicholas Metropolis, Arianna W. Rosenbluth, Marshall N. Rosenbluth, Augusta H. Teller, and Edward Teller. Equation of state calculations by fast computing machines. The journal of chemical physics, 21(6):1087–1092, 1953.
  27. 27.Denis Paperno, German Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernandez. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Association for Computational Linguistics (ACL), 2016.
  28. 28.Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. Grokking: Generalization beyond overfitting on small algorithmic datasets. In ICLR MATH AI Workshop, 2021.
  29. 29.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 2019.
  30. 30.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. In International Conference on Learning Representations (ICLR), 2017.
  31. 31.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. Multitask prompted training enables zero-shot task generalization, 2021.
  32. 32.Timo Schick and Hinrich Schütze. Exploiting cloze questions for few shot text classification and natural language inference. In European Association for Computational Linguistics (EACL), 2021.
  33. 33.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Eliciting knowledge from language models using automatically generated prompts. In Empirical Methods in Natural Language Processing (EMNLP), 2020.
  34. 34.Ingo Steinwart. How to compare different loss functions and their risks. Constructive Approximation, 26, 2007.
  35. 35.A. W. van der Vaart. Asymptotic statistics. Cambridge University Press, 1998.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  37. 37.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  38. 38.Colin Wei, Sang Michael Xie, and Tengyu Ma. Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning. arXiv, 2021a.
  39. 39.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. Finetuned language models are zero-shot learners. arXiv, 2021b.
  40. 40.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. HuggingFace’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  41. 41.Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. In International Conference on Learning Representations (ICLR), 2017.
  42. 42.Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning (ICML), 2021.
  43. 43.Bernardo Ávila Pires and Csaba Szepesvári. Multiclass classification calibration functions. arXiv, 2016.

Citation

MLA
Xie, S. M., et al. “An Explanation of In-context Learning as Implicit Bayesian Inference”. arXiv, 2021, http://arxiv.org/abs/2111.02080v6.
APA
Xie, S. M., Raghunathan, A., Liang, P., & Ma, T. (2021). An Explanation of In-context Learning as Implicit Bayesian Inference. arXiv. http://arxiv.org/abs/2111.02080v6
Chicago
Xie, S. M., A. Raghunathan, P. Liang, and T. Ma. 2021. “An Explanation of In-context Learning as Implicit Bayesian Inference”. arXiv. http://arxiv.org/abs/2111.02080v6.
Harvard
Xie, S.M. et al. (2021) “An Explanation of In-context Learning as Implicit Bayesian Inference”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.02080v6.
Vancouver
1. Xie SM, Raghunathan A, Liang P, Ma T (2021) An Explanation of In-context Learning as Implicit Bayesian Inference. arXiv

BibTeX

@article{xie2021explanation,
  title = {An Explanation of In-context Learning as Implicit Bayesian Inference},
  author = {Xie, Sang Michael and Raghunathan, Aditi and Liang, Percy and Ma, Tengyu},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.02080v6},
  eprint = {2111.02080}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors