In-Context Language Learning: Architectures and Algorithms

Ekin AkyürekBailin WangYoon KimJacob Andreas

article2024ICML104 citations

Reveals that Transformers perform in-context formal language learning by computing smoothed n-gram statistics via specialized attention heads, providing an architectural insight that reduces natural language modeling perplexity when hard-wired into sequence models.

Listen

Modern artificial intelligence relies heavily on the ability of large language models to adapt on the fly to new tasks simply by processing contextual examples, a process known as in-context learning. However, current research into how this mechanism works has predominantly relied on overly simplistic synthetic benchmarks, such as linear regression or basic key-value retrieval, which fail to capture the complex generative dynamics of real language processing.

The article addresses this gap by introducing a new evaluation framework called in-context language learning to investigate how neural architectures infer entire generative languages from contextual string samples. Specifically, the article evaluates how various sequence models learn regular formal languages generated by random finite automata and demonstrates how identifying their internal learning mechanisms can directly inform better model design.

To conduct this evaluation, the researchers created a synthetic benchmark comprising thousands of problem instances derived from randomly generated probabilistic automata. Across this benchmark, the article tested ten sequence architectures spanning standard attention-based Transformers, linear attention variants, recurrent networks, and convolutional models. The investigation combined behavioral accuracy tests, probabilistic error measurements, internal representation probing, and mechanistic attention analyses. Finally, the authors evaluated their structural insights by training 340-million-parameter language models on seven billion tokens of natural web text.

The investigation produced four central findings. First, standard Transformers substantially outperformed all recurrent and convolutional models in learning formal languages from context, whereas alternative architectures frequently failed to surpass simple statistical baselines. Second, internal probing and attention visualizations revealed that successful Transformers achieve this by developing specialized attention circuits that track and normalize short sequence pattern frequencies, functioning as higher-order induction heads. Third, explicitly hard-wiring these pattern-matching layers into recurrent and linear attention models elevated their formal language performance to Transformer-level accuracy. Fourth, integrating these dedicated layers into full-scale language models trained on real text consistently boosted performance, reducing test perplexity by up to 6.7 percent in 340-million-parameter Transformers.

These findings indicate that the ability to track local pattern statistics within context is a fundamental driver of language model performance. Rather than requiring models to expend capacity learning basic statistical induction from scratch, neural architectures can be explicitly designed with dedicated pattern-tracking layers to achieve lower perplexity and better data efficiency without increasing parameter counts. This presents a viable architectural pathway to reduce the computational expense of large-scale pre-training while enhancing generative quality.

Engineering and research teams designing next-generation language models should consider integrating dedicated multi-order pattern heads into both Transformer and recurrent architectures. For long-context and recurrent systems, adopting these heads allows models to maintain efficient inference while closing the capability gap with full attention networks. Prior to wide-scale deployment, teams should conduct pilot pre-training runs across larger multi-billion parameter configurations and expand evaluations into richer grammatical structures, such as context-free languages.

Confidence in these findings is high for regular formal languages and small-scale language models up to 340 million parameters. However, decision-makers should maintain caution regarding very large scales, as the empirical validation on natural text was restricted to seven billion tokens, and regular languages do not encapsulate the full hierarchical complexity of human language.

Akyürek et al (2024).pdf
  • Paper: Evaluating the World Model Implicit in a Generative Model, Keyon Vafa et al. (2024). Its finite-automata framework for testing whether models learn coherent state structure provides the formal-language evaluation context this paper develops for in-context learning.
  • Paper: Overcoming a Theoretical Limitation of Self-Attention, David Chiang et al. (2022). Its analysis of Transformers recognizing formal languages clarifies architectural capabilities and limitations relevant to this paper’s automata-based comparisons.
  • Paper: Transformers Learn In-Context by Gradient Descent, Johannes von Oswald et al. (2023). Its account of Transformers implementing gradient-descent-like learning in context supplies a key mechanistic precedent for this paper’s investigation of learned in-context circuits.
  • Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). Its meta-training approach explicitly teaches models to infer tasks from demonstrations, grounding the broader question of how architectures learn from context.
Cover for In-Context Language Learning: Architectures and Algorithms

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Neural sequence modeling
  • 2.2. In-context learning
  • 2.3. Formal Languages
  • 3. REGBENCH: A Benchmark Dataset for In-Context Language Learning
  • 3.1. Sampling languages
  • 3.2. Sampling strings
  • 3.3. REGBENCH Dataset
  • 4. Which Model Classes Learn to Perform ICLL Efficiently?
  • 4.1. Setup
  • 4.2. Neural sequence models
  • 4.3. Baseline learning algorithms
  • 4.4. Metrics
  • 4.5. Results
  • 5. What Algorithmic Solutions do In-Context Language Learners Implement?
  • 5.1. Transformers form in-context n-gram heads
  • 5.2. Transformers represent in-context n-gram counts better than other models
  • 5.3. Transformer predictions resemble n-gram models with learned reweighting
  • 6. How Can Findings About ICLL Inform the Design of Neural Sequence Models?
  • 7. Conclusion
  • Impact Statement
  • Acknowledgements
  • References
  • A. Model Architectures
  • A.1. Overview
  • A.2. Modeling Overview
  • A.3. Transformers with Self Attention (Vaswani et al., 2017)
  • A.4. Transformers with Linear Attention (Katharopoulos et al., 2020)
  • A.5. RetNet (Sun et al., 2023; Qin et al., 2022)
  • A.6. GLA (Yang et al., 2023)
  • A.7. LSTM (Hochreiter & Schmidhuber, 1997)
  • A.8. RWKV (Peng et al., 2023)
  • A.9. S4 (Gu et al., 2022b)
  • A.10. H3 (Fu et al., 2023)
  • A.11. Hyena (Poli et al., 2023)
  • A.12. Mamba (Gu & Dao, 2023)
  • B. Optimization & Hyperparameter Search
  • C. Algorithms
  • C.1. In-context N-gram Language Model
  • C.2. In-context Baum-Welch HMM Language Model
  • C.3. Masking A to enforce state transitions
  • C.4. Masking π to start at the initial state
  • D. Learned MLP Reweighting
  • E. Probing Experiments
  • F. Implementations of N-gram Layers
  • G. Language Model Experiments

Knowls

  1. Knowl 1 — In-context language learning as adaptation to an unknown language

    definition

    In-context language learning (ICLL) is next-token prediction in which a sequence model receives a finite collection of example strings sampled from an unknown language and must predict continuations that are also valid under that language. The examples are provided in the same context, so the learner must use them to estimate a context-dependent distribution over strings rather than identify a language from a fixed, pre-known set. This paper studies ICLL for probabilistic regular languages generated by finite automata.

  2. Knowl 2 — REGBENCH samples new probabilistic regular languages for each instance

    experimental setup

    REGBENCH creates each language by sampling an automaton, then sampling strings from its induced probabilistic language. An automaton has a start state and 4–12 additional states; its alphabet contains 4–18 symbols selected from a shared set of 18. For each non-start state, the generator chooses 1–4 outgoing labeled edges, with labels sampled without replacement from the language-specific alphabet and destinations chosen uniformly without replacement from the other non-start states. Unselected alphabet symbols lead to a dead state in the associated deterministic automaton. The deterministic automaton is minimized, then converted to a probabilistic automaton with no terminal states: each selected edge from a state with oo selected edges has probability 1/o1/o, and all other transitions have probability zero.

    A string is generated by choosing its length uniformly from 1 to 50 and sampling that many transitions from the automaton. A problem instance contains 10–20 such strings from one automaton, with average total length about 382 symbols. Training and test instances use disjoint automata, so evaluation tests adaptation to previously unseen languages.

  3. Knowl 3 — Transformers outperform recurrent and convolutional models on regular ICLL

    empirical result

    Across REGBENCH training-set sizes and both evaluation metrics, Transformers substantially outperform the tested recurrent and convolutional sequence models. The comparison included Transformers, Linear Transformers and RetNet; recurrent LSTM, RWKV, GLA and Mamba models; and convolutional S4, H3 and Hyena models. Most non-Transformer models perform worse than the in-context n-gram and Baum–Welch baselines except in the high-data regime. In contrast, associative recall produces less separation between model classes. No architecture achieves non-trivial REGBENCH performance with only one layer; Transformer accuracy rises monotonically with depth, whereas other architectures begin to overfit as depth increases.

  4. Knowl 4 — REGBENCH training and evaluation measure next-token language adaptation

    experimental setup

    Models are trained by maximizing next-token log likelihood over the serialized strings and their prefixes: for parameters θ\theta, the objective is ∑d∈Dtrain∑ilog⁡pθ(xi∣d<i)\sum_{d\in D_{\mathrm{train}}}\sum_i \log p_\theta(x_i\mid d_{<i}), where dd is a problem instance and xix_i is its symbol at position ii. The evaluation uses 500 test instances and training subsets ranging from 150 to 40,000 instances. Greedy accuracy is the fraction of evaluated positions where the model's highest-probability next symbol is in the support of the true language's next-symbol distribution, conditioned on the current partial string. Total variation distance (TVD) compares the full predicted next-symbol distribution with that true conditional distribution; lower TVD is better. The study also compares with associative recall using vocabulary size 40, sequence length 382, and a 500-instance test set.

  5. Knowl 5 — Transformer attention implements higher-order induction-like n-gram heads

    empirical result

    Attention visualizations of an 8-layer, 1-head Transformer trained on 2,500 REGBENCH instances show a sequence of computations consistent with higher-order induction. Heads in layers 2 and 3 attend to the immediately previous token, allowing later representations to incorporate the identities of preceding tokens. A layer-5 head then attends to tokens that followed occurrences of the same two-symbol context as the current context—for example, when the input ends in “nh,” it attends to tokens appearing after “nh” elsewhere in the context. The head selects by matching the n-gram, not by selecting tokens generated from the same underlying automaton state.

  6. Knowl 6 — Transformer representations encode n-gram frequencies better than other architectures

    empirical result

    Probes trained on hidden representations from REGBENCH models find that Transformers make higher-order n-gram frequencies more decodable than the tested recurrent and convolutional models. Transformers also support more accurate decoding of n-gram existence and equivalence of the underlying automaton states. The advantage does not extend to unnormalized n-gram counts: count probes do not show a meaningful Transformer advantage over other architectures. The probing results therefore distinguish the encoding of normalized frequencies and existence information from the encoding of raw counts.

  7. Knowl 7 — Transformer predictions resemble learned mixtures of in-context n-gram statistics

    empirical result

    A learned n-gram reweighting predictor represents each context using the empirical next-symbol distributions for contexts matching the most recent one symbol, the most recent two symbols, and the unigram context, then maps the concatenated features through a one-hidden-layer MLP. The paper tests versions using raw counts and normalized distributions. In pairwise next-token TVD comparisons, the normalized-feature version is the closest match to the 12-layer Transformer. Larger Transformers are more similar to n-gram-based predictors than to the ground-truth automaton predictor or the Baum–Welch predictor; the 2-layer Transformer is more similar to a 2-gram baseline than to a 3-gram baseline. This supports the interpretation that Transformer ICLL predictions use n-gram statistics with learned reweighting.

  8. Knowl 8 — Static n-gram attention heads retrieve representations after matching contexts

    model/method

    The paper defines an n-gram head (NGH) that uniformly pools hidden representations at prior positions whose following contexts match the current n-symbol context, then combines that pooled representation with the current hidden state through learned linear maps. For token positions ii and jj, define Aij(n)A^{(n)}_{ij} to be zero unless j<ij<i and xi−k=xj−k−1x_{i-k}=x_{j-k-1} for every k∈{1,…,n}k\in\{1,\ldots,n\}; among positions satisfying those conditions, Aij(n)A^{(n)}_{ij} is uniform and sums to one. If no position matches, the pooled contribution is zero. For a hidden width dd, the head output is

    NGH⁡(n)(h)i=W1hi+W2∑jAij(n)hj,\operatorname{NGH}^{(n)}(h)_i = W_1 h_i + W_2\sum_j A^{(n)}_{ij}h_j,

    where hi∈Rdh_i\in\mathbb{R}^d is the current hidden representation and W1,W2∈Rd×dW_1,W_2\in\mathbb{R}^{d\times d} are learned matrices. One head has 2d22d^2 parameters and can be inserted as a standalone layer. The paper also notes that a trie can store and query in-context n-grams, allowing recurrent-form models to use these heads with little inference overhead.

  9. Knowl 9 — Adding n-gram heads brings RetNet and GLA to Transformer-level ICLL

    empirical result

    On REGBENCH with 2,500 training instances, adding static n-gram heads improves both RetNet and GLA; each entry below is TVD (lower is better) and greedy accuracy (higher is better). A single unigram-context head improves GLA substantially and RetNet modestly, while adding higher-order heads produces large further gains. The three-head hybrids exceed the Transformer baseline in accuracy, and approach or improve its TVD.

    RetNet: baseline 0.392 / 0.800; with head order 1, 0.310 / 0.814; orders 1 and 2, 0.229 / 0.925; orders 1, 2, and 3, 0.217 / 0.940. GLA: baseline 0.624 / 0.526; with order 1, 0.302 / 0.819; orders 1 and 2, 0.211 / 0.929; orders 1, 2, and 3, 0.207 / 0.946. The Transformer baseline is 0.203 / 0.926.

  10. Knowl 10 — N-gram heads reduce perplexity in 340M-parameter language models

    empirical result

    The authors trained equal-sized 340M-parameter language models on 7 billion SlimPajama tokens and compared test perplexity with and without n-gram-head blocks. Adding heads improves all three reported base architectures, with the largest gain for the Transformer. RetNet falls from 16.55 to 15.86 with order-1, -2, and -3 head blocks (4.2%); GLA falls from 15.65 to 15.54 with three order-1 blocks (0.7%), or to 15.24 with order-1, -2, and -3 blocks (2.6%); and the Llama Transformer falls from 16.96 to 15.82 with order-1, -2, and -3 blocks (6.7%), an absolute reduction of 1.14 perplexity points. The comparison places one head bundle after the second layer and another before the output, replacing the corresponding original layers. Using multiple n-gram orders matters: the reported all-order configuration reduces perplexity four times more than using only order-1 layers.

Coverage note — Supplementary implementation details for probing, baseline HMM fitting, and hyperparameter searches are omitted because they support rather than add to the paper’s main task definition, findings, and n-gram-head contribution.

References

  1. 1.Akyurek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. What learning algorithm is in-context learning? Investigations with linear models. In Proceedings of the International Conference on Learning Representations, 2023.
  2. 2.Angluin, D. Identifying languages from stochastic examples. Yale University. Department of Computer Science, 1988.
  3. 3.Arora, S., Eyuboglu, S., Timalsina, A., Johnson, I., Poli, M., Zou, J., Rudra, A., and Re, C. Zoology: Measuring and improving recall in efficient language models. ArXiv preprint, abs/2312.04927, 2023.
  4. 4.Belinkov, Y. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1), 2022.
  5. 5.Bhattamishra, S., Ahuja, K., and Goyal, N. On the ability and limitations of Transformers to recognize formal languages. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2020.
  6. 6.Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viegas, F., and Wattenberg, M. An interpretability illusion for BERT. ArXiv preprint, abs/2104.07143, 2021.
  7. 7.Chan, S., Santoro, A., Lampinen, A., Wang, J., Singh, A., Richemond, P., McClelland, J., and Hill, F. Data distributional properties drive emergent in-context learning in Transformers. Advances in Neural Information Processing Systems, 2022.
  8. 8.Chen, S. F. and Goodman, J. An empirical study of smoothing techniques for language modeling. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Santa Cruz, California, USA, 1996.
  9. 9.Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F. Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers. In Findings of the Association for Computational Linguistics, 2023.
  10. 10.Dempster, A. P., Laird, N. M., and Rubin, D. B. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1), 1977.
  11. 11.Drozdov, A., Scharli, N., Akyurek, E., Scales, N., Song, X., Chen, X., Bousquet, O., and Zhou, D. Compositional semantic parsing with large language models. In Proceedings of the International Conference on Learning Representations, 2023.
  12. 12.Dupont, P., Denis, F., and Esposito, Y. Links between probabilistic automata and hidden Markov models: Probability distributions, learning models and induction algorithms. Pattern Recognition, 38(9), 2005.
  13. 13.Elman, J. L. Finding structure in time. Cognitive science, 14(2), 1990.
  14. 14.Finlayson, M., Richardson, K., Sabharwal, A., and Clark, P. What makes instruction learning hard? An investigation and a new challenge in a synthetic environment. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2022.
  15. 15.Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Re, C. Hungry hungry hippos: Towards language modeling with state space models. In Proceedings of the International Conference on Learning Representations, 2023.
  16. 16.Garg, S., Tsipras, D., Liang, P. S., and Valiant, G. What can Transformers learn in-context? A case study of simple function classes. Advances in Neural Information Processing Systems, 2022.
  17. 17.Gers, F. A. and Schmidhuber, E. LSTM recurrent networks learn simple context-free and context-sensitive languages. IEEE Transactions on Neural Networks, 12(6), 2001.
  18. 18.Gold, E. M. Language identification in the limit. Information and Control, 10(5), 1967.
  19. 19.Gu, A. and Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. ArXiv preprint, abs/2312.00752, 2023.
  20. 20.Gu, A., Goel, K., Gupta, A., and Re, C. On the parameterization and initialization of diagonal state space models. Advances in Neural Information Processing Systems, 35, 2022a.
  21. 21.Gu, A., Goel, K., and Re, C. Efficiently modeling long sequences with structured state spaces. In Proceedings of the International Conference on Learning Representations, 2022b.
  22. 22.Hahn, M. and Goyal, N. A theory of emergent in-context learning as implicit structure induction. ArXiv preprint, abs/2303.07971, 2023.
  23. 23.Hewitt, J., Hahn, M., Ganguli, S., Liang, P., and Manning, C. D. RNNs can generate bounded hierarchical languages with optimal memory. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2020.
  24. 24.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation, 9(8), 1997.
  25. 25.Hopcroft, J. An n log n algorithm for minimizing states in a finite automaton. In Theory of Mchines and Computations. Elsevier, 1971.
  26. 26.Hua, W., Dai, Z., Liu, H., and Le, Q. V. Transformer quality in linear time. In Procedings of the International Conference on Machine Learning, 2022.
  27. 27.Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. Transformers are RNNs: Fast autoregressive Transformers with linear attention. In Proceedings of the International Conference on Machine Learning, 2020.
  28. 28.Lee, I., Jiang, N., and Berg-Kirkpatrick, T. Exploring the relationship between model architecture and in-context learning ability. ArXiv preprint, abs/2310.08049, 2023.
  29. 29.Mehta, H., Gupta, A., Cutkosky, A., and Neyshabur, B. Long range language modeling via gated state spaces. ArXiv preprint, abs/2206.13947, 2022.
  30. 30.Merrill, W. On the linguistic capacity of real-time counter automata. ArXiv preprint, abs/2004.06866, 2020.
  31. 31.Merrill, W. and Sabharwal, A. The parallelism tradeoff: Limitations of log-precision transformers. Transactions of the Association for Computational Linguistics, 11, 2023.
  32. 32.Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Rethinking the role of demonstrations: What makes in-context learning work? ArXiv preprint, abs/2202.12837, 2022.
  33. 33.Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al. In-context learning and induction heads. ArXiv preprint, abs/2209.11895, 2022.
  34. 34.Pauls, A. and Klein, D. Faster and smaller n-gram language models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 2011.
  35. 35.Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al. RWKV: Reinventing RNNs for the transformer era. ArXiv preprint, abs/2305.13048, 2023.
  36. 36.Pitt, L. Probabilistic inductive inference. Journal of the ACM (JACM), 36(2), 1989.
  37. 37.Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., and Re, C. Hyena hierarchy: Towards larger convolutional language models. ArXiv preprint, abs/2302.10866, 2023.
  38. 38.Qin, Z., Han, X., Sun, W., Li, D., Kong, L., Barnes, N., and Zhong, Y. The devil in linear transformer. ArXiv preprint, abs/2210.10340, 2022.
  39. 39.Rabiner, L. R. A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE, 77(2), 1989.
  40. 40.Shazeer, N. GLU variants improve Transformer. ArXiv preprint, abs/2002.05202, 2020.
  41. 41.Shi, X., Padhi, I., and Knight, K. Does string-based neural MT learn source syntax? In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Austin, Texas, 2016.
  42. 42.Shin, R. and Van Durme, B. Few-shot semantic parsing with language models trained on code. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, United States, 2022.
  43. 43.Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama, 2023. URL https://huggingface.co/datasets/cerebras/SlimPajama-627B.
  44. 44.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  45. 45.Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F. Retentive network: A successor to transformer for large language models. ArXiv preprint, abs/2307.08621, 2023.
  46. 46.Suzgun, M., Gehrmann, S., Belinkov, Y., and Shieber, S. M. Memory-augmented recurrent neural networks can learn generalized Dyck languages. ArXiv preprint, abs/1911.03329, 2019.
  47. 47.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. ArXiv preprint, abs/2302.13971, 2023.
  48. 48.Valiant, L. G. A theory of the learnable. Communications of the ACM, 27(11):1134–1142, 1984.
  49. 49.van der Poel, S., Lambert, D., Kostyszyn, K., Gao, T., Verma, R., Andersen, D., Chau, J., Peterson, E., Clair, C. S., Fodor, P., et al. MLRegTest: A benchmark for the machine learning of regular languages. ArXiv preprint, abs/2304.07687, 2023.
  50. 50.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems, 2017.
  51. 51.von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. Transformers learn in-context by gradient descent. In Proceedings of the International Conference on Machine Learning, 2023a.
  52. 52.von Oswald, J., Niklasson, E., Schlegel, M., Kobayashi, S., Zucchet, N., Scherrer, N., Miller, N., Sandler, M., Vladymyrov, M., Pascanu, R., et al. Uncovering mesa-optimization algorithms in Transformers. ArXiv preprint, abs/2309.05858, 2023b.
  53. 53.Wen, K., Li, Y., Liu, B., and Risteski, A. Transformers are uninterpretable with myopic methods: A case study with bounded Dyck grammars. ArXiv preprint, abs/2312.01429, 2023. URL https://arxiv.org/abs/2312.01429.
  54. 54.Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. An explanation of in-context learning as implicit Bayesian inference. In Proceedings of the International Conference on Learning Representations, 2022.
  55. 55.Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y. Gated linear attention Transformers with hardware-efficient training. ArXiv preprint, abs/2312.06635, 2023.
  56. 56.Zhai, S., Talbott, W., Srivastava, N., Huang, C., Goh, H., Zhang, R., and Susskind, J. An attention free Transformer. ArXiv preprint, abs/2105.14103, 2021.
  57. 57.Zhang, B. and Sennrich, R. Root mean square layer normalization. Advances in Neural Information Processing Systems, 2019.

Citation

MLA
Akyürek, E., et al. “In-Context Language Learning: Architectures and Algorithms”. arXiv, 2024, https://doi.org/10.48550/arxiv.2401.12973.
APA
Akyürek, E., Wang, B., Kim, Y., & Andreas, J. (2024). In-Context Language Learning: Architectures and Algorithms. arXiv. https://doi.org/10.48550/arxiv.2401.12973
Chicago
Akyürek, E., B. Wang, Y. Kim, and J. Andreas. 2024. “In-Context Language Learning: Architectures and Algorithms”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2401.12973.
Harvard
Akyürek, E. et al. (2024) “In-Context Language Learning: Architectures and Algorithms”. arXiv. Available at: https://doi.org/10.48550/arxiv.2401.12973.
Vancouver
1. Akyürek E, Wang B, Kim Y, Andreas J (2024) In-Context Language Learning: Architectures and Algorithms. https://doi.org/10.48550/arxiv.2401.12973

BibTeX

@misc{https://doi.org/10.48550/arxiv.2401.12973,
  doi = {10.48550/ARXIV.2401.12973},
  url = {https://arxiv.org/abs/2401.12973},
  author = {Akyürek, Ekin and Wang, Bailin and Kim, Yoon and Andreas, Jacob},
  keywords = {Computation and Language (cs.CL), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {In-Context Language Learning: Architectures and Algorithms},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/