Memorizing Transformers

Yuhuai WuMarkus Norman RabeDeLesley HutchinsChristian Szegedy

article2022ICLR228 citations

Extends Transformer language models with an approximate k-nearest-neighbors memory over past key-value representations, enabling models to scale test-time context up to 262,000 tokens and immediately apply new knowledge without updating weights.

Listen

Modern language models struggle to process long documents because their attention context is typically constrained to short sequences. Consequently, acquiring new information normally requires expensive model retraining or weight updates, making it difficult for models to reference distant concepts such as earlier chapters in books, function declarations in software repositories, or previously proven lemmas in mathematics.

The article demonstrates an architectural extension that enables language models to memorize and access representations of past inputs at runtime. The objective is to evaluate whether augmenting standard transformers with approximate nearest-neighbor retrieval over a large, non-differentiable external memory improves language modeling performance on long-context tasks without requiring prohibitive computational overhead.

To evaluate this capability, the authors integrated an approximate nearest-neighbor search mechanism into a single layer of a standard decoder-only transformer. This layer queries a memory cache of past key-value pairs that are retained across sequential chunks of long documents without backpropagating gradients into the memory store. The researchers tested the system across five long-text domains—generic web text (C4), mathematics papers (arXiv), books (PG-19), open-source code (GitHub), and formal proofs (Isabelle)—across memory capacities ranging from 1,536 tokens to 262,000 tokens and model sizes up to 8 billion parameters.

The evaluation yielded several key findings. First, adding external memory consistently reduced language perplexity across all datasets and architectures, with performance steadily improving as memory size increased up to 262,000 tokens. Second, an 8-thousand-token memory allowed a model to match the perplexity of a standard baseline with roughly five times more parameters, offering major parameter efficiency. Third, existing pretrained transformers could be fine-tuned to use external memory within just 20,000 steps—representing 4% of initial pretraining time—recovering 85% of the performance gap compared to models built with memory from scratch. Qualitative inspections confirmed that the models primarily used the memory to look up distant definitions, citations, and variable names, correctly locating referenced lemmas in formal proof tests 80% of the time.

These findings indicate that non-differentiable memory mechanisms can dramatically expand context windows with modest computational costs, providing a practical alternative to training increasingly massive models. Memory lookups increase single-step training time only moderately (from 0.20 seconds to 0.25 seconds for an 8-thousand-token cache on tested hardware), while enabling immediate knowledge acquisition without changing model weights. Additionally, storing context in external memory allows sensitive or proprietary text to be cleared after processing, presenting privacy and compliance advantages over permanently embedding information into static neural network weights.

Organizations developing long-context applications should consider adopting nearest-neighbor memory augmentation and fine-tuning their existing models rather than retraining from scratch. While the findings provide high confidence regarding perplexity improvements on structured, long-form datasets, the memory exhibits diminishing returns once capacity exceeds average document length, and early-stage training with large memories can encounter stability challenges. Teams implementing this architecture should stabilize initial training with smaller memory caches before expanding capacity, and future efforts should investigate how to scale these retrieval mechanisms over permanent external knowledge repositories.

arXiv: 2203.08913
  • Paper: Generalization through Memorization: Nearest Neighbor Language Models, Urvashi Khandelwal et al. (2020). Introduces kNN-LM and explicit non-parametric nearest-neighbor memory lookups over key-value representations, establishing the foundational paradigm directly built upon and generalized in Memorizing Transformers.
  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). Provides the seminal architecture for differentiable memory-augmented networks and multi-hop attention over explicit key-value storage that underpins external memory mechanisms in language modeling.
  • Paper: Memory Networks, Jason Weston et al. (2014). Pioneered the core conceptual framework of augmenting neural models with an explicit, compartmentalized external memory to overcome fixed hidden state capacity.
  • Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). Establishes modern pre-training and dense retrieval mechanics over large external document indexes to augment parametric language model representations.
  • Paper: Pointer Sentinel Mixture Models, Stephen Merity et al. (2016). Introduces pointer-based continuous memory lookup over past representations to dynamically retrieve and reproduce long-context information during language modeling.
Cover for Memorizing Transformers

Abstract

Language models typically need to be trained or finetuned in order to acquire new knowledge, which involves updating their weights. We instead envision language models that can simply read and memorize new data at inference time, thus acquiring new knowledge immediately. In this work, we extend language models with the ability to memorize the internal representations of past inputs. We demonstrate that an approximate kNN lookup into a non-differentiable memory of recent (key, value) pairs improves language modeling across various benchmarks and tasks, including generic webtext (C4), math papers (arXiv), books (PG-19), code (Github), as well as formal theorems (Isabelle). We show that the performance steadily improves when we increase the size of memory up to 262K tokens. On benchmarks including code and mathematics, we find that the model is capable of making use of newly defined functions and theorems during test time.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 kkNN-augmented Attention Layer
  • 3.2 Distributional Shift
  • 3.3 Approximate kkNN
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Experimental Method
  • 4.3 Effect of external memory
  • 4.4 Scaling to larger models
  • 4.5 Finetuning on Larger Memories
  • 4.6 Information Retrieval Patterns
  • 5 Conclusion
  • References
  • A Length of inputs
  • A.1 Ablation studies
  • B What does the model retrieve from memory?
  • B.1 More Retrieving examples in formal theorem proving corpus

Knowls

  1. Knowl 1 — kNN-Augmented Attention Architecture for Memorizing Transformers

    model/method

    The Memorizing Transformer is an autoregressive decoder-only language model that augments standard attention with an approximate kk-nearest-neighbor (kNNk\text{NN}) lookup into a large, non-differentiable external memory of past key-value representations.

    Documents are partitioned into sequential subsequences of fixed length (e.g., 512 tokens) and processed sequentially without shuffling. For standard layers, dense causal self-attention is applied over the local context, optionally prepended with cached keys and values from the immediately preceding training step (a Transformer-XL style cache). In one designated layer (typically in the upper-middle of the transformer stack, such as layer 9 of 12), a kNNk\text{NN}-augmented attention mechanism is introduced.

    For each query vector in the local sequence, an approximate kNNk\text{NN} search retrieves the top-kk most similar key-value pairs (Kmem,Vmem)(K_{\text{mem}}, V_{\text{mem}}) from an external memory cache containing up to MM previously computed pairs (where MM can range from 1,536 to 262,144 tokens). Standard dot-product attention followed by softmax is computed over the retrieved kk key-value pairs to produce a memory context vector VmV_m. Simultaneously, dense attention over the local context produces a local context vector VcV_c. The output representations are then combined via a learned per-head gate and passed through the subsequent feed-forward network (FFN).

  2. Knowl 2 — Learned Gating Mechanism for Combining Memory and Local Attention

    equation

    In the kNNk\text{NN}-augmented attention layer, the output vector from external memory retrieval VmV_m and the output vector from dense local context attention VcV_c are combined into the final layer attention representation VaV_a using a learned per-head gating parameter:

    g=σ(bg)g = \sigma(b_g)

    Va=Vm⊙g+Vc⊙(1−g)V_a = V_m \odot g + V_c \odot (1 - g)

    where:

    • bg∈Rb_g \in \mathbb{R} is a learned per-head scalar bias parameter.
    • σ(x)=11+e−x\sigma(x) = \frac{1}{1 + e^{-x}} denotes the standard sigmoid activation function, mapping the bias to a gating weight g∈(0,1)g \in (0, 1).
    • ⊙\odot denotes element-wise multiplication.
    • VmV_m is the attention vector resulting from softmax dot-product attention over the top-kk key-value pairs retrieved from external memory.
    • VcV_c is the attention vector resulting from standard dense attention over the local sequence (and optional local recurrence cache).

    Over the course of training, individual attention heads learn distinct biases bgb_g, with most heads specializing to attend almost exclusively to the external long-term memory.

  3. Knowl 3 — Non-Differentiable External Memory Management and Key-Query Normalization

    model/method

    The external memory stores a FIFO cache of the prior MM (key,value)(key, value) pairs computed at the kNNk\text{NN}-augmented attention layer across previous consecutive training chunks of a document. Memory is cleared at document boundaries.

    Two critical design choices govern memory management:

    1. Non-Differentiable Cache: Gradients are not backpropagated into the external memory. This allows previously computed (key,value)(key, value) vectors to be stored directly and reused without recomputing representations across the entire historical sequence at every training step, enabling scaling up to 262,144 tokens on a single accelerator device.
    2. Key and Query Normalization: Because the model parameters evolve across training steps, older cached keys become subject to distributional shift (staleness) relative to newly generated queries. To mitigate this staleness and stabilize training, L2L_2 normalization is applied to all query and key vectors prior to attention and kNNk\text{NN} index matching.
  4. Knowl 4 — Language Modeling Perplexity Improvements Across Long-Document Benchmarks

    data/table

    The addition of external memory significantly reduces token-level perplexity across diverse domains (mathematics, source code, formal proofs, web articles, and books). Models are 12-layer decoder-only transformers (~200M parameters) trained for 500k steps (except Isabelle, trained for 100k steps) with k=32k=32.

    Context Memory XL cache arXiv PG19 C4(4K+) GitHub Isabelle
    512 None None 3.29 13.71 17.20 3.05 3.09
    2048 None None 2.69 12.37 14.81 2.22 2.39
    512 None 512 2.67 12.34 15.38 2.26 2.46
    2048 None 2048 2.42 11.88 14.03 2.10 2.16
    512 1536 None 2.61 12.50 14.97 2.20 2.33
    512 8192 None 2.49 12.29 14.42 2.09 2.19
    512 8192 512 2.37 11.93 14.04 2.03 2.08
    512 65K 512 2.31 11.62 14.04 1.87 2.06
    2048 8192 2048 2.33 11.84 13.80 1.98 2.06
    2048 65K 2048 2.26 11.37 13.64 1.80 1.99

    Adding non-differentiable memory to a single layer yields perplexity gains superior to extending full dense context across all layers. For example, a context of 512 with an 8,192 memory (perplexity 2.49 on arXiv) outperforms a full dense context of 2048 without memory (perplexity 2.69).

  5. Knowl 5 — Scaling Memory Capacity up to 262K Tokens via Two-Stage Pretraining and Finetuning

    data/table

    Training directly from scratch with very large memory sizes (≥131K\ge 131\text{K}) can introduce optimization instability due to distributional shift in stored keys. A two-stage training procedure—pretraining the model for 500k steps with a moderate memory size (8K8\text{K} or 65K65\text{K}) followed by finetuning for 20k steps with an expanded memory—enables stable scaling up to 262K tokens.

    Evaluation on the arXiv Math dataset demonstrates continuous perplexity improvements up to 262K tokens:

    Context Pretrain Memory Fine-tune Memory Perplexity
    512 8192 None 2.37
    512 65K None 2.31
    512 8192 65K 2.32
    512 8192 131K 2.30
    512 8192 262K 2.26
    2048 8192 None 2.33
    2048 65K None 2.26
    2048 65K 131K 2.23
    2048 65K 262K 2.21
  6. Knowl 6 — Parameter Efficiency of Memory Augmentation Compared to Model Scaling

    empirical result

    Augmenting a transformer with external memory produces perplexity improvements comparable to scaling model parameters by more than 5×5\times.

    On the arXiv Math benchmark:

    • A 200M parameter Memorizing Transformer with an 8K external memory achieves a perplexity of ~2.42, matching the perplexity of a 1B parameter vanilla Transformer (~2.42).
    • Scaling to 1B and 8B parameter models maintains this advantage: an 8B Memorizing Transformer with 8K memory reaches ~1.91 perplexity, significantly lower than an 8B vanilla Transformer (~2.12).
    • Computational overhead on TPUv3 is modest: for a ~200M parameter model, training step time increases from 0.20s (vanilla) to 0.25s with 8K memory, and to 0.60s with 65K memory.
  7. Knowl 7 — Memory Adaptation via Finetuning of Pretrained Vanilla Transformers

    empirical result

    A standard vanilla Transformer pretrained without external memory can rapidly learn to utilize external memory through subsequent finetuning without requiring pretraining from scratch.

    When a 1B parameter vanilla Transformer pretrained on arXiv Math is finetuned with a 65K external memory layer:

    • Within 20k steps (only 4% of the original 500k pretraining duration), the finetuned model closes 85% of the perplexity performance gap relative to a 1B Memorizing Transformer trained from scratch.
    • Within 100k finetuning steps, the finetuned model fully closes the gap, reaching identical perplexity.
  8. Knowl 8 — Structural Ablations on Memory Layer Index, Neighborhood Size, and Layer Count

    empirical result

    Ablation experiments on the 12-layer Memorizing Transformer (context 512, XL cache 512, memory 8192) show the sensitivity of perplexity to architectural hyperparameters on arXiv Math:

    1. Memory Layer Placement: Inserting the kNNk\text{NN} attention layer into the middle/upper-middle of the transformer stack performs best:
      • Layer 3: 2.40 perplexity
      • Layer 6: 2.36 perplexity
      • Layer 9: 2.37 perplexity
      • Layer 12: 2.43 perplexity
    2. Number of Retrieved Neighbors (kk): Model quality is robust to the number of retrieved neighbors. Retrieving k=32k=32 neighbors yields a perplexity of 2.38, which is virtually identical to retrieving k=128k=128 (2.37) or k=256k=256 (2.37), while substantially reducing computational retrieval cost.
    3. Multiple Memory Layers: Adding more than one kNNk\text{NN}-augmented attention layer across the transformer stack does not provide meaningful performance gains over a single well-placed memory layer.
  9. Knowl 9 — Information Retrieval Patterns and Definition Grounding in Long Contexts

    empirical result

    Qualitative and token-level loss difference analysis (Δi=cross-entropy8192(xi)−cross-entropy32K(xi)\Delta_i = \text{cross-entropy}_{8192}(x_i) - \text{cross-entropy}_{32\text{K}}(x_i)) reveals that the performance gains of external memory are sparse and driven by large accuracy improvements on specific token classes across long distances:

    • Rare Identifiers: Memory is primarily used to look up rare entities such as citation keys/bibitems, function names, and variable definitions introduced thousands of tokens earlier in a document.
    • Formal Theorem Definitions: In formal mathematics (Isabelle), when predicting mathematical objects and lemma names, the kNNk\text{NN} layer directly retrieves the definition body of the corresponding lemma from earlier proof context (confirmed in 8 out of 10 sampled manual test cases).

Coverage note — Omitted routine implementation details such as the specific SentencePiece tokenizer subword construction, standard Adafactor optimizer learning rate sweep lists, and raw data pipeline URL lists.

References

  1. 1.Joshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. ETC: encoding long and structured inputs in transformers. In EMNLP, 2020.
  2. 2.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. Program synthesis with large language models. CoRR, abs/2108.07732, 2021. URL https://arxiv.org/abs/2108.07732.
  3. 3.Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The long-document transformer. CoRR, abs/2004.05150, 2020. URL https://arxiv.org/abs/2004.05150.
  4. 4.James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/google/jax.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In NeurIPS, 2020.
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021. URL https://arxiv.org/abs/2107.03374.
  7. 7.Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J. Colwell, and Adrian Weller. Rethinking attention with performers. In ICLR, 2021.
  8. 8.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.
  9. 9.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. Transformer-XL: Attentive language models beyond a fixed-length context. In ACL, 2019.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In ACL, 2019.
  11. 11.Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar. Addressing some limitations of transformers with feedback memory. arXiv preprint arXiv:2002.09402, 2020.
  12. 12.Angela Fan, Claire Gardent, Chloé Braud, and Antoine Bordes. Augmenting transformers with KNN-based composite memory for dialog. Transactions of the Association for Computational Linguistics, 9:82–99, 2021.
  13. 13.Edouard Grave, Armand Joulin, and Nicolas Usunier. Improving neural language models with a continuous cache. In ICLR, 2017.
  14. 14.Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. Accelerating large-scale inference with anisotropic vector quantization. In ICML, 2020.
  15. 15.Ankit Gupta, Guy Dar, Shaya Goodman, David Ciprut, and Jonathan Berant. Memory-efficient transformers via top-k attention. CoRR, abs/2106.06899, 2021. URL https://arxiv.org/abs/2106.06899.
  16. 16.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. Retrieval augmented language model pre-training. In ICML, 2020.
  17. 17.Christopher Hahn, Frederik Schmitt, Jens U. Kreber, Markus Norman Rabe, and Bernd Finkbeiner. Teaching temporal logics to neural networks. In ICLR, 2021.
  18. 18.Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee. Flax: A neural network library and ecosystem for JAX, 2020. URL http://github.com/google/flax.
  19. 19.Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. Query-key normalization for transformers. In EMNLP, 2020.
  20. 20.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 2021.
  21. 21.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. Generalization through memorization: Nearest neighbor language models. In ICLR, 2020.
  22. 22.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. In ICLR, 2020.
  23. 23.Taku Kudo and John Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In EMNLP, 2018.
  24. 24.Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. Large memory layers with product keys. In NeurIPS, 2019.
  25. 25.Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida Wang, and Luke Zettlemoyer. Pre-training via paraphrasing. In NeurIPS, 2020a.
  26. 26.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In NeurIPS, 2020b.
  27. 27.Wenda Li, Lei Yu, Yuhuai Wu, and Lawrence C. Paulson. Isarstep: a benchmark for high-level mathematical reasoning. In ICLR, 2021.
  28. 28.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals. Competition-level code generation with alphacode. DeepMind, 2022.
  29. 29.Stanislas Polu and Ilya Sutskever. Generative language modeling for automated theorem proving. CoRR, abs/2009.03393, 2020. URL https://arxiv.org/abs/2009.03393.
  30. 30.Markus Norman Rabe, Dennis Lee, Kshitij Bansal, and Christian Szegedy. Mathematical reasoning via self-supervised skip-tree training. In ICLR, 2021.
  31. 31.Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. In ICLR, 2020.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 2020.
  33. 33.Hongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang, Jure Leskovec, Dale Schuurmans, and Bo Dai. Combiner: Full attention transformer with sparse computation cost. CoRR, abs/2107.05768, 2021. URL https://arxiv.org/abs/2107.05768.
  34. 34.Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. Efficient content-based sparse attention with routing transformers. Transactions of the Association for Computational Linguistics, 9:53–68, 2021.
  35. 35.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In ICML, 2018.
  36. 36.Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. Augmenting self-attention with persistent memory. arXiv preprint arXiv:1907.01470, 2019.
  37. 37.Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur Szlam, Jason Weston, and Angela Fan. Not all memories are created equal: Learning to forget by expiring. In ICML, 2021.
  38. 38.Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, and Mohit Iyyer. Do long-range language models actually use long-range context? In EMNLP, 2021.
  39. 39.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. arXiv preprint arXiv:2009.06732, 2020.
  40. 40.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long range arena: A benchmark for efficient transformers. In ICLR, 2021.
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  42. 42.Qingxiang Wang, Chad Brown, Cezary Kaliszyk, and Josef Urban. Exploration of neural machine translation in autoformalization of mathematics in mizar. In International Conference on Certified Programs and Proofs, 2020a.
  43. 43.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020b.
  44. 44.Xun Wang, Haozhi Zhang, Weilin Huang, and Matthew R. Scott. Cross-batch memory for embedding learning. In CVPR, 2020c.
  45. 45.Ronald J. Williams and Jing Peng. An efficient gradient-based algorithm for on-line training of recurrent network trajectories. Neural Computation, 1990.
  46. 46.Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. Adaptive semiparametric language models. ACL, 9:362–373, 2021.
  47. 47.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big bird: Transformers for longer sequences. In NeurIPS, 2020.
  48. 48.Yury Zemlyanskiy, Joshua Ainslie, Michiel de Jong, Philip Pham, Ilya Eckstein, and Fei Sha. Readtwice: Reading very large documents with memories. In ACL: Human Language Technologies, 2021.
  49. 49.Zhenhai Zhu and Radu Soricut. H-transformer-1d: Fast one-dimensional hierarchical attention for sequences. In ACL, 2021.

Citation

MLA
Wu, Y., et al. “Memorizing Transformers”. arXiv, 2022, http://arxiv.org/abs/2203.08913v1.
APA
Wu, Y., Rabe, M. N., Hutchins, D., & Szegedy, C. (2022). Memorizing Transformers. arXiv. http://arxiv.org/abs/2203.08913v1
Chicago
Wu, Y., M. N. Rabe, D. Hutchins, and C. Szegedy. 2022. “Memorizing Transformers”. arXiv. http://arxiv.org/abs/2203.08913v1.
Harvard
Wu, Y. et al. (2022) “Memorizing Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.08913v1.
Vancouver
1. Wu Y, Rabe MN, Hutchins D, Szegedy C (2022) Memorizing Transformers. arXiv

BibTeX

@article{wu2022memorizing,
  title = {Memorizing Transformers},
  author = {Wu, Yuhuai and Rabe, Markus N. and Hutchins, DeLesley and Szegedy, Christian},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.08913v1},
  eprint = {2203.08913}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission