Memorizing Transformers
Yuhuai WuMarkus Norman RabeDeLesley HutchinsChristian Szegedy
Extends Transformer language models with an approximate k-nearest-neighbors memory over past key-value representations, enabling models to scale test-time context up to 262,000 tokens and immediately apply new knowledge without updating weights.
Modern language models struggle to process long documents because their attention context is typically constrained to short sequences. Consequently, acquiring new information normally requires expensive model retraining or weight updates, making it difficult for models to reference distant concepts such as earlier chapters in books, function declarations in software repositories, or previously proven lemmas in mathematics.
The article demonstrates an architectural extension that enables language models to memorize and access representations of past inputs at runtime. The objective is to evaluate whether augmenting standard transformers with approximate nearest-neighbor retrieval over a large, non-differentiable external memory improves language modeling performance on long-context tasks without requiring prohibitive computational overhead.
To evaluate this capability, the authors integrated an approximate nearest-neighbor search mechanism into a single layer of a standard decoder-only transformer. This layer queries a memory cache of past key-value pairs that are retained across sequential chunks of long documents without backpropagating gradients into the memory store. The researchers tested the system across five long-text domains—generic web text (C4), mathematics papers (arXiv), books (PG-19), open-source code (GitHub), and formal proofs (Isabelle)—across memory capacities ranging from 1,536 tokens to 262,000 tokens and model sizes up to 8 billion parameters.
The evaluation yielded several key findings. First, adding external memory consistently reduced language perplexity across all datasets and architectures, with performance steadily improving as memory size increased up to 262,000 tokens. Second, an 8-thousand-token memory allowed a model to match the perplexity of a standard baseline with roughly five times more parameters, offering major parameter efficiency. Third, existing pretrained transformers could be fine-tuned to use external memory within just 20,000 steps—representing 4% of initial pretraining time—recovering 85% of the performance gap compared to models built with memory from scratch. Qualitative inspections confirmed that the models primarily used the memory to look up distant definitions, citations, and variable names, correctly locating referenced lemmas in formal proof tests 80% of the time.
These findings indicate that non-differentiable memory mechanisms can dramatically expand context windows with modest computational costs, providing a practical alternative to training increasingly massive models. Memory lookups increase single-step training time only moderately (from 0.20 seconds to 0.25 seconds for an 8-thousand-token cache on tested hardware), while enabling immediate knowledge acquisition without changing model weights. Additionally, storing context in external memory allows sensitive or proprietary text to be cleared after processing, presenting privacy and compliance advantages over permanently embedding information into static neural network weights.
Organizations developing long-context applications should consider adopting nearest-neighbor memory augmentation and fine-tuning their existing models rather than retraining from scratch. While the findings provide high confidence regarding perplexity improvements on structured, long-form datasets, the memory exhibits diminishing returns once capacity exceeds average document length, and early-stage training with large memories can encounter stability challenges. Teams implementing this architecture should stabilize initial training with smaller memory caches before expanding capacity, and future efforts should investigate how to scale these retrieval mechanisms over permanent external knowledge repositories.
- Paper: Generalization through Memorization: Nearest Neighbor Language Models, Urvashi Khandelwal et al. (2020). Introduces kNN-LM and explicit non-parametric nearest-neighbor memory lookups over key-value representations, establishing the foundational paradigm directly built upon and generalized in Memorizing Transformers.
- Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). Provides the seminal architecture for differentiable memory-augmented networks and multi-hop attention over explicit key-value storage that underpins external memory mechanisms in language modeling.
- Paper: Memory Networks, Jason Weston et al. (2014). Pioneered the core conceptual framework of augmenting neural models with an explicit, compartmentalized external memory to overcome fixed hidden state capacity.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). Establishes modern pre-training and dense retrieval mechanics over large external document indexes to augment parametric language model representations.
- Paper: Pointer Sentinel Mixture Models, Stephen Merity et al. (2016). Introduces pointer-based continuous memory lookup over past representations to dynamically retrieve and reproduce long-context information during language modeling.
- Paper: Titans: Learning to Memorize at Test Time, Ali Behrouz et al. (2024). Extends test-time memory augmentation by formulating long-term memory as a neural meta-model updated online during inference across multi-million-token contexts.
- Paper: Test-time regression: a unifying framework for designing sequence models with associative memory, Ke Alexander Wang et al. (2025). Unifies test-time memorization and associative recall across attention mechanisms and recurrent sequence layers through a formal regression framework.
- Paper: Memory Caching: RNNs with Growing Memory, Ali Behrouz et al.. Generalizes segment-level memory caching and growing context retrieval to recurrent neural network architectures.
- Paper: $δ$-mem: Efficient Online Memory for Large Language Models, Jingdi Lei et al. (2026). Develops a compact online state memory add-on for frozen language models to maintain history across long interactions without full context growth.
- Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). Treats dynamic knowledge acquisition and long-horizon memorization as a dedicated modular model operating alongside frozen executive language models.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). Investigates training-free, in-context retrieval mechanisms as an alternative to non-differentiable external memory layers for large language models.
- Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). Explores end-to-end test-time gradient learning to compress long context into model weights rather than indexing key-value memory stores.
