Random-Access Infinite Context Length for Transformers
Amirkeivan MohtashamiMartin Jaggi
Introduces landmark attention, a mechanism that uses dedicated tokens to retrieve relevant context blocks directly within attention, reducing memory and computation to allow fine-tuned models like LLaMA 7B to scale inference past 32k tokens.
Modern large language models struggle to process long documents because their standard attention mechanism requires computing resources and memory that grow quadratically with input length. Existing remedies—such as compressing past text through recurrent memory or searching external databases with separate retrieval algorithms—either lose crucial fine-grained details or fail to integrate cleanly with the core model architecture. There is a critical operational need for scalable methods that allow models to access extensive context windows without sacrificing precision or inflating infrastructure costs.
The article demonstrates a new architecture, called landmark attention, which enables transformer models to access arbitrary context lengths at inference time while maintaining full random-access flexibility to individual tokens. To evaluate this approach, the researchers trained language models from scratch on standard benchmarks, including the PG-19 books dataset and arXiv math papers, and fine-tuned a pre-trained LLaMA 7B model. The approach divides long inputs into fixed-size blocks (e.g., 50 tokens) and inserts a special landmark token at the end of each block. The attention mechanism is trained using a grouped softmax method, allowing the landmark token to act as a representative gate: if a token requires information from a past block, it attends to the block’s landmark token, seamlessly retrieving only the most relevant blocks into active memory.
The evaluation yielded several key findings. First, landmark attention matches the language modeling perplexity of recurrent architectures like Transformer-XL while processing substantially fewer tokens per step, reducing computational operations by roughly a factor of the block size (approximately 50x in tested configurations). Second, models trained on short sequence lengths (512 tokens) successfully extrapolate to inputs of 4,096 tokens and beyond without retraining. Third, fine-tuning LLaMA 7B on sequence lengths of only 512 tokens enabled the model to process contexts exceeding 32,000 tokens, achieving a 98% success rate on targeted information-retrieval tests across 50 trials, matching the operational context reach of leading proprietary models like GPT-4. Finally, by caching only landmark tokens on the accelerator and offloading non-active memory blocks to system RAM, the method dramatically lowers memory requirements.
These findings indicate that organizations can significantly cut compute hardware costs and expand context windows without retraining massive foundation models from scratch. The native attention-gating mechanism also enhances operational interpretability by revealing exactly which past text blocks the model accessed to generate an output. However, offloading memory blocks between processor and system memory introduces data transfer overhead, and the current positional indexing scheme relies on approximations for distant tokens. Technical teams should consider piloting landmark attention on existing open-source models for long-form analysis, with next steps focused on integrating the method with optimized hardware kernels (such as FlashAttention) and testing complex, domain-specific retrieval workloads.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). It introduces the segment-level recurrence and memory caching baseline that Landmark Attention directly compares against and aims to improve upon for long-context processing.
- Paper: Memorizing Transformers, Yuhuai Wu et al. (2022). It establishes external memory retrieval mechanisms for transformers that Landmark Attention seeks to replace with native, attention-integrated block selection.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). It provides foundational principles for sparse and block-based attention mechanisms that combine local context with global landmark-style tokens.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). It examines sequence length extrapolation challenges in transformers, setting up the positional and scaling issues addressed by landmark-based attention.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). It details the GPU memory hierarchy and IO-aware block tiling fundamentals essential for understanding how landmark tokens interface with system hardware.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). It introduces the foundational Transformer self-attention architecture whose quadratic memory constraints necessitate landmark token compression.
- Paper: Ring Attention with Blockwise Transformers for Near-Infinite Context, Hao Liu et al. (2024). It expands on blockwise long-context processing by distributing block attention over a ring of devices to achieve near-infinite sequence lengths without token selection.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). It investigates position-dependent retrieval failure modes in extended-context models, providing critical evaluation criteria for landmark-based retrieval systems.
- Paper: Efficient Streaming Language Models with Attention Sinks, Guangxuan Xiao et al. (2023). It discovers the attention sink phenomenon, providing a complementary mechanism for maintaining streaming context stability alongside landmark-based block selection.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). It provides a comprehensive multilingual benchmark to rigorously evaluate models extended to long context windows via methods like Landmark Attention.
- Paper: Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers, Sotiris Anagnostidis et al. (2023). It explores dynamic context pruning inside autoregressive transformers as an alternative strategy to blockwise landmark retrieval for reducing memory overhead.
- Paper: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention, Jingyang Yuan et al. (2025). It advances block-compressed sparse attention by training hierarchical coarse-to-fine token selection natively on hardware.
- Paper: XAttention: Block Sparse Attention with Antidiagonal Scoring, Ruyi Xu et al. (2025). It presents a training-free block-sparse attention alternative using antidiagonal scoring to dynamically identify relevant context blocks.
