Random-Access Infinite Context Length for Transformers

Amirkeivan MohtashamiMartin Jaggi

article2023NeurIPS207 citations

Introduces landmark attention, a mechanism that uses dedicated tokens to retrieve relevant context blocks directly within attention, reducing memory and computation to allow fine-tuned models like LLaMA 7B to scale inference past 32k tokens.

Listen

Modern large language models struggle to process long documents because their standard attention mechanism requires computing resources and memory that grow quadratically with input length. Existing remedies—such as compressing past text through recurrent memory or searching external databases with separate retrieval algorithms—either lose crucial fine-grained details or fail to integrate cleanly with the core model architecture. There is a critical operational need for scalable methods that allow models to access extensive context windows without sacrificing precision or inflating infrastructure costs.

The article demonstrates a new architecture, called landmark attention, which enables transformer models to access arbitrary context lengths at inference time while maintaining full random-access flexibility to individual tokens. To evaluate this approach, the researchers trained language models from scratch on standard benchmarks, including the PG-19 books dataset and arXiv math papers, and fine-tuned a pre-trained LLaMA 7B model. The approach divides long inputs into fixed-size blocks (e.g., 50 tokens) and inserts a special landmark token at the end of each block. The attention mechanism is trained using a grouped softmax method, allowing the landmark token to act as a representative gate: if a token requires information from a past block, it attends to the block’s landmark token, seamlessly retrieving only the most relevant blocks into active memory.

The evaluation yielded several key findings. First, landmark attention matches the language modeling perplexity of recurrent architectures like Transformer-XL while processing substantially fewer tokens per step, reducing computational operations by roughly a factor of the block size (approximately 50x in tested configurations). Second, models trained on short sequence lengths (512 tokens) successfully extrapolate to inputs of 4,096 tokens and beyond without retraining. Third, fine-tuning LLaMA 7B on sequence lengths of only 512 tokens enabled the model to process contexts exceeding 32,000 tokens, achieving a 98% success rate on targeted information-retrieval tests across 50 trials, matching the operational context reach of leading proprietary models like GPT-4. Finally, by caching only landmark tokens on the accelerator and offloading non-active memory blocks to system RAM, the method dramatically lowers memory requirements.

These findings indicate that organizations can significantly cut compute hardware costs and expand context windows without retraining massive foundation models from scratch. The native attention-gating mechanism also enhances operational interpretability by revealing exactly which past text blocks the model accessed to generate an output. However, offloading memory blocks between processor and system memory introduces data transfer overhead, and the current positional indexing scheme relies on approximations for distant tokens. Technical teams should consider piloting landmark attention on existing open-source models for long-form analysis, with next steps focused on integrating the method with optimized hardware kernels (such as FlashAttention) and testing complex, domain-specific retrieval workloads.

arXiv: 2305.16300
Cover for Random-Access Infinite Context Length for Transformers

Abstract

While Transformers have shown remarkable success in natural language processing, their attention mechanism's large memory requirements have limited their ability to handle longer contexts. Prior approaches, such as recurrent memory or retrieval-based augmentation, have either compromised the random-access flexibility of attention (i.e., the capability to select any token in the entire context) or relied on separate mechanisms for relevant context retrieval, which may not be compatible with the model's attention. In this paper, we present a novel approach that allows access to the complete context while retaining random-access flexibility, closely resembling running attention on the entire context. Our method uses a landmark token to represent each block of the input and trains the attention to use it for selecting relevant blocks, enabling retrieval of blocks directly through the attention mechanism instead of by relying on a separate mechanism. Our approach seamlessly integrates with specialized data structures and the system's memory hierarchy, enabling processing of arbitrarily long context lengths. We demonstrate that our method can obtain comparable performance with Transformer-XL while significantly reducing the number of retrieved tokens in each step. Finally, we show that fine-tuning LLaMA 7B with our method successfully extends its context length capacity to over 32k tokens, allowing for inference at the context lengths of GPT-4. We release the implementation of landmark attention and the code to reproduce our experiments at https://github.com/epfml/landmark-attention/.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Training Landmark Tokens
  • 3.2 Inference
  • 3.2.1 Positional Encoding
  • 3.3 Memory & Computation
  • 4 Experiments
  • 4.1 Language Modeling
  • 4.2 Fine-Tuning Pre-Trained Models
  • 5 Future Work
  • 6 Conclusion
  • Acknowledgment
  • References
  • A Grouped Softmax Example
  • B Dataset Description
  • C Number of Unique Retrieved Blocks
  • D Context Miss Token
  • E Positional Augmentation
  • F Additional Extensions and Details
  • G Offloading KV Cache to CPU

Knowls

  1. Knowl 1 — Landmark Attention Token Representation and Block Formulation

    model/method

    Landmark Attention partitions an input sequence into contiguous, non-overlapping blocks of fixed length ℓblock\ell_{\text{block}}. To represent each block within the attention mechanism without relying on separate retrieval modules or heuristic pooling, a dedicated special token—termed a landmark token—is appended immediately after the final token of every block.

    During self-attention across transformer layers, landmark tokens are processed and updated alongside ordinary sequence tokens. The key vector computed for a landmark token acts as a learned representative vector and gating mechanism for its corresponding block. Consequently, the attention weight assigned to any regular token in an earlier block is scaled by the attention score assigned to that block's landmark token, enabling random-access retrieval directly through the self-attention operation.

  2. Knowl 2 — Grouped Softmax for Landmark Attention

    equation

    To enable landmark tokens to gate attention to their respective blocks, the standard softmax function is replaced by Grouped Softmax. Given an input vector v∈Rℓseqv \in \mathbb{R}^{\ell_{\text{seq}}} and an integer group assignment vector g∈Nℓseqg \in \mathbb{N}^{\ell_{\text{seq}}}, Grouped Softmax applies softmax normalization separately over indices that share identical group IDs:

    σG(v,g)x:=GroupedSoftmax(v,g)x:=evx∑y:gy=gxevy\sigma_G(v, g)_x := \text{GroupedSoftmax}(v, g)_x := \frac{e^{v_x}}{\sum_{y: g_y = g_x} e^{v_y}}

    For an input sequence of length ℓseq\ell_{\text{seq}} containing landmark tokens, let pip_i be the sequence index of the landmark token associated with the block containing the ii-th token (with pi:=ip_i := i if token ii is itself a landmark token, and pi:=ℓseqp_i := \ell_{\text{seq}} if the final block is incomplete). When computing attention from query token ii to key token jj, the group assignment Gi,jG_{i,j} is defined as:

    Gi,j:={pjif pj≠j(places normal tokens in their own block groups)−1if pi=j(ignores the landmark token of token i’s own block)piif pi≠j∧pj=j(places other blocks’ landmarks in token i’s group)G_{i,j} := \begin{cases} p_j & \text{if } p_j \neq j \quad (\text{places normal tokens in their own block groups}) \\ -1 & \text{if } p_i = j \quad (\text{ignores the landmark token of token } i\text{'s own block}) \\ p_i & \text{if } p_i \neq j \land p_j = j \quad (\text{places other blocks' landmarks in token } i\text{'s group}) \end{cases}

    This grouping forces query token ii to distribute a single probability mass across its own block tokens and the landmark tokens of preceding blocks.

  3. Knowl 3 — Landmark Gated Attention Weight Computation

    equation

    In Landmark Attention, given a query token Qi∈RdheadQ_i \in \mathbb{R}^{d_{\text{head}}}, key matrix K∈Rℓseq×dheadK \in \mathbb{R}^{\ell_{\text{seq}} \times d_{\text{head}}}, head dimension dheadd_{\text{head}}, and group assignment vector GiG_i, the intermediate grouped attention score Si,jS_{i,j} is computed as:

    Si,j:=GroupedSoftmax(Qi⊤Kdhead,Gi)jS_{i,j} := \text{GroupedSoftmax}\left(\frac{Q_i^\top K}{\sqrt{d_{\text{head}}}}, G_i\right)_j

    The final attention weight AttWeight(Q,K)i,j\text{AttWeight}(Q, K)_{i,j} allocated from token ii to token jj is obtained by gating token-level scores with block landmark scores:

    AttWeight(Q,K)i,j:={0if pj=jSi,jif Gi,j=Gi,i∧pj≠jSi,j⋅Si,pjif Gi,j≠Gi,i∧pj≠j\text{AttWeight}(Q, K)_{i,j} := \begin{cases} 0 & \text{if } p_j = j \\ S_{i,j} & \text{if } G_{i,j} = G_{i,i} \land p_j \neq j \\ S_{i,j} \cdot S_{i,p_j} & \text{if } G_{i,j} \neq G_{i,i} \land p_j \neq j \end{cases}

    where pjp_j denotes the landmark token index corresponding to token jj's block. Under this definition, landmark tokens themselves receive zero final attention weight in the value aggregation step, and the attention weights over all normal tokens strictly sum to 1.

  4. Knowl 4 — Landmark-Based KV Cache Retrieval at Inference

    model/method

    During inference over arbitrarily long contexts, the input stream is segmented into chunks of length ℓlocal\ell_{\text{local}}, with a landmark token inserted every ℓblock\ell_{\text{block}} tokens. Each attention layer maintains a key-value (KV) cache of preceding blocks.

    1. For each query token in the current chunk, attention scores are computed solely against the cached landmark key vectors.
    2. For each query token (or query head), the kk highest-scoring landmark blocks are identified.
    3. The full KV vectors for tokens in the selected kk blocks are retrieved and prepended to the local chunk's KV matrix.
    4. Grouped Softmax and landmark gating are computed over the combined set of retrieved and local tokens.

    Because non-landmark KV pairs are only accessed upon retrieval, unselected block KV tensors can be swapped out to host CPU memory or disk, keeping only landmark vectors resident in GPU memory. This reduces computation and memory access overhead by a factor proportional to block size ℓblock\ell_{\text{block}} (e.g., approximately 50x for ℓblock=50\ell_{\text{block}} = 50) compared to full-context attention.

  5. Knowl 5 — Stingy Position Mapping for Length Extrapolation

    model/method

    Transformers utilizing positional encoding mechanisms such as Rotary Position Embeddings (RoPE) suffer performance degradation when evaluated on position indices beyond those observed during training. To allow retrieval from unbounded context lengths without altering positional embeddings or restricting receptive fields to local windows, Stingy Position Mapping maps block positions to a bounded index prefix.

    A prefix index window of length (k+1)⋅(ℓblock+1)(k+1) \cdot (\ell_{\text{block}} + 1) is reserved at the sequence start, where kk is the number of retrieved blocks and ℓblock\ell_{\text{block}} is the block size. Tokens within the active chunk are indexed starting immediately after this prefix window. For cached blocks in memory:

    • The most recent kk memory blocks are mapped to the last kk block slots within the prefix window.
    • All older memory blocks are assigned position indices corresponding to the first block of the prefix window.
    • When kk blocks are retrieved at inference, retrieved blocks among the most recent kk are positioned on the right of the allocated prefix, while older retrieved blocks are mapped to the left, separated by an empty block buffer while preserving relative chronological order.

    This ensures that recent tokens maintain accurate local relative positions while older tokens remain retrievable via semantic key matching without encountering out-of-distribution position indices.

  6. Knowl 6 — Language Modeling Perplexity on PG-19 and arXiv Math

    data/table

    Language modeling perplexity was evaluated on the PG-19 book corpus and arXiv Math datasets using 12-layer GPT-2 architectures trained with context length ℓseq=512\ell_{\text{seq}} = 512 and landmark block size ℓblock=50\ell_{\text{block}} = 50, compared against standard attention baselines and Transformer-XL trained on sequences of length 2048 (window size 256, effective context 512).

    Model / Setting Eval Length ℓlocal\ell_{\text{local}} Memory Blocks kk PG-19 PPL arXiv PPL
    Baseline 512 512 None - 16.12 4.01
    Baseline 512 360 None - 16.76 4.31
    Landmark Attention 512 250 10 2 16.23 4.01
    Transformer-XL 2048 256 (XL cache 256) None - 14.72 -
    Landmark Attention 2048 250 40 2 15.14 3.43
    Landmark Attention 2048 350 40 2 15.07 3.41
    Landmark Attention 2048 300 40 3 14.97 3.36
    Landmark Attention 2048 250 20 4 15.02 3.37
    Landmark Attention 2048 250 40 4 14.92 3.35
    Transformer-XL 4096 256 (XL cache 256) None - 14.55 -
    Landmark Attention 4096 250 40 4 14.79 3.19
    Landmark Attention 4096 250 80 2 15.00 3.29
    Landmark Attention 4096 250 80 4 14.72 3.18

    When evaluated at context lengths of 2048 and 4096, Landmark Attention models trained on sequence lengths of only 512 tokens successfully retrieve prior context to decrease perplexity (from 16.12 down to 14.72 on PG-19, and from 4.01 down to 3.18 on arXiv Math), matching Transformer-XL while maintaining random-access attention.

  7. Knowl 7 — Passkey Retrieval Context Extension on LLaMA 7B

    empirical result

    LLaMA 7B (pretrained with a default context limit of 2048 tokens) was fine-tuned for 15,000 steps using landmark attention on a subset of RedPajama using a context length of 512 tokens and block size ℓblock=50\ell_{\text{block}} = 50.

    In a passkey retrieval evaluation—where a target 5-digit passkey is embedded at a random position inside irrelevant repeating filler text—the base LLaMA 7B achieves 100% retrieval accuracy within its native 2048 context length but experiences complete failure (0% accuracy or Out-Of-Memory) on inputs exceeding 2500 tokens.

    In contrast, the landmark-fine-tuned LLaMA 7B (evaluating in chunks of 250 tokens, retrieving top k=4k=4 or k=5k=5 landmark blocks, and offloading non-landmark KV caches to CPU) maintains high retrieval accuracy across context lengths up to 32,070 tokens. Across 50 randomized trials at ~32k token prompts, the landmark-augmented model achieves 98% passkey retrieval accuracy.

  8. Knowl 8 — Granularity of Cache Block Retrieval and Efficiency Trade-offs

    data/table

    Block selection granularity at inference can be configured at different levels of flexibility: allowing independent block retrieval per head and per token, sharing retrieved blocks across tokens within each head, or sharing retrieved blocks across heads. The impact on language modeling perplexity was evaluated on PG-19 using input chunk size ℓlocal=250\ell_{\text{local}} = 250:

    Per-Head Selection Per-Token Selection Eval Length kk PG-19 Perplexity
    ✓ ✓ 2048 2 15.14
    ✓ ✓ 2048 4 14.92
    ✓ ✓ 4096 4 14.72
    ✓ ×\times 2048 2 15.48
    ✓ ×\times 2048 4 15.10
    ✓ ×\times 4096 4 14.95
    ×\times ✓ 2048 2 15.44
    ×\times ✓ 2048 4 15.04
    ×\times ✓ 4096 4 14.89

    Restricting retrieval to be shared across tokens within each head (per-head, token-shared) substantially reduces memory transfer bandwidth (accessing fewer than 10 unique blocks out of 32 possible per chunk) with a perplexity penalty of only 0.18–0.23 points. Increasing the number of retrieved blocks kk from 2 to 4 under the less flexible scheme compensates for this difference.

  9. Knowl 9 — Context Miss Token for Adaptive Memory Retrieval

    model/method

    To adaptively decide whether long-term memory retrieval is required for a given token, a Context Miss Token (CMT) is placed at sequence position −1-1. During training, a subset of landmark tokens LCMTL_{\text{CMT}} is designated as CMT-controlled (each landmark is assigned to LCMTL_{\text{CMT}} independently with probability PCMT=0.5P_{\text{CMT}} = 0.5).

    The Grouped Softmax group assignment Gi,jG_{i,j} and final attention weights are defined as:

    Gi,j:={piif j=−1pjif pj≠j−1if pi=jpiif j∉LCMT∧pj=j−2if j∈LCMTG_{i,j} := \begin{cases} p_i & \text{if } j = -1 \\ p_j & \text{if } p_j \neq j \\ -1 & \text{if } p_i = j \\ p_i & \text{if } j \notin L_{\text{CMT}} \land p_j = j \\ -2 & \text{if } j \in L_{\text{CMT}} \end{cases}

    AttWeight(Q,K)i,j:={0if pj=j∨j=−1Si,jif Gi,j=Gi,iSi,j⋅Si,pjif j∉LCMTSi,j⋅Si,pj⋅Si,−1if j∈LCMT\text{AttWeight}(Q, K)_{i,j} := \begin{cases} 0 & \text{if } p_j = j \lor j = -1 \\ S_{i,j} & \text{if } G_{i,j} = G_{i,i} \\ S_{i,j} \cdot S_{i,p_j} & \text{if } j \notin L_{\text{CMT}} \\ S_{i,j} \cdot S_{i,p_j} \cdot S_{i,-1} & \text{if } j \in L_{\text{CMT}} \end{cases}

    At inference, setting cached landmarks to LCMTL_{\text{CMT}} allows the CMT attention score Si,−1S_{i,-1} to act as a retrieval threshold: if Si,−1<τS_{i,-1} < \tau, memory retrieval is skipped. On PG-19 at evaluation length 2048 (k=4k=4), applying threshold τ=0.3\tau = 0.3 drops 57% of memory retrieval operations while increasing perplexity from 16.38 to only 16.43 (baseline model without CMT: 16.28).

  10. Knowl 10 — Positional Augmentation via Landmark Jumps

    model/method

    To train transformers capable of context extrapolation without KV cache mechanisms, positional augmentation introduces random discrete skips at landmark token boundaries. Instead of standard sequential indexing (1,2,…,T1, 2, \dots, T), the position indices of all tokens following each landmark token are incremented by a random integer sampled uniformly from {1,…,pjump}\{1, \dots, p_{\text{jump}}\}.

    For a training context length of 512 with landmark blocks of 50 tokens (yielding 10–11 landmarks per sequence) and jump parameter pjump=100p_{\text{jump}} = 100, the model encounters simulated position indices up to 11×100+512=161211 \times 100 + 512 = 1612. When evaluated on PG-19 by processing full inputs in a single forward pass without caching, models trained with pjump=100p_{\text{jump}} = 100 maintain decreasing perplexity up to sequence lengths of 1400 tokens, whereas standard models (pjump=1p_{\text{jump}} = 1) suffer perplexity degradation before reaching 1024 tokens.

Coverage note — Conceptual adaptations for masked language modeling (layer-by-layer chunking) and Triton/FlashAttention implementation details were omitted as secondary implementation notes without primary empirical validation.

References

  1. 1.Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, and Yuxiong He. DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale, June 2022. arXiv:2207.00032 [cs].
  2. 2.Anthropic. claude-v1.3-100k, Blog post ‘Introducing 100K Context Windows‘. https://www.anthropic.com/index/100k-context-windows, 2023.
  3. 3.Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The long-document transformer.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  5. 5.Aydar Bulatov, Yuri Kuratov, and Mikhail S. Burtsev. Recurrent memory transformer.
  6. 6.Mikhail S. Burtsev, Yuri Kuratov, Anton Peganov, and Grigory V. Sapunov. Memory transformer.
  7. 7.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating Long Sequences with Sparse Transformers, 2019. URL https://arxiv.org/abs/1904.10509.
  8. 8.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. Rethinking Attention with Performers, November 2022. arXiv:2009.14794 [cs, stat].
  9. 9.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. Transformer-XL: Attentive language models beyond a fixed-length context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2978–2988, Florence, Italy, 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1285. URL https://aclanthology.org/P19-1285.
  10. 10.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, June 2022.
  11. 11.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. REALM: Retrieval-augmented language model pre-training.
  12. 12.Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy. Transformer Language Models without Positional Encodings Still Learn Positional Information, 2022. URL https://arxiv.org/abs/2203.16634.
  13. 13.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. Atlas: Few-shot learning with retrieval augmented language models.
  14. 14.Zhengbao Jiang, Luyu Gao, Jun Araki, Haibo Ding, Zhiruo Wang, Jamie Callan, and Graham Neubig. Retrieval as attention: End-to-end learning of retrieval and reading within a single transformer.
  15. 15.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with GPUs.
  16. 16.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.550. URL https://aclanthology.org/2020.emnlp-main.550.
  17. 17.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. Generalization through memorization: Nearest neighbor language models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=HklBjCEKvH.
  18. 18.Omar Khattab, Christopher Potts, and Matei Zaharia. Relevance-guided supervision for OpenQA with ColBERT. Transactions of the Association for Computational Linguistics, 9:929–944, 2021. doi: 10.1162/tacl_a_00405. URL https://aclanthology.org/2021.tacl-1.55.
  19. 19.Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=rkgNKkHtvB.
  20. 20.Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4365–4374, Hong Kong, China, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1445. URL https://aclanthology.org/D19-1445.
  21. 21.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012.
  22. 22.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7.
  23. 23.Pedro Henrique Martins, Zita Marinho, and André F. T. Martins. ∞\infty-former: Infinite memory transformer.
  24. 24.Jesse Mu, Xiang Lisa Li, and Noah Goodman. Learning to Compress Prompts with Gist Tokens, 2023. URL https://arxiv.org/abs/2304.08467.
  25. 25.OpenAI. GPT-4 Technical Report. arXiv, 2023.
  26. 26.Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. Efficiently Scaling Transformer Inference, November 2022. arXiv:2211.05102 [cs].
  27. 27.Ofir Press, Noah A. Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=R8sQPpGCv0.
  28. 28.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  29. 29.Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. Compressive transformers for long-range sequence modelling. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=SylKikSYDH.
  30. 30.Hongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang, Jure Leskovec, Dale Schuurmans, and Bo Dai. Combiner: Full Attention Transformer with Sparse Computation Cost, October 2021. arXiv:2107.05768 [cs].
  31. 31.Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. ColBERTv2: Effective and efficient retrieval via lightweight late interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3715–3734, Seattle, United States, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.272. URL https://aclanthology.org/2022.naacl-main.272.
  32. 32.Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. High-throughput Generative Inference of Large Language Models with a Single GPU, March 2023. arXiv:2303.06865 [cs].
  33. 33.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. RoFormer: Enhanced transformer with rotary position embedding.
  34. 34.Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. Adaptive Attention Span in Transformers, August 2019. arXiv:1905.07799 [cs, stat].
  35. 35.Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, and Mohit Iyyer. Do long-range language models actually use long-range context? In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 807–822, Online and Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.62. URL https://aclanthology.org/2021.emnlp-main.62.
  36. 36.Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei. A Length-Extrapolatable Transformer, 2022. URL https://arxiv.org/abs/2212.10554.
  37. 37.Philippe Tillet, Hsiang-Tsung Kung, and David Cox. Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, pages 10–19, 2019.
  38. 38.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and Efficient Foundation Language Models, 2023. URL https://arxiv.org/abs/2302.13971.
  39. 39.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  40. 40.Qingyang Wu, Zhenzhong Lan, Kun Qian, Jing Gu, Alborz Geramifard, and Zhou Yu. Memformer: A memory-augmented transformer for sequence modeling.
  41. 41.Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, and Christian Szegedy. Memorizing transformers. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=TrjbxzRcnf-.
  42. 42.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big bird: Transformers for longer sequences. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/c8512d142a2d849725f31a9a7a361ab9-Abstract.html.

Citation

MLA
Mohtashami, A., and M. Jaggi. “Random-Access Infinite Context Length for Transformers”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 54567–85, https://proceedings.neurips.cc/paper_files/paper/2023/file/ab05dc8bf36a9f66edbff6992ec86f56-Paper-Conference.pdf.
APA
Mohtashami, A., & Jaggi, M. (2023). Random-Access Infinite Context Length for Transformers. Advances in Neural Information Processing Systems, 36, 54567–54585. https://proceedings.neurips.cc/paper_files/paper/2023/file/ab05dc8bf36a9f66edbff6992ec86f56-Paper-Conference.pdf
Chicago
Mohtashami, A., and M. Jaggi. 2023. “Random-Access Infinite Context Length for Transformers”. Advances in Neural Information Processing Systems 36: 54567–85. https://proceedings.neurips.cc/paper_files/paper/2023/file/ab05dc8bf36a9f66edbff6992ec86f56-Paper-Conference.pdf.
Harvard
Mohtashami, A. and Jaggi, M. (2023) “Random-Access Infinite Context Length for Transformers”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 54567–54585. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/ab05dc8bf36a9f66edbff6992ec86f56-Paper-Conference.pdf.
Vancouver
1. Mohtashami A, Jaggi M (2023) Random-Access Infinite Context Length for Transformers. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 54567–54585

BibTeX

@inproceedings{mohtashami2023random,
  title = {Random-Access Infinite Context Length for Transformers},
  author = {Mohtashami, Amirkeivan and Jaggi, Martin},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {54567-54585},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/ab05dc8bf36a9f66edbff6992ec86f56-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors