A Length-Extrapolatable Transformer

Yutao SunLi DongBarun PatraShuming MaShaohan HuangAlon BenhaimVishrav ChaudharyXia SongFuru Wei

article2023ACL185 citations

Proposes the Length-Extrapolatable Transformer, which combines a decay-augmented relative position embedding with blockwise causal attention to continually lower perplexity on sequences far longer than those seen during training.

Listen

Modern natural language processing relies heavily on Transformer architectures, but pre-training these models on long sequences is prohibitively expensive in terms of computational resources and memory. Standard architectures suffer severe performance degradation when processing text inputs longer than the maximum length encountered during training. Enabling models to generalize effectively from short training sequences to longer deployment inputs—a capability known as length extrapolation—is essential for practical applications that involve long documents, transcripts, and books.

The article aims to design, evaluate, and demonstrate an enhanced architecture, named the Length-Extrapolatable (LEX) Transformer. This framework trains on short sequences while maintaining high accuracy and continuously improving performance when processing significantly longer sequences during inference.

To address this challenge, the authors defined a formal metric called attention resolution to evaluate how effectively an attention mechanism distinguishes relative token positions. They then developed the LEX Transformer by integrating two key components: an Extrapolatable Position Embedding (XPOS), which adds exponential decay to vector rotations to stabilize position calculations, and blockwise causal attention, which segments attention into manageable blocks during inference. The authors validated their approach by pre-training medium-scale models from scratch on sequences of 1,024 tokens and evaluating perplexity across sequence lengths extending up to 8,192 tokens across benchmark datasets including PG22, QMSum, arXiv, and NarrativeQA.

The evaluation produced several critical findings. First, the LEX Transformer was the only architecture tested whose perplexity consistently decreased as input lengths scaled up to 8,192 tokens, verifying that the model successfully leverages longer context. Second, baseline Transformer models degraded severely on long sequences, with perplexity on the PG22 dataset escalating from 30.54 at 1,024 tokens to over 12,700 at 8,192 tokens. Third, alternative relative position approaches struggled at extended lengths; rotary position embeddings suffered catastrophic degradation (reaching a perplexity of 458.83 at 8,192 tokens), while attention with linear biases saw perplexity rise from 26.01 at 2,048 tokens to 32.8 at 8,192 tokens. Finally, LEX outperformed existing methods on short, in-distribution texts (interpolation) while achieving the highest attention resolution scores across both standard and extended lengths.

These results demonstrate that engineering teams can deploy long-context language processing capabilities without incurring the heavy financial, hardware, and operational costs of training models directly on long sequences. The findings resolve a long-standing trade-off in sequence modeling by showing that robust long-range capability does not require sacrificing accuracy on standard-length inputs.

Engineering and research teams developing generative language models should consider adopting XPOS position embeddings and blockwise attention as drop-in upgrades. When implementing long-context inference, organizations can choose between blockwise causal attention for cache efficiency and sliding-window attention for slightly higher accuracy at the expense of computational throughput. Further work is recommended to adapt and validate these mechanisms in bidirectional encoder architectures.

The study's primary limitation is its focus on causal language models, leaving bidirectional architectures (such as masked language models) requiring further adaptation. Additionally, XPOS introduces approximately a 6 percent inference overhead compared to basic absolute position embeddings, although this is partially offset by accelerated training convergence. Confidence in the core findings remains high due to consistent empirical improvements across multiple diverse datasets.

Cover for A Length-Extrapolatable Transformer

Abstract

Position modeling plays a critical role in Transformers. In this paper, we focus on length extrapolation, i.e., training on short texts while evaluating longer sequences. We define attention resolution as an indicator of extrapolation. Then we propose two designs to improve the above metric of Transformers. Specifically, we introduce a relative position embedding to explicitly maximize attention resolution. Moreover, we use blockwise causal attention during inference for better efficiency. The proposed architecture is named Length-Extrapolatable (LEX) Transformer. We evaluate different Transformer variants on language modeling. Experimental results show that our model achieves better performance in both interpolation and extrapolation settings. The code will be available at https://aka.ms/LeX-Transformer.

Table of Contents

  • 1 Introduction
  • 2 Design Principles of Transformers for Position Modeling
  • 2.1 Order Variance
  • 2.2 Translation Invariance
  • 2.3 Length Extrapolation
  • 3 A Length-Extrapolatable Transformer
  • 3.1 Attention Resolution
  • 3.2 Improve Resolution by Position Encoding
  • 3.3 Blockwise Causal Attention
  • 4 Experiments
  • 4.1 Pre-training
  • 4.2 Language Modeling
  • 4.3 Measuring Resolution
  • 4.4 Ablation Studies
  • 4.4.1 Rotation Computation
  • 4.4.2 Blockwise Causal Attention
  • 5 Related Work
  • 5.1 Long-Sequence Transformers
  • 5.2 Position Modeling
  • 5.2.1 Absolute Position Embedding
  • 5.2.2 Relative Position Embedding
  • 6 Conclusion
  • Limitations
  • References
  • A Additional Experiments
  • B Hyperparameters for Pre-Training

Knowls

  1. Knowl 1 — Extrapolatable Position Embedding

    model/method

    Extrapolatable Position Embedding (xPos) is a relative position encoding method that augments rotary position embeddings (RoPE) with dimension-specific exponential decay to suppress high-frequency oscillations across long token distances.

    Let dd denote the hidden dimension of an attention head, and let nn denote the position index of a token. In complex representation for dimension coordinate pair i∈{0,1,…,d/2−1}i \in \{0, 1, \dots, d/2 - 1\}, the position embedding transformation parameters are given by λi=ξi+iθi∈C\lambda_i = \xi_i + \mathrm{i}\theta_i \in \mathbb{C}, where θi=10000−2i/d\theta_i = 10000^{-2i/d}. For query vector q∈Rdq \in \mathbb{R}^d and key vector k∈Rdk \in \mathbb{R}^d at position nn, the transformations are:

    fq(q,n)i=qieξin+iθinf_q(q, n)_i = q_i e^{\xi_i n + \mathrm{i}\theta_i n}

    fk(k,n)i=kie−ξin+iθinf_k(k, n)_i = k_i e^{-\xi_i n + \mathrm{i}\theta_i n}

    In real-valued 2D block matrix form, for coordinates (2i,2i+1)(2i, 2i+1):

    fq(q,n)2i:2i+1=ζ^in/B(cos⁡(nθi)−sin⁡(nθi)sin⁡(nθi)cos⁡(nθi))(q2iq2i+1)f_q(q, n)_{2i:2i+1} = \hat{\zeta}_i^{n/B} \begin{pmatrix} \cos(n\theta_i) & -\sin(n\theta_i) \\ \sin(n\theta_i) & \cos(n\theta_i) \end{pmatrix} \begin{pmatrix} q_{2i} \\ q_{2i+1} \end{pmatrix}

    fk(k,n)2i:2i+1=ζ^i−n/B(cos⁡(nθi)−sin⁡(nθi)sin⁡(nθi)cos⁡(nθi))(k2ik2i+1)f_k(k, n)_{2i:2i+1} = \hat{\zeta}_i^{-n/B} \begin{pmatrix} \cos(n\theta_i) & -\sin(n\theta_i) \\ \sin(n\theta_i) & \cos(n\theta_i) \end{pmatrix} \begin{pmatrix} k_{2i} \\ k_{2i+1} \end{pmatrix}

    where ζ^i=eξi=2i/d+γ1+γ∈[0,1]\hat{\zeta}_i = e^{\xi_i} = \frac{2i/d + \gamma}{1 + \gamma} \in [0, 1] is a per-dimension decay factor, γ=0.4\gamma = 0.4 is a smoothing constant, and B=512B = 512 is a base scale parameter introduced to prevent floating-point underflow and overflow in 16-bit precision.

  2. Knowl 2 — Attention Resolution Metric

    definition

    Attention resolution R(s)R(s) is a quantitative metric that measures an attention mechanism's ability to distinguish relative token positions monotonically across sequence distances.

    Let s[n]s[n] denote the expected attention score between two tokens separated by relative distance n≥0n \ge 0:

    s[n]=E0≤i≤N[qi+nkiTd]s[n] = \mathbb{E}_{0 \le i \le N}\left[ \frac{q_{i+n} k_i^T}{\sqrt{d}} \right]

    where qi+n∈Rdq_{i+n} \in \mathbb{R}^d is the query at position i+ni+n, ki∈Rdk_i \in \mathbb{R}^d is the key at position ii, dd is the attention head dimension, and NN is the context length.

    The attention resolution R(s)R(s) is defined as:

    R(s)=∑i=0Nes[i](es[i]−es[i+1])(∑j=0Nes[j])2R(s) = \sum_{i=0}^N \frac{e^{s[i]} (e^{s[i]} - e^{s[i+1]})}{\left(\sum_{j=0}^N e^{s[j]}\right)^2}

    A strictly decreasing attention score expectation s[i]>s[i+1]s[i] > s[i+1] yields positive resolution contributions. Higher values of R(s)R(s) indicate superior position discriminability, whereas oscillations in s[n]s[n] over long distances degrade R(s)R(s).

  3. Knowl 3 — Blockwise Causal Attention for Inference Length Extrapolation

    model/method

    Blockwise Causal Attention (BCA) is a constrained local masking mechanism employed during inference to bound relative attention distances and maintain high attention resolution when evaluating sequences longer than the pre-training context window.

    During training, the Transformer is trained with standard full causal attention on sequences of length ll. During inference on sequences of length L>lL > l, the query sequence is divided into contiguous blocks of length l/2l/2. Each query token in block bb attends only to keys and values belonging to its own block bb and the immediately preceding block b−1b-1. Context information propagates across the sequence through the recurrent caching and reuse of overlapping key and value representations, ensuring that no query computes relative position offsets exceeding ll.

  4. Knowl 4 — Length-Extrapolatable Transformer Architecture

    model/method

    The Length-Extrapolatable (LeX) Transformer is an autoregressive sequence model designed to achieve monotonic perplexity reduction as evaluation lengths extend far beyond the training context length.

    LeX Transformer integrates two primary components:

    1. Extrapolatable Position Embedding (xPos), which applies a combined 2D rotation and dimension-dependent exponential decay to queries and keys during both pre-training and inference.
    2. Blockwise Causal Attention (BCA), which restricts the self-attention window during inference to the current block of length l/2l/2 and the preceding block of length l/2l/2, where ll is the pre-training sequence length.

    Unlike vanilla Transformers or standard rotary embeddings whose perplexity diverges when evaluating inputs longer than the pre-training window, LeX Transformer maintains translation invariance and order variance while monotonically decreasing perplexity up to at least 8×8\times the pre-training sequence length.

  5. Knowl 5 — Attention Computation with xPos

    algorithm

    Attention with xPos applies vectorized rotations and exponential decay to query and key tensors prior to computing multi-head scaled dot-product attention.

    function AttentionWithXPos(Q, K, V, M, d, l, gamma, B)
        Input: Query Q in R^{h x l x d}, Key K in R^{h x l x d}, Value V in R^{h x l x d}
        Input: Causal mask M in R^{l x l}, head dimension d, length l, gamma = 0.4, B = 512
        Output: Context output in R^{h x l x d}
        for i = 0 to d - 1 do
            θi←10000−2⌊i/2⌋/d\theta_i \leftarrow 10000^{-2 \lfloor i/2 \rfloor / d}
            ζ^i←2⌊i/2⌋/d+γ1+γ\hat{\zeta}_i \leftarrow \frac{2 \lfloor i/2 \rfloor / d + \gamma}{1 + \gamma}
        end for
        for m = 0 to l - 1 do
            for n = 0 to d - 1 do
                Cm,n←cos⁡(mθn)C_{m,n} \leftarrow \cos(m \theta_n)
                Sm,n←sin⁡(mθn)S_{m,n} \leftarrow \sin(m \theta_n)
                Tm,n←ζ^nm/BT_{m,n} \leftarrow \hat{\zeta}_n^{m / B}
            end for
        end for
        rot(X) = [-X[:, :, 1], X[:, :, 0], -X[:, :, 3], X[:, :, 2], ...]
        Q←Q⊙(C⊙T)+rot(Q)⊙(S⊙T)Q \leftarrow Q \odot (C \odot T) + \text{rot}(Q) \odot (S \odot T)
        K←K⊙(C⊙T−1)+rot(K)⊙(S⊙T−1)K \leftarrow K \odot (C \odot T^{-1}) + \text{rot}(K) \odot (S \odot T^{-1})
        output←softmax(QKTd⊙M)V\text{output} \leftarrow \text{softmax}\left(\frac{Q K^T}{\sqrt{d}} \odot M\right) V
        return output
  6. Knowl 6 — Language Modeling Perplexity under Length Interpolation and Extrapolation

    data/table

    Language models trained with context length 1024 were evaluated on language modeling datasets across interpolation lengths (256, 512, 1024) and extrapolation lengths (2048, 4096, 8192). LeX Transformer is the only architecture where perplexity decreases monotonically as sequence length grows to 8192.

    Length 256 512 1024 2048 4096 8192
    Setting Interpolation Extrapolation
    PG22
    Transformer 38.10 33.50 30.54 132.46 1446.95 12747.41
    Alibi 34.25 30.01 27.34 26.01 28.46 32.80
    Roformer 33.27 29.20 26.68 68.86 235.71 458.83
    LEX Transformer (Ours) 33.18 29.11 26.59 25.53 25.07 24.89
    QMSum
    Transformer 24.25 18.81 16.05 86.56 1196.92 10781.38
    Alibi 22.85 17.74 15.17 13.97 15.36 18.37
    Roformer 22.66 17.65 15.12 36.54 146.61 331.56
    LEX Transformer (Ours) 22.01 17.24 14.85 13.92 13.56 13.48

    Vanilla Transformer and RoFormer suffer from severe perplexity explosion during extrapolation (>450>450 and >10000>10000 at length 8192). ALiBi extrapolates moderately well up to length 2048 but begins degrading at 4096 and 8192. LeX Transformer outperforms all baselines at every evaluated length and continues to improve perplexity up to 8192 tokens.

  7. Knowl 7 — Empirical Attention Resolution across Positional Encodings

    data/table

    Attention resolution R(s)R(s) was empirically measured across Transformer variants at sequence length 1024 (in-distribution) and length 2048 (extrapolation), averaged over test inputs and layers.

    Model 1024 (Interpolation) 2048 (Extrapolation)
    Transformer 0.87 0.28
    Alibi 0.81 0.88
    Roformer 0.91 0.08
    LEX (Ours) 0.98 1.08
    −- BCA 0.98 0.54

    RoFormer's attention resolution collapses from 0.91 to 0.08 during extrapolation due to severe rotation oscillation. LeX Transformer achieves the highest attention resolution at both 1024 (0.98) and 2048 (1.08). Ablating Blockwise Causal Attention (BCA) reduces LeX's extrapolation resolution from 1.08 to 0.54, confirming that BCA significantly enhances the model's position discrimination at extended lengths.

  8. Knowl 8 — Ablation of Rotation and Exponential Decay Components in xPos

    data/table

    Ablation experiments evaluated on the PG22 dataset demonstrate that both vector rotation and dimension-specific exponential decay are required for effective length interpolation and extrapolation.

    Methods 1024 (Interpolation) 8192 (Extrapolation)
    LEX 26.59 24.89
    w/o Rotation 37.11 34.50
    ζ=0\zeta = 0 (RoPE) 26.68 26.16
    Scalar ζ\zeta 26.85 25.10

    Removing vector rotation (thetai=0\\theta_i = 0) causes a substantial degradation in perplexity (37.11 at 1024 and 34.50 at 8192). Setting decay to zero (ζ=0\zeta = 0, which reduces to RoPE) hurts extrapolation performance relative to LeX (26.16 vs. 24.89). Replacing the vector ζ^\hat{\zeta} with a single scalar decay value ζ=γ/(1+γ)\zeta = \gamma / (1 + \gamma) performs worse than the per-dimension decay vector (25.10 vs. 24.89).

  9. Knowl 9 — Interaction of Position Encodings with Blockwise and Sliding Window Attention

    data/table

    Evaluating various position embeddings with and without windowed causal attention on PG22 at sequence lengths 2048 and 8192 reveals the complementary relationship between relative embeddings and causal masking.

    Methods 2048 8192
    Absolute 132.46 12747.41
    Absolute + BCA 322.73 28787.01
    ROPE 68.86 458.83
    ROPE + BCA 26.37 26.16
    Alibi 26.01 32.80
    Alibi + BCA 27.53 31.82
    xPOS 27.29 63.99
    xPOS + BCA 25.53 24.89
    xPOS + Sliding Window 25.33 24.61

    Blockwise Causal Attention (BCA) prevents perplexity explosion for both RoPE (dropping from 458.83 to 26.16 at 8192) and xPos (dropping from 63.99 to 24.89 at 8192). While Sliding Window attention achieves marginally lower perplexity (24.61 vs. 24.89), BCA is chosen for LeX Transformer due to being approximately 1.5×1.5\times faster and more cache-friendly.

  10. Knowl 10 — Language Model Pre-Training Setup

    experimental setup

    Language models were pre-trained from scratch on 16 NVIDIA V100 GPUs using a medium GPT-3 architecture: 24 layers, hidden dimension d=1024d = 1024, feed-forward intermediate dimension 4096, and 16 attention heads. Models were trained on a 0.5M token batch size (512 sequences of length 1024) for 50,000 steps using the GPT-2 tokenizer.

    The training corpus comprises a subset of The Pile (Books3, OpenWebText2, Stack Exchange, PubMed Abstracts, Wikipedia, PG-19, BookCorpus2, NIH ExPorter, and Pile-CC). The optimizer is Adam (β1=0.9,β2=0.98,ϵ=10−6\beta_1 = 0.9, \beta_2 = 0.98, \epsilon = 10^{-6}) with a peak learning rate of 3×10−43 \times 10^{-4}, 20,000 warmup steps, polynomial learning rate decay, gradient clipping of 2.0, and weight decay of 0.01.

  11. Knowl 11 — Limitations of xPos and LeX Transformer

    limitation

    The theoretical formulation and implementation of xPos rely on the causal autoregressive language modeling assumption where query position expectations dominate key position expectations (E(∠q)≤E(∠k)\mathbb{E}(\angle q) \le \mathbb{E}(\angle k)). Consequently, applying xPos to bidirectional architectures, such as masked language models (e.g., BERT), requires additional theoretical and architectural adaptations.

    In addition, computing the xPos transformations introduces approximately 6% inference time overhead compared to standard absolute position embeddings.

Coverage note — Omitted Table 6 (additional language modeling evaluation on arXiv and NarrativeQA) as it repeats the same comparative findings reported in Table 2 on PG22 and QMSum.

References

  1. 1.Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  3. 3.Ta-Chung Chi, Ting-Han Fan, Peter J Ramadge, and Alexander I Rudnicky. 2022a. Kerple: Kernelized relative positional embedding for length extrapolation. arXiv preprint arXiv:2205.09921.
  4. 4.Ta-Chung Chi, Ting-Han Fan, and Alexander I Rudnicky. 2022b. Receptive field alignment enables transformer length extrapolation. arXiv preprint arXiv:2212.10356.
  5. 5.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. URL https://openai.com/blog/sparse-transformers.
  6. 6.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. 2020. Rethinking attention with performers. arXiv preprint arXiv:2009.14794.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek B Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Oliveira Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. ArXiv, abs/2204.02311.
  8. 8.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
  11. 11.Philipp Dufter, Martin Schmitt, and Hinrich Schütze. 2022. Position information in transformers: An overview. Computational Linguistics, 48(3):733–763.
  12. 12.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  13. 13.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654.
  14. 14.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural Computation, 9:1735–1780.
  15. 15.DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam Neyshabur. 2022. Block-recurrent Transformers. In Advances in Neural Information Processing Systems.
  16. 16.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR.
  17. 17.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, San Diego, CA.
  18. 18.Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328.
  19. 19.Shuming Ma, Hongyu Wang, Shaohan Huang, Wenhui Wang, Zewen Chi, Li Dong, Alon Benhaim, Barun Patra, Vishrav Chaudhary, Xia Song, and Furu Wei. 2022a. TorchScale: Transformers at scale. CoRR, abs/2211.13184.
  20. 20.Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. 2022b. Mega: Moving average equipped gated attention. arXiv preprint arXiv:2209.10655.
  21. 21.Ofir Press, Noah A Smith, and Mike Lewis. 2021. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409.
  22. 22.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR.
  23. 23.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  24. 24.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  25. 25.Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, et al. 2022. Scrolls: Standardized comparison over long language sequences. arXiv preprint arXiv:2201.03533.
  26. 26.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155.
  27. 27.Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. 2021. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864.
  28. 28.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA, pages 6000–6010.
  29. 29.Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen. 2020a. On position embeddings in bert. In International Conference on Learning Representations.
  30. 30.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020b. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768.
  31. 31.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. 2022. Image as a foreign language: BEiT pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442.
  32. 32.Wenhan Xiong, Barlas Oğuz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Wen-tau Yih, and Yashar Mehdad. 2021. Simple local attentions remain competitive for long-context tasks. arXiv preprint arXiv:2112.07210.
  33. 33.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  34. 34.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020. Big bird: Transformers for longer sequences. Advances in Neural Information Processing Systems, 33:17283–17297.
  35. 35.Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, et al. 2021. Qmsum: A new benchmark for query-based multidomain meeting summarization. arXiv preprint arXiv:2104.05938.

Citation

MLA
Sun, Y., et al. “A Length-Extrapolatable Transformer”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 14590–604, https://doi.org/10.18653/v1/2023.acl-long.816.
APA
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., & Wei, F. (2023). A Length-Extrapolatable Transformer. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14590–14604. https://doi.org/10.18653/v1/2023.acl-long.816
Chicago
Sun, Y., L. Dong, B. Patra, et al. 2023. “A Length-Extrapolatable Transformer”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14590–604. https://doi.org/10.18653/v1/2023.acl-long.816.
Harvard
Sun, Y. et al. (2023) “A Length-Extrapolatable Transformer”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14590–14604. Available at: https://doi.org/10.18653/v1/2023.acl-long.816.
Vancouver
1. Sun Y, Dong L, Patra B, Ma S, Huang S, Benhaim A, Chaudhary V, Song X, Wei F (2023) A Length-Extrapolatable Transformer. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14590–14604

BibTeX

@inproceedings{sun-etal-2023-length,
    title = "A Length-Extrapolatable Transformer",
    author = "Sun, Yutao  and
      Dong, Li  and
      Patra, Barun  and
      Ma, Shuming  and
      Huang, Shaohan  and
      Benhaim, Alon  and
      Chaudhary, Vishrav  and
      Song, Xia  and
      Wei, Furu",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.816/",
    doi = "10.18653/v1/2023.acl-long.816",
    pages = "14590--14604"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/