Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis

Ta-Chung ChiTing-Han FanAlexander RudnickyPeter J. Ramadge

article2023ACL55 citationsOutstanding Paper Award

Explains why transformer length extrapolation succeeds by analyzing empirical receptive fields and introduces Sandwich, a parameter-free relative positional embedding that effectively utilizes context beyond the training sequence length.

Listen

Training modern transformer language models on very long sequences requires substantial compute and memory resources due to quadratic cost scaling. To keep training costs practical while enabling models to handle long documents at testing time, developers rely on length extrapolation, which allows a model trained on short text sequences to maintain performance on much longer sequences. While specific positional embedding designs like Attention with Linear Biases (ALiBi) have achieved widespread empirical success in enabling extrapolation, the underlying mechanisms governing why these techniques succeed or fail have remained poorly understood.

The article evaluates why transformer models extrapolate to longer sequences by analyzing their receptive field—the span of surrounding text that actually influences a model's predictions. The investigation demonstrates the structural connection between linear relative positional biases and windowed attention, explains prior failure modes of common positional embeddings, and introduces an improved positional embedding design.

To conduct this evaluation, the researchers introduced a cumulative normalized gradient tool to measure a model's empirical receptive field across various configurations. They trained 12-layer, 162-million parameter models on sequences of 512 tokens using three diverse datasets—academic papers from ArXiv, open web text from OpenWebText2, and code from GitHub—and tested their language modeling perplexity across extended lengths ranging up to 8,192 tokens.

The analysis yielded four major findings. First, popular methods like ALiBi succeed primarily because steep slopes restrict the model to a narrow, windowed receptive field; as long as the training sequence length fully covers this empirical receptive field, the model extrapolates smoothly by ignoring distant tokens rather than integrating them. Second, standard absolute and rotary positional embeddings fail at extrapolation because they overfit to the exact position indices seen during training and fail to constrain their receptive field. Third, the authors introduced "Sandwich," a new parameter-free relative positional embedding derived from simplified sinusoidal embeddings. Unlike ALiBi, Sandwich genuinely leverages context beyond the initial training length, achieving lower (better) perplexity on longer test sequences across multiple benchmarks. Fourth, Sandwich exhibits a logarithmic decaying temporal bias pattern similar to parameterized designs like KERPLE and T5, which provides a beneficial averaging and denoising effect over distant tokens.

These findings provide clear practical implications for designing and deploying large-scale language models. Relying on ALiBi enables predictable compute costs and efficient inference caching without performance degradation, but it leaves long-range document context largely unexploited. By contrast, adopting logarithmically decaying positional designs like Sandwich allows organizations to train cost-effectively on short sequences while gaining the actual performance benefits of longer contextual inputs during deployment, all without adding learnable parameters or runtime overhead.

For engineering and research teams developing foundation models, the article recommends adopting relative positional embeddings that incorporate logarithmic distance decay. When resource efficiency and training simplicity are paramount, teams should utilize parameter-free designs like Sandwich over linear-decay designs. If maximum task adaptability is required and additional computational overhead is acceptable, learnable counterparts such as KERPLE or T5 present viable alternatives. Further work should explore adding learnable compression ratios to Sandwich and testing these designs across broader architectures.

Decision-makers should interpret these results with an awareness of specific architectural limitations. The empirical evaluation focused on 162-million parameter models, and while the underlying receptive field mechanics are general, absolute scaling dynamics in multi-billion parameter regimes warrant ongoing validation. Furthermore, the logarithmic decay mechanism inherently introduces a recency bias that favors nearby context; while this aligns well with natural language text, caution is advised when applying these architectures to non-linguistic tasks where all sequence elements possess equal structural importance.

arXiv: 2212.10356
Cover for Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis

Abstract

Length extrapolation permits training a transformer language model on short sequences that preserves perplexities when tested on substantially longer sequences. A relative positional embedding design, ALiBi, has had the widest usage to date. We dissect ALiBi via the lens of receptive field analysis empowered by a novel cumulative normalized gradient tool. The concept of receptive field further allows us to modify the vanilla Sinusoidal positional embedding to create Sandwich, the first parameter-free relative positional embedding design that truly length information uses longer than the training sequence. Sandwich shares with KERPLE and T5 the same logarithmic decaying temporal bias pattern with learnable relative positional embeddings; these elucidate future extrapolatable positional embedding design.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Length Extrapolation
  • 2.2 Positional Embeddings
  • 2.3 Windowed and Sparse Attention
  • 2.4 Receptive Field
  • 3 Background and Notations
  • 3.1 Transformer Language Model
  • 3.2 ALiBi
  • 3.3 Windowed Attention
  • 3.4 Evaluation of Length Extrapolation
  • 4 ALiBi and Windowed Attention
  • 4.1 Slope Shift (Shift all h by ∆ )
  • 4.2 Slope Equalization (Same h for all heads)
  • 4.3 Windowed Attention (Size w )
  • 4.4 Other Observations
  • 5 Receptive Field Measurement
  • 5.1 Quantifying Empirical Receptive Field
  • 5.2 Fixing Failed Cases
  • 5.3 Analyses of Sinusoidal and Rotary
  • 6 A New RPE for Length Extrapolation
  • 6.1 Introduction to Sandwich
  • 6.2 Experiments and Discussion
  • 6.3 Connection to KERPLE and T5
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgment
  • References
  • A Results on OpenWebText2
  • B Efficient Inference
  • C Scientific Artifacts
  • D Implementation Details
  • E Python Implementation of Sandwich

Knowls

  1. Knowl 1 — Sandwich Relative Positional Embedding

    model/method

    Sandwich is a parameter-free relative positional embedding (RPE) for Transformer language models derived from the inner product of sinusoidal positional embeddings. For a query at sequence position mm and key at position nn (m≥nm \ge n), the inner product between sinusoidal position vectors pm,pn∈Rdˉp_m, p_n \in \mathbb{R}^{\bar{d}} simplifies to:

    pm⊤pn=∑i=1dˉ/2[sin⁡(m100002i/dˉ)sin⁡(n100002i/dˉ)+cos⁡(m100002i/dˉ)cos⁡(n100002i/dˉ)]=∑i=1dˉ/2cos⁡(m−n100002i/dˉ)p_m^\top p_n = \sum_{i=1}^{\bar{d}/2} \left[ \sin\left(\frac{m}{10000^{2i/\bar{d}}}\right) \sin\left(\frac{n}{10000^{2i/\bar{d}}}\right) + \cos\left(\frac{m}{10000^{2i/\bar{d}}}\right) \cos\left(\frac{n}{10000^{2i/\bar{d}}}\right) \right] = \sum_{i=1}^{\bar{d}/2} \cos\left(\frac{m - n}{10000^{2i/\bar{d}}}\right)

    where dˉ\bar{d} is a hyperparameter decoupling the positional embedding dimension from the hidden dimension dd (default dˉ=128\bar{d} = 128). The maximal value of pm⊤pnp_m^\top p_n occurs at m−n=0m - n = 0, where pm⊤pm=dˉ/2p_m^\top p_m = \bar{d}/2.

    To construct an attention bias that decays monotonically with relative distance and prevents logit divergence during length extrapolation, the temporal bias added to the pre-softmax attention logit for head h∈{1,…,H}h \in \{1, \dots, H\} is defined as:

    Biasm,n(h)=pm⊤pn−dˉ/2ch\text{Bias}^{(h)}_{m,n} = \frac{p_m^\top p_n - \bar{d}/2}{c_h}

    where ch=h⋅8Hc_h = h \cdot \frac{8}{H} denotes the head-specific compression ratio. The resulting self-attention probability matrix for head hh is computed as:

    am,n(h)=exp⁡((⟨qm(h),kn(h)⟩+pm⊤pn−dˉ/2ch)/d/H)∑i=1mexp⁡((⟨qm(h),ki(h)⟩+pm⊤pi−dˉ/2ch)/d/H)a_{m,n}^{(h)} = \frac{\exp\left(\left(\langle q_m^{(h)}, k_n^{(h)} \rangle + \frac{p_m^\top p_n - \bar{d}/2}{c_h}\right) / \sqrt{d/H}\right)}{\sum_{i=1}^m \exp\left(\left(\langle q_m^{(h)}, k_i^{(h)} \rangle + \frac{p_m^\top p_i - \bar{d}/2}{c_h}\right) / \sqrt{d/H}\right)}

    where qm(h),kn(h)∈Rd/Hq_m^{(h)}, k_n^{(h)} \in \mathbb{R}^{d/H} are query and key representations.

  2. Knowl 2 — Cumulative Normalized Gradient and Empirical Receptive Field (ERF)

    definition

    The Empirical Receptive Field (ERF) quantifies the effective input context that contributes to the prediction of a target token at sequence position LexL_{ex} in an autoregressive Transformer.

    Let gm∈Rdg_m \in \mathbb{R}^d be the gradient of the predicted target token output with respect to the input embedding eme_m at position m∈{1,…,Lex−1}m \in \{1, \dots, L_{ex}-1\}. The normalized gradient sms_m at position mm is defined as:

    sm=∥gm∥2∑n=1Lex∥gn∥2s_m = \frac{\|g_m\|_2}{\sum_{n=1}^{L_{ex}} \|g_n\|_2}

    where ∥⋅∥2\|\cdot\|_2 denotes the Euclidean norm.

    The cumulative normalized gradient cmc_m measures the total gradient contribution from position mm up to the target token LexL_{ex}:

    cm=∑n=mLexsn,0≤cm≤1c_m = \sum_{n=m}^{L_{ex}} s_n, \quad 0 \le c_m \le 1

    The empirical receptive field (ERF) of the model is defined as the minimum lookback context containing at least 99% of the total normalized gradient:

    ERF=Lex−min⁡{m∣cm>0.99}+1\text{ERF} = L_{ex} - \min\{m \mid c_m > 0.99\} + 1

  3. Knowl 3 — Empirical Receptive Field Coverage Principle for Length Extrapolation

    theoretical result

    An autoregressive Transformer language model pretrained on sequence length LtrL_{tr} successfully performs length extrapolation to evaluation sequence lengths Lex≫LtrL_{ex} \gg L_{tr} without perplexity divergence if and only if its empirical receptive field (ERF) is bounded by and covered during training:

    ERF≤Ltr\text{ERF} \le L_{tr}

    For an RR-layer Transformer with local windowed attention size ww, the theoretical receptive field (TRF) is wRw R. When TRF≤Ltr\text{TRF} \le L_{tr}, the model extrapolates trivially because all accessible token interactions were observed during training. When TRF>Ltr\text{TRF} > L_{tr}, the model still extrapolates successfully if ERF≤Ltr\text{ERF} \le L_{tr}. Conversely, if ERF>Ltr\text{ERF} > L_{tr}, evaluating on Lex>LtrL_{ex} > L_{tr} causes perplexity to plateau or explode. Increasing LtrL_{tr} to satisfy Ltr≥ERFL_{tr} \ge \text{ERF} recovers stable extrapolation for previously failing configurations.

  4. Knowl 4 — Sandwich Temporal Bias Vectorized Construction

    algorithm

    The parameter-free multi-head temporal bias matrix for Sandwich positional embeddings across a maximum sequence length LL and attention head count HH is constructed as follows:

    import numpy as np
    
    def compute_sandwich_bias(seq_len: int, num_heads: int, bar_d: int = 128, base: float = 1e4) -> np.ndarray:
        positions = np.arange(seq_len)[:, None]
        i = np.arange(bar_d // 2)
        
        # Generate sine and cosine positional embeddings
        angles = positions / (base ** (2 * i / bar_d))
        pos_embs = np.concatenate([np.sin(angles), np.cos(angles)], axis=-1)
        
        # Inner product yields distance-dependent cosine sums, shifted by bar_d / 2
        sandwich = np.matmul(pos_embs, pos_embs.T) - (bar_d / 2.0)
        
        # Compute head compression ratios: h = n * 8 / H for n in 1..H
        compression_ratios = np.arange(1, num_heads + 1) * (8.0 / num_heads)
        
        # Broadcast across heads: shape (num_heads, seq_len, seq_len)
        multi_head_bias = sandwich[None, :, :] / compression_ratios[:, None, None]
        return multi_head_bias
    

    This tensor requires no learnable parameters and can be precomputed once before training and inference.

  5. Knowl 5 — Logarithmic Temporal Bias and Equivalence to KERPLE and T5

    theoretical result

    The temporal attention bias matrix of Sandwich positional embeddings with dˉ=128\bar{d}=128 follows a logarithmic decay curve with respect to relative token distance ∣m−n∣|m - n|. A least-squares logarithmic fit to the temporal bias yields:

    Bias(m,n)≈−0.825⋅log⁡(1+∣m−n∣)−0.8\text{Bias}(m, n) \approx -0.825 \cdot \log(1 + |m - n|) - 0.8

    Because the softmax operation is shift-invariant, constant offsets are eliminated by normalization. This connects Sandwich directly to other relative positional embedding architectures:

    1. KERPLE: The generalized logarithmic kernel of KERPLE is formulated as c−r1log⁡(1+r2∣m−n∣)c - r_1 \log(1 + r_2 |m - n|). Setting r1=0.825r_1 = 0.825 and r2=1.0r_2 = 1.0 turns Sandwich into a parameter-free realization of KERPLE.
    2. T5 Relative Bias: T5 implements logarithmic binning where nearby relative distances receive distinct learned biases while distant positions are assigned to shared, constant bins.

    The logarithmic plateau at large ∣m−n∣|m - n| assigns uniform negative biases to distant tokens, allowing the model to perform averaging and denoising over historical tokens beyond LtrL_{tr} rather than strictly masking them out.

  6. Knowl 6 — Fixed-Target Evaluation Protocol for Length Extrapolation

    experimental setup

    To evaluate length extrapolation without confounding factors from position-dependent token prediction difficulty (the "early token" effect), perplexity is computed strictly on a fixed target set of final tokens across varying context lengths.

    From an evaluation dataset, N=1000N = 1000 text segments of length LexL_{ex} are sampled. For each segment i∈{1,…,N}i \in \{1, \dots, N\}, the model predicts the probability pip_i of only the final (LexL_{ex}-th) token given a context prefix of length L−1L - 1 where L∈[Ltr,Lex]L \in [L_{tr}, L_{ex}]:

    PPL=exp⁡(1N∑i=1N−log⁡pi)\text{PPL} = \exp\left(\frac{1}{N} \sum_{i=1}^N -\log p_i\right)

    This guarantees that identical tokens are evaluated across all evaluation sequence lengths LexL_{ex}, ensuring that changes in perplexity reflect the utility of the extended context.

  7. Knowl 7 — Gradient Concentration and Overfitting in Absolute and Rotary Positional Embeddings

    empirical result

    Cumulative normalized gradient analysis demonstrates why Sinusoidal absolute positional embeddings (APE) and Rotary Position Embeddings (RoPE) fail to extrapolate to evaluation lengths Lex>LtrL_{ex} > L_{tr}:

    1. Sinusoidal APE Overfitting: When predicting the target token at position Lex=2048L_{ex} = 2048, a Sinusoidal Transformer trained with Ltr=512L_{tr} = 512 exhibits a sharp gradient concentration spike precisely at position 2048−512=15362048 - 512 = 1536. When trained with Ltr=128L_{tr} = 128, the gradient spike shifts exactly to position 2048−128=19202048 - 128 = 1920. This confirms that self-attention layers overfit to the exact set of absolute positional vectors observed during training.
    2. Rotary Embeddings: RoPE rotates query and key embeddings using position-specific rotation matrices. Its cumulative gradient does not decay within the most recent LtrL_{tr} tokens, violating the ERF coverage condition (ERF≤Ltr\text{ERF} \le L_{tr}) and leading to severe perplexity degradation as LexL_{ex} increases.

    By discarding absolute position-to-content cross-terms and keeping only relative displacement cosine sums, relative positional embeddings such as Sandwich eliminate positional overfitting.

  8. Knowl 8 — Language Modeling Perplexity Comparison across Evaluation Context Lengths

    data/table

    12-layer, 12-head, 768-hidden dimension (162M parameter) Transformer language models were trained on Ltr=512L_{tr} = 512 for 50,000 steps across OpenWebText2, ArXiv, and GitHub datasets, and evaluated across extended sequence lengths Lex∈{512,1024,2048,4096,8192}L_{ex} \in \{512, 1024, 2048, 4096, 8192\} averaged over 5 random seeds.

    Dataset Method Lex=512L_{ex}=512 Lex=1024L_{ex}=1024 Lex=2048L_{ex}=2048 Lex=4096L_{ex}=4096 Lex=8192L_{ex}=8192
    OpenWebText2 Sandwich 23.5±3.823.5 \pm 3.8 23.0±3.623.0 \pm 3.6 23.3±3.523.3 \pm 3.5 23.8±3.323.8 \pm 3.3 24.7±3.424.7 \pm 3.4
    ALiBi 22.8±3.322.8 \pm 3.3 23.3±3.423.3 \pm 3.4 23.5±3.323.5 \pm 3.3 23.5±3.323.5 \pm 3.3 23.5±3.323.5 \pm 3.3
    Sinusoidal 26.0±1.026.0 \pm 1.0 1416814168 2037020370 4200342003 6786967869
    Rotary 23.0±3.423.0 \pm 3.4 6161 9696 232232 343343
    KERPLE 22.6±3.522.6 \pm 3.5 22.0±3.322.0 \pm 3.3 21.9±3.121.9 \pm 3.1 22.1±2.922.1 \pm 2.9 22.3±2.922.3 \pm 2.9
    T5 22.6±3.622.6 \pm 3.6 22.2±3.322.2 \pm 3.3 23.0±3.123.0 \pm 3.1 26.8±3.226.8 \pm 3.2 38.6±7.238.6 \pm 7.2
    ArXiv Sandwich 5.27±0.335.27 \pm 0.33 5.05±0.335.05 \pm 0.33 5.02±0.345.02 \pm 0.34 5.15±0.395.15 \pm 0.39 5.28±0.445.28 \pm 0.44
    ALiBi 5.25±0.335.25 \pm 0.33 5.41±0.365.41 \pm 0.36 5.58±0.405.58 \pm 0.40 5.58±0.405.58 \pm 0.40 5.58±0.405.58 \pm 0.40
    Sinusoidal 5.85.8 10701070 17841784 1805018050 4410044100
    Rotary 5.25±0.335.25 \pm 0.33 16.0216.02 33.7633.76 71.9671.96 111111
    KERPLE 5.22±0.375.22 \pm 0.37 4.95±0.344.95 \pm 0.34 4.83±0.354.83 \pm 0.35 4.84±0.344.84 \pm 0.34 4.90±0.334.90 \pm 0.33
    T5 5.16±0.375.16 \pm 0.37 4.91±0.354.91 \pm 0.35 4.92±0.354.92 \pm 0.35 5.35±0.365.35 \pm 0.36 6.74±0.906.74 \pm 0.90
    GitHub Sandwich 2.88±0.122.88 \pm 0.12 2.71±0.092.71 \pm 0.09 2.69±0.112.69 \pm 0.11 2.73±0.122.73 \pm 0.12 2.79±0.152.79 \pm 0.15
    ALiBi 2.83±0.112.83 \pm 0.11 2.97±0.112.97 \pm 0.11 3.01±0.103.01 \pm 0.10 3.01±0.103.01 \pm 0.10 3.01±0.103.01 \pm 0.10
    Sinusoidal 4.04.0 83428342 91799179 1101711017 1127011270
    Rotary 2.82±0.112.82 \pm 0.11 3.86±0.253.86 \pm 0.25 5.94±0.645.94 \pm 0.64 11.1±1.5511.1 \pm 1.55 20.2±2.7520.2 \pm 2.75
    KERPLE 2.81±0.142.81 \pm 0.14 2.67±0.102.67 \pm 0.10 2.65±0.102.65 \pm 0.10 2.70±0.092.70 \pm 0.09 2.75±0.082.75 \pm 0.08
    T5 2.76±0.142.76 \pm 0.14 2.61±0.082.61 \pm 0.08 2.65±0.052.65 \pm 0.05 2.91±0.122.91 \pm 0.12 3.68±0.503.68 \pm 0.50

    Sandwich achieves lower perplexities at Lex>Ltr=512L_{ex} > L_{tr} = 512 than at Lex=512L_{ex} = 512 on all three datasets (e.g., dropping from 5.275.27 to 5.025.02 on ArXiv), showing effective use of context beyond LtrL_{tr}. ALiBi perplexity degrades and then plateaus, while Sinusoidal and Rotary diverge.

  9. Knowl 9 — ALiBi Attention Slope Sensitivity and Windowing Equivalence

    empirical result

    Systematic manipulation of ALiBi slope parameters (1/2h1/2^h) reveals the structural mechanism underlying its length extrapolation:

    1. Slope Shift: Shifting head slopes by a constant Δ\Delta (h=n⋅8H+Δh = n \cdot \frac{8}{H} + \Delta) causes perplexity explosion when Δ≥6\Delta \ge 6 on evaluation lengths Lex>Ltr=512L_{ex} > L_{tr} = 512. Diversity of slopes across heads is not sufficient to guarantee extrapolation; the magnitude of the slope is the governing factor.
    2. Slope Equalization: Setting identical slopes across all heads (h∈{0,2,4,6,8}h \in \{0, 2, 4, 6, 8\}) demonstrates that only steep slopes (h≤4h \le 4, where 1/2h≥1/161/2^h \ge 1/16) extrapolate successfully. Steep linear slopes impose an effective window that suppresses attention weights between distant token pairs.
    3. Equivalence to Windowed Attention: Explicit local windowed attention with window size w∈[40,320]w \in [40, 320] exhibits identical behavior to ALiBi: models with w≤100w \le 100 maintain stable perplexities across Lex∈[512,8192]L_{ex} \in [512, 8192] (e.g., perplexity 7.027.02 at w=100w=100 on ArXiv), while models with w≥160w \ge 160 diverge when Lex>LtrL_{ex} > L_{tr}.
    4. Saturation Behavior: In ALiBi, perplexities do not decrease when Lex>LtrL_{ex} > L_{tr}; they rise from Lex=512L_{ex}=512 to Lex=2048L_{ex}=2048 and remain flat thereafter, showing that ALiBi ignores tokens outside its empirical receptive field rather than utilizing them.
  10. Knowl 10 — Fixed-Window Key-Value Caching for Linear-Complexity Inference

    model/method

    Because relative positional embeddings such as ALiBi and Sandwich have an empirical receptive field bounded within a finite window wˉ\bar{w}, autoregressive generation at long evaluation lengths Lex≫LtrL_{ex} \gg L_{tr} can be computed with O(wˉ×Lex)O(\bar{w} \times L_{ex}) time complexity instead of standard O(Lex2)O(L_{ex}^2).

    For a fixed cache window size wˉ=2048\bar{w} = 2048:

    1. Key and value vectors km,vmk_m, v_m are cached for the first wˉ\bar{w} generated tokens (m∈[1,wˉ]m \in [1, \bar{w}]).
    2. When generating token wˉ+1\bar{w} + 1, the earliest cached vectors k1,v1k_1, v_1 are discarded, maintaining an active cache of size wˉ\bar{w}.
    3. Because relative positional biases depend only on relative position differences within the sliding window, previously computed keys and values do not need to be re-encoded.

    With wˉ=2048\bar{w} = 2048, ALiBi maintains constant perplexities up to Lex=16384L_{ex} = 16384 (23.523.5 on OpenWebText2, 5.595.59 on ArXiv, 3.013.01 on GitHub), and Sandwich maintains stable perplexities (24.124.1 on OpenWebText2, 5.355.35 on ArXiv, 2.812.81 on GitHub).

  11. Knowl 11 — Recency Bias Limitation of Decaying Relative Positional Embeddings

    limitation

    Relative positional embeddings with distance-decaying temporal biases (including ALiBi, Sandwich, KERPLE, and T5) impose an intrinsic architectural recency bias that concentrates attention weights primarily on the most recent tokens.

    While this recency bias matches the local dependency structure of natural language, it fails on algorithmic tasks where distant and recent tokens are equally critical. In bit-string parity prediction—determining whether an input binary sequence contains an even or odd number of ones—every bit across the sequence has equal importance. Distance-decaying relative positional embeddings downweight early tokens, preventing Transformers from extrapolating length on parity prediction and other non-recency tasks.

Coverage note — No substantial contributed material was omitted from the knowls.

References

  1. 1.Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020. ETC: Encoding long and structured inputs in transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 268–284, Online. Association for Computational Linguistics.
  2. 2.Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Shivanshu Purohit, Tri Songz, Wang Phil, and Samuel Weinbach. 2021. GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch.
  3. 3.Cem Anil, Yuhuai Wu, Anders Johan Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Venkatesh Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022. Exploring length generalization in large language models. In Advances in Neural Information Processing Systems.
  4. 4.Andre Araujo, Wade Norris, and Jack Sim. 2019. Computing receptive fields of convolutional neural networks. Distill. Https://distill.pub/2019/computing-receptive-fields.
  5. 5.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  6. 6.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer.
  7. 7.Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A Hasegawa-Johnson, and Thomas S Huang. 2017. Dilated recurrent neural networks. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  8. 8.Ta-Chung Chi, Ting-Han Fan, Peter J Ramadge, and Alexander I Rudnicky. 2022. Kerple: Kernelized relative positional embedding for length extrapolation. arXiv preprint arXiv:2205.09921.
  9. 9.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. 2017. Deformable convolutional networks. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 764–773.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  11. 11.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  12. 12.Ankit Gupta and Jonathan Berant. 2020. GMAT: global memory augmentation for transformers. CoRR, abs/2006.03274.
  13. 13.Wenjie Luo, Yujia Li, Raquel Urtasun, and Richard Zemel. 2016. Understanding the effective receptive field in deep convolutional neural networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, page 4905–4913, Red Hook, NY, USA. Curran Associates Inc.
  14. 14.Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernocky, and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In Interspeech, volume 2, pages 1045–1048. Makuhari.
  15. 15.Tomas Mikolov and Geoffrey Zweig. 2012. Context dependent recurrent neural network language model. In 2012 IEEE Spoken Language Technology Workshop (SLT), pages 234–239.
  16. 16.Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. Cite arxiv:1609.03499.
  17. 17.Ofir Press. 2022. The use case for relative position embeddings.
  18. 18.Ofir Press, Noah Smith, and Mike Lewis. 2022. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations.
  19. 19.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  20. 20.Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. 2021. Do vision transformers see like convolutional neural networks? In Advances in Neural Information Processing Systems.
  21. 21.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  22. 22.Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. 2021. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864.
  23. 23.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  24. 24.Hang Yan, Bocao Deng, Xiaonan Li, and Xipeng Qiu. 2019. Tener: adapting transformer encoder for named entity recognition. arXiv preprint arXiv:1911.04474.
  25. 25.Fisher Yu and Vladlen Koltun. 2016. Multi-scale context aggregation by dilated convolutions. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.
  26. 26.Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big bird: Transformers for longer sequences. CoRR, abs/2007.14062.
  27. 27.Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014. Recurrent neural network regularization. Cite arxiv:1409.2329.

Citation

MLA
Chi, T.-C., et al. “Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 13522–37, https://doi.org/10.18653/v1/2023.acl-long.756.
APA
Chi, T.-C., Fan, T.-H., Rudnicky, A., & Ramadge, P. (2023). Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13522–13537. https://doi.org/10.18653/v1/2023.acl-long.756
Chicago
Chi, T.-C., T.-H. Fan, A. Rudnicky, and P. Ramadge. 2023. “Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13522–37. https://doi.org/10.18653/v1/2023.acl-long.756.
Harvard
Chi, T.-C. et al. (2023) “Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13522–13537. Available at: https://doi.org/10.18653/v1/2023.acl-long.756.
Vancouver
1. Chi T-C, Fan T-H, Rudnicky A, Ramadge P (2023) Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13522–13537

BibTeX

@inproceedings{chi-etal-2023-dissecting,
    title = "Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis",
    author = "Chi, Ta-Chung  and
      Fan, Ting-Han  and
      Rudnicky, Alexander  and
      Ramadge, Peter",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.756/",
    doi = "10.18653/v1/2023.acl-long.756",
    pages = "13522--13537"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/