Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis
Ta-Chung ChiTing-Han FanAlexander RudnickyPeter J. Ramadge
Explains why transformer length extrapolation succeeds by analyzing empirical receptive fields and introduces Sandwich, a parameter-free relative positional embedding that effectively utilizes context beyond the training sequence length.
Training modern transformer language models on very long sequences requires substantial compute and memory resources due to quadratic cost scaling. To keep training costs practical while enabling models to handle long documents at testing time, developers rely on length extrapolation, which allows a model trained on short text sequences to maintain performance on much longer sequences. While specific positional embedding designs like Attention with Linear Biases (ALiBi) have achieved widespread empirical success in enabling extrapolation, the underlying mechanisms governing why these techniques succeed or fail have remained poorly understood.
The article evaluates why transformer models extrapolate to longer sequences by analyzing their receptive field—the span of surrounding text that actually influences a model's predictions. The investigation demonstrates the structural connection between linear relative positional biases and windowed attention, explains prior failure modes of common positional embeddings, and introduces an improved positional embedding design.
To conduct this evaluation, the researchers introduced a cumulative normalized gradient tool to measure a model's empirical receptive field across various configurations. They trained 12-layer, 162-million parameter models on sequences of 512 tokens using three diverse datasets—academic papers from ArXiv, open web text from OpenWebText2, and code from GitHub—and tested their language modeling perplexity across extended lengths ranging up to 8,192 tokens.
The analysis yielded four major findings. First, popular methods like ALiBi succeed primarily because steep slopes restrict the model to a narrow, windowed receptive field; as long as the training sequence length fully covers this empirical receptive field, the model extrapolates smoothly by ignoring distant tokens rather than integrating them. Second, standard absolute and rotary positional embeddings fail at extrapolation because they overfit to the exact position indices seen during training and fail to constrain their receptive field. Third, the authors introduced "Sandwich," a new parameter-free relative positional embedding derived from simplified sinusoidal embeddings. Unlike ALiBi, Sandwich genuinely leverages context beyond the initial training length, achieving lower (better) perplexity on longer test sequences across multiple benchmarks. Fourth, Sandwich exhibits a logarithmic decaying temporal bias pattern similar to parameterized designs like KERPLE and T5, which provides a beneficial averaging and denoising effect over distant tokens.
These findings provide clear practical implications for designing and deploying large-scale language models. Relying on ALiBi enables predictable compute costs and efficient inference caching without performance degradation, but it leaves long-range document context largely unexploited. By contrast, adopting logarithmically decaying positional designs like Sandwich allows organizations to train cost-effectively on short sequences while gaining the actual performance benefits of longer contextual inputs during deployment, all without adding learnable parameters or runtime overhead.
For engineering and research teams developing foundation models, the article recommends adopting relative positional embeddings that incorporate logarithmic distance decay. When resource efficiency and training simplicity are paramount, teams should utilize parameter-free designs like Sandwich over linear-decay designs. If maximum task adaptability is required and additional computational overhead is acceptable, learnable counterparts such as KERPLE or T5 present viable alternatives. Further work should explore adding learnable compression ratios to Sandwich and testing these designs across broader architectures.
Decision-makers should interpret these results with an awareness of specific architectural limitations. The empirical evaluation focused on 162-million parameter models, and while the underlying receptive field mechanics are general, absolute scaling dynamics in multi-billion parameter regimes warrant ongoing validation. Furthermore, the logarithmic decay mechanism inherently introduces a recency bias that favors nearby context; while this aligns well with natural language text, caution is advised when applying these architectures to non-linguistic tasks where all sequence elements possess equal structural importance.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). This paper introduces Attention with Linear Biases (ALiBi), the primary length extrapolation baseline dissected and contextualized by the receptive field analysis in the source.
- Paper: RoFormer: Enhanced Transformer with Rotary Position Embedding, Jianlin Su et al. (2024). This paper presents Rotary Position Embedding (RoPE), whose extrapolation failures and receptive field dynamics are directly evaluated and compared against in the source.
- Paper: Self-Attention with Relative Position Representations, Peter Shaw et al. (2018). This work establishes the foundation of relative positional representations in self-attention that underlies the parameter-free relative embedding designs analyzed in the source.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). This foundational paper develops relative positional encodings and segment-level recurrence to model sequences beyond fixed training context lengths.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). This foundational text introduces standard sinusoidal positional encodings and the multi-head attention architecture whose length extrapolation behavior is analyzed.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). This paper formalizes techniques for tracking layer-wise attention flow and token attribution, providing methodological context for empirical receptive field analysis.
- Paper: Exploring Length Generalization in Large Language Models, Cem Anil et al. (2022). This study establishes key empirical failure modes of standard transformers when performing length generalization out of training distribution.
- Paper: A Length-Extrapolatable Transformer, Yutao Sun et al. (2023). This work develops the LEX Transformer and XPOS, implementing exponential distance decay on rotary embeddings to achieve genuine length extrapolation.
- Paper: LeRoPE: Learnable RoPE Frequencies Improve Language Modeling, Petros Karypis et al. (2026). This paper extends the study of positional frequency decay by introducing learnable frequency multipliers to improve rotary position encodings across sequence lengths.
- Paper: Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings, Yoav Gelberg et al. (2025). This study offers an alternative strategy to the positional embedding extrapolation problem by investigating the complete removal and post-training recalibration of positional embeddings.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). This paper evaluates the practical utilization of long contexts across language models, revealing positional biases and middle-context neglect in extended sequences.
- Paper: Ring Attention with Blockwise Transformers for Near-Infinite Context, Hao Liu et al. (2024). This work addresses system-level scaling to massive context windows by distributing blockwise causal attention across hardware devices in a ring topology.
- Paper: Qwen2.5 Technical Report, Qwen et al. (2024). This report details practical length extrapolation and progressive context extension strategies deployed in large-scale frontier open-weight language models.
- Paper: Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling, Liliang Ren et al. (2025). This architecture pairs sliding-window attention with recurrent state-space models to extrapolate effectively to sequence lengths far beyond initial training limits.
- Paper: Prefix Sliding for efficient test-time scaling, Niklas Muennighoff et al. (2026). This work explores test-time context management by retaining a static prefix alongside a sliding window to mitigate long-range attention distractions.
