A Length-Extrapolatable Transformer
Yutao SunLi DongBarun PatraShuming MaShaohan HuangAlon BenhaimVishrav ChaudharyXia SongFuru Wei
Proposes the Length-Extrapolatable Transformer, which combines a decay-augmented relative position embedding with blockwise causal attention to continually lower perplexity on sequences far longer than those seen during training.
Modern natural language processing relies heavily on Transformer architectures, but pre-training these models on long sequences is prohibitively expensive in terms of computational resources and memory. Standard architectures suffer severe performance degradation when processing text inputs longer than the maximum length encountered during training. Enabling models to generalize effectively from short training sequences to longer deployment inputs—a capability known as length extrapolation—is essential for practical applications that involve long documents, transcripts, and books.
The article aims to design, evaluate, and demonstrate an enhanced architecture, named the Length-Extrapolatable (LEX) Transformer. This framework trains on short sequences while maintaining high accuracy and continuously improving performance when processing significantly longer sequences during inference.
To address this challenge, the authors defined a formal metric called attention resolution to evaluate how effectively an attention mechanism distinguishes relative token positions. They then developed the LEX Transformer by integrating two key components: an Extrapolatable Position Embedding (XPOS), which adds exponential decay to vector rotations to stabilize position calculations, and blockwise causal attention, which segments attention into manageable blocks during inference. The authors validated their approach by pre-training medium-scale models from scratch on sequences of 1,024 tokens and evaluating perplexity across sequence lengths extending up to 8,192 tokens across benchmark datasets including PG22, QMSum, arXiv, and NarrativeQA.
The evaluation produced several critical findings. First, the LEX Transformer was the only architecture tested whose perplexity consistently decreased as input lengths scaled up to 8,192 tokens, verifying that the model successfully leverages longer context. Second, baseline Transformer models degraded severely on long sequences, with perplexity on the PG22 dataset escalating from 30.54 at 1,024 tokens to over 12,700 at 8,192 tokens. Third, alternative relative position approaches struggled at extended lengths; rotary position embeddings suffered catastrophic degradation (reaching a perplexity of 458.83 at 8,192 tokens), while attention with linear biases saw perplexity rise from 26.01 at 2,048 tokens to 32.8 at 8,192 tokens. Finally, LEX outperformed existing methods on short, in-distribution texts (interpolation) while achieving the highest attention resolution scores across both standard and extended lengths.
These results demonstrate that engineering teams can deploy long-context language processing capabilities without incurring the heavy financial, hardware, and operational costs of training models directly on long sequences. The findings resolve a long-standing trade-off in sequence modeling by showing that robust long-range capability does not require sacrificing accuracy on standard-length inputs.
Engineering and research teams developing generative language models should consider adopting XPOS position embeddings and blockwise attention as drop-in upgrades. When implementing long-context inference, organizations can choose between blockwise causal attention for cache efficiency and sliding-window attention for slightly higher accuracy at the expense of computational throughput. Further work is recommended to adapt and validate these mechanisms in bidirectional encoder architectures.
The study's primary limitation is its focus on causal language models, leaving bidirectional architectures (such as masked language models) requiring further adaptation. Additionally, XPOS introduces approximately a 6 percent inference overhead compared to basic absolute position embeddings, although this is partially offset by accelerated training convergence. Confidence in the core findings remains high due to consistent empirical improvements across multiple diverse datasets.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). This paper establishes the core problem of length extrapolation in Transformers and introduces attention bias mechanisms (ALiBi) that the LeX Transformer directly builds upon and compares against.
- Paper: Self-Attention with Relative Position Representations, Peter Shaw et al. (2018). It introduces the foundational concept of relative position encodings in self-attention, which LeX adapts to formulate its high-resolution positional embeddings.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). It introduces relative positional encodings and segment-level recurrence for evaluating beyond fixed context windows, establishing essential concepts for sequence extrapolation.
- Paper: RoFormer: Enhanced Transformer with Rotary Position Embedding, Jianlin Su et al. (2024). It provides key theoretical grounding for relative position modeling via vector rotation (RoPE), representing a major baseline and foundation for modern extrapolation research.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). It defines the standard multi-head self-attention architecture and original absolute positional encodings that motivate the need for length-extrapolatable designs.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). It systematically categorizes efficient attention patterns, including blockwise and causal mechanisms used by LeX for inference efficiency.
- Paper: Ring Attention with Blockwise Transformers for Near-Infinite Context, Hao Liu et al. (2024). It extends blockwise attention computation across distributed device topologies to scale context lengths to millions of tokens without memory bottlenecks.
- Paper: Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings, Yoav Gelberg et al. (2025). It explores an alternative paradigm for length generalization by recalibrating and removing positional embeddings post-training to prevent out-of-distribution failure.
- Paper: LeRoPE: Learnable RoPE Frequencies Improve Language Modeling, Petros Karypis et al. (2026). It refines relative positional encoding frequencies by making them learnable parameters to improve language model scaling and context utilization.
- Paper: World Model on Million-Length Video And Language With Blockwise RingAttention, Hao Liu 0055 et al. (2025). It scales blockwise attention architectures to multimodal video and million-token language contexts using progressive stage training.
- Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). It addresses ultra-long context extrapolation by framing sequence processing as test-time training rather than purely static architectural positional design.
