LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
Yiran DingLi Lyna ZhangChengruidong ZhangYuanyuan XuNing ShangJiahang XuFan YangMao Yang
Presents LongRoPE, an efficient positional interpolation method that extends pre-trained large language model context windows up to two million tokens using only one thousand fine-tuning steps while preserving original short-context performance.
Modern large language models struggle to process long sequences because standard positional embeddings degrade when faced with sequence lengths beyond those seen during pre-training. Current extension techniques typically reach around 128k tokens, limited by the scarcity of million-token training texts, extreme compute costs, and performance drops on shorter inputs. The article addresses these bottlenecks by demonstrating a method called LongRoPE, which scales the effective context window of pre-trained models up to 2,048k (over two million) tokens while preserving original short-context capabilities.
To achieve this, the article utilizes an evolutionary search algorithm to discover non-uniform scaling factors for rotary position embeddings across two dimensions: different embedding channels and initial sequence token positions. Instead of directly fine-tuning models on million-token texts, the authors implement a progressive scaling workflow. They first search for optimal scaling factors to reach a 256k token window, fine-tune the model within this length for only 1,000 steps across 16 GPUs, and then perform a secondary non-fine-tuned search on the resulting model to extend its reach to 2,048k tokens. An additional search dynamically readjusts scaling factors during inference when processing standard 4k and 8k sequences.
The evaluation reveals several key findings. First, non-uniform positional scaling enables up to an 8x context extension without any fine-tuning, whereas previous methods experience severe perplexity spikes beyond a 2x extension. Second, on benchmark models such as LLaMA-2-7B and Mistral-7B, LongRoPE successfully reaches 2,048k tokens and achieves over 90% accuracy on passkey retrieval tasks across the full context window. Third, the dynamic short-window readjustment preserves model performance on standard evaluation suites, matching or slightly exceeding baseline accuracy on benchmarks like TruthfulQA.
These results demonstrate that organizations can deploy ultra-long context capabilities without prohibitive training compute or architectural redesigns. Because the method relies strictly on embedding rescaling, it integrates into existing Transformer serving infrastructure with minimal code changes. However, practitioners should note that inference compute and memory still scale heavily with sequence length, and evaluations beyond 256k were conducted on a limited set of long-text samples and synthetic retrieval tasks. Teams seeking to process large code repositories or document libraries should consider running targeted pilot validations on domain-specific long texts before broad production deployment.
- Paper: RoFormer: Enhanced Transformer with Rotary Position Embedding, Jianlin Su et al. (2024). Introduces Rotary Position Embedding (RoPE), the foundational positional representation that LongRoPE modifies and interpolates to extend context windows.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). Establishes core concepts of length extrapolation and the challenges transformers face when evaluating on contexts longer than their training sequences.
- Paper: A Length-Extrapolatable Transformer, Yutao Sun et al. (2023). Analyzes position embedding stabilization and attention resolution for length extrapolation beyond standard pre-training limits.
- Paper: Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis, Ta-Chung Chi et al. (2023). Provides fundamental receptive-field analysis explaining why standard rotary embeddings break under length extrapolation.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). Introduces standard multi-task benchmarks for evaluating language model performance and degradation across extended context lengths.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). Demonstrates critical failure modes and position biases in long-context language models, motivating architectural improvements in position encoding.
- Paper: LeRoPE: Learnable RoPE Frequencies Improve Language Modeling, Petros Karypis et al. (2026). Extends RoPE frequency modification strategies by making individual frequency multipliers learnable parameters during training.
- Paper: Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings, Yoav Gelberg et al. (2025). Explores an alternative paradigm for long-context scaling by recalibrating and entirely removing rotary position embeddings post-pretraining.
- Paper: Qwen2.5 Technical Report, Qwen et al. (2024). Applies multi-stage progressive context scaling and length extrapolation to push open-source language models to million-token contexts in production.
- Paper: World Model on Million-Length Video And Language With Blockwise RingAttention, Hao Liu 0055 et al. (2025). Builds upon million-token context expansion techniques to scale multimodal sequence modeling across massive video and text inputs.
- Paper: Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA, Minzheng Wang et al. (2024). Evaluates whether models with ultra-long context windows can reliably synthesize and reason over complex, multi-document enterprise inputs.
- Paper: DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, DeepSeek AI (2026). Demonstrates native million-token context intelligence using specialized hybrid attention architectures designed for ultra-long sequence efficiency.
