LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Yiran DingLi Lyna ZhangChengruidong ZhangYuanyuan XuNing ShangJiahang XuFan YangMao Yang

article2024ICML309 citations

Presents LongRoPE, an efficient positional interpolation method that extends pre-trained large language model context windows up to two million tokens using only one thousand fine-tuning steps while preserving original short-context performance.

Listen

Modern large language models struggle to process long sequences because standard positional embeddings degrade when faced with sequence lengths beyond those seen during pre-training. Current extension techniques typically reach around 128k tokens, limited by the scarcity of million-token training texts, extreme compute costs, and performance drops on shorter inputs. The article addresses these bottlenecks by demonstrating a method called LongRoPE, which scales the effective context window of pre-trained models up to 2,048k (over two million) tokens while preserving original short-context capabilities.

To achieve this, the article utilizes an evolutionary search algorithm to discover non-uniform scaling factors for rotary position embeddings across two dimensions: different embedding channels and initial sequence token positions. Instead of directly fine-tuning models on million-token texts, the authors implement a progressive scaling workflow. They first search for optimal scaling factors to reach a 256k token window, fine-tune the model within this length for only 1,000 steps across 16 GPUs, and then perform a secondary non-fine-tuned search on the resulting model to extend its reach to 2,048k tokens. An additional search dynamically readjusts scaling factors during inference when processing standard 4k and 8k sequences.

The evaluation reveals several key findings. First, non-uniform positional scaling enables up to an 8x context extension without any fine-tuning, whereas previous methods experience severe perplexity spikes beyond a 2x extension. Second, on benchmark models such as LLaMA-2-7B and Mistral-7B, LongRoPE successfully reaches 2,048k tokens and achieves over 90% accuracy on passkey retrieval tasks across the full context window. Third, the dynamic short-window readjustment preserves model performance on standard evaluation suites, matching or slightly exceeding baseline accuracy on benchmarks like TruthfulQA.

These results demonstrate that organizations can deploy ultra-long context capabilities without prohibitive training compute or architectural redesigns. Because the method relies strictly on embedding rescaling, it integrates into existing Transformer serving infrastructure with minimal code changes. However, practitioners should note that inference compute and memory still scale heavily with sequence length, and evaluations beyond 256k were conducted on a limited set of long-text samples and synthetic retrieval tasks. Teams seeking to process large code repositories or document libraries should consider running targeted pilot validations on domain-specific long texts before broad production deployment.

Cover for LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Abstract

Large context window is a desirable feature in large language models (LLMs). However, due to high fine-tuning costs, scarcity of long texts, and catastrophic values introduced by new token positions, current extended context windows are limited to around 128k tokens. This paper introduces LongRoPE that, for the first time, extends the context window of pre-trained LLMs to an impressive 2048k tokens, with up to only 1k fine-tuning steps at within 256k training lengths, while maintaining performance at the original short context window. This is achieved by three key innovations: (i) we identify and exploit two forms of non-uniformities in positional interpolation through an efficient search, providing a better initialization for fine-tuning and enabling an 8x extension in non-fine-tuning scenarios; (ii) we introduce a progressive extension strategy that first fine-tunes a 256k length LLM and then conducts a second positional interpolation on the fine-tuned extended LLM to achieve a 2048k context window; (iii) we readjust LongRoPE on 8k length to recover the short context window performance. Extensive experiments on LLaMA2 and Mistral across various tasks demonstrate the effectiveness of our method. Models extended via LongRoPE retain the original architecture with minor modifications to the positional embedding, and can reuse most pre-existing optimizations.

Table of Contents

  • 1 Introduction
  • 2 Non-uniformity in Positional Interpolation
  • 2.1 Preliminary
  • 2.2 Study on Non-uniform Positional Interpolation
  • 3 LongRoPE
  • 3.1 Problem Formulation
  • 3.2 Searching the Non-uniform Position Interpolation
  • 3.3 Extending LLM Context Window to 2048K
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Main Results
  • 4.3 Ablation Results
  • 5 Related Works
  • 6 Conclusion
  • References
  • A Appendix
  • A.1 Settings
  • A.2 Additional details on fine-tuning
  • A.3 Additional details on the search

Knowls

  1. Knowl 1 — Non-Uniform Positional Interpolation in LongRoPE

    model/method

    Rotary Position Embedding (RoPE) encodes a token at position index nn into a dd-dimensional representation using rotation frequencies hetai=heta−2i/d heta_i = heta^{-2i/d}, where i imes 2 eta = heta^{2/d}, and heta heta is the base frequency (default 1000010000). When extending the context window from original length LL to L′L' with extension ratio s=L′/Ls = L'/L, standard Linear Position Interpolation (PI) scales all rotary angles down by a single factor ss.

    LongRoPE introduces non-uniform positional interpolation along two axes: RoPE embedding dimensions and token sequence positions. The position- and dimension-dependent rescaling operator I( rac{1}{\lambda_i}, \hat{n}) is defined as:

    I(1λi,n^)={1,if n<n^1λi,if n≥n^I\left(\frac{1}{\lambda_i}, \hat{n}\right) = \begin{cases} 1, & \text{if } n < \hat{n} \\ \frac{1}{\lambda_i}, & \text{if } n \ge \hat{n} \end{cases}

    where i∈{0,1,…,d/2−1}i \in \{0, 1, \dots, d/2 - 1\} indexes the RoPE dimension pair, λi∈[1.0,1.25s]\lambda_i \in [1.0, 1.25s] denotes the dimension-specific rescale factor, and n^∈{0,1,2,4,8,12,16,20,24,28,32,64,128,256}\hat{n} \in \{0, 1, 2, 4, 8, 12, 16, 20, 24, 28, 32, 64, 128, 256\} is the starting token threshold. For the initial n^\hat{n} tokens (which act as attention sinks and carry large attention scores), no positional interpolation is applied (I=1I=1), preserving original rotary angles. For tokens at positions n≥n^n \ge \hat{n}, each dimension ii is rescaled by its searched factor 1/λi1/\lambda_i:

    RoPE(n)i=[cos⁡(I(1λi,n^)nβi),sin⁡(I(1λi,n^)nβi)]\text{RoPE}(n)_i = \left[ \cos\left(I\left(\frac{1}{\lambda_i}, \hat{n}\right) \frac{n}{\beta^i}\right), \sin\left(I\left(\frac{1}{\lambda_i}, \hat{n}\right) \frac{n}{\beta^i}\right) \right]

    This formulation minimizes positional information loss on critical dimensions and key starting token positions.

  2. Knowl 2 — Evolutionary Search Algorithm for Non-Uniform RoPE Rescale Factors

    algorithm

    To identify optimal RoPE rescale factors (λ0,λ1,…,λd/2−1)(\lambda_0, \lambda_1, \dots, \lambda_{d/2-1}) and starting token threshold n^\hat{n} that minimize next-token prediction perplexity L\mathcal{L} over validation text samples XX of length ≥L′\ge L', LongRoPE uses an evolutionary search algorithm enhanced by two domain-specific constraints:

    1. Monotonically Non-Decreasing Constraint: Guided by Neural Tangent Kernel (NTK) theory—which posits that high-frequency (lower) dimensions require less interpolation while low-frequency (higher) dimensions tolerate more—candidate parameter vectors are strictly constrained such that λi≤λi+1\lambda_i \le \lambda_{i+1} for all i∈{0,…,d/2−2}i \in \{0, \dots, d/2 - 2\}. Candidates violating this constraint are pruned prior to evaluation.
    2. Informed Population Initialization: The initial population of size PP includes the parameter vectors corresponding to Linear Position Interpolation (PI), Neural Tangent Kernel (NTK) scaling, and YaRN, with the remaining P−3P-3 individuals generated by mutating these seeds with probability pp.
    Input: Target LLM, input samples XX of length ≥L′\ge L', population size PP, mutation size N1N_1, crossover size N2N_2, iterations TT, mutation probability pp
    Output: Rescale factors (λ0,…,λd/2−1)(\lambda_0, \dots, \lambda_{d/2-1}) and threshold n^\hat{n} with lowest validation perplexity
    Topk = empty_set
    Population = Initialize_Population_With_Seeds(P, p)
    for iteration = 1 to T do
        for each individual in Population do
            Perplexity = Compute_Perplexity(LLM, individual, X)
        end for
        Topk = Update_Topk(Topk, Population, k=32)
        P_mutation = Mutate_With_Monotonicity_Constraint(Topk, N_1, p)
        P_crossover = Crossover_With_Monotonicity_Constraint(Topk, N_2)
        Population = P_mutation union P_crossover union Topk
    end for
    return individual in Topk with lowest Perplexity

    For context lengths up to 256k256\text{k}, search hyperparameters are P=64P=64, N1=16N_1=16, N2=16N_2=16, p=0.3p=0.3, T=40T=40, selecting the top 32 individuals on 5 validation samples from PG19. For lengths over 512k512\text{k}, population and offspring sizes are halved, evaluated on 3 validation samples from Books3.

  3. Knowl 3 — Progressive Context Extension Strategy to 2048k Tokens

    model/method

    Directly fine-tuning an LLM from a short pre-trained context window (e.g., 4k4\text{k}) to extreme context lengths (e.g., 2048k2048\text{k}, a 512×512\times extension) suffers from severe training instability, GPU memory constraints, and data scarcity. LongRoPE overcomes this via a multi-phase progressive extension pipeline requiring at most 1000 fine-tuning steps within a maximum training length of 256k256\text{k}:

    1. Pre-Trained Stage Search: Using evolutionary search on the pre-trained base model, non-uniform RoPE rescale factors are determined for intermediate context targets of 128k128\text{k} (32×32\times) and 256k256\text{k} (64×64\times).
    2. Staged Fine-Tuning: The base LLM is first fine-tuned for 400 steps with the 128k128\text{k} searched RoPE parameters on 128k128\text{k} sequence chunks. Fine-tuning then continues from the 128k128\text{k} checkpoint for an additional 600 steps using the 256k256\text{k} RoPE parameters on 256k256\text{k} sequence chunks (1000 total fine-tuning steps).
    3. Secondary Search to 2048k2048\text{k}: Taking advantage of the fact that LongRoPE enables an 8×8\times extension without fine-tuning, a secondary evolutionary search is performed directly on the 256k256\text{k} fine-tuned model to discover rescale factors for 2048k2048\text{k} (8×8\times extension over 256k256\text{k}, 512×512\times over base 4k4\text{k}), requiring no additional fine-tuning.
  4. Knowl 4 — Short Context Window Performance Recovery via Dynamic Rescaling

    model/method

    When an LLM's positional embeddings are scaled up by large factors (e.g., 512×512\times to support 2048k2048\text{k} tokens), positions in the original short context window (such as 4k4\text{k} and 8k8\text{k}) become compressed into an excessively narrow subspace, impairing short-sequence language modeling and benchmark accuracy.

    To recover short-context capabilities without modifying model weights:

    1. An auxiliary evolutionary search is executed on the final extended model to discover specialized RoPE rescale factors optimized specifically for short contexts (4k4\text{k} and 8k8\text{k}). The search space restricts the maximum allowable scale factor λi\lambda_i to smaller upper bounds reflecting the lesser degree of required interpolation.
    2. At inference time, the model dynamically swaps its RoPE rescale parameters based on the current input length: if the input sequence length is less than or equal to 8k8\text{k} tokens, the model applies the searched short-context RoPE factors; otherwise, it uses the 2048k2048\text{k} rescale factors.
  5. Knowl 5 — Extreme Context Language Modeling Perplexity on Books3 up to 2048k Tokens

    data/table

    Perplexity was evaluated on 20 randomly selected Books3 books (each >2048k>2048\text{k} tokens) using a sliding window of 256k256\text{k} tokens across evaluation lengths from 8k8\text{k} to 2048k2048\text{k} (1k=10241\text{k} = 1024 tokens). LongRoPE models extended from 128k128\text{k} and 256k256\text{k} fine-tuned checkpoints scale effectively up to 2048k2048\text{k} without perplexity explosion, whereas baseline methods fail rapidly beyond their training lengths:

    Model Target Method 8k 16k 32k 64k 128k 256k 512k 1024k 2048k
    LongLoRA 100k PI 6.99 6.80 6.66 6.59 20.57 246.45 >10^3 >10^4 >10^4
    Code LLaMA 100k NTK 7.68 7.49 7.38 7.88 9.80 98.30 >10^3 >10^4 >10^4
    YaRN (s=16) 64k YaRN 6.33 6.20 6.11 6.06 >10^4 >10^4 >10^4 >10^4 >10^4
    YaRN (s=32) 128k YaRN 6.38 6.25 6.16 6.11 6.12 >10^4 >10^4 >10^4 >10^4
    LongRoPE-LLaMA2 (ft=128k) 2048k LongRoPE 6.55 6.35 6.24 6.18 6.17 6.17 6.36 6.83 7.80
    LongRoPE-LLaMA2 (ft=256k) 2048k LongRoPE 6.81 6.66 6.31 6.27 6.21 6.17 6.17 6.35 7.08
    Mistral v0.1 8k - 6.32 66.61 >10^2 >10^3 >10^3 >10^3 - - -
    YaRN-Mistral (s=16) 64k YaRN 6.59 6.48 6.42 6.45 104.15 727.20 >10^3 >10^4 >10^4
    YaRN-Mistral (s=32) 128k YaRN 6.70 6.63 6.65 6.72 6.85 99.90 >10^3 >10^4 >10^4
    LongRoPE-Mistral (ft=128k) 2048k LongRoPE 6.64 6.48 6.39 6.45 6.64 7.08 7.71 8.93 12.78
    LongRoPE-Mistral (ft=256k) 2048k LongRoPE 6.63 6.48 6.38 6.43 6.68 7.15 7.98 9.42 13.71

    LongRoPE-LLaMA2-7B (ft=256k) achieves monotonic perplexity reduction from 8k8\text{k} (6.816.81) down to 256k256\text{k} and 512k512\text{k} (6.176.17), with only mild perplexity increases at 1024k1024\text{k} (6.356.35) and 2048k2048\text{k} (7.087.08).

  6. Knowl 6 — Million-Token Passkey Retrieval Accuracy

    empirical result

    Synthetic passkey retrieval was evaluated using a prompt where a random 5-digit number is embedded inside repeated filler sentences ('The grass is green. The sky is blue. The sun is yellow. Here we go. There and back again.') at a position uniformly sampled across the target context window. Models were tested over 10 iterations per evaluation context length.

    Baseline models (including LongLoRA-100k, Code LLaMA-100k, and YaRN-128k) experience an abrupt drop to 0% retrieval accuracy once evaluation lengths exceed 128k128\text{k} tokens.

    In contrast:

    • LongRoPE-LLaMA2-2048k (fine-tuned at 256k256\text{k}) maintains ≥90%\ge 90\% retrieval accuracy across the entire evaluation spectrum from 4k4\text{k} to 2048k2048\text{k} tokens.
    • LongRoPE-Mistral-2048k (fine-tuned at 128k128\text{k}) sustains 100%100\% accuracy from 4k4\text{k} up to 1800k1800\text{k} tokens, decreasing to 60%60\% accuracy at 2048k2048\text{k} tokens.
  7. Knowl 7 — Standard Benchmark Performance Retention within Original Context Window

    data/table

    LongRoPE-2048k models were evaluated on the Hugging Face Open LLM Leaderboard tasks within standard short contexts: 25-shot ARC-Challenge, 10-shot HellaSwag, 5-shot MMLU, and 0-shot TruthfulQA.

    Model Window ARC-c HellaSwag MMLU TruthfulQA Avg.
    Original LLaMA2-7B 4k 53.1 78.6 46.6 39.0 54.3
    Together 32k 47.6 76.1 43.3 39.2 51.6
    Code LLaMA 100k 42.4 64.8 40.1 37.1 46.1
    YaRN (s=16) 64k 52.4 78.7 42.4 38.2 52.9
    YaRN (s=32) 128k 52.2 78.5 41.8 37.4 52.5
    LongRoPE-LLaMA2-2048k (ft=128k) 2048k 52.9 76.5 43.4 38.8 52.9
    LongRoPE-LLaMA2-2048k (ft=256k) 2048k 51.0 75.3 39.6 37.3 50.8
    Original Mistral-7B 8k 60.6 83.2 63.6 42.6 62.5
    MistralLite 16k 59.2 81.6 50.4 38.3 57.4
    YaRN-Mistral (s=16) 64k 59.3 81.3 61.3 42.5 61.1
    YaRN-Mistral (s=32) 128k 59.0 80.5 60.5 42.5 60.6
    LongRoPE-Mistral-2048k (ft=128k) 2048k 59.0 81.2 61.3 43.1 61.2
    LongRoPE-Mistral-2048k (ft=256k) 2048k 59.2 80.9 61.1 42.2 60.9

    With dynamic short-context recovery enabled, LongRoPE-2048k models retain near-original accuracy across all tasks. For Mistral-7B, LongRoPE-2048k (ft=128k) achieves a TruthfulQA score of 43.1%43.1\%, exceeding the original base model (42.6%42.6\%).

  8. Knowl 8 — Ablation of Positional Interpolation Components and 8x Non-Fine-Tuned Extension

    empirical result

    Ablation experiments isolating the two non-uniformity components (RoPE dimension-wise λi\lambda_i and token position threshold n^\hat{n}) show distinct contributions to context extension without fine-tuning:

    1. Non-Uniformity in RoPE Dimensions: Extending non-fine-tuned LLaMA2-7B to 16k16\text{k} and 32k32\text{k} context sizes on PG19, linear interpolation (PI) yields perplexities of 14.8814.88 (16k16\text{k}) and 136.30136.30 (32k32\text{k}). Searching dimension-wise λi\lambda_i alone dramatically lowers perplexity to 7.287.28 (16k16\text{k}) and 13.0013.00 (32k32\text{k}).
    2. Non-Uniformity in Initial Token Positions: Adding the search for starting token threshold n^\hat{n} without interpolation further reduces PG19 perplexity to 7.227.22 (16k16\text{k}) and 11.5111.51 (32k32\text{k}).
    3. 8x Extension without Fine-Tuning: Exploiting both non-uniformities enables extending base models by an 8×8\times factor (from 4k4\text{k} to 32k32\text{k} on LLaMA2-7B) without fine-tuning, whereas standard linear interpolation, dynamic NTK, and YaRN experience sharp perplexity spikes past 2×2\times extension without fine-tuning.
    4. Diminishing Impact at Extreme Lengths: When extending the fine-tuned LLaMA2-256k model to 2048k2048\text{k} on Books3 without further fine-tuning, dimension-wise search achieves a perplexity of 7.087.08 (compared to 20.1720.17 for PI), while adding initial token threshold preservation yields the identical perplexity (7.087.08), indicating that preserving starting token angles becomes negligible at extreme multi-million token scales.
  9. Knowl 9 — Efficacy of Secondary Interpolation and Short-Context Recovery

    data/table

    Ablation studies confirm the standalone effectiveness of (1) secondary positional interpolation on the 256k256\text{k} fine-tuned model and (2) short-context parameter readjustment.

    Secondary Interpolation Performance on Books3 (LLaMA2-7B ft=256k without further fine-tuning):

    Model Extension Method 512k 1024k 2048k
    LLaMA2-7B (ft=256k) PI 6.60 8.73 20.17
    LLaMA2-7B (ft=256k) YaRN 6.39 6.79 8.27
    LLaMA2-7B (ft=256k) LongRoPE 6.17 6.35 7.08

    Impact of Short-Context RoPE Readjustment on Proof-Pile Perplexity and Standard LLM Benchmark Average:

    Model With Recovery Proof-Pile 4k Proof-Pile 8k LLM Bench Avg. (%)
    LLaMA2-7B-2048k (ft=128k) No 4.16 3.72 49.3
    LLaMA2-7B-2048k (ft=128k) Yes 3.71 3.50 52.9
    LLaMA2-7B-2048k (ft=256k) No 4.51 3.82 47.9
    LLaMA2-7B-2048k (ft=256k) Yes 3.85 3.65 50.8

    LongRoPE secondary search restrains perplexity growth at 2048k2048\text{k} (7.087.08 vs 20.1720.17 for PI and 8.278.27 for YaRN). Short-context recovery improves average 4-task benchmark accuracy by +3.6%+3.6\% (ft=128k) and +2.9%+2.9\% (ft=256k).

  10. Knowl 10 — Comparison of Fine-Tuning Pathways for 256k Context Extension

    data/table

    To evaluate the progression to a 256k256\text{k} context window, three fine-tuning pathways on LLaMA2-7B were evaluated on Proof-Pile test perplexity:

    1. Training directly on 128k128\text{k} text length with 256k256\text{k} RoPE factors from base LLaMA2-7B.
    2. Direct fine-tuning on 256k256\text{k} text length with 256k256\text{k} RoPE factors from base LLaMA2-7B.
    3. Progressive two-stage fine-tuning: fine-tuning on 128k128\text{k} text length for 400 steps, followed by fine-tuning on 256k256\text{k} text length for 600 steps with updated 256k256\text{k} RoPE factors.
    Method (Training Length, Window, Init Model) 32768 65536 98304 131072 262144
    (128k, 256k, LLaMA2-7B) 9.75 6.56 5.15 5.19 2.21
    (256k, 256k, LLaMA2-7B) 4.51 2.87 2.53 2.39 1.95
    (256k, 256k, LLaMA2-7B [ft=128k]) 2.66 2.38 2.28 2.26 1.87

    Progressive fine-tuning from a 128k128\text{k} intermediate checkpoint achieves lower perplexity across all evaluation intervals (1.871.87 at 262k262\text{k} evaluation context length vs. 1.951.95 and 2.212.21), while converging substantially faster and maintaining lower initial loss during the 256k256\text{k} training phase.

Coverage note — None was omitted. All major methodological contributions (non-uniform interpolation, evolutionary search algorithm with monotonicity constraint, progressive context extension, short-context recovery) and primary empirical findings (perplexity scaling up to 2048k, passkey retrieval, standard benchmarks, and ablations) are fully covered.

References

  1. 1.Long-data collections, 2024. URL https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T.
  2. 2.Amazon. Mistrallite, 2023. URL https://huggingface.co/amazon/MistralLite.
  3. 3.Azerbayev, Z., Ayers, E., and Piotrowski, B. Proofpile, 2022. URL https://github.com/zhangir-azerbayev/ProofNet.
  4. 4.Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pp. 2206–2240. PMLR, 2022.
  5. 5.Chen, S., Wong, S., Chen, L., and Tian, Y. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023a.
  6. 6.Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J. Longlora: Efficient fine-tuning of long-context large language models. arXiv:2309.12307, 2023b.
  7. 7.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018.
  8. 8.Computer, T. Redpajama: An open source recipe to reproduce llama training dataset, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  9. 9.Dao, T. FlashAttention-2: Faster attention with better parallelism and work partitioning. 2023.
  10. 10.Face, H. Open llm leaderboard, 2024. URL https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard.
  11. 11.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C. The Pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  12. 12.Guo, Z., Zhang, X., Mu, H., Heng, W., Liu, Z., Wei, Y., and Sun, J. Single path one-shot neural architecture search with uniform sampling. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVI 16, pp. 544–560. Springer, 2020.
  13. 13.Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S. Lm-infinite: Simple on-the-fly length generalization for large language models. arXiv preprint arXiv:2308.16137, 2023.
  14. 14.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  15. 15.Huang, X., Zhang, L. L., Cheng, K.-T., Yang, F., and Yang, M. Fewer is more: Boosting llm reasoning with reinforced context pruning. 2023. URL https://api.semanticscholar.org/CorpusID:266210460.
  16. 16.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. Advances in neural information processing systems, 31, 2018.
  17. 17.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7b, 2023.
  18. 18.Lin, S., Hilton, J., and Evans, O. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  19. 19.Lin, Z., Miao, Y., Liu, G., Shi, X., Zhang, Q., Yang, F., Maleki, S., Zhu, Y., Cao, X., Li, C., et al. Superscaler: Supporting flexible dnn parallelization via a unified abstraction. arXiv preprint arXiv:2301.08984, 2023.
  20. 20.Liu, X., Yan, H., Zhang, S., An, C., Qiu, X., and Lin, D. Scaling laws of rope-based extrapolation. arXiv preprint arXiv:2310.05209, 2023.
  21. 21.LocalLLaMA. Dynamically scaled rope further increases performance of long context llama with zero fine-tuning, 2023a. URL https://www.reddit.com/r/LocalLLaMA/comments/14mrgpr/dynamically_scaled_rope_further_increases/.
  22. 22.LocalLLaMA. Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degration, 2023b. URL https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/.
  23. 23.Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P. Self-refine: Iterative refinement with self-feedback, 2023.
  24. 24.Mohtashami, A. and Jaggi, M. Landmark attention: Random-access infinite context length for transformers. arXiv preprint arXiv:2305.16300, 2023.
  25. 25.OpenAI, :, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mely, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B. Gpt-4 technical report, 2023.
  26. 26.Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pp. 1–22, 2023.
  27. 27.Peng, B., Quesnelle, J., Fan, H., and Shippole, E. Yarn: Efficient context window extension of large language models. arXiv preprint arXiv:2309.00071, 2023.
  28. 28.Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P. Compressive transformers for long-range sequence modelling. arXiv preprint, 2019. URL https://arxiv.org/abs/1911.05507.
  29. 29.Ratner, N., Levine, Y., Belinkov, Y., Ram, O., Abend, O., Karpas, E., Shashua, A., Leyton-Brown, K., and Shoham, Y. Parallel context windows improve in-context learning of large language models. arXiv preprint arXiv:2212.10947, 2022.
  30. 30.Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Defossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G. Code llama: Open foundation models for code, 2023.
  31. 31.Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021.
  32. 32.Tancik, M., Srinivasan, P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J., and Ng, R. Fourier features let networks learn high frequency functions in low dimensional domains. Advances in Neural Information Processing Systems, 33: 7537–7547, 2020.
  33. 33.Together, 2023. URL https://huggingface.co/togethercomputer/LLaMA-2-7B-32K.
  34. 34.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models, 2023.
  35. 35.Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłos, P. Focused transformer: Contrastive training for context scaling. 2023.
  36. 36.Wang, W., Dong, L., Cheng, H., Liu, X., Yan, X., Gao, J., and Wei, F. Augmenting language models with long-term memory. arXiv preprint arXiv:2306.07174, 2023.
  37. 37.Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. Efficient streaming language models with attention sinks. arXiv, 2023.
  38. 38.Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. Effective long-context scaling of foundation models, 2023.
  39. 39.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  40. 40.Zhang, P., Liu, Z., Xiao, S., Shao, N., Ye, Q., and Dou, Z. Soaring from 4k to 400k: Extending llm’s context with activation beacon. arXiv preprint arXiv:2401.03462, 2024.
  41. 41.Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S. Pose: Efficient context window extension of llms via positional skip-wise training, 2023.

Citation

MLA
Ding, Y., et al. “LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens”. arXiv, 2024, http://arxiv.org/abs/2402.13753v1.
APA
Ding, Y., Zhang, L. L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., & Yang, M. (2024). LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens. arXiv. http://arxiv.org/abs/2402.13753v1
Chicago
Ding, Y., L. L. Zhang, C. Zhang, et al. 2024. “LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens”. arXiv. http://arxiv.org/abs/2402.13753v1.
Harvard
Ding, Y. et al. (2024) “LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.13753v1.
Vancouver
1. Ding Y, Zhang LL, Zhang C, Xu Y, Shang N, Xu J, Yang F, Yang M (2024) LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens. arXiv

BibTeX

@article{ding2024longrope,
  title = {LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens},
  author = {Ding, Yiran and Zhang, Li Lyna and Zhang, Chengruidong and Xu, Yuanyuan and Shang, Ning and Xu, Jiahang and Yang, Fan and Yang, Mao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.13753v1},
  eprint = {2402.13753}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/