Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation

Zichong LiChen LiangLiliang RenTuo ZhaoYelong ShenWeizhu Chen

article2026arXiv0 citations

Proposes RoPE-Perturbed Self-Distillation, a training regularizer that enforces prediction consistency across perturbed position embeddings to eliminate positional bias and improve long-context retrieval and extrapolation in large language models.

Listen

Large language models are increasingly deployed in real-world applications that require processing vast amounts of information at once, such as analyzing full codebases, synthesizing multi-document research, and powering search-based question answering. A common strategy to enable these capabilities is fine-tuning models on longer sequences. However, existing models often exhibit positional brittleness: their accuracy fluctuates widely depending on where critical evidence appears within the input text, frequently dropping when key facts are placed in the middle of long documents.

The article demonstrates that standard rotary position embedding mechanisms cause models to over-rely on brittle positional artifacts rather than core semantic meaning. To resolve this, the article introduces and evaluates RoPE-Perturbed Self-Distillation, a training regularizer designed to improve positional robustness and ensure reliable retrieval regardless of evidence placement.

The evaluated approach introduces a two-view training technique applied to open-source models, specifically Llama-3-8B and Qwen-3-4B, across context lengths ranging from 32,000 to 256,000 tokens. During training, the model processes each text sequence twice: once under standard positional numbering, and once under perturbed positional indices created by skipping or shifting positions. By minimizing a statistical difference (divergence) between the two resulting prediction streams, the model learns to align the perturbed predictions with the stable standard view, forcing it to focus on semantic content rather than exact numerical placement. Evaluations were conducted across standardized long-context benchmarks, realistic mixed-length training schedules, and subsequent instruction fine-tuning stages.

The evaluation produced several key findings. First, the proposed method substantially improved overall long-context accuracy, raising average benchmark scores on the RULER suite by up to 12.04 percentage points on Llama-3-8B at 64,000 tokens and by 2.71 percentage points on Qwen-3-4B at 256,000 tokens after supervised instruction tuning. Second, the method directly alleviated the performance drop in the middle of inputs, producing a much more uniform accuracy curve across all possible fact positions. Third, the models showed superior length extrapolation, outperforming standard baselines when tested on sequence lengths two to four times beyond their trained context window. Finally, these long-range improvements were achieved without degrading standard short-context reasoning and knowledge benchmarks.

These findings indicate that positional brittleness is not an inevitable limitation of long-context architectures, but a training deficiency that can be corrected at the objective level. Organizations deploying long-context models can reduce the operational risk of missed information in large-scale retrieval-augmented pipelines and multi-document workflows. Although the technique introduces an approximate 1.6-times computational overhead per training step due to the additional forward pass, wall-clock compute comparisons show that standard fine-tuning plateaus early, meaning this regularizer delivers significantly higher performance per training hour and improves data efficiency on scarce long-form datasets.

Engineering teams adapting language models to long context windows should adopt structure-preserving position perturbations, such as skip-based shifts, during continued pretraining. Teams can also combine this regularizer with complementary training techniques, like token-loss reweighting, which showed the highest combined benchmark scores in the study. Before broader deployment across non-standard architectures, practitioners should pilot the approach on their specific domain workflows, as different tasks show varying sensitivities to order-preserving versus non-order-preserving perturbations.

The findings are supported with high confidence through multiple benchmark evaluations, seed stability tests, and ablation studies across two distinct model families. However, certain limitations remain: the article specifically evaluates rotary position embeddings in decoder-only transformer models. While the underlying principle of index invariance is broadly applicable, further empirical validation is necessary before generalizing these results to non-rotary positional frameworks or unstructured document pipelines.

arXiv: 2604.14339
Cover for Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation

Abstract

Large language models (LLMs) increasingly operate in settings that require reliable long-context understanding, such as retrieval-augmented generation and multi-document reasoning. A common strategy is to fine-tune pretrained short-context models at the target sequence length. However, we find that standard long-context adaptation can remain brittle: model accuracy depends strongly on the absolute placement of relevant evidence, exhibiting high positional variance even when controlling for task format and difficulty.

We propose RoPE-Perturbed Self-Distillation, a training regularizer that improves positional robustness. The core idea is to form alternative "views" of the same training sequence by perturbing its RoPE indices -- effectively moving parts of the context to different positions -- and to train the model to produce consistent predictions across views via self-distillation. This encourages reliance on semantic signals instead of brittle position dependencies. Experiments on long-context adaptation of Llama-3-8B and Qwen-3-4B demonstrate consistent gains on long-context benchmarks, including up to 12.04% improvement on RULER-64K for Llama-3-8B and 2.71% on RULER-256K for Qwen-3-4B after SFT, alongside improved length extrapolation beyond the training context window.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Problem setting and notation
  • 2.2 RoPE-perturbed views via skip-based index shifts
  • 2.3 Self-distillation between perturbed and standard views
  • 2.4 Cyclic-shift perturbation as a working variant
  • 3 Experiments
  • 3.1 Experimental Setup
  • 3.2 Main Results on Long-Context Benchmarks
  • 3.3 Length Extrapolation via YaRN
  • 3.4 Does the Gain Survive Instruction SFT?
  • 3.5 Short-Context Performance
  • 3.6 Ablations and Analysis
  • 3.7 Compute-Matched Training Efficiency
  • 4 Related Work
  • 5 Conclusion
  • References
  • A Experimental Details
  • A.1 Training protocol and compute
  • A.2 Baselines and implementation notes
  • B Extended Results
  • B.1 Qwen RULER task breakdowns at 256K
  • B.2 Mix-length Training
  • B.3 Order-sensitive NIAH: Multi-value ordinal retrieval
  • B.4 Evaluation Variance Across Random Seeds
  • C Extended Related work discussion

Knowls

  1. Knowl 1 — RoPE-Perturbed Self-Distillation Objective

    model/method

    RoPE-Perturbed Self-Distillation is a long-context training regularizer designed to reduce the sensitivity of autoregressive Transformers to token position assignments by enforcing prediction consistency across perturbed Rotary Position Embedding (RoPE) index assignments.

    Let x0:L−1=(x0,x1,…,xL−1)x_{0:L-1} = (x_0, x_1, \dots, x_{L-1}) be a token sequence of length LL drawn from a training corpus D\mathcal{D}. An autoregressive model parameterized by θ\theta predicts the next token given RoPE index vector r=(r0,…,rL−1)∈ZLr = (r_0, \dots, r_{L-1}) \in \mathbb{Z}^L under a causal mask:

    pθ(⋅∣x<i;r),i=0,…,L−1p_\theta(\cdot \mid x_{<i}; r), \quad i = 0, \dots, L-1

    The standard view uses unperturbed RoPE indices ri(0,0)=ir_i^{(0,0)} = i for all i∈{0,…,L−1}i \in \{0, \dots, L-1\} and is optimized using the causal language modeling (CLM) loss:

    LCLM(θ)=Ex0:L−1∼D[−∑i=0L−1log⁡pθ(xi∣x<i;r(0,0))]\mathcal{L}_{\mathrm{CLM}}(\theta) = \mathbb{E}_{x_{0:L-1} \sim \mathcal{D}} \left[ -\sum_{i=0}^{L-1} \log p_\theta\left(x_i \mid x_{<i}; r^{(0,0)}\right) \right]

    A perturbed view is constructed by sampling a split point s∈{0,…,L−1}s \in \{0, \dots, L-1\} and a skip offset y∈{1,…,Y}y \in \{1, \dots, Y\} (with default Y=LY = L) from a distribution q(s,y)q(s, y), defining the skip-perturbed RoPE indices:

    ri(s,y)={i,i<si+y,i≥sr_i^{(s,y)} = \begin{cases} i, & i < s \\ i + y, & i \ge s \end{cases}

    Because indices for i<si < s are unchanged, the regularizer computes the reverse Kullback-Leibler (KL) divergence strictly on suffix targets i∈[s,L−1]i \in [s, L-1]:

    ℓi(s,y)(θ)=KL(pθ(⋅∣x<i;r(s,y))  ∥  sg(pθ(⋅∣x<i;r(0,0))))\ell_i^{(s,y)}(\theta) = \mathrm{KL}\left(p_\theta\left(\cdot \mid x_{<i}; r^{(s,y)}\right) \;\Big\|\; \mathrm{sg}\left(p_\theta\left(\cdot \mid x_{<i}; r^{(0,0)}\right)\right)\right)

    where sg(⋅)\mathrm{sg}(\cdot) denotes the stop-gradient operator, treating the standard forward pass as a stable teacher. The full distillation regularizer is the expectation over suffix positions and the data distribution:

    LKL(θ)=Ex0:L−1∼D,(s,y)∼q(s,y)[1L−s∑i=sL−1ℓi(s,y)(θ)]\mathcal{L}_{\mathrm{KL}}(\theta) = \mathbb{E}_{x_{0:L-1} \sim \mathcal{D}, (s,y) \sim q(s,y)} \left[ \frac{1}{L-s} \sum_{i=s}^{L-1} \ell_i^{(s,y)}(\theta) \right]

    The total training loss balances both objectives via weight λ\lambda (set by default to λ=1\lambda = 1):

    Ltotal(θ)=LCLM(θ)+λLKL(θ)\mathcal{L}_{\mathrm{total}}(\theta) = \mathcal{L}_{\mathrm{CLM}}(\theta) + \lambda \mathcal{L}_{\mathrm{KL}}(\theta)

    At inference time, the model is evaluated exclusively with standard, unperturbed indices (ri=ir_i = i), requiring no additional forward passes or architectural modifications.

  2. Knowl 2 — RoPE Index Perturbation Schemes

    model/method

    The RoPE-Perturbed Self-Distillation framework generates alternative views of an input sequence by applying transformations to RoPE positional index vectors while leaving the underlying token order and causal mask unchanged.

    Two primary index perturbation schemes are formulated:

    1. Skip-based Perturbation: Given split point s∈{0,…,L−1}s \in \{0, \dots, L-1\} and skip offset y∈{1,…,Y}y \in \{1, \dots, Y\}, the indices are assigned as:

    ri(s,y)={i,i<si+y,i≥sr_i^{(s,y)} = \begin{cases} i, & i < s \\ i + y, & i \ge s \end{cases}

    This perturbation increases the RoPE distance between prefix [0,s−1][0, s-1] and suffix [s,L−1][s, L-1] by yy, preserving the monotonic ordering of tokens within each segment and across the boundary.

    1. Cyclic-shift Perturbation: Given a shift offset u∈{0,…,L−1}u \in \{0, \dots, L-1\} sampled from q(u)q(u), the RoPE indices are cyclically shifted:

    ricyc(u)=(i+u) mod L,i=0,…,L−1r_i^{\mathrm{cyc}(u)} = (i + u) \bmod L, \quad i = 0, \dots, L-1

    Unlike skip shifts, cyclic shifts map tokens from the original prefix to the end of the RoPE index range, significantly altering relative positional offsets and serving as a more aggressive stress test of position invariance.

    Alternative perturbation variants include Random Dilation (scaling RoPE indices by a random factor in [0.5,2.0][0.5, 2.0]), Chunked Permutation (permuting large blocks of RoPE indices), and NoPE (setting all RoPE indices to zero in the perturbed pass).

  3. Knowl 3 — Long-Context Adaptation Benchmark Performance on RULER and HELMET

    data/table

    When adapting pretrained instruction models to longer context windows—Llama-3-8B-Instruct adapted to 64K context (4B tokens) and Qwen-3-4B adapted to 256K context (8B tokens)—RoPE-perturbed self-distillation outperforms standard causal language modeling (CLM) fine-tuning and related baselines (LongCE, PoSE) on the RULER and HELMET benchmarks.

    Llama-3-8B-Instruct RULER HELMET (64K)
    Method 32K 64K RAG ICL LongQA Rerank Avg.
    Standard 85.23 57.25 59.08 86.30 30.30 1.51 44.30
    LongCE 86.11 52.51 54.71 85.80 31.64 11.20 45.84
    PoSE 87.20 67.47 57.22 86.30 30.23 8.85 45.65
    Ours (cyclic shift) 87.33 71.30 58.34 86.80 30.63 8.70 46.12
    Ours (skip) 87.87 69.29 59.12 87.65 32.41 12.34 47.88
    Qwen-3-4B RULER HELMET (128K)
    Method 128K 256K RAG ICL LongQA Rerank Avg.
    Standard 74.45 67.01 51.54 86.08 56.02 13.57 51.80
    LongCE 76.87 68.01 52.88 87.24 56.00 15.53 52.90
    PoSE 75.47 65.48 52.04 87.16 55.47 12.51 51.79
    Ours (cyclic shift) 78.37 67.91 53.00 86.52 54.48 13.58 51.89
    Ours (skip) 77.51 68.10 53.96 87.40 57.23 13.39 52.99
    Ours (skip) + LongCE 77.50 68.95 54.25 87.56 57.23 15.32 53.59

    On Llama-3-8B-Instruct at 64K context, the skip-based self-distillation objective improves RULER accuracy by +12.04% over Standard fine-tuning (69.29% vs. 57.25%) and HELMET average by +3.58% (47.88% vs. 44.30%). On Qwen-3-4B at 256K context, combining skip-based distillation with LongCE yields the highest aggregate scores (68.95% on RULER-256K and 53.59% on HELMET-128K).

  4. Knowl 4 — Context Length Extrapolation with YaRN

    data/table

    Models adapted with RoPE-perturbed self-distillation exhibit superior length extrapolation capabilities when evaluated beyond their training context length using YaRN scaling (2×2\times and 4×4\times context extension).

    Llama-3-8B-Instruct (64K trained) Qwen-3-4B (256K trained)
    Method 128K 256K 512K 1M
    Standard 42.03 23.90 58.11 45.31
    LongCE 46.54 26.91 60.52 47.53
    PoSE 58.21 43.18 58.00 44.62
    Ours (cyclic shift) 60.89 45.33 60.27 49.56
    Ours (skip) 59.43 44.74 60.73 50.41
    Ours (skip) + LongCE – – 60.13 50.02

    Standard CLM fine-tuning degrades significantly under extrapolation (e.g., Llama-3-8B drops to 23.90% on RULER at 256K). Both skip and cyclic perturbation variants maintain higher performance (44.74% and 45.33% respectively on Llama at 256K, and 50.41% on Qwen at 1M), showing that invariance regularizers improve generalization beyond the trained context limit.

  5. Knowl 5 — Preservation of Long-Context Adaptation Gains After Supervised Fine-Tuning

    data/table

    The performance gains achieved by RoPE-perturbed self-distillation during long-context continued pretraining persist after post-training supervised fine-tuning (SFT).

    Qwen-3-4B models adapted to 256K context were subjected to one epoch of full-parameter instruction SFT on the Tulu-v3 dataset and evaluated across RULER (128K and 256K), HELMET (128K), and LongBench-v2 (LBv2):

    Method RULER HELMET LBv2
    128K 256K
    Standard 75.52 64.44 52.61 27.63
    LongCE 76.95 65.92 53.36 31.73
    PoSE 76.22 63.88 52.53 28.83
    Ours (skip) 77.28 66.86 53.48 29.40
    Ours (skip) + LongCE 78.68 67.15 54.02 32.22

    After SFT, Ours (skip) retains a +2.42% lead over Standard fine-tuning on RULER 256K (66.86% vs. 64.44%) and outperforms Standard on HELMET (53.48% vs. 52.61%) and LongBench-v2 (29.40% vs. 27.63%). The combined Ours (skip) + LongCE checkpoint achieves the highest overall accuracy across all three benchmark suites.

  6. Knowl 6 — Order-Sensitive Needle-In-A-Haystack Evaluation (NIAH-Position)

    empirical result

    To test whether positional invariance regularization harms tasks that require tracking relative sequence order, models were evaluated on NIAH-position at 64K context. In this task, the same key appears multiple times in a synthetic haystack paired with distinct values, and the query asks for an ordinally specified occurrence (e.g., first, second, third, or last occurrence).

    Method 64K NIAH-position Accuracy (%)
    Baseline (Standard CLM) 44.0
    LongCE 56.2
    PoSE 48.2
    Ours (cyclic shift) 42.0
    Ours (skip) 59.4

    The skip-based self-distillation method achieves 59.4% accuracy, improving upon the standard CLM baseline (44.0%) and all comparison baselines. Because skip perturbations shift RoPE indices without modifying token sequence order, the model preserves ordinal relations. In contrast, cyclic-shift perturbation drops to 42.0% because cyclic wrapping disrupts relative ordinal positioning.

  7. Knowl 7 — Ablation of Self-Distillation Components, Perturbation Types, and Consistency Controls

    data/table

    Ablation experiments conducted on Llama-3-8B-Instruct evaluated on RULER at 64K context (reported at an early 200-step checkpoint) demonstrate the impact of loss direction, stochasticity, perturbation structure, and generic two-pass baselines:

    Objective Variant / Control Avg. (%) Perturbation Scheme Avg. (%)
    Baseline (Standard CLM) 47.9 Baseline (Standard CLM) 47.9
    Ours (Skip, Reverse KL, λ=1\lambda=1) 59.2 Ours (Skip-based shift) 59.2
    w/ forward KL 57.8 Random dilation 53.4
    w/ CLM on perturbed view 56.5 Chunked permutation 46.2
    w/o random skip (fixed s=y=32Ks=y=32\text{K}) 55.1 NoPE (zero positional indices) 29.5
    Dropout consistency 48.9
    Attention noise consistency 49.6

    Key takeaways:

    1. Distillation loss: Reverse KL (59.2%) outperforms forward KL (57.8%) and optimizing CLM on the perturbed pass (56.5%).
    2. Stochasticity: Randomly sampling (s,y)(s, y) outperforms a deterministic fixed skip (55.1%).
    3. Generic two-pass controls: Adding consistency over dropout passes (48.9%) or attention noise (49.6%) yields marginal gains over baseline, demonstrating that index perturbation specifically drives the improvement.
    4. Perturbation structure: Preserving local and monotonic ordering (skip-based shift: 59.2%) substantially outperforms order-disrupting schemes like chunked permutation (46.2%) and index removal / NoPE (29.5%).
  8. Knowl 8 — Short-Context Benchmark Evaluation Across Adaptation Methods

    data/table

    Evaluating models on standard downstream short-context benchmarks confirms that RoPE-perturbed self-distillation does not degrade general language understanding capabilities.

    Model Method MMLU HellaSwag Winogrande OpenBookQA Avg.
    Llama-3-8B Base 64.7 75.6 71.4 43.0 63.7
    Standard 61.7 78.7 72.5 43.4 64.1
    LongCE 62.9 78.7 72.2 43.8 64.4
    PoSE 62.2 78.7 72.0 43.6 64.1
    Ours (cyclic shift) 61.4 78.4 72.4 44.2 64.1
    Ours (skip) 61.2 78.3 72.2 43.6 63.8
    Qwen-3-4B Base 68.4 68.4 65.8 40.6 60.8
    Standard 68.6 70.6 67.8 40.4 61.8
    LongCE 68.4 71.1 68.2 40.4 62.0
    PoSE 68.6 71.1 68.8 40.6 62.3
    Ours (cyclic shift) 68.2 70.3 68.0 40.6 61.8
    Ours (skip) 68.5 71.2 68.9 39.4 62.0
    Ours (skip) + LongCE 68.0 69.9 67.1 40.6 61.4

    Across both Llama-3-8B and Qwen-3-4B, models trained with RoPE-perturbed self-distillation achieve downstream task averages comparable to standard long-context fine-tuning and the original base checkpoints.

Coverage note — None was omitted; all primary methods, formulas, benchmark results (RULER, HELMET, YaRN extrapolation, post-SFT LongBench-v2, NIAH-position), ablations, and efficiency analyses from the paper are represented.

References

  1. 1.An, C., Zhang, J., Zhong, M., Li, L., Gong, S., Luo, Y., Xu, J., and Kong, L. Why does the effective context length of llms fall short? arXiv preprint arXiv:2410.18745, 2024.
  2. 2.Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H. Self-rag: Learning to retrieve, generate, and critique through self-reflection. 2024.
  3. 3.Bai, Y., Tu, S., Zhang, J., Peng, H., Wang, X., Lv, X., Cao, S., Xu, J., Hou, L., Dong, Y., Tang, J., and Li, J. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. arXiv preprint arXiv:2412.15204, 2024.
  4. 4.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  5. 5.Bulatov, A., Kuratov, Y., Kapushev, Y., and Burtsev, M. S. Scaling transformer to 1m tokens and beyond with rmt. arXiv preprint arXiv:2304.11062, 2023.
  6. 6.Ding, Y., Zhang, L. L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M. Longrope: Extending LLM context window beyond 2 million tokens. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=ONOtpXLqqw.
  7. 7.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Roziere, B., ` Biron, B., Tang, B., Chern, B., Caucheteux, C., Nayak, C., Bi, C., Marra, C., McConnell, C., Keller, C., Touret, C., Wu, C., Wong, C., Ferrer, C. C., Nikolaidis, C., Allonsius, D., Song, D., Pintz, D., Livshits, D., Esiobu, D., Choudhary, D., Mahajan, D., Garcia-Olano, D., Perino, D., Hupkes, D., Lakomkin, E., AlBadawy, E., Lobanova, E., Dinan, E., Smith, E. M., Radenovic, F., Zhang, F., Synnaeve, G., Lee, G., Anderson, G. L., Nail, G., Mialon, G., Pang, G., Cucurell, G., Nguyen, H., Korevaar, H., Xu, H., Touvron, H., Zarov, I., Ibarra, I. A., Kloumann, I. M., Misra, I., Evtimov, I., Copet, J., Lee, J., Geffert, J., Vranes, J., Park, J., Mahadeokar, J., Shah, J., van der Linde, J., Billock, J., Hong, J., Lee, J., Fu, J., Chi, J.,
  8. 8.Huang, J., Liu, J., Wang, J., Yu, J., Bitton, J., Spisak, J., Park, J., Rocca, J., Johnstun, J., Saxe, J., Jia, J., Alwala, K. V., Upasani, K., Plawiak, K., Li, K., Heafield, K., Stone, K., and et al. The llama 3 herd of models. CoRR, abs/2407.21783, 2024.
  9. 9.Fang, L., Wang, Y., Liu, Z., Zhang, C., Jegelka, S., Gao, J., Ding, B., and Wang, Y. What is wrong with perplexity for long-context language modeling? In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=fL4qWkSmtM.
  10. 10.Fu, Y., Panda, R., Niu, X., Yue, X., Hajishirzi, H., Kim, Y., and Peng, H. Data engineering for scaling language models to 128k context. arXiv preprint arXiv:2402.10171, 2024.
  11. 11.Gao, T., Wettig, A., Yen, H., and Chen, D. How to train long-context language models (effectively). In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T. (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, pp. 7376–7399. Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025.acl-long.366/.
  12. 12.Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Guo, Q., Wang, M., and Wang, H. Retrieval-augmented generation for large language models: A survey. CoRR, abs/2312.10997, 2023.
  13. 13.Han, C., Wang, Q., Peng, H., Xiong, W., Chen, Y., Ji, H., and Wang, S. Lm-infinite: Zero-shot extreme length generalization for large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 3991–4008, 2024.
  14. 14.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021.
  15. 15.Hsieh, C., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., and Ginsburg, B. RULER: what’s the real context size of your long-context language models? CoRR, abs/2404.06654, 2024.
  16. 16.Hu, Z., Liu, Y., Zhao, J., Wang, S., WangYan, W., Shen, W., Gu, Q., Tuan, L. A., Ng, S. K., Jiang, Z., et al. Longrecipe: Recipe for efficient long context generalization in large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11857–11870, 2025.
  17. 17.Jiang, Z., Ma, X., and Chen, W. Longrag: Enhancing retrieval-augmented generation with long-context llms. arXiv preprint arXiv:2406.15319, 2024.
  18. 18.Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R. SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=VTF8yNQM66.
  19. 19.Jin, H., Han, X., Yang, J., Jiang, Z., Chang, C.-Y., and Hu, X. Growlength: Accelerating llms pretraining by progressively growing training length. arXiv preprint arXiv:2310.00576, 2023.
  20. 20.Kazemnejad, A., Padhi, I., Natesan Ramamurthy, K., Das, P., and Reddy, S. The impact of positional encoding on length generalization in transformers. Advances in Neural Information Processing Systems, 36:24892–24928, 2023.
  21. 21.Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, S., Gu, Y., Malik, S., Graf, V., Hwang, J. D., Yang, J., Bras, R. L., Tafjord, O., Wilhelm, C., Soldaini, L., Smith, N. A., Wang, Y., Dasigi, P., and Hajishirzi, H. Tulu 3: ¨ Pushing frontiers in open language model post-training. 2024.
  22. 22.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. Transactions of the association for computational linguistics, 12:157–173, 2024.
  23. 23.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? A new dataset for open book question answering. In EMNLP, pp. 2381–2391. Association for Computational Linguistics, 2018.
  24. 24.Peng, B., Quesnelle, J., Fan, H., and Shippole, E. Yarn: Efficient context window extension of large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=wHBfxhZu1u.
  25. 25.Ruoss, A., Deletang, G., Genewein, T., Grau-Moya, J., ´ Csordas, R., Bennani, M., Legg, S., and Veness, J. Ran- ´ domized positional encodings boost length generalization of transformers. arXiv preprint arXiv:2305.16843, 2023.
  26. 26.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. In AAAI, pp. 8732–8740. AAAI Press, 2020.
  27. 27.Su, J., Ahmed, M. H. M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024. doi: 10.1016/J.NEUCOM.2023.127063. URL https://doi.org/10.1016/j.neucom.2023.127063.
  28. 28.Team, Q. et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2(3), 2024.
  29. 29.Wang, S., Kobyzev, I., Lu, P., Rezagholizadeh, M., and Liu, B. Resonance rope: Improving context length generalization of large language models. arXiv preprint arXiv:2403.00071, 2024.
  30. 30.Wu, T., Zhao, Y., and Zheng, Z. An efficient recipe for long context extension via middle-focused positional encoding. Advances in Neural Information Processing Systems, 37: 56349–56373, 2024a.
  31. 31.Wu, W., Wang, Y., Fu, Y., Yue, X., Zhu, D., and Li, S. Long context alignment with short instructions and synthesized positions. arXiv preprint arXiv:2405.03939, 2024b.
  32. 32.Wu, Y., Rabe, M. N., Hutchins, D., and Szegedy, C. Memorizing transformers. arXiv preprint arXiv:2203.08913, 2022.
  33. 33.Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453, 2023.
  34. 34.Xiaomi, L.-C. Mimo-v2-flash technical report, 2026. URL https://arxiv.org/abs/2601.02780.
  35. 35.Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. Effective long-context scaling of foundation models. In Duh, K., Gomez, H., and Bethard, S. (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 4643–4663, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.naacl-long.260. URL https://aclanthology.org/2024.naacl-long.260/.
  36. 36.Xu, C., Ping, W., Xu, P., Liu, Z., Wang, B., Shoeybi, M., Li, B., and Catanzaro, B. From 128k to 4m: Efficient training of ultra-long context large language models. arXiv preprint arXiv:2504.06214, 2025.
  37. 37.Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. Qwen3 technical report. CoRR, abs/2505.09388, 2025a.
  38. 38.Yang, B., Venkitesh, B., Talupuru, D., Lin, H., Cairuz, D., Blunsom, P., and Locatelli, A. Rope to nope and back again: A new hybrid attention strategy. arXiv preprint arXiv:2501.18795, 2025b.
  39. 39.Yen, H., Gao, T., Hou, M., Ding, K., Fleischer, D., Izsak, P., Wasserblat, M., and Chen, D. Helmet: How to evaluate long-context language models effectively and thoroughly. In International Conference on Learning Representations (ICLR), 2025.
  40. 40.Yu, Y., Jiang, H., Luo, X., Wu, Q., Lin, C.-Y., Li, D., Yang, Y., Huang, Y., and Qiu, L. Mitigate position bias in llms via scaling a single hidden states channel. In Findings of the Association for Computational Linguistics: ACL 2025, pp. 6092–6111, 2025.
  41. 41.Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al. Big bird: Transformers for longer sequences. Advances in neural information processing systems, 33:17283–17297, 2020.
  42. 42.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In ACL (1), pp. 4791–4800. Association for Computational Linguistics, 2019.
  43. 43.Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Re, C., Barrett, C., et al. H2o: ´ Heavy-hitter oracle for efficient generative inference of large language models. Advances in Neural Information Processing Systems, 36:34661–34710, 2023.
  44. 44.Zhang, Z., Chen, R., Liu, S., Yao, Z., Ruwase, O., Chen, B., Wu, X., Wang, Z., et al. Found in the middle: How language models use long contexts better via plug-and-play positional encoding. Advances in Neural Information Processing Systems, 37:60755–60775, 2024.
  45. 45.Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S. Pose: Efficient context window extension of llms via positional skip-wise training. arXiv preprint arXiv:2309.10400, 2023.

Citation

MLA
Li, Z., et al. “Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation”. arXiv, 2026, http://arxiv.org/abs/2604.14339v1.
APA
Li, Z., Liang, C., Ren, L., Zhao, T., Shen, Y., & Chen, W. (2026). Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation. arXiv. http://arxiv.org/abs/2604.14339v1
Chicago
Li, Z., C. Liang, L. Ren, T. Zhao, Y. Shen, and W. Chen. 2026. “Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation”. arXiv. http://arxiv.org/abs/2604.14339v1.
Harvard
Li, Z. et al. (2026) “Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.14339v1.
Vancouver
1. Li Z, Liang C, Ren L, Zhao T, Shen Y, Chen W (2026) Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation. arXiv

BibTeX

@article{li2026shuffle,
  title = {Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation},
  author = {Li, Zichong and Liang, Chen and Ren, Liliang and Zhao, Tuo and Shen, Yelong and Chen, Weizhu},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.14339v1},
  eprint = {2604.14339}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/