Training-Free Long-Context Scaling of Large Language Models

Chenxin AnFei HuangJun ZhangShansan GongXipeng QiuChang ZhouLingpeng Kong

article2024ICML72 citations

Presents Dual Chunk Attention, a training-free framework that decomposes attention into intra-chunk and inter-chunk modules to scale the context window of large language models like LLaMA-2 70B beyond 100k tokens while retaining practical task accuracy.

Listen

Modern large language models struggle to process extended text sequences that exceed the specific window length encountered during their initial training. While extending this capacity through further training on long sequences works well, it requires massive computing resources and proprietary datasets that are inaccessible or prohibitively expensive for most organizations. Common training-free workarounds either truncate earlier context—losing critical long-range connections—or alter position calculations in ways that degrade model coherence and accuracy.

The article demonstrates and evaluates Dual Chunk Attention, a training-free framework that expands the input capacity of existing language models without requiring retraining. The objective is to enable models to retain both local detail and long-range context across sequences that are significantly longer than their original design limits.

The authors implemented the approach across open-source model architectures ranging from 7 billion to 70 billion parameters, including base and instruction-tuned variants. They evaluated the framework on standard language modeling tasks up to 192,000 tokens, passkey retrieval tests, and established benchmarks for long-form question answering and summarization. Instead of altering positional values across the entire input, the method divides sequences into manageable chunks smaller than the original training window. It then computes relationships using three distinct modules: within the same chunk, across different chunks, and between immediately adjacent chunks to preserve local sentence flow.

The results highlight significant performance gains. First, the method allows models designed for a 4,000-token limit to process over 32,000 tokens with almost no degradation in coherence, while larger 70-billion-parameter models scale smoothly beyond 100,000 tokens. Second, on practical question-answering and summarization tasks, the training-free 70-billion model achieved performance comparable to—and in several cases exceeding—models that underwent resource-intensive continued training. Third, the framework reached 94% of the performance of proprietary commercial models like GPT-3.5-16k. Finally, the approach integrates directly with accelerated computation tools like Flash Attention, maintaining standard memory consumption and inference speeds.

These findings provide strong strategic value for enterprise deployment. Organizations can dramatically scale their document analysis, customer support history, and information retrieval pipelines simply by updating their inference procedures, completely avoiding the high capital costs, energy consumption, and project timelines tied to extensive model retraining. The framework also stacks on top of models that have already undergone prior context tuning, offering a path to reach up to 192,000 tokens efficiently.

Decision-makers should consider piloting this inference-only patch as an immediate, low-cost upgrade for long-document tasks rather than funding expensive fine-tuning runs. For organizations requiring maximum possible accuracy, light fine-tuning using long conversation data can still be layered on top to capture additional incremental gains. Future internal pilots should benchmark latency across target production sequence lengths and evaluate specific task performance.

Confidence in these findings is high across standard research benchmarks, passkey retrieval tests, and base model architectures. However, decision-makers should note that smaller models (such as the 13-billion-parameter variant) still exhibit noticeable comprehension gaps on deeply nuanced reasoning across extended documents compared to 70-billion-parameter configurations. Readers should validate performance on complex domain-specific tasks that require complete global document synthesis.

An et al (2024).pdf

No sufficiently relevant recommendations were found.

No sufficiently relevant recommendations were found.

Cover for Training-Free Long-Context Scaling of Large Language Models

Abstract

The ability of Large Language Models (LLMs) to process and generate coherent text is markedly weakened when the number of input tokens exceeds their pretraining length. Given the expensive overhead of finetuning large-scale models with longer sequences, we propose Dual Chunk Attention (DCA), which enables Llama2 70B to support context windows of more than 100k tokens without continual training. By decomposing the attention computation for long sequences into chunk-based modules, DCA manages to effectively capture the relative positional information of tokens within the same chunk (Intra-Chunk) and across distinct chunks (Inter-Chunk), as well as integrates seamlessly with Flash Attention. In addition to its impressive extrapolation capability, DCA achieves performance on practical long-context tasks that is comparable to or even better than that of finetuned models. When compared with proprietary models, our training-free 70B model attains 94% of the performance of gpt-3.5-16k, indicating it is a viable open-source alternative. All code and data used in this work are released at https://github.com/HKUNLP/ChunkLlama.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Positional Encoding
  • 2.2. Extrapolation of RoPE
  • 3. Method
  • 3.1. Intra-Chunk Attention
  • 3.2. Inter-Chunk Attention
  • 3.3. Successive-Chunk Attention
  • 3.4. Normalization
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Long-Sequence Language Modeling
  • 4.3. Practical Tasks
  • 4.4. Analysis
  • 5. Conclusion
  • Acknowledgements
  • References
  • A. Appendix
  • A.1. Passkey retrieval
  • A.2. More Examples
  • A.3. Flash Attention
  • A.4. In-Context Examples Selection
  • A.5. Performance on Unseen Data

Knowls

  1. Knowl 1 — Dual Chunk Attention reuses pretrained RoPE positions to handle long sequences

    model/method

    Dual Chunk Attention (DCA) extends a causal language model’s context without updating its weights. It divides a sequence into chunks of size ss, smaller than the model’s pretrained context length cc, and assigns query positions according to whether a key is in the same chunk, the immediately preceding chunk, or a more distant chunk. For a sequence of length LL, let ii and jj be zero-based query and key indices with 0≤j≤i<L0\leq j\leq i<L, and let r=i mod sr=i\bmod s. The key position is Pk[j]=j mod sP_k[j]=j\bmod s. The query positions are Pqintra[i]=i mod sP_q^{\mathrm{intra}}[i]=i\bmod s, Pqinter[i]=c−1P_q^{\mathrm{inter}}[i]=c-1, and Pqsucc[i]=s+rP_q^{\mathrm{succ}}[i]=s+r when r<wr<w, otherwise c−1c-1, where ww is the successive-chunk local-window size and 0<w≤c−s0<w\leq c-s. The paper suggests setting w=c−sw=c-s.

    The relative position supplied to rotary positional encoding is selected by the distance between the chunks containing the query and key:

    M[i,j]={Pqintra[i]−Pk[j],⌊i/s⌋−⌊j/s⌋=0,Pqsucc[i]−Pk[j],⌊i/s⌋−⌊j/s⌋=1,Pqinter[i]−Pk[j],⌊i/s⌋−⌊j/s⌋>1.M[i,j]=\begin{cases} P_q^{\mathrm{intra}}[i]-P_k[j], & \lfloor i/s\rfloor-\lfloor j/s\rfloor=0,\\ P_q^{\mathrm{succ}}[i]-P_k[j], & \lfloor i/s\rfloor-\lfloor j/s\rfloor=1,\\ P_q^{\mathrm{inter}}[i]-P_k[j], & \lfloor i/s\rfloor-\lfloor j/s\rfloor>1. \end{cases}

    Here, M[i,j]M[i,j] is the relative position used for the attention interaction between query ii and key jj. Same-chunk interactions retain their within-chunk positions; interactions with more distant chunks use the pretrained maximum query position; and interactions with the immediately preceding chunk use positions that preserve a local window. With these assignments, the relative positions remain within the pretrained range. The query and key are rotated using the selected positions, and causal softmax attention is computed over the allowed keys.

  2. Knowl 2 — Evaluation setup and practical DCA configuration

    experimental setup

    The evaluations apply DCA at inference to Llama2 models of 7B, 13B, and 70B parameters, including chat variants, and also test Llama3, Together-32k, and CodeLlama. The implementation replaces the original Llama attention inference code and is compatible with Flash Attention 2. The chunk size is typically three quarters of the pretrained context length; for a 4k-context Llama2, the paper uses s=3072s=3072. The number of chunks grows with input length.

    For PG19 language modeling, contexts range from 4k to 192k tokens. The 7B and 13B evaluations use a sliding window of 256 tokens; the 70B evaluation uses 2048 tokens, reduced to half the input length for contexts above 96k. Few-shot tests use zero shots on NarrativeQA, one shot on QMSum, and two shots each on QuALITY and Qasper; prompts are capped at 16,384 tokens and longer prompts are truncated from the left. Zero-shot chat evaluations use the four closed-ended L-Eval tasks TOFEL, QuALITY, Coursera, and SFiction. For inference, 7B and 13B models fit on one A100 80GB GPU with Flash Attention 2; the 70B models use two A100 GPUs for inputs up to 16k tokens.

  3. Knowl 3 — DCA preserves PG19 perplexity well beyond the pretrained context on Llama2

    empirical result

    On the PG19 validation set, DCA substantially improves training-free extrapolation over unchunked Llama2 and the tested PI- and NTK-based position-scaling baselines. Perplexity (PPL) is reported below for evaluation contexts from 4k to 64k tokens. The 7B DCA model’s PPL at 32k is only 0.02 above its 4k value; the 70B model remains close to its 4k PPL through 64k. The smaller DCA models show larger degradation at 64k.

    Model 4k 8k 16k 32k 64k
    Llama2 7B 7.87 >102 >102 >102 >102
    Llama2-PI 7B 7.87 9.19 15.11 >102 >102
    Llama2-PI-Yarn 7B 7.87 8.80 11.75 42.42 >102
    Llama2-NTK 7B 7.87 11.98 26.12 58.91 >102
    Llama2-NTK-Yarn 7B 7.87 8.06 9.82 11.74 41.57
    CHUNKLLAMA2 7B 7.87 7.67 7.64 7.89 15.87
    CHUNKLLAMA2 13B 7.15 6.95 6.99 7.90 15.14
    CHUNKLLAMA2 70B 5.24 5.18 5.21 5.30 5.59
  4. Knowl 4 — DCA extends existing long-context models to 192k tokens

    empirical result

    On PG19, DCA also extends models that already use longer-context positional encodings: CodeLlama with NTK-based RoPE and Together-32k with positional interpolation (PI). Llama2 70B and Llama3 models are also evaluated at long contexts. The table reports PPL at evaluation lengths of 4k, 32k, 64k, 96k, 128k, 160k, and 192k tokens. The original unchunked Llama2 70B and Llama3 70B exceed 102 PPL at 32k, while their DCA versions remain below 8 through 192k. ChunkCodeLlama and ChunkTogether also remain below 10 PPL at 192k.

    Model 4k 32k 64k 96k 128k 160k 192k
    Llama2 70B 5.24 >102 >102 >102 >102 >102 >102
    CHUNKLLAMA2 70B 5.24 5.30 5.59 5.80 6.12 6.52 7.05
    Llama3 70B 5.36 >102 >102 >102 >102 >102 >102
    CHUNKLLAMA3 70B 5.36 5.14 5.14 5.21 5.32 5.40 5.45
    CodeLlama 7B 8.93 8.36 8.65 9.14 9.87 15.68 24.78
    ChunkCodeLlama 7B 8.93 8.36 8.13 8.33 8.66 9.30 9.83
    Together 7B 8.21 7.64 >102 >102 >102 >102 >102
    ChunkTogether 7B 8.21 7.64 7.59 7.64 7.67 7.74 7.83

    This demonstrates that DCA can be combined with PI- or NTK-based long-context models rather than replacing their positional encoding; applying it to those models requires choosing a chunk size.

  5. Knowl 5 — Training-free DCA is competitive on few-shot long-document tasks

    empirical result

    On NarrativeQA, Qasper, QuALITY, and QMSum, the paper evaluates base models with zero-shot, two-shot, two-shot, and one-shot prompts, respectively. The metrics are F1 for NarrativeQA and Qasper, exact match (EM) for QuALITY, and ROUGE for QMSum. The following results compare DCA models with the strongest relevant 70B trained baselines and the Llama2 Long 70B proprietary baseline; DCA models use no additional training.

    Model NarrativeQA F1 Qasper F1 QuALITY EM QMSum R-g Average
    Longlora 70B 34.2 29.0 69.9 15.6 37.2
    CHUNKLLAMA2 7B 20.0 28.2 35.6 14.7 24.6
    CHUNKLLAMA2 13B 26.3 29.3 47.9 15.2 29.7
    CHUNKLLAMA2 70B 32.5 29.6 73.2 16.0 37.8
    CHUNKLLAMA3 8B 27.4 30.5 52.6 15.4 31.5
    CHUNKLLAMA3 70B 33.7 33.1 75.4 16.0 39.5
    Llama2 Long 70B 30.9 35.7 79.7 16.5 40.7

    CHUNKLLAMA2 70B obtains an average of 37.8, compared with 37.2 for Longlora 70B, which was trained for long context. The DCA result is competitive with that baseline without extra training, although it remains below Llama2 Long 70B’s average of 40.7.

  6. Knowl 6 — DCA improves zero-shot chat evaluation, especially at 70B scale

    empirical result

    On four closed-ended L-Eval tasks, prompts range from 3k to 27k tokens and scores are exact match percentages. The task-specific input ranges are TOFEL 3k–5k, QuALITY 4k–9k, Coursera 5k–17k, and SFiction 6k–27k. The table compares DCA chat models with selected open-source and proprietary systems. CHUNKLLAMA2-Chat 70B averages 63.20, about 94% of GPT-3.5-16k’s 67.03 average. CHUNKLLAMA3-Instruct 70B averages 79.89, above the reported Claude 1.3 average of 72.52.

    Model TOFEL QuALITY Coursera SFiction Average
    Longlora-Chat 70B 71.37 55.45 44.76 67.96 59.88
    CHUNKLLAMA2-Chat 7B 57.62 35.14 32.12 61.72 46.64
    CHUNKLLAMA2-Chat 13B 66.54 43.06 41.56 57.03 52.04
    CHUNKLLAMA2-Chat 70B 82.15 60.39 48.54 61.72 63.20
    CHUNKLLAMA3-Instruct 8B 83.27 63.86 56.24 70.31 68.42
    CHUNKLLAMA3-Instruct 70B 84.75 82.17 76.88 75.78 79.89
    GPT3.5-16k-0613 78.43 61.38 63.51 64.84 67.03
    Claude1.3-100k 83.64 60.03 73.76 72.65 72.52
  7. Knowl 7 — DCA supports long-distance passkey retrieval

    empirical result

    In the passkey-retrieval test, a random five-digit key is embedded at document depths distributed uniformly through otherwise nonsensical text. For each tested depth, the evaluation uses 20 different passkeys. On Llama2 13B with a 4k pretrained context, CHUNKLLAMA2 achieves 100% retrieval accuracy at all tested depths through an 18k-token input. Its accuracy remains high through 32k tokens. The paper also reports that DCA-enhanced existing long-context models maintain 90% retrieval accuracy at context lengths up to 192k tokens. These results test whether the models can retrieve information placed away from the end of a long prompt, not just whether they maintain low language-modeling perplexity.

  8. Knowl 8 — Ablations identify distinct roles for the three DCA attention ranges

    empirical result

    The paper compares intra-chunk attention alone, intra- plus inter-chunk attention, and the full combination including successive-chunk attention on language modeling and passkey retrieval at input lengths from 8k to 32k. Intra-chunk attention alone retains low perplexity but discards information from earlier chunks, impairing passkey retrieval. Adding inter-chunk attention improves retrieval, including at 12k, but loses locality between adjacent chunks and increases perplexity. Adding successive-chunk attention restores local positional relationships across adjacent chunks, yielding both low perplexity and high retrieval accuracy. The ablation supports the separate functions of within-chunk precision, distant-chunk information access, and near-boundary locality.

  9. Knowl 9 — DCA retains Flash Attention-like inference efficiency

    empirical result

    Inference time and GPU memory are compared for PyTorch self-attention, Flash Attention, and DCA integrated with Flash Attention on Llama2 7B. Tests use a single NVIDIA A100 80GB GPU, long prompts from NarrativeQA, and the average of 20 trials. Without Flash Attention, the maximum input length on one GPU is approximately 12k–16k tokens. The authors report that DCA’s GPU-memory use and inference speed are similar to the original Flash Attention implementation, without considerable added overhead. This compatibility matters because ordinary PyTorch self-attention does not support the same long inputs under the tested hardware constraints.

  10. Knowl 10 — DCA’s Flash Attention implementation separates attention by chunk range

    algorithm

    For a query at absolute position ii, DCA can compute three partial attention results using Flash Attention: keys in the query’s own chunk, keys in the immediately preceding chunk, and keys in earlier chunks. Let ss be the chunk size and n=⌊i/s⌋n=\lfloor i/s\rfloor. The respective key ranges contain i−nsi-ns tokens, ss tokens, and s(n−1)s(n-1) tokens; the within-chunk calculation is causal, while the two previous-chunk calculations are not. The query is rotary-embedded separately with its intra-chunk, successive-chunk, or inter-chunk position, and the keys are rotary-embedded with their repeated within-chunk key positions. The three partial attention outputs are combined using their softmax normalization statistics so that the result is normalized over all allowed keys, equivalent to one attention over the union of the three ranges. The stated attention-computation costs for a query are O(i−ns)O(i-ns), O(s)O(s), and O(s(n−1))O(s(n-1)), respectively.

  11. Knowl 11 — Evaluation on unseen paper text exposes remaining comprehension errors

    limitation

    To probe possible benchmark data contamination, the authors use the paper’s own LaTeX text as a test input, omitting its title, abstract, and conclusion; the tokenized input is 19,388 tokens. On simple questions, CHUNKLLAMA2 70B answers accurately, while the 13B model makes errors, including misidentifying the corpora used for the separate long-dialogue finetuning experiments. On more demanding questions about DCA, the 70B model can explain the motivation for inter-chunk and successive-chunk attention, but still has difficulty with questions requiring a global understanding of the method. Thus, the paper’s own test indicates that long-context access does not by itself guarantee reliable document-level comprehension.

Coverage note — The separate long-dialogue finetuning recipe and the in-context example-selection ablation are omitted because they support comparison and prompt analysis rather than the central training-free DCA contribution.

References

  1. 1.An, C., Gong, S., Zhong, M., Li, M., Zhang, J., Kong, L., and Qiu, X. L-eval: Instituting standardized evaluation for long context language models. arXiv preprint arXiv:2307.11088, 2023.
  2. 2.Anthropic. Introducing 100K Context Windows, 2023. URL https://www.anthropic.com/index/100k-context-windows.
  3. 3.Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T. Qwen technical report, 2023.
  4. 4.Chen, G., Li, X., Meng, Z., Liang, S., and Bing, L. Clex: Continuous length extrapolation for large language models, 2023a.
  5. 5.Chen, S., Wong, S., Chen, L., and Tian, Y. Extending context window of large language models via positional interpolation, 2023b.
  6. 6.Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J. Longlora: Efficient fine-tuning of long-context large language models. arXiv:2309.12307, 2023c.
  7. 7.Chi, T.-C., Fan, T.-H., Rudnicky, A. I., and Ramadge, P. J. Dissecting transformer length extrapolation via the lens of receptive field analysis, 2023.
  8. 8.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.
  9. 9.Chowdhury, J. R. and Caragea, C. Monotonic location attention for length generalization, 2023.
  10. 10.Computer, T. Redpajama: an open dataset for training large language models, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  11. 11.Dao, T. Flashattention-2: Faster attention with better parallelism and work partitioning, 2023.
  12. 12.Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C. Flashattention: Fast and memory-efficient exact attention with io-awareness. In NeurIPS, 2022.
  13. 13.Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4599–4610, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.365. URL https://aclanthology.org/2021.naacl-main.365.
  14. 14.Han, C., Wang, Q., Xiong, W., Chen, Y., Ji, H., and Wang, S. Lm-infinite: Simple on-the-fly length generalization for large language models, 2023.
  15. 15.He, Z., Feng, G., Luo, S., Yang, K., He, D., Xu, J., Zhang, Z., Yang, H., and Wang, L. Two stones hit one bird: Bilevel positional encoding for better length extrapolation, 2024.
  16. 16.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models, 2021.
  17. 17.Jin, H., Han, X., Yang, J., Jiang, Z., Liu, Z., Chang, C.-Y., Chen, H., and Hu, X. Llm maybe longlm: Self-extend llm context window without tuning, 2024.
  18. 18.Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S. The impact of positional encoding on length generalization in transformers, 2023.
  19. 19.Kociský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E. The NarrativeQA reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328, 2018. doi: 10.1162/tacl_a_00023. URL https://aclanthology.org/Q18-1023.
  20. 20.Lee, G., Hartmann, V., Park, J., Papailiopoulos, D., and Lee, K. Prompted llms as chatbot modules for long open-domain conversation. In Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.findings-acl.277. URL http://dx.doi.org/10.18653/v1/2023.findings-acl.277.
  21. 21.Li, D., Shao, R., Xie, A., Sheng, Y., Zheng, L., Gonzalez, J. E., Stoica, I., Ma, X., and Zhang, H. How long can open-source llms truly promise on context length. 2023a.
  22. 22.Li, S., You, C., Guruganesh, G., Ainslie, J., Ontanon, S., Zaheer, M., Sanghai, S., Yang, Y., Kumar, S., and Bhojanapalli, S. Functional interpolation for relative positions improves long context transformers, 2023b.
  23. 23.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts, 2023a.
  24. 24.Liu, X., Yan, H., Zhang, S., An, C., Qiu, X., and Lin, D. Scaling laws of rope-based extrapolation, 2023b.
  25. 25.LMSYS. Vicuna: An open-source chatbot impressing gpt-4 with 90 URL https://lmsys.org/blog/2023-03-30-vicuna/.
  26. 26.LocalLLaMA. Dynamically scaled rope further increases performance of long context llama with zero fine-tuning, July 2023a. URL https://www.reddit.com/r/LocalLLaMA/comments/14mrgpr/dynamically_scaled_rope_further_increases/.
  27. 27.LocalLLaMA. Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., June 2023b. URL https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/.
  28. 28.Lv, K., Liu, X., Guo, Q., Yan, H., He, C., Qiu, X., and Lin, D. Longwanjuan: Towards systematic measurement for long text quality, 2024.
  29. 29.Mohtashami, A. and Jaggi, M. Landmark attention: Random-access infinite context length for transformers. arXiv preprint arXiv:2305.16300, 2023.
  30. 30.MosaicML. Introducing mpt-30b: Raising the bar for open-source foundation models, 2023a. URL www.mosaicml.com/blog/mpt-30b. Accessed: 2023-06-22.
  31. 31.MosaicML. Introducing mpt-7b: A new standard for open-source, ly usable llms, 2023b. URL www.mosaicml.com/blog/mpt-7b.
  32. 32.OpenAI. Gpt-4 technical report, 2023.
  33. 33.Pang, R. Y., Parrish, A., Joshi, N., Nangia, N., Phang, J., Chen, A., Padmakumar, V., Ma, J., Thompson, J., He, H., and Bowman, S. QuALITY: Question answering with long input texts, yes! In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 5336–5358, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.391. URL https://aclanthology.org/2022.naacl-main.391.
  34. 34.Peng, B., Quesnelle, J., Fan, H., and Shippole, E. Yarn: Efficient context window extension of large language models, 2023.
  35. 35.Press, O., Smith, N. A., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation, 2022.
  36. 36.Qin, Z., Sun, W., Li, D., Shen, X., Sun, W., and Zhong, Y. Lightning attention-2: A free lunch for handling unlimited sequence lengths in large language models. ArXiv, abs/2401.04658, 2024. URL https://api.semanticscholar.org/CorpusID:266900042.
  37. 37.Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P. Compressive transformers for long-range sequence modelling. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=SylKikSYDH.
  38. 38.Ratner, N., Levine, Y., Belinkov, Y., Ram, O., Magar, I., Abend, O., Karpas, E., Shashua, A., Leyton-Brown, K., and Shoham, Y. Parallel context windows for large language models, 2023.
  39. 39.Robertson, S., Zaragoza, H., et al. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389, 2009.
  40. 40.Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Defossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G. Code llama: Open foundation models for code, 2023.
  41. 41.Rula, A. and D’Souza, J. Procedural text mining with large language models, 2023.
  42. 42.Ruoss, A., Deletang, G., Genewein, T., Grau-Moya, J., Csordas, R., Bennani, M., Legg, S., and Veness, J. Randomized positional encodings boost length generalization of transformers, 2023.
  43. 43.Saad-Falcon, J., Barrow, J., Siu, A., Nenkova, A., Yoon, D. S., Rossi, R. A., and Dernoncourt, F. Pdftriage: Question answering over long, structured documents, 2023.
  44. 44.Song, K., Wang, X., Cho, S., Pan, X., and Yu, D. Zebra: Extending context window with layerwise grouped local-global attention, 2023.
  45. 45.Su, J. Rectified rotary position embeddings. https://github.com/bojone/rerope, 2023.
  46. 46.Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding, 2022.
  47. 47.Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F. A length-extrapolatable transformer, 2022.
  48. 48.Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  49. 49.Together. Llama-2-7b-32k-instruct — and fine-tuning for llama-2 models with together api, 2023. URL https://together.ai/blog/llama-2-7b-32k-instruct.
  50. 50.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023a.
  51. 51.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  52. 52.Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłos, P. Focused transformer: Contrastive training for context scaling, 2023.
  53. 53.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need, 2017.
  54. 54.Wang, L., Yang, N., and Wei, F. Learning to retrieve in-context examples for large language models, 2024.
  55. 55.Wei, J., Kim, S., Jung, H., and Kim, Y.-H. Leveraging large language models to power chatbots for collecting user self-reported data, 2023.
  56. 56.Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., Han, S., and Sun, M. Infllm: Unveiling the intrinsic capacity of llms for understanding extremely long sequences with training-free memory, 2024.
  57. 57.Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. Efficient streaming language models with attention sinks, 2023.
  58. 58.Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. Effective long-context scaling of foundation models. CoRR, abs/2309.16039, 2023. doi: 10.48550/ARXIV.2309.16039. URL https://doi.org/10.48550/arXiv.2309.16039.
  59. 59.Ye, J., Wu, Z., Feng, J., Yu, T., and Kong, L. Compositional exemplars for in-context learning. arXiv preprint arXiv:2302.05698, 2023.
  60. 60.Zhang, J., Jiang, S., Feng, J., Zheng, L., and Kong, L. Linear attention via orthogonal memory. ArXiv, abs/2312.11135, 2023. URL https://api.semanticscholar.org/CorpusID:266359128.
  61. 61.Zhang, P., Liu, Z., Xiao, S., Shao, N., Ye, Q., and Dou, Z. Soaring from 4k to 400k: Extending llm’s context with activation beacon. ArXiv, abs/2401.03462, 2024. URL https://api.semanticscholar.org/CorpusID:266844488.
  62. 62.Zhong, M., Yin, D., Yu, T., Zaidi, A., Mutuma, M., Jha, R., Awadallah, A. H., Celikyilmaz, A., Liu, Y., Qiu, X., and Radev, D. QMSum: A new benchmark for query-based multi-domain meeting summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 5905–5921, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.472. URL https://aclanthology.org/2021.naacl-main.472.
  63. 63.Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S. Pose: Efficient context window extension of llms via positional skip-wise training, 2023.

Citation

MLA
An, C., et al. “Training-Free Long-Context Scaling of Large Language Models”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.17463.
APA
An, C., Huang, F., Zhang, J., Gong, S., Qiu, X., Zhou, C., & Kong, L. (2024). Training-Free Long-Context Scaling of Large Language Models. arXiv. https://doi.org/10.48550/arxiv.2402.17463
Chicago
An, C., F. Huang, J. Zhang, et al. 2024. “Training-Free Long-Context Scaling of Large Language Models”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.17463.
Harvard
An, C. et al. (2024) “Training-Free Long-Context Scaling of Large Language Models”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.17463.
Vancouver
1. An C, Huang F, Zhang J, Gong S, Qiu X, Zhou C, Kong L (2024) Training-Free Long-Context Scaling of Large Language Models. https://doi.org/10.48550/arxiv.2402.17463

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.17463,
  doi = {10.48550/ARXIV.2402.17463},
  url = {https://arxiv.org/abs/2402.17463},
  author = {An, Chenxin and Huang, Fei and Zhang, Jun and Gong, Shansan and Qiu, Xipeng and Zhou, Chang and Kong, Lingpeng},
  keywords = {Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Training-Free Long-Context Scaling of Large Language Models},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/