Data Engineering for Scaling Language Models to 128K Context
Yao FuRameswar PandaXinyao NiuXiang YueHannaneh HajishirziYoon KimHao Peng
Demonstrates that continually pretraining language models on just 1 to 5 billion domain-balanced, length-upsampled tokens enables accurate 128K context retrieval matching GPT-4 performance under accessible academic budgets.
Scaling the context length of large language models to 128,000 tokens enables critical real-world applications, such as multi-document question answering, repository-scale code analysis, and autonomous agent workflows. However, existing open-source models often fail to accurately retrieve information across such massive inputs, leaving a significant capability gap compared to leading proprietary systems. Expanding context capacity through full retraining or massive continuous pretraining has historically been seen as prohibitively expensive due to high computational demands and the quadratic cost of sequence attention.
The article demonstrates an efficient data engineering strategy for extending standard open-source language models to a 128,000-token context window. Specifically, it evaluates how data quantity and composition affect a model's ability to locate and retrieve arbitrary information across long contexts.
The authors conducted continual pretraining experiments using standard open-source base models with 7 billion and 13 billion parameters, originally trained on 4,000-token sequences. Utilizing an academic-scale hardware setup of 8 to 16 graphics processing units alongside modern memory-optimization techniques, the models were trained on 64,000 to 80,000 token chunk sizes. The evaluation evaluated precise retrieval via the Needle-in-a-Haystack test up to 128,000 tokens, general knowledge retention via massive multitask language benchmarks, and downstream long-form book comprehension.
The findings establish that long-range retrieval is an inherent capability largely acquired during initial pretraining rather than a new skill requiring heavy training from scratch. First, training on only 1 to 5 billion tokens is sufficient to unlock full retrieval across 128,000 tokens, achieving an 88.0% accuracy on the retrieval benchmark for a 7-billion parameter model and 90.0% for a 13-billion model—matching or exceeding proprietary benchmark performance. Second, performance saturates around 5 billion tokens; scaling training data to 10 billion tokens caused overfitting and reduced length generalization. Third, data composition requires both length upsampling and domain balance; simply increasing long-text domains like books improves single-domain metrics while severely harming other domains such as coding. Upsampling long sequences within each individual data source while preserving the original domain mixture delivers the best performance.
These results demonstrate that organizations do not need hundreds of billions of tokens or multi-million-dollar computing budgets to achieve frontier-level long-context capabilities. Continual pretraining can be completed in approximately 5 to 7 days on standard academic computing clusters, representing roughly 1% of the compute budget required by prior approaches. Crucially, this approach maintains general short-context task accuracy while closing the gap with top proprietary models.
Organizations scaling context lengths should adopt a lightweight continual pretraining stage using per-source length upsampling rather than massive retraining or single-domain data inflation. Long-context adaptation should be treated as a targeted, separate stage following general, coding, or mathematical pretraining. Future efforts should focus on supervised instruction tuning and multi-step reasoning for extended sequences, as well as exploring advanced sequence parallelism to scale contexts beyond 128,000 tokens.
These conclusions are supported by controlled evaluations on 7-billion and 13-billion parameter architectures. Confidence in precise retrieval up to 128,000 tokens is high; however, readers should note that the underlying base models remain unaligned and require further instruction tuning before deployment in complex conversational and interactive environments.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). This paper identifies the critical limitation where models fail to utilize information located in the middle of long contexts, motivating the data engineering and continual pretraining strategies developed in the source.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). This benchmark provides the foundational multitask evaluation framework for measuring comprehension and degradation across extended sequence lengths.
- Paper: Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora, Xisen Jin et al. (2022). This study analyzes continual pretraining trade-offs and domain adaptation challenges in language models that underpin continual context scaling.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). This work introduces key length extrapolation dynamics and positional bias considerations when extending sequence processing beyond initial training lengths.
- Paper: LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens, Yiran Ding et al. (2024). This work extends context window scaling beyond the 128K frontier to over two million tokens using evolutionary positional search and progressive multi-stage adaptation.
- Paper: Qwen2.5 Technical Report, Qwen et al. (2024). This technical report demonstrates large-scale execution of progressive multi-stage training schedules and long-dependency data recipes to reach 1-million-token contexts.
- Paper: World Model on Million-Length Video And Language With Blockwise RingAttention, Hao Liu 0055 et al. (2025). This paper extends progressive text context expansion up to one million tokens and broadens the paradigm to multimodal video sequences using RingAttention.
- Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). This work explores an alternative test-time training paradigm to scale effective context handling to 128K tokens without requiring standard full-attention pretraining recipes.
- Paper: Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models, Mosh Levy et al. (2024). This study evaluates the degradation of reasoning capabilities across expanded context lengths even when underlying task complexity remains unchanged.
- Paper: Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time, Michael Y. Hu et al. (2026). This research generalizes data mixture optimization across continual training and midtraining stages using dynamic multi-domain adapter interpolation.
