Data Engineering for Scaling Language Models to 128K Context

Yao FuRameswar PandaXinyao NiuXiang YueHannaneh HajishirziYoon KimHao Peng

article2024ICML233 citations

Demonstrates that continually pretraining language models on just 1 to 5 billion domain-balanced, length-upsampled tokens enables accurate 128K context retrieval matching GPT-4 performance under accessible academic budgets.

Listen

Scaling the context length of large language models to 128,000 tokens enables critical real-world applications, such as multi-document question answering, repository-scale code analysis, and autonomous agent workflows. However, existing open-source models often fail to accurately retrieve information across such massive inputs, leaving a significant capability gap compared to leading proprietary systems. Expanding context capacity through full retraining or massive continuous pretraining has historically been seen as prohibitively expensive due to high computational demands and the quadratic cost of sequence attention.

The article demonstrates an efficient data engineering strategy for extending standard open-source language models to a 128,000-token context window. Specifically, it evaluates how data quantity and composition affect a model's ability to locate and retrieve arbitrary information across long contexts.

The authors conducted continual pretraining experiments using standard open-source base models with 7 billion and 13 billion parameters, originally trained on 4,000-token sequences. Utilizing an academic-scale hardware setup of 8 to 16 graphics processing units alongside modern memory-optimization techniques, the models were trained on 64,000 to 80,000 token chunk sizes. The evaluation evaluated precise retrieval via the Needle-in-a-Haystack test up to 128,000 tokens, general knowledge retention via massive multitask language benchmarks, and downstream long-form book comprehension.

The findings establish that long-range retrieval is an inherent capability largely acquired during initial pretraining rather than a new skill requiring heavy training from scratch. First, training on only 1 to 5 billion tokens is sufficient to unlock full retrieval across 128,000 tokens, achieving an 88.0% accuracy on the retrieval benchmark for a 7-billion parameter model and 90.0% for a 13-billion model—matching or exceeding proprietary benchmark performance. Second, performance saturates around 5 billion tokens; scaling training data to 10 billion tokens caused overfitting and reduced length generalization. Third, data composition requires both length upsampling and domain balance; simply increasing long-text domains like books improves single-domain metrics while severely harming other domains such as coding. Upsampling long sequences within each individual data source while preserving the original domain mixture delivers the best performance.

These results demonstrate that organizations do not need hundreds of billions of tokens or multi-million-dollar computing budgets to achieve frontier-level long-context capabilities. Continual pretraining can be completed in approximately 5 to 7 days on standard academic computing clusters, representing roughly 1% of the compute budget required by prior approaches. Crucially, this approach maintains general short-context task accuracy while closing the gap with top proprietary models.

Organizations scaling context lengths should adopt a lightweight continual pretraining stage using per-source length upsampling rather than massive retraining or single-domain data inflation. Long-context adaptation should be treated as a targeted, separate stage following general, coding, or mathematical pretraining. Future efforts should focus on supervised instruction tuning and multi-step reasoning for extended sequences, as well as exploring advanced sequence parallelism to scale contexts beyond 128,000 tokens.

These conclusions are supported by controlled evaluations on 7-billion and 13-billion parameter architectures. Confidence in precise retrieval up to 128,000 tokens is high; however, readers should note that the underlying base models remain unaligned and require further instruction tuning before deployment in complex conversational and interactive environments.

Cover for Data Engineering for Scaling Language Models to 128K Context

Abstract

We study the continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in particular the ability to utilize information at arbitrary input locations, is a capability that is mostly already acquired through large-scale pretraining, and that this capability can be readily extended to contexts substantially longer than seen during training (e.g., 4K to 128K) through lightweight continual pretraining on appropriate data mixture. We investigate the quantity and quality of the data for continual pretraining: (1) for quantity, we show that 500 million to 5 billion tokens are enough to enable the model to retrieve information anywhere within the 128K context; (2) for quality, our results equally emphasize domain balance and length upsampling. Concretely, we find that naïvely upsampling longer data on certain domains like books, a common practice of existing work, gives suboptimal performance, and that a balanced domain mixture is important. We demonstrate that continual pretraining of the full model on 1B-5B tokens of such data is an effective and affordable strategy for scaling the context length of language models to 128K. Our recipe outperforms strong open-source long-context models and closes the gap to frontier models like GPT-4 128K.

Citation

MLA
Fu, Y., et al. “Data Engineering for Scaling Language Models to 128K Context”. arXiv, 2024, http://arxiv.org/abs/2402.10171v1.
APA
Fu, Y., Panda, R., Niu, X., Yue, X., Hajishirzi, H., Kim, Y., & Peng, H. (2024). Data Engineering for Scaling Language Models to 128K Context. arXiv. http://arxiv.org/abs/2402.10171v1
Chicago
Fu, Y., R. Panda, X. Niu, et al. 2024. “Data Engineering for Scaling Language Models to 128K Context”. arXiv. http://arxiv.org/abs/2402.10171v1.
Harvard
Fu, Y. et al. (2024) “Data Engineering for Scaling Language Models to 128K Context”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.10171v1.
Vancouver
1. Fu Y, Panda R, Niu X, Yue X, Hajishirzi H, Kim Y, Peng H (2024) Data Engineering for Scaling Language Models to 128K Context. arXiv

BibTeX

@article{fu2024data,
  title = {Data Engineering for Scaling Language Models to 128K Context},
  author = {Fu, Yao and Panda, Rameswar and Niu, Xinyao and Yue, Xiang and Hajishirzi, Hannaneh and Kim, Yoon and Peng, Hao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.10171v1},
  eprint = {2402.10171}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/