Adapting Language Models to Compress Contexts
Alexis ChevalierAlexander WettigAnirudh AjithDanqi Chen
Introduces AutoCompressors, an unsupervised method for fine-tuning pre-trained language models to recursively compress long contexts into compact summary vectors that serve as soft prompts, significantly extending effective context windows up to 30,720 tokens while lowering inference costs for in-context learning and retrieval tasks.
Transformer-based language models are central to modern artificial intelligence applications, but they face severe practical constraints due to finite context windows and steep computational costs when processing long texts. Standard attention mechanisms scale quadratically with sequence length, making the ingestion of long documents prohibitively expensive in memory and compute. The article addresses this bottleneck by evaluating whether pre-trained language models can be adapted into "AutoCompressors" that recursively compress long context sequences into compact "summary vectors" (short soft prompts) to extend context capacity and accelerate inference.
The authors develop an unsupervised fine-tuning approach that equips models like OPT (1.3B and 2.7B parameters) and Llama-2 (7B parameters) with special summary tokens. Documents are divided into segments, compressed into summary vectors, and accumulated across segments as soft prompts for future text. To enable efficient training on a single 80GB GPU, the framework incorporates summary accumulation, randomized segment lengths, and backpropagation through time with stopped gradients. The study evaluates long-range language modeling on sequences up to 30,720 tokens, in-context learning across 11 benchmark tasks, retrieval-augmented generation on multi-billion token corpora, and unsupervised document re-ranking.
The findings show that AutoCompressors effectively compress long contexts while retaining critical factual and semantic information. First, in long-context evaluations, AutoCompressors successfully leveraged contexts up to 28,000 tokens to consistently reduce perplexity, outperforming baseline Recurrent Memory Transformers. Second, for in-context learning, substituting plain-text examples with summary vectors yielded higher accuracy than 150 plain tokens on 8 out of 11 tasks, and outperformed 750 plain tokens on 8 tasks while substantially lowering token processing requirements. Third, when applied to retrieval-augmented modeling, fusing pre-computed summary vectors achieved 1.5 times the perplexity gain of plain-text passages and delivered a 1.7-fold throughput increase over traditional multi-passage ensembling. Finally, in passage re-ranking, caching summary vectors established a Pareto-optimal trade-off between retrieval recall and computational throughput.
These results demonstrate that pre-computing and caching summary vectors offers a scalable, cost-effective method to expand context windows and speed up high-volume inference workflows. Organizations running large-scale retrieval or few-shot classification systems can lower runtime latency and operational costs without training massive architectures from scratch. Next steps supported by the article include evaluating the approach on larger foundation models, refining training to better capture granular information that full attention retains, and optimizing methods for aggregating high volumes of summary vectors.
Readers should note certain limitations: testing was restricted to models up to 7B parameters, and summary vectors exhibited a slight performance gap compared to full attention over very long spans due to information loss during compression. Nevertheless, the empirical findings provide high confidence that context compression via summary vectors is a viable and efficient enhancement for production language model pipelines.
- Paper: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context, Zihang Dai et al. (2019). Introduces segment-level recurrence and cached hidden states across chunks, providing the foundational mechanism for segment-based context compression in transformers.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). Establishes attention with linear biases to enable input length extrapolation, directly addressing the length generalization challenges central to long-context language modeling.
- Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Presents localized windowed and global attention patterns that serve as core reference baselines for scaling transformers to long documents.
- Paper: Improving language models by retrieving from trillions of tokens, Sebastian Borgeaud et al. (2022). Demonstrates how chunked cross-attention and external retrieval can scale context capacity, motivating the retrieval-augmented evaluations in AutoCompressors.
- Paper: Memorizing Transformers, Yuhuai Wu et al. (2022). Shows how appending external memory caches across document chunks reduces perplexity, laying groundwork for accumulating summary vectors across sequence segments.
- Paper: General-purpose, long-context autoregressive modeling with Perceiver AR, Curtis Hawthorne et al. (2022). Pioneers the compression of large input contexts into compact latent representations to decouple context length from deep attention layers.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). Provides a comprehensive taxonomy and efficiency analysis of sub-quadratic attention and recurrence mechanisms that context compression seeks to improve upon.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). Extends long-context prompt compression to downstream black-box models by dynamically budgeting and reconstructing prompt tokens to combat position bias.
- Paper: Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference, Piotr Nawrot et al. (2024). Advances the compression of sequence memory during inference by retrofitting models to dynamically accumulate key-value representations on the fly.
- Paper: REFRAG: Rethinking RAG based Decoding, Xiaoqiang Lin et al. (2025). Applies context-chunk compression specifically to retrieval-augmented generation pipelines to drastically cut time-to-first-token latency.
- Paper: DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model, Zhihong Shao et al. (2024). Employs multi-head latent attention to compress key-value caches directly into compact latent vectors within a large-scale mixture-of-experts model.
- Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). Generalizes context compression by framing long-context processing as test-time learning that compresses sequence history directly into model weights.
- Paper: Titans: Learning to Memorize at Test Time, Ali Behrouz et al. (2024). Explores architectural memory extensions that compress and retain long-term historical context using test-time memory updates alongside short-term attention.
- Paper: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention, Jingyang Yuan et al. (2025). Combines hierarchical block compression of past tokens with native sparse attention to scale long-context processing efficiently on hardware.
- Paper: PolarQuant: Quantizing KV Caches with Polar Transformation, Insu Han et al. (2025). Addresses inference cache bloat by quantizing long-context key-value embeddings into compact polar coordinates.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). Builds an operating-system-style virtual context manager that paginates and evicts compressed context memory to support unbounded conversational lengths.
- Paper: Prefix Sliding for efficient test-time scaling, Niklas Muennighoff et al. (2026). Investigates prefix and sliding window strategies for bounded-memory context management during long-chain test-time reasoning.
