Compressing Context to Enhance Inference Efficiency of Large Language Models
Yucheng LiBo DongFrank GuerinChenghua Lin
Proposes Selective Context, an approach that uses language model surprisal to identify and prune redundant content from long input prompts, reducing inference memory usage by 36% and generation latency by 32% with minimal impact on output quality.
Modern large language models struggle to process extensive documents and prolonged conversations efficiently. Because the computational memory and processing time of standard model architectures grow quadratically with input length, handling long contexts causes high operational costs, increased latency, and frequent loss of information due to context truncation. Current efforts to address this issue mainly rely on modifying neural network architectures or training specialized compressed representations, which can be computationally intensive and difficult to generalize.
The article demonstrates and evaluates an alternative, model-agnostic technique called Selective Context. The main objective is to establish whether identifying and pruning redundant text from the input context itself can substantially enhance inference efficiency without degrading generation quality.
The researchers assessed the approach through empirical experiments on three datasets spanning academic papers from arXiv, BBC News articles, and multi-turn dialogues from ShareGPT.com. All test data were curated from content created after March 2023 to ensure the models had never seen the material during pre-training. Using a smaller base causal language model, the method calculates the self-information—or statistical surprisal—of lexical units such as tokens, phrases, or sentences. Units falling below a chosen percentile threshold of informativeness are pruned, and the compressed text is then processed by larger target models, including GPT-3.5, GPT-4, LLaMA, and Vicuna, across tasks such as summarization, question answering, conversation, and context reconstruction.
The key findings reveal substantial efficiency gains alongside preserved output quality. First, Selective Context achieved a 50% reduction in context length, which yielded an approximate 36% decrease in graphics memory usage and a 32% reduction in generation latency, while incurring only a minimal drop of 0.023 in BERTScore semantic similarity and 0.038 in factual unfaithfulness. Second, pruning at the noun-phrase level proved to be the optimal granularity, consistently outperforming token-level and sentence-level filtering. Third, the method significantly outperformed a random deletion baseline, where a 50% Selective Context compression retained higher output fidelity than random deletion at only 20%. Finally, human and manual evaluations showed that instruction-tuned models remained notably robust against compressed prompts, and models occasionally responded with brief non-answers rather than generating factual hallucinations when critical information was omitted.
These results demonstrate that natural language inputs contain significant inherent redundancy and overlap with background knowledge already stored in model parameters. In practice, this means organizations can process significantly longer inputs within existing hardware constraints, lower computational hosting expenses, and shorten turnaround times for real-time applications without retraining underlying models. Because the technique operates strictly at the input data level, it can also be combined with existing model-level optimization techniques.
For operational deployment, teams should consider adopting phrase-level context compression for summarization and question answering pipelines where efficiency is paramount. Practitioners should implement moderate compression ratios, roughly between 20% and 50%, to capture the majority of latency and cost savings while avoiding the steeper quality degradation observed at aggressive pruning rates above 65%. Further engineering work should focus on integrating dynamic threshold selection tailored to specific document types and developing advanced dependency-tree parsing to refine phrase boundary detection.
While confidence in the core efficiency and quality outcomes is high across the tested open-source and proprietary models, decision-makers should note certain limitations. The experimental implementation used simple noun phrase chunking without verb phrase processing, and the fixed-percentile pruning strategy does not yet adapt automatically to inputs with unusually dense information. Highly critical workflows requiring absolute verbatim recall should conduct focused pilot evaluations before applying aggressive compression.
- Paper: LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Huiqiang Jiang et al. (2023). Read LLMLingua first to see the prompt-compression approach that Selective Context develops for reducing inference cost while preserving task quality.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). LongLLMLingua extends prompt compression with query-aware ranking, context reordering, and adaptive budgets to improve long-context accuracy as well as efficiency.
- Paper: Adapting Language Models to Compress Contexts, Alexis Chevalier et al. (2023). AutoCompressors carry context reduction beyond input pruning by learning compact summary vectors that retain information across long sequences.
