Compressing Context to Enhance Inference Efficiency of Large Language Models

Yucheng LiBo DongFrank GuerinChenghua Lin

article2023EMNLP247 citations

Proposes Selective Context, an approach that uses language model surprisal to identify and prune redundant content from long input prompts, reducing inference memory usage by 36% and generation latency by 32% with minimal impact on output quality.

Listen

Modern large language models struggle to process extensive documents and prolonged conversations efficiently. Because the computational memory and processing time of standard model architectures grow quadratically with input length, handling long contexts causes high operational costs, increased latency, and frequent loss of information due to context truncation. Current efforts to address this issue mainly rely on modifying neural network architectures or training specialized compressed representations, which can be computationally intensive and difficult to generalize.

The article demonstrates and evaluates an alternative, model-agnostic technique called Selective Context. The main objective is to establish whether identifying and pruning redundant text from the input context itself can substantially enhance inference efficiency without degrading generation quality.

The researchers assessed the approach through empirical experiments on three datasets spanning academic papers from arXiv, BBC News articles, and multi-turn dialogues from ShareGPT.com. All test data were curated from content created after March 2023 to ensure the models had never seen the material during pre-training. Using a smaller base causal language model, the method calculates the self-information—or statistical surprisal—of lexical units such as tokens, phrases, or sentences. Units falling below a chosen percentile threshold of informativeness are pruned, and the compressed text is then processed by larger target models, including GPT-3.5, GPT-4, LLaMA, and Vicuna, across tasks such as summarization, question answering, conversation, and context reconstruction.

The key findings reveal substantial efficiency gains alongside preserved output quality. First, Selective Context achieved a 50% reduction in context length, which yielded an approximate 36% decrease in graphics memory usage and a 32% reduction in generation latency, while incurring only a minimal drop of 0.023 in BERTScore semantic similarity and 0.038 in factual unfaithfulness. Second, pruning at the noun-phrase level proved to be the optimal granularity, consistently outperforming token-level and sentence-level filtering. Third, the method significantly outperformed a random deletion baseline, where a 50% Selective Context compression retained higher output fidelity than random deletion at only 20%. Finally, human and manual evaluations showed that instruction-tuned models remained notably robust against compressed prompts, and models occasionally responded with brief non-answers rather than generating factual hallucinations when critical information was omitted.

These results demonstrate that natural language inputs contain significant inherent redundancy and overlap with background knowledge already stored in model parameters. In practice, this means organizations can process significantly longer inputs within existing hardware constraints, lower computational hosting expenses, and shorten turnaround times for real-time applications without retraining underlying models. Because the technique operates strictly at the input data level, it can also be combined with existing model-level optimization techniques.

For operational deployment, teams should consider adopting phrase-level context compression for summarization and question answering pipelines where efficiency is paramount. Practitioners should implement moderate compression ratios, roughly between 20% and 50%, to capture the majority of latency and cost savings while avoiding the steeper quality degradation observed at aggressive pruning rates above 65%. Further engineering work should focus on integrating dynamic threshold selection tailored to specific document types and developing advanced dependency-tree parsing to refine phrase boundary detection.

While confidence in the core efficiency and quality outcomes is high across the tested open-source and proprietary models, decision-makers should note certain limitations. The experimental implementation used simple noun phrase chunking without verb phrase processing, and the fixed-percentile pruning strategy does not yet adapt automatically to inputs with unusually dense information. Highly critical workflows requiring absolute verbatim recall should conduct focused pilot evaluations before applying aggressive compression.

Cover for Compressing Context to Enhance Inference Efficiency of Large Language Models

Abstract

Large language models (LLMs) achieved remarkable performance across various tasks. However, they face challenges in managing long documents and extended conversations, due to significantly increased computational requirements, both in memory and inference time, and potential context truncation when the input exceeds the LLM’s fixed context length. This paper proposes a method called Selective Context that enhances the inference efficiency of LLMs by identifying and pruning redundancy in the input context to make the input more compact. We test our approach using common data sources requiring long context processing: arXiv papers, news articles, and long conversations, on tasks of summarisation, question answering, and response generation. Experimental results show that Selective Context significantly reduces memory cost and decreases generation latency while maintaining comparable performance compared to that achieved when full context is used. Specifically, we achieve a 50% reduction in context cost, resulting in a 36% reduction in inference memory usage and a 32% reduction in inference time, while observing only a minor drop of .023 in BERTscore and .038 in faithfulness on four downstream applications, indicating that our method strikes a good balance between efficiency and performance. Code and data are available at https://github.com/liyucheng09/Selective_Context.

Table of Contents

  • Abstract — removed placeholder
  • 1 Introduction
  • 2 Self-Information
  • 3 Method
  • 3.1 Computing Self-Information
  • 3.2 Merging into Lexical Units
  • 3.3 Selective Retention of Informative Context
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Models
  • 4.3 Tasks and Metrics
  • 4.4 Experimental Settings
  • 5 Results
  • 5.1 Overview
  • 5.2 Faithfulness
  • 5.3 Tasks
  • 5.4 Scaling and Instruct-Tuning
  • 5.5 Lexical Units
  • 5.6 Case Study
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Dataset Statistics
  • B Example of Selective Context on Long Context
  • C The Previous Version of Selective Context

Knowls

  1. Knowl 1 — Selective Context scores and prunes lexical units by informativeness

    model/method

    Selective Context compresses an input context by assigning token-level self-information scores with a causal language model, grouping tokens into lexical units, and discarding low-scoring units. The implementation can filter individual tokens, NLTK-tokenized sentences, or spaCy noun phrases; it does not merge verb phrases. To avoid the observed tendency for later units in a long input to receive lower scores, the paper computes scores sentence by sentence rather than scoring the entire context at once. For a lexical unit, it sums the scores of its constituent tokens. Given a desired content-reduction fraction, it computes the corresponding percentile of unit scores and retains units whose scores are at least that percentile; for example, the 50th percentile removes about half the units. Retained units are kept in their original order to form the compressed context.

  2. Knowl 2 — Self-information supplies additive scores for tokens and lexical units

    equation

    For a token sequence x1,…,xTx_1,\ldots,x_T and a causal language model MM, the self-information of token xix_i is its negative conditional log probability, measured in bits. For a contiguous lexical unit uu spanning tokens xax_a through xbx_b, its score is the sum of the token scores:

    I(xi)=−log⁡2PM(xi∣x1,…,xi−1),I(u)=∑i=abI(xi).I(x_i)=-\log_2 P_M(x_i\mid x_1,\ldots,x_{i-1}), \qquad I(u)=\sum_{i=a}^{b} I(x_i).

    Here PMP_M is the next-token probability assigned by MM, and a,b,ia,b,i are token indices with 1≤a≤b≤T1\leq a\leq b\leq T. The method treats low-scoring units as less informative and therefore more likely to be inferable from the remaining context.

  3. Knowl 3 — Evaluation covers long documents, conversations, and four generation tasks

    experimental setup

    The evaluation uses 408 arXiv articles, 294 BBC news articles, and 470 ShareGPT conversations. The average document lengths are 864.85 tokens for arXiv, 732.54 for BBC, and 689.32 for ShareGPT; arXiv inputs include the first two sections, BBC inputs use the full article, and ShareGPT is used for conversation response generation. The tasks are original-context reconstruction, summarization, question answering, and conversation response generation. For question answering, questions and reference answers are first generated from the full context; for summarization and reconstruction, full-context outputs or the original context serve as references; for conversation, the model answers the final user query using the pruned history. Tested generators are GPT-3.5, GPT-4, LLaMA-7B/13B/30B, and Vicuna-7B/13B. Self-information is computed with LLaMA-7B for LLaMA and Vicuna models and with OpenAI Curie for OpenAI models. The tested reduction fractions are 0.2, 0.35, 0.5, 0.65, and 0.8. The study uses BLEU, METEOR, ROUGE, and BERTScore, comparing compressed-context outputs with full-context outputs; it uses four full-context references for temperature-0.7 decoding and one for greedy decoding.

  4. Knowl 4 — Output similarity declines gradually as more context is removed

    empirical result

    In the aggregate comparison with full-context generation, the reported full-context scores are BLEU 0.347, ROUGE-1 0.571, and BERTScore-F1 0.909. At content-reduction fractions 0.2, 0.35, 0.5, 0.65, and 0.8, respectively, the corresponding scores are: BLEU 0.295, 0.243, 0.179, 0.127, and 0.070; ROUGE-1 0.540, 0.504, 0.449, 0.391, and 0.311; and BERTScore-F1 0.902, 0.897, 0.887, 0.877, and 0.863. The scores are averaged over the tested models. Thus, reductions of 0.2 or 0.35 produce relatively small semantic-similarity drops, while larger reductions increasingly degrade similarity.

  5. Knowl 5 — Informativeness-based pruning beats random deletion at matched reduction rates

    empirical result

    Under greedy decoding, Selective Context yields higher scores than random deletion at the same content-reduction fraction. For example, at reduction fractions 0.2, 0.5, and 0.8, Selective Context versus random deletion scores are: BLEU 0.527/0.437, 0.350/0.283, and 0.160/0.156; ROUGE-1 0.714/0.666, 0.642/0.576, and 0.464/0.450; and BERTScore-F1 0.931/0.899, 0.900/0.873, and 0.866/0.847. The full comparison reports the same advantage for Selective Context across the evaluated metrics at each matched reduction rate, showing that selecting units by self-information is more effective than deleting the same amount of content at random.

  6. Knowl 6 — Summarization and question answering tolerate pruning better than reconstruction

    empirical result

    Across summarization, question answering, original-context reconstruction, and conversation response generation, reconstruction shows the steepest performance decline as the content-reduction fraction increases. In the task-wise BERTScore comparisons, summarization and question answering remain above reconstruction and conversation across the tested reductions. The paper therefore identifies summarization and answer generation as particularly suitable uses of Selective Context, while reconstructing the complete original input is more sensitive to removed content.

  7. Knowl 7 — Reported efficiency gains accompany moderate quality changes

    empirical result

    The paper reports that reducing context cost by 50% yields a 36% reduction in inference memory use and a 32% reduction in inference time, with an average BERTScore drop of 0.023 and faithfulness drop of 0.038 across its reported downstream applications. In one Vicuna-13B summary-generation example, full-context inference used 77,695 MB of CUDA memory and took 110.8 ms per generated token; with a 0.5 reduction fraction, the reported values were 61,885 MB and 76.3 ms per token. Constructing that selective context took 46.1 ms. The single-example measurements and the paper's aggregate efficiency claims describe different reported evaluation results.

  8. Knowl 8 — Faithfulness is mostly preserved, but unfaithful tuples increase at the strongest pruning

    empirical result

    The paper manually audits 1,000 question-answer pairs from GPT-3.5-turbo, with 200 pairs at each reduction fraction. OpenIE tuples are extracted from generated answers and checked for entailment by the full-context reference answer; answers of Sorry, I don't know are counted separately and excluded from the unfaithfulness calculation. For reduction fractions 0.2, 0.35, 0.5, 0.65, and 0.8, the numbers of refusals are 0, 6, 4, 19, and 27; average answer lengths are 156.5, 136.0, 140.2, 131.2, and 103.7; and the unfaithful-tuple rates are 0.027, 0.050, 0.038, 0.051, and 0.086. The rate is not monotonic across intermediate settings, but is highest at the 0.8 reduction fraction.

  9. Knowl 9 — Phrase-level filtering performs best among the tested unit granularities

    empirical result

    The paper compares token-, phrase-, and sentence-level Selective Context using BERTScore-F1. Phrase-level filtering consistently performs best, token-level filtering ranks next, and sentence-level filtering is less stable. This supports using phrase-sized units rather than deleting isolated tokens or whole sentences; the implemented phrase units are noun phrases.

  10. Knowl 10 — Instruct tuning helps more than model scaling in the tested summary evaluation

    empirical result

    Three student evaluators judged 1,150 generated summaries from LLaMA and Vicuna models for reasonableness, with about 55 summaries per model and reduction fraction. The authors report no clear improvement in acceptance as model size increases within the tested LLaMA or Vicuna families. Vicuna, which is instruction-tuned from LLaMA, has higher acceptance than the LLaMA models under selective context. The paper also observes that LLaMA outputs often fail to follow instructions when context is compressed; it presents instruction tuning as a possible aid to robustness, rather than establishing a general scaling result.

  11. Knowl 11 — Phrase boundaries and context-dependent pruning thresholds limit the method

    limitation

    Selective Context depends on the quality of lexical-unit boundaries: the experiments use spaCy noun-phrase merging and do not handle verb phrases because the authors report no mature solution for identifying them. The authors speculate that dependency-tree-based filtering could improve boundary identification, but do not demonstrate that improvement. The pruning percentile that works best also varies with the task and input context, so a fixed setting need not be optimal; developing a tool to select that threshold is left as future work.

Coverage note — Illustrative appendix passages and explanations of standard evaluation metrics are omitted because they add no separate contributed method or result beyond the findings represented here.

References

  1. 1.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  2. 2.Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712.
  5. 5.Razvan Bunescu and Oseremen O Uduehi. 2022. Distribution-based measures of surprise for creative language: Experiments with humor and metaphor. In Proceedings of the 3rd Workshop on Figurative Language Processing (FLP), pages 68–78.
  6. 6.Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. arXiv preprint arXiv:2305.14788.
  7. 7.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  8. 8.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509.
  9. 9.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv, abs/1803.05457.
  10. 10.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Xiaodong Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. ArXiv, abs/2009.03300.
  11. 11.Yucheng Li. 2023. Unlocking context constraints of llms: Enhancing context efficiency of llms with self-information-based content filtering. ArXiv, abs/2304.12102.
  12. 12.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  13. 13.Jesse Mu, Xiang Lisa Li, and Noah Goodman. 2023. Learning to compress prompts with gist tokens. arXiv preprint arXiv:2304.08467.
  14. 14.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  15. 15.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  16. 16.Claude E Shannon. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423.
  17. 17.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  18. 18.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  19. 19.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. ArXiv, abs/2004.04228.
  20. 20.Ernst-Jan C. Wit and Marie Gillette. 1999. What is linguistic redundancy. University of Chicago.
  21. 21.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  22. 22.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.

Citation

MLA
Li, Y., et al. “Compressing Context to Enhance Inference Efficiency of Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 6342–53, https://doi.org/10.18653/v1/2023.emnlp-main.391.
APA
Li, Y., Dong, B., Guerin, F., & Lin, C. (2023). Compressing Context to Enhance Inference Efficiency of Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6342–6353. https://doi.org/10.18653/v1/2023.emnlp-main.391
Chicago
Li, Y., B. Dong, F. Guerin, and C. Lin. 2023. “Compressing Context to Enhance Inference Efficiency of Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6342–53. https://doi.org/10.18653/v1/2023.emnlp-main.391.
Harvard
Li, Y. et al. (2023) “Compressing Context to Enhance Inference Efficiency of Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6342–6353. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.391.
Vancouver
1. Li Y, Dong B, Guerin F, Lin C (2023) Compressing Context to Enhance Inference Efficiency of Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6342–6353

BibTeX

@inproceedings{li-etal-2023-compressing,
    title = "Compressing Context to Enhance Inference Efficiency of Large Language Models",
    author = "Li, Yucheng  and
      Dong, Bo  and
      Guerin, Frank  and
      Lin, Chenghua",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.391/",
    doi = "10.18653/v1/2023.emnlp-main.391",
    pages = "6342--6353"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/