LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Huiqiang JiangQianhui WuChin-Yew LinYuqing YangLili Qiu
Proposes LLMLingua, a coarse-to-fine prompt compression framework that uses smaller language models to reduce prompt length by up to 20x while preserving key reasoning and context capabilities across black-box large language models.
Modern large language model applications increasingly rely on extended inputs—such as detailed instructions, multi-step chain-of-thought demonstrations, and retrieved context—which can span thousands of tokens. Processing these large inputs significantly escalates cloud computation costs, increases operational latency, and risks exceeding system context windows. Although existing acceleration techniques modify internal model parameters, they cannot be readily applied to proprietary, black-box systems accessed exclusively via commercial application programming interfaces.
The article demonstrates and evaluates LLMLingua, a prompt compression method designed to shorten long inputs before they are sent to large language models. The primary objective is to substantially accelerate model processing and reduce API expenses without fine-tuning the underlying target model or sacrificing overall output quality and reasoning capability.
The authors implemented a three-part framework that evaluates token informativeness using a smaller, aligned language model. First, a budget controller dynamically preserves critical instructions and questions while selectively pruning redundant examples at the sentence or demonstration level. Second, an iterative token-level compression algorithm evaluates and drops low-information tokens while preserving context dependencies. Third, instruction tuning aligns the probability distributions of the small compression model and the target large model. The approach was evaluated across four diverse benchmarks covering mathematical reasoning, symbolic logic, multi-turn dialogue, and scientific summarization using systems including GPT-3.5-Turbo and Claude.
The evaluation produced four central findings. First, the method achieved up to a 20x compression ratio on complex mathematical reasoning tasks while experiencing only a minor performance decline of roughly 1.5 points in accuracy, significantly outperforming prior compression techniques. Second, the compressed inputs enabled end-to-end processing speedups ranging from 1.7x to 5.7x across tested configurations. Third, shortening the input prompts naturally shortened the length of the generated outputs, yielding compounding computational and financial savings; for instance, evaluation costs dropped from 0.50 on reasoning benchmarks and 0.20 on summarization tasks. Fourth, ablation experiments confirmed that iterative token filtering and dynamic budget allocation are both critical to preventing catastrophic logic loss during compression.
These findings indicate that organizations deploying large language models can immediately lower operational expenses and improve response times without altering target model architectures or renegotiating infrastructure contracts. The results also counter the common assumption that highly compressed text necessarily degrades complex reasoning, proving that modern frontier models can successfully interpret non-fluent, semantically dense prompts.
Organizations handling high volumes of long-form context should consider piloting prompt compression pipelines ahead of their primary model calls, particularly for repetitive few-shot demonstrations and extended background context. Because the method operates as a preprocessing layer, technical teams can adopt it modularly alongside existing prompt engineering strategies.
Decision-makers should note certain operational boundaries. Accuracy declines steeply when pushing compression beyond extreme thresholds, such as 25x to 30x, or when applied to highly intricate spatial reasoning tasks. Minor discrepancies between tokenization schemes across different models can also slightly skew compression budgets. Nevertheless, within recommended compression ratios, the evidence provides high confidence that the method offers substantial cost and latency reductions for production language model deployments.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Understanding how soft prompt tuning and scale interact provides foundational context for prompt-level adaptation methods used in LLM prompt compression.
- Paper: P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks, Xiao Liu et al. (2022). This paper establishes continuous prompt tuning across model scales, a core concept underlying efficient prompt parameterization and manipulation.
- Paper: RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning, Mingkai Deng et al. (2022). This work introduces techniques for discrete prompt optimization, informing how discrete token selection and compression algorithms can be structured.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). LongLLMLingua directly extends LLMLingua to address long-context challenges like position bias and lost-in-the-middle phenomena via question-aware prompt compression.
- Paper: Adapting Language Models to Compress Contexts, Alexis Chevalier et al. (2023). This work explores adapting language models into recursive summary compressors, offering a learned model-based alternative to token-pruning prompt compression.
- Paper: Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching, Simon A. Aytes et al. (2025). Sketch-of-Thought applies cognitive sketching principles to condense intermediate reasoning prompts and reduce inference token overhead.
- Paper: Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference, Piotr Nawrot et al. (2024). This research continues the goal of inference acceleration by compressing intermediate key-value states dynamically rather than compressing raw input prompts.
- Paper: REFRAG: Rethinking RAG based Decoding, Xiaoqiang Lin et al. (2025). REFRAG applies context compression specifically to retrieval-augmented generation pipelines to accelerate time-to-first-token during decoding.
