LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Huiqiang JiangQianhui WuChin-Yew LinYuqing YangLili Qiu

article2023EMNLP342 citations

Proposes LLMLingua, a coarse-to-fine prompt compression framework that uses smaller language models to reduce prompt length by up to 20x while preserving key reasoning and context capabilities across black-box large language models.

Listen

Modern large language model applications increasingly rely on extended inputs—such as detailed instructions, multi-step chain-of-thought demonstrations, and retrieved context—which can span thousands of tokens. Processing these large inputs significantly escalates cloud computation costs, increases operational latency, and risks exceeding system context windows. Although existing acceleration techniques modify internal model parameters, they cannot be readily applied to proprietary, black-box systems accessed exclusively via commercial application programming interfaces.

The article demonstrates and evaluates LLMLingua, a prompt compression method designed to shorten long inputs before they are sent to large language models. The primary objective is to substantially accelerate model processing and reduce API expenses without fine-tuning the underlying target model or sacrificing overall output quality and reasoning capability.

The authors implemented a three-part framework that evaluates token informativeness using a smaller, aligned language model. First, a budget controller dynamically preserves critical instructions and questions while selectively pruning redundant examples at the sentence or demonstration level. Second, an iterative token-level compression algorithm evaluates and drops low-information tokens while preserving context dependencies. Third, instruction tuning aligns the probability distributions of the small compression model and the target large model. The approach was evaluated across four diverse benchmarks covering mathematical reasoning, symbolic logic, multi-turn dialogue, and scientific summarization using systems including GPT-3.5-Turbo and Claude.

The evaluation produced four central findings. First, the method achieved up to a 20x compression ratio on complex mathematical reasoning tasks while experiencing only a minor performance decline of roughly 1.5 points in accuracy, significantly outperforming prior compression techniques. Second, the compressed inputs enabled end-to-end processing speedups ranging from 1.7x to 5.7x across tested configurations. Third, shortening the input prompts naturally shortened the length of the generated outputs, yielding compounding computational and financial savings; for instance, evaluation costs dropped from 5.20to5.20 to 0.50 on reasoning benchmarks and 1.30to1.30 to 0.20 on summarization tasks. Fourth, ablation experiments confirmed that iterative token filtering and dynamic budget allocation are both critical to preventing catastrophic logic loss during compression.

These findings indicate that organizations deploying large language models can immediately lower operational expenses and improve response times without altering target model architectures or renegotiating infrastructure contracts. The results also counter the common assumption that highly compressed text necessarily degrades complex reasoning, proving that modern frontier models can successfully interpret non-fluent, semantically dense prompts.

Organizations handling high volumes of long-form context should consider piloting prompt compression pipelines ahead of their primary model calls, particularly for repetitive few-shot demonstrations and extended background context. Because the method operates as a preprocessing layer, technical teams can adopt it modularly alongside existing prompt engineering strategies.

Decision-makers should note certain operational boundaries. Accuracy declines steeply when pushing compression beyond extreme thresholds, such as 25x to 30x, or when applied to highly intricate spatial reasoning tasks. Minor discrepancies between tokenization schemes across different models can also slightly skew compression budgets. Nevertheless, within recommended compression ratios, the evidence provides high confidence that the method offers substantial cost and latency reductions for production language model deployments.

arXiv: 2310.05736microsoft/LLMLingua
Cover for LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Abstract

Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens. To accelerate model inference and reduce cost, this paper presents LLMLingua, a coarse-to-fine prompt compression method that involves a budget controller to maintain semantic integrity under high compression ratios, a token-level iterative compression algorithm to better model the interdependence between compressed contents, and an instruction tuning based method for distribution alignment between language models. We conduct experiments and analysis over four datasets from different scenarios, i.e., GSM8K, BBH, ShareGPT, and Arxiv-March23; showing that the proposed approach yields state-of-the-art performance and allows for up to 20x compression with little performance loss.¹

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Efficient LLMs
  • 2.2 Out-of-Distribution (OoD) Detection
  • 2.3 LLMs as a Compressor
  • 3 Problem Formulation
  • 4 Methodology
  • 4.1 Budget Controller
  • 4.2 Iterative Token-level Prompt Compression
  • 4.3 Distribution Alignment
  • 5 Experiments
  • 5.1 Settings
  • 5.2 Main Results
  • 5.3 Analysis on Reasoning & ICL Tasks.
  • 5.4 Ablation
  • 5.5 Discussion
  • 6 Conclusion
  • Limitations
  • References
  • A Experiment Details
  • A.1 Dataset Details
  • A.2 Other Implementation Details
  • B Economic Cost
  • C Instructions used in GPT-4 Generation
  • D Recovering Compressed Prompts with Large Language Model
  • E Cases Study

Knowls

  1. Knowl 1 — Problem Formulation of Prompt Compression

    definition

    Prompt compression is formulated as the task of generating a shortened prompt x~={x~i}i=1L~\tilde{x} = \{\tilde{x}_i\}_{i=1}^{\tilde{L}} from an original prompt x=(xins,xdems,xque)x = (x^{\text{ins}}, x^{\text{dems}}, x^{\text{que}}), where xins={xiins}i=1Linsx^{\text{ins}} = \{x_i^{\text{ins}}\}_{i=1}^{L_{\text{ins}}} represents the instruction, xdems={xidems}i=1Ldemsx^{\text{dems}} = \{x_i^{\text{dems}}\}_{i=1}^{L_{\text{dems}}} represents the set of in-context demonstrations, and xque={xique}i=1Lquex^{\text{que}} = \{x_i^{\text{que}}\}_{i=1}^{L_{\text{que}}} represents the question or input query. The total original prompt length is L=Lins+Ldems+LqueL = L_{\text{ins}} + L_{\text{dems}} + L_{\text{que}}, and L~\tilde{L} denotes the number of tokens in x~\tilde{x}.

    The compression rate is defined as τ=L~/L∈[0,1]\tau = \tilde{L} / L \in [0, 1], and the compression ratio is 1/τ1/\tau. Given an autoregressive target large language model (LLM), let xGx_G denote the output tokens generated from original prompt xx, and let x~G\tilde{x}_G denote the output tokens generated from compressed prompt x~\tilde{x}. The prompt compression objective is to minimize the Kullback-Leibler (KL) divergence between the predictive output distributions under budget constraints:

    min⁡x~,τKL(P(x~G∣x~)∥P(xG∣x))\min_{\tilde{x}, \tau} \text{KL}(P(\tilde{x}_G \mid \tilde{x}) \parallel P(x_G \mid x))

  2. Knowl 2 — Budget Controller for Demonstration-Level Compression and Budget Allocation

    algorithm

    The Budget Controller is a coarse-grained prompt compression module that dynamically allocates compression budgets to the instruction xinsx^{\text{ins}}, demonstrations xdemsx^{\text{dems}}, and question xquex^{\text{que}}, while filtering demonstrations at the demonstration level using perplexity evaluated by a small language model MsM_s.

    Given target overall compression rate τ\tau and pre-defined compression rates for instruction and question (τins\tau_{\text{ins}} and τque\tau_{\text{que}}, set to 0.85 and 0.90, respectively), the target compression rate for demonstrations is computed as:

    τdems=τL−(τinsLins+τqueLque)Ldems\tau_{\text{dems}} = \frac{\tau L - (\tau_{\text{ins}} L_{\text{ins}} + \tau_{\text{que}} L_{\text{que}})}{L_{\text{dems}}}

    Demonstrations are ranked in descending order of perplexity computed by MsM_s. Demonstrations with higher perplexities (higher information content) are sequentially added to selected demonstration set D\mathcal{D} until adding an additional demonstration causes the total tokens L~D\tilde{L}_{\mathcal{D}} to exceed k⋅τdemsLdemsk \cdot \tau_{\text{dems}} L_{\text{dems}}, where kk is a granular control coefficient (default k=2k=2). The remaining unspent demonstration budget is reallocated to the instruction and question:

    Δτ=k⋅τdemsLdems−L~DLins+Lque\Delta\tau = \frac{k \cdot \tau_{\text{dems}} L_{\text{dems}} - \tilde{L}_{\mathcal{D}}}{L_{\text{ins}} + L_{\text{que}}}

    Input: Small LM MsM_s; original prompt x=(xins,xdems,xque)x = (x^{\text{ins}}, x^{\text{dems}}, x^{\text{que}}); target rate τ\tau; pre-defined rates τins,τque\tau_{\text{ins}}, \tau_{\text{que}}; granular coefficient kk.
    Output: Filtered demonstration set D\mathcal{D}; additional budget Δτ\Delta\tau.
    1: Initialize demonstration set D=∅\mathcal{D} = \emptyset.
    2: Compute demonstration compression rate τdems=τL−(τinsLins+τqueLque)Ldems\tau_{\text{dems}} = \frac{\tau L - (\tau_{\text{ins}} L_{\text{ins}} + \tau_{\text{que}} L_{\text{que}})}{L_{\text{dems}}}.
    3: Calculate perplexity for each demonstration in xdemsx^{\text{dems}} using MsM_s.
    4: Sort demonstrations in descending order of perplexity: (x(1)dem,…,x(N)dem)(x^{\text{dem}}_{(1)}, \dots, x^{\text{dem}}_{(N)}).
    5: for i=1i = 1 to NN do
    6: if L~D+Length(x(i)dem)>k⋅τdemsLdems\tilde{L}_{\mathcal{D}} + \text{Length}(x^{\text{dem}}_{(i)}) > k \cdot \tau_{\text{dems}} L_{\text{dems}} then
    7: break
    8: end if
    9: Append x(i)demx^{\text{dem}}_{(i)} to D\mathcal{D}.
    10: end for
    11: Calculate leftover budget Δτ=k⋅τdemsLdems−L~DLins+Lque\Delta\tau = \frac{k \cdot \tau_{\text{dems}} L_{\text{dems}} - \tilde{L}_{\mathcal{D}}}{L_{\text{ins}} + L_{\text{que}}}.
    12: return D,Δτ\mathcal{D}, \Delta\tau
  3. Knowl 3 — Iterative Token-Level Prompt Compression Algorithm

    algorithm

    Iterative Token-level Prompt Compression (ITPC) performs fine-grained token filtering while addressing the conditional independence deficiency of non-iterative perplexity pruning. In standard pruning, token self-information assumes independence (p(x~)≈∏p(xi∣x<i)p(\tilde{x}) \approx \prod p(x_i \mid x_{<i})), ignoring the removal of prior tokens.

    ITPC partitions the prompt x′=(xins,xD,xque)x' = (x^{\text{ins}}, x^{\mathcal{D}}, x^{\text{que}}) into segments S={s1,s2,…,sm}S = \{s_1, s_2, \dots, s_m\} (default segment length is 100 tokens). It computes conditional token probabilities for segment sjs_j conditioned on the preserved tokens from all preceding segments s~<j\tilde{s}_{<j}:

    p(s~j)≈∏i=1Ls,j+∑k=1j−1L~s,kp(sj,i∣sj,<i,s~<j)p(\tilde{s}_j) \approx \prod_{i=1}^{L_{s,j} + \sum_{k=1}^{j-1} \tilde{L}_{s,k}} p(s_{j,i} \mid s_{j,<i}, \tilde{s}_{<j})

    where sj,is_{j,i} is the ii-th token of segment jj, and Ls,jL_{s,j} and L~s,j\tilde{L}_{s,j} are original and compressed lengths of segment jj.

    For each segment sjs_j, its target compression rate is assigned based on prompt component origin:

    τsj={τins+Δτ,if sj∈xinsτdems,if sj∈xDτque+Δτ,if sj∈xque\tau_{s_j} = \begin{cases} \tau_{\text{ins}} + \Delta\tau, & \text{if } s_j \in x^{\text{ins}} \\ \tau_{\text{dems}}, & \text{if } s_j \in x^{\mathcal{D}} \\ \tau_{\text{que}} + \Delta\tau, & \text{if } s_j \in x^{\text{que}} \end{cases}

    A probability threshold γj\gamma_j is dynamically chosen such that retaining tokens satisfying p(sj,i)>γjp(s_{j,i}) > \gamma_j achieves compression rate τsj\tau_{s_j}:

    s~j={sj,i∣p(sj,i)>γj}\tilde{s}_j = \{s_{j,i} \mid p(s_{j,i}) > \gamma_j\}

    Input: Small LM MsM_s; prompt x′=(xins,xD,xque)x' = (x^{\text{ins}}, x^{\mathcal{D}}, x^{\text{que}}); rates τins,τdems,τque\tau_{\text{ins}}, \tau_{\text{dems}}, \tau_{\text{que}}; budget adjustment Δτ\Delta\tau.
    Output: Compressed prompt x~\tilde{x}.
    1: Initialize token set T=∅\mathcal{T} = \emptyset.
    2: Divide x′x' into segments S={s1,s2,…,sm}S = \{s_1, s_2, \dots, s_m\}.
    3: for j=1j = 1 to mm do
    4: Compute conditional probabilities p(sj,i∣sj,<i,s~<j)p(s_{j,i} \mid s_{j,<i}, \tilde{s}_{<j}) using MsM_s.
    5: Determine component rate τsj\tau_{s_j} and threshold γj\gamma_j matching τsj\tau_{s_j}.
    6: Select tokens s~j={sj,i∣p(sj,i)>γj}\tilde{s}_j = \{s_{j,i} \mid p(s_{j,i}) > \gamma_j\}.
    7: Append s~j\tilde{s}_j to T\mathcal{T}.
    8: end for
    9: Concatenate all preserved tokens T\mathcal{T} to yield x~\tilde{x}.
    10: return x~\tilde{x}
  4. Knowl 4 — Small-to-Large Language Model Distribution Alignment

    model/method

    To mitigate the probability distribution discrepancy between the small language model MsM_s (used to estimate token perplexity) and the target black-box large language model (LLM), distribution alignment is conducted via instruction tuning.

    Starting from a pre-trained small model MsM_s with parameters θMs\theta_{M_s}, the model is fine-tuned on instruction-response pairs (xi,yiLLM)(x_i, y_i^{\text{LLM}}) where outputs yiLLMy_i^{\text{LLM}} are generated by the target LLM (e.g., using the Stanford Alpaca dataset). The training objective optimizes the cross-entropy loss L\mathcal{L} over NN instruction tuning pairs:

    min⁡θMsE[1N∑i=1NL(xi,yiLLM;θMs)]\min_{\theta_{M_s}} \mathbb{E} \left[ \frac{1}{N} \sum_{i=1}^N \mathcal{L}(x_i, y_i^{\text{LLM}}; \theta_{M_s}) \right]

    Aligning the compression model's distribution enables more accurate identification of tokens that are informative to the target black-box LLM.

  5. Knowl 5 — Computational Cost Formulation and Latency Speedup of LLMLingua

    theoretical result

    The overall computational cost cc of compressing a prompt and executing inference with a target LLM is given by:

    c=(L+kLτ+Lτ)⋅csmall+Lτ⋅cLLMsc = \left(L + k \frac{L}{\tau} + \frac{L}{\tau}\right) \cdot c_{\text{small}} + \frac{L}{\tau} \cdot c_{\text{LLMs}}

    where csmallc_{\text{small}} and cLLMsc_{\text{LLMs}} denote the per-token computational loads of the small language model MsM_s and the target LLM, respectively. LL tokens are processed by the Budget Controller for demonstration perplexity evaluation, kL/τk L / \tau tokens are evaluated to determine token perplexities in ITPC, and L/τL / \tau tokens are processed for conditioned perplexity calculation in ITPC using key-value (KV) caching.

    Assuming parameter-proportional compute scaling (e.g., Alpaca-7B vs. a 175B LLM yields csmall≈7/175 cLLMs=1/25 cLLMsc_{\text{small}} \approx 7/175 \, c_{\text{LLMs}} = 1/25 \, c_{\text{LLMs}}), at an overall compression ratio of 1/τ=51/\tau = 5 (so τ=0.2\tau = 0.2), total computational cost evaluates to:

    c≈0.264⋅LcLLMs≈14LcLLMsc \approx 0.264 \cdot L c_{\text{LLMs}} \approx \frac{1}{4} L c_{\text{LLMs}}

    This delivers an approximate 4×4\times computational resource reduction. On a single NVIDIA V100-32G GPU evaluated on GSM8K, end-to-end inference latency is reduced from 8.6 seconds (uncompressed 1x baseline) to 4.9s (1.7×1.7\times speedup at 2x compression), 2.3s (3.3×3.3\times speedup at 5x compression), and 1.3s (5.7×5.7\times speedup at 10x compression), while LLMLingua's own compression runtime accounts for 0.8s, 0.3s, and 0.2s, respectively.

  6. Knowl 6 — Empirical Benchmark Performance on Reasoning and Contextual Tasks

    data/table

    LLMLingua was evaluated across reasoning benchmarks (GSM8K and BIG-bench Hard / BBH using Exact Match EM) and contextual understanding benchmarks (ShareGPT conversation and Arxiv-March23 summarization using BLEU, ROUGE-1/2/L, and BERTScore F1) against full-shot prompts, Selective-Context, Sentence Selection, and GPT-4 generation.

    Method GSM8K BBH
    EM Tokens 1/τ1/\tau EM Tokens 1/τ1/\tau
    Full-shot 78.85 2,366 - 70.07 774 -
    1-shot constraint
    Selective-Context 53.98 452 5x 54.27 276 3x
    GPT4 Generation 71.87 496 5x 27.13 260 3x
    Ours (LLMLingua) 79.08 446 5x 70.11 288 3x
    half-shot constraint
    Sentence Selection 72.33 230 10x 39.56 175 4x
    Selective-Context 52.99 218 11x 54.02 155 5x
    GPT4 Generation 68.61 223 11x 27.09 161 5x
    Ours (LLMLingua) 77.41 171 14x 61.60 171 5x
    quarter-shot constraint
    Sentence Selection 66.67 195 12x 46.00 109 7x
    Selective-Context 44.20 157 15x 47.37 108 7x
    GPT4 Generation 56.33 188 20x 26.81 101 8x
    Ours (LLMLingua) 77.33 117 20x 56.85 110 7x

    On GSM8K, LLMLingua achieves 79.08 EM at 5x compression (exceeding the full-shot baseline of 78.85) and maintains 77.33 EM at 20x compression (a drop of only 1.52 EM points, while outperforming Selective-Context by 33.13 points). On ShareGPT at 3.3x compression, LLMLingua scores 19.55 BLEU and 87.70 BERTScore F1 (vs. Selective-Context's 15.79 BLEU and 87.12 BS F1). On Arxiv-March23 at 9x compression, LLMLingua achieves 13.45 BLEU and 89.03 BERTScore F1 (vs. Selective-Context's 12.23 BLEU and 88.16 BS F1).

  7. Knowl 7 — Component-Wise Ablation Analysis of LLMLingua

    empirical result

    Ablation experiments on GSM8K under the 1-shot constraint quantify the contribution of each module in LLMLingua:

    1. Full LLMLingua: 79.08 EM (439 tokens, 5x compression).
    2. w/o Iterative Token-level Prompt Compression (ITPC) (performing token pruning in a single pass without prefix conditioning): Exact Match drops by 6.15 points to 72.93 EM (453 tokens), indicating that ignoring conditional dependencies between preserved tokens loses essential reasoning logic and low-frequency keywords.
    3. w/o Budget Controller (applying ITPC uniformly across all prompt parts): Performance drops by 5.46 points to 73.62 EM (486 tokens), demonstrating the necessity of coarse-grained demonstration filtering.
    4. w/o Dynamic Compression Ratio (uniform budget per component): Performance drops to 77.26 EM (457 tokens), confirming that instructions and questions are more sensitive to token pruning than demonstrations.
    5. w/ Random Selection in Budget Controller (random demonstration pruning instead of perplexity-based ranking): Performance drops to 72.78 EM (477 tokens), showing that small LM perplexity effectively identifies informative demonstration examples.
    6. w/o Distribution Alignment (using raw pre-trained LLaMA-7B without Alpaca tuning): Performance drops to 78.62 EM (452 tokens), indicating a 0.46–0.56 EM gain from instruction tuning alignment.
    7. w/ Remove Stop Words (rule-based stop-word removal using NLTK): Yields 76.27 EM with 1,882 tokens (only 1.3x compression), showing simple linguistic heuristics fail to achieve high compression rates.
  8. Knowl 8 — Generalizability Across Target LLMs and Compressor Model Scales

    empirical result

    LLMLingua generalizes across distinct target black-box LLMs and small compressor model architectures:

    1. Target LLM Generalization (Claude-v1.3 on GSM8K): When tested on Anthropic's Claude-v1.3 using Alpaca-7B as the compressor, LLMLingua achieves 83.51 EM with 439 tokens (5x compression ratio) and 82.61 EM with 171 tokens (14x compression ratio), both outperforming the simple prompt baseline (81.80 EM with 691 tokens, 3x ratio).
    2. Compressor Model Scale Variation (GPT2-Alpaca 124M on GSM8K): When replacing Alpaca-7B with a fine-tuned GPT2-small (124M parameters) on GPT-3.5-Turbo, LLMLingua achieves 77.02 EM at 5x compression (447 tokens), 76.42 EM at 14x compression (173 tokens), and 76.27 EM at 18x compression (128 tokens). Compared to Alpaca-7B, performance drops by 2.06, 0.99, and 1.06 EM points, demonstrating that while larger compression models provide better distribution alignment, lightweight 124M models remain viable.
  9. Knowl 9 — Effect of Prompt Compression on Generated Output Length and Reconstructibility

    empirical result

    Experiments show two key behavioral properties of LLMs interacting with compressed prompts:

    1. Reduction in Generated Text Length: As prompt compression ratio increases (from 1x to 20x), the length of output texts generated by target LLMs systematically decreases across GSM8K, BBH, ShareGPT, and Arxiv datasets. This demonstrates that prompt compression reduces computational costs during both the prompt prefill stage and the autoregressive decoding stage.
    2. Emergent Prompt Reconstructibility: Advanced black-box LLMs (specifically GPT-4) can successfully reconstruct full, grammatically coherent original multi-step reasoning prompts from highly compressed prompts (e.g., recovering 9-step Chain-of-Thought reasoning from a 17x compressed prompt produced by Alpaca-7B). In contrast, less capable models such as GPT-3.5-Turbo fail to accurately reconstruct compressed prompts, indicating that compressed prompt decoding is an emergent capability of frontier models.
  10. Knowl 10 — Performance Degradation at Extreme Compression and Tokenizer Mismatches

    limitation

    LLMLingua exhibits two primary limitations:

    1. Performance Cliff at Extreme Compression Ratios: While LLMLingua preserves task performance up to approximately 20x compression (e.g., maintaining 77.33 EM on GSM8K), compression ratios beyond 25x–30x lead to steep performance degradation across all tasks due to severe loss of semantic and structural reasoning information.
    2. Tokenizer Discrepancy: Small compression models and black-box target LLMs frequently use different tokenizers and vocabularies (e.g., LLaMA tokenizer vs. OpenAI tiktoken). Sub-token boundary differences can cause underestimation or misalignment of actual token counts and compression rates.

Coverage note — Specific natural language prompts used to instruct GPT-4 as a compression baseline (Appendix C) and qualitative raw text outputs (Appendices D and E) were omitted as they serve as auxiliary case demonstrations rather than core methodological contributions.

References

  1. 1.
    1. Sharegpt. https://sharegpt.com/.
  2. 2.Udit Arora, William Huang, and He He. 2021. Types of out-of-distribution texts and how to detect them. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10687–10701, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  3. 3.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. 2023. Token merging: Your vit but faster. In The Eleventh International Conference on Learning Representations.
  4. 4.Harrison Chase. 2022. LangChain.
  5. 5.Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. ArXiv preprint, abs/2305.14788.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. ArXiv preprint, abs/2110.14168.
  8. 8.Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al. 2023. Language modeling is compression. ArXiv preprint, abs/2309.10668.
  9. 9.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. GPT3.int8(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems.
  10. 10.Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning.
  11. 11.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2023. OPTQ: Accurate quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations.
  12. 12.Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot. 2023a. Chain-of-thought hub: A continuous effort to measure large language models' reasoning performance. ArXiv preprint, abs/2305.17306.
  13. 13.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2023b. Complexity-based prompting for multi-step reasoning. In The Eleventh International Conference on Learning Representations.
  14. 14.Tao Ge, Jing Hu, Li Dong, Shaoguang Mao, Yan Xia, Xun Wang, Si-Qing Chen, and Furu Wei. 2022. Extensible prompts for language models. ArXiv preprint, abs/2212.00616.
  15. 15.Tao Ge, Jing Hu, Xun Wang, Si-Qing Chen, and Furu Wei. 2023. In-context autoencoder for context compression in a large language model. ArXiv preprint, abs/2307.06945.
  16. 16.Henry Gilbert, Michael Sandborn, Douglas C Schmidt, Jesse Spencer-Smith, and Jules White. 2023. Semantic compression with large language models. ArXiv preprint, abs/2304.12512.
  17. 17.Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020. Power-bert: Accelerating BERT inference via progressive word-vector elimination. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 3690–3699. PMLR.
  18. 18.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  19. 19.Gyuwan Kim and Kyunghyun Cho. 2021. Length-adaptive transformer: Train once with length drop, use anytime with search. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6501–6511, Online. Association for Computational Linguistics.
  20. 20.Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer. 2022. Learned token pruning for transformers. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 784–794.
  21. 21.Yucheng Li. 2023. Unlocking context constraints of llms: Enhancing context efficiency of llms with self-information-based content filtering. ArXiv preprint, abs/2304.12102.
  22. 22.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  23. 23.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  24. 24.Kimberly T Mai, Toby Davies, and Lewis D Griffin. 2022. Self-supervised losses for one-class textual anomaly detection. ArXiv preprint, abs/2204.05695.
  25. 25.Ali Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. 2022. AdapLeR: Speeding up inference by adaptive length reduction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1–15, Dublin, Ireland. Association for Computational Linguistics.
  26. 26.Jesse Mu, Xiang Lisa Li, and Noah Goodman. 2023. Learning to compress prompts with gist tokens. ArXiv preprint, abs/2304.08467.
  27. 27.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  28. 28.Richard Clark Pasco. 1976. Source coding algorithms for fast data compression. Ph.D. thesis, Citeseer.
  29. 29.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. 2021. Dynamicvit: Efficient vision transformers with dynamic token sparsification. In Advances in Neural Information Processing Systems.
  30. 30.Jorma J Rissanen. 1976. Generalized kraft inequality and arithmetic coding. IBM Journal of research and development, 20(3):198–203.
  31. 31.Claude E Shannon. 1951. Prediction and entropy of printed english. Bell system technical journal, 30(1):50–64.
  32. 32.Ilya Sutskever. 2023. A theory of unsupervised learning. https://simons.berkeley.edu/talks/ilya-sutskever-openai-2023-08-14.
  33. 33.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, , and Jason Wei. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. ArXiv preprint, abs/2210.09261.
  34. 34.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  35. 35.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  36. 36.David Wingate, Mohammad Shoeybi, and Taylor Sorensen. 2022. Prompt compression and contrastive conditioning for controllability and toxicity reduction in language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5621–5634, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  37. 37.Qianhui Wu, Huqiang Jiang, Haonan Yin, Börje F. Karlsson, and Chin-Yew Lin. 2023. Multi-level knowledge distillation for out-of-distribution detection in text. In Proceedings of the 61th Annual Meeting of the Association for Computational Linguistics (Long Papers).
  38. 38.Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning.
  39. 39.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. ArXiv preprint, abs/2304.12244.
  40. 40.Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023. Inference with reference: Lossless acceleration of large language models. ArXiv preprint, abs/2304.04487.
  41. 41.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 5754–5764.
  42. 42.Lei Zhang, Yuge Zhang, Kan Ren, Dongsheng Li, and Yuqing Yang. 2023. Mlcopilot: Unleashing the power of large language models in solving machine learning tasks. ArXiv preprint, abs/2304.14979.
  43. 43.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  44. 44.Wangchunshu Zhou, Yuchen Eleanor Jiang, Ryan Cotterell, and Mrinmaya Sachan. 2023. Efficient prompting via dynamic in-context learning. ArXiv preprint, abs/2305.11170.

Citation

MLA
Jiang, H., et al. “LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 13358–76, https://doi.org/10.18653/v1/2023.emnlp-main.825.
APA
Jiang, H., Wu, Q., Lin, C.-Y., Yang, Y., & Qiu, L. (2023). LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13358–13376. https://doi.org/10.18653/v1/2023.emnlp-main.825
Chicago
Jiang, H., Q. Wu, C.-Y. Lin, Y. Yang, and L. Qiu. 2023. “LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13358–76. https://doi.org/10.18653/v1/2023.emnlp-main.825.
Harvard
Jiang, H. et al. (2023) “LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13358–13376. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.825.
Vancouver
1. Jiang H, Wu Q, Lin C-Y, Yang Y, Qiu L (2023) LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13358–13376

BibTeX

@inproceedings{jiang-etal-2023-llmlingua,
    title = "{LLML}ingua: Compressing Prompts for Accelerated Inference of Large Language Models",
    author = "Jiang, Huiqiang  and
      Wu, Qianhui  and
      Lin, Chin-Yew  and
      Yang, Yuqing  and
      Qiu, Lili",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.825/",
    doi = "10.18653/v1/2023.emnlp-main.825",
    pages = "13358--13376"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/