LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

Huiqiang JiangQianhui WuXufang LuoDongsheng LiChin-Yew LinYuqing YangLili Qiu

article2024ACL399 citationsBest Paper Award

Proposes a question-aware prompt compression method that improves large language model performance by up to 21.4% while reducing token counts by up to four times and mitigating position bias in long-context processing.

Listen

Modern artificial intelligence applications increasingly rely on long prompts containing thousands of words to support tasks such as document search, complex reasoning, and multi-turn interactions. However, feeding lengthy contexts into large language models introduces three major operational hurdles: high computational and financial costs, slower response times, and degraded accuracy. Models frequently get distracted by irrelevant details or suffer from position bias, also known as the "lost in the middle" phenomenon, where critical data placed in the middle of a prompt is overlooked.

The article demonstrates and evaluates LongLLMLingua, a prompt compression framework designed to improve how large language models perceive key information. The primary objective is to simultaneously reduce prompt token size, cut costs, speed up processing latency, and increase answering accuracy across various long-context tasks.

To achieve this, the authors developed a multi-stage approach using a smaller, cost-effective language model to filter and structure prompts before passing them to target models like GPT-3.5-Turbo and LongChat-13B. The technique applies coarse-to-fine compression by evaluating document relevance based on the specific question, reordering documents so key facts avoid the model's middle blind spot, dynamically allocating compression budgets, and using a post-processing algorithm to reconstruct any incomplete entities in the generated answer. This framework was tested across five standardized benchmarks covering multi-document question answering, code completion, multi-hop reasoning, and long-dependency analysis.

The experimental findings demonstrate substantial improvements across cost, speed, and accuracy metrics. LongLLMLingua compressed prompts by 2x to 6x while matching or exceeding the accuracy of uncompressed inputs. In multi-document question answering, the compressed prompts improved model accuracy by up to 21.4% with four times fewer tokens, overcoming the drop in performance typically caused by misplaced context. Economically, prompt compression reduced operational inference costs by 52.6% to 94.0% across the evaluated datasets, while accelerating end-to-end processing speeds by 1.4x to 2.6x.

These results show that smaller contexts with higher key-information density yield more reliable outputs than simply maximizing context length. For enterprise deployment, adopting prompt compression can lower API operating expenses, mitigate middle-context hallucination risks, and improve user turnaround times without sacrificing task performance.

Stakeholders and engineering teams deploying long-context workflows should consider integrating question-aware compression pipelines prior to querying primary language models. Next steps should focus on extending the technique from question-specific compression to broader task-aware frameworks, enabling prompt caching to lower computational overhead. Readers should note that because compression is tailored to individual questions, contexts cannot currently be pre-cached across different queries, and highly convoluted multi-hop dependencies may introduce marginal loss during initial document filtering.

arXiv: 2310.06839microsoft/LLMLingua
Cover for LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

Abstract

In long context scenarios, large language models (LLMs) face three main challenges: higher computational cost, performance reduction, and position bias. Research indicates that LLM performance hinges on the density and position of key information in the input prompt. Inspired by these findings, we propose LongLLMLingua for prompt compression towards improving LLMs' perception of the key information to simultaneously address the three challenges. Our extensive evaluation across various long context scenarios demonstrates that LongLLMLingua not only enhances performance but also significantly reduces costs and latency. For instance, in the NaturalQuestions benchmark, LongLLMLingua boosts performance by up to 21.4% with around 4x fewer tokens in GPT-3.5-Turbo, leading to substantial cost savings. It achieves a 94.0% cost reduction in the LooGLE benchmark. Moreover, when compressing prompts of about 10k tokens at ratios of 2x-6x, LongLLMLingua can accelerate end-to-end latency by 1.4x-2.6x.

Table of Contents

  • 1 Introduction
  • 2 Problem Formulation
  • 3 Preliminary: LLMLingua
  • 4 LongLLMLingua
  • 4.1 How to improve key information density in the prompt?
  • 4.2 How to reduce information loss in the middle?
  • 4.3 How to achieve adaptive granular control during compression?
  • 4.4 How to improve the integrity of key information?
  • 5 Experiments
  • 6 Related Works
  • 7 Conclusion
  • Limitation
  • References
  • A Derivation Of Question-Aware Fine-Grained Compression
  • B Experiment Details
  • B.1 Dataset Details
  • B.2 Other Implementation Details
  • C Additional Experimental Results
  • C.1 Empirical Study of Question-aware Fine-grained Compression
  • C.2 Ablation in LongBench
  • C.3 LongBench Using LongChat-13b-16k
  • C.4 ZeroSCROLLS
  • C.5 MuSiQue
  • C.6 LooGLE
  • D Economic Cost
  • E Ablation Analysis
  • F Cases Study

Knowls

  1. Knowl 1 — LongLLMLingua Prompt Compression Framework

    model/method

    LongLLMLingua is a prompt compression method designed for long-context scenarios in large language models (LLMs). Given an input prompt x=(xins,x1doc,…,xKdoc,xque)x = (x^{\text{ins}}, x_1^{\text{doc}}, \dots, x_K^{\text{doc}}, x^{\text{que}}) consisting of instructions xinsx^{\text{ins}}, KK context documents {xkdoc}k=1K\{x_k^{\text{doc}}\}_{k=1}^K, and a query xquex^{\text{que}}, the objective is to find a compressed token subsequence x~\tilde{x} that minimizes downstream output divergence while satisfying a compression budget:

    min⁡x~Dϕ(y,y~)+λ∥x~∥0\min_{\tilde{x}} D_\phi(y, \tilde{y}) + \lambda \|\tilde{x}\|_0

    where yy and y~\tilde{y} are the text generations produced by the target LLM from xx and x~\tilde{x} respectively, DϕD_\phi is a distance metric (such as Kullback-Leibler divergence), and λ\lambda balances the compression ratio.

    To address computational cost, information density degradation from irrelevant text, and position bias ("lost-in-the-middle"), LongLLMLingua uses a smaller language model MS\mathcal{M}_S (such as LLaMA-2-7B-Chat) across four integrated steps:

    1. Question-Aware Coarse-Grained Compression: Evaluates document importance scores rkr_k using question perplexity conditioned on each document and retains the top K′K' documents.
    2. Document Reordering: Reorders the retained documents by descending importance score to place the most relevant information at the beginning of the prompt.
    3. Question-Aware Fine-Grained Compression with Dynamic Budgets: Computes token importance using contrastive perplexity (conditional pointwise mutual information) and dynamically allocates token budgets to documents based on their coarse relevance rank.
    4. Subsequence Recovery: Post-processes the target LLM's output by mapping compressed or truncated entity strings back to their original full surface forms from the source prompt.
  2. Knowl 2 — Question-Aware Coarse-Grained Document Importance Metric

    equation

    In the coarse-grained compression stage of LongLLMLingua, the importance score rkr_k of the kk-th document xkdoc={xk,idoc}i=1Nkx_k^{\text{doc}} = \{x_{k,i}^{\text{doc}}\}_{i=1}^{N_k} (with token length NkN_k) relative to the question xquex^{\text{que}} is computed using the negative log-likelihood of the question conditioned on that document:

    rk=−1Nc∑i=1Nclog⁡p(xique,restrict∣xkdoc),k∈{1,2,…,K}r_k = -\frac{1}{N_c} \sum_{i=1}^{N_c} \log p(x_i^{\text{que,restrict}} \mid x_k^{\text{doc}}), \quad k \in \{1, 2, \dots, K\}

    where xque,restrictx^{\text{que,restrict}} is the sequence formed by concatenating the question xquex^{\text{que}} with a restrictive statement xrestrictx^{\text{restrict}} (specifically, "We can get the answer to this question in the given documents."), xique,restrictx_i^{\text{que,restrict}} is the ii-th token of that concatenated sequence, and NcN_c is its total token count.

    Conditioning the question on the document rather than the document on the question (p(xkdoc∣xque)p(x_k^{\text{doc}} \mid x^{\text{que}})) prevents irrelevant background information within lengthy documents from diluting the relevance signal. The restrictive prompt serves as a regularizer that reduces hallucinations in the smaller compression model. Documents with higher rkr_k scores are retained for subsequent fine-grained compression.

  3. Knowl 3 — Contrastive Perplexity for Token-Level Question-Aware Compression

    equation

    To evaluate token importance during fine-grained compression, LongLLMLingua defines the importance score sis_i for each token xix_i in the retained documents {xkdoc}k=1K′\{x_k^{\text{doc}}\}_{k=1}^{K'} via contrastive perplexity, which measures the distribution shift induced by conditioning on the question xquex^{\text{que}}:

    si=perplexity(xi∣x<i)−perplexity(xi∣xque,x<i)s_i = \text{perplexity}(x_i \mid x_{<i}) - \text{perplexity}(x_i \mid x^{\text{que}}, x_{<i})

    where x<ix_{<i} denotes the sequence of preceding tokens.

    Expressed in terms of log-probabilities and the ground-truth distribution q(xi)q(x_i):

    si=q(xi)log⁡p(xi∣xque,x<i)p(xi∣x<i)s_i = q(x_i) \log \frac{p(x_i \mid x^{\text{que}}, x_{<i})}{p(x_i \mid x_{<i})}

    By Bayes' theorem, p(xque∣xi,x<i)=p(xque)p(xi∣xque,x<i)p(xi∣x<i)p(x^{\text{que}} \mid x_i, x_{<i}) = p(x^{\text{que}}) \frac{p(x_i \mid x^{\text{que}}, x_{<i})}{p(x_i \mid x_{<i})}. Because p(xque)p(x^{\text{que}}) and q(xi)q(x_i) are constant with respect to token selection, si∝p(xque∣xi,x<i)s_i \propto p(x^{\text{que}} \mid x_i, x_{<i}). This formulation is mathematically equivalent to conditional pointwise mutual information (PMI) and allows the compression model to identify question-sensitive tokens in a single inference pass.

  4. Knowl 4 — Dynamic Compression Budget Allocation

    equation

    Rather than enforcing a uniform token retention ratio across all retained documents, LongLLMLingua adaptively assigns document-specific compression budgets τkdoc\tau_k^{\text{doc}} guided by the coarse-grained importance scores {rk}k=1K′\{r_k\}_{k=1}^{K'}.

    Given the initial document budget τdoc\tau^{\text{doc}}, the count of retained documents K′K', the descending rank index I(rk)∈{0,1,…,K′−1}I(r_k) \in \{0, 1, \dots, K'-1\} of document kk based on its coarse score rkr_k, and a dynamic allocation budget parameter δτ\delta\tau, the retention budget τi\tau_i for each token xix_i in document xkdocx_k^{\text{doc}} is defined as:

    τi=τkdoc,∀xi∈xkdoc\tau_i = \tau_k^{\text{doc}}, \quad \forall x_i \in x_k^{\text{doc}} τkdoc=max⁡(min⁡((1−2I(rk)K′)δτ+τdoc,1),0)\tau_k^{\text{doc}} = \max\left(\min\left(\left(1 - \frac{2I(r_k)}{K'}\right)\delta\tau + \tau^{\text{doc}}, 1\right), 0\right)

    This linear scheduler allocates higher token retention budgets (lower compression ratios) to higher-ranked documents, concentrating prompt capacity on regions containing key evidence while aggressively compressing lower-ranked context.

  5. Knowl 5 — Question-Aware Document Reordering for Position Bias Mitigation

    model/method

    Large language models suffer from positional bias ("lost-in-the-middle"), exhibiting peak retrieval and comprehension accuracy when relevant information is positioned at the start or end of the input context, and deteriorating when relevant information is situated in the middle.

    Following coarse-grained scoring, LongLLMLingua reorders the retained documents {xkdoc}k=1K′\{x_k^{\text{doc}}\}_{k=1}^{K'} according to their importance scores {rk}k=1K′\{r_k\}_{k=1}^{K'}:

    (xins,x1doc,…,xK′doc,xque)→rk(xins,xr1doc,…,xrK′doc,xque)(x^{\text{ins}}, x_1^{\text{doc}}, \dots, x_{K'}^{\text{doc}}, x^{\text{que}}) \xrightarrow{r_k} (x^{\text{ins}}, x_{r_1}^{\text{doc}}, \dots, x_{r_{K'}}^{\text{doc}}, x^{\text{que}})

    where (r1,r2,…,rK′)(r_1, r_2, \dots, r_{K'}) is the permutation sorting documents by descending coarse relevance rkr_k. Placing the most question-relevant documents at the front of the prompt context aligns with LLM attention patterns and minimizes middle-context information loss.

  6. Knowl 6 — Token-Level Subsequence Recovery Algorithm

    algorithm

    Token-level compression can drop sub-word tokens belonging to entity names, numbers, or dates, causing the target LLM to output corrupted entity names. The subsequence recovery algorithm reconstructs the original, uncorrupted strings in the LLM's response using the subsequence mapping between the original prompt xx, compressed prompt x~\tilde{x}, and generated output yy.

    Input: Original prompt xx, compressed prompt x~\tilde{x}, LLM generation response yy
    Output: Recovered response token list yrecy_{\text{rec}}
    yrec←[]y_{\text{rec}} \leftarrow []
    l←0l \leftarrow 0
    while l<len(y)l < \text{len}(y) do
        if token substring yl∈x~y_l \in \tilde{x} then
            Find the longest substring y~key,l={yl,yl+1,…,yr}∈x~\tilde{y}_{\text{key},l} = \{y_l, y_{l+1}, \dots, y_r\} \in \tilde{x}
            Find the maximum common shortest subsequence xi,j={xi,xi+1,…,xj}x_{i,j} = \{x_i, x_{i+1}, \dots, x_j\} in xx corresponding to y~key,l\tilde{y}_{\text{key},l}
            Append xi,jx_{i,j} to yrecy_{\text{rec}}
            l←r+1l \leftarrow r + 1
        else
            Append token yly_l to yrecy_{\text{rec}}
            l←l+1l \leftarrow l + 1
        end if
    end while
    return yrecy_{\text{rec}}

    Subsequence and prefix search operations over xx and x~\tilde{x} are accelerated using prefix trees or sequence automata.

  7. Knowl 7 — Multi-Document QA Performance on NaturalQuestions

    empirical result

    LongLLMLingua was evaluated on the 20-document NaturalQuestions multi-document QA dataset (average prompt length of 2,946 tokens) across five ground-truth document placements (1st, 5th, 10th, 15th, and 20th position) using GPT-3.5-Turbo and LongChat-13B-16k under 2×2\times and 4×4\times compression budgets.

    Key results include:

    • Under a 4×4\times constraint using GPT-3.5-Turbo (prompt compressed to 748 tokens, a 3.9×3.9\times reduction), LongLLMLingua achieved accuracies of 75.0% (1st), 71.8% (5th), 71.2% (10th), 71.2% (15th), and 74.7% (20th), reaching 75.5% with document reordering.
    • The original uncompressed prompt scored 75.7% (1st), 57.3% (5th), 54.1% (10th), 55.4% (15th), and 63.1% (20th), demonstrating severe degradation when the answer was in the middle of the prompt.
    • At the 10th position, LongLLMLingua outperformed the original prompt by 17.1 percentage points without reordering (71.2% vs. 54.1%) and by 21.4 percentage points with reordering (75.5% vs. 54.1%).
    • Question-agnostic compression baselines failed severely under long contexts: Selective-Context achieved 24.7% accuracy at the 10th position and LLMLingua achieved 23.5%, falling below zero-shot performance (56.1%).
  8. Knowl 8 — Long-Context Benchmark Results on LongBench

    data/table

    LongLLMLingua was evaluated across the 6 task categories of the LongBench benchmark (average original prompt length 10,289 tokens) under 3,000-token (3×3\times) and 2,000-token (5–6×5\text{--}6\times) constraints using GPT-3.5-Turbo.

    Method Constraint SingleDoc MultiDoc Summ. FewShot Synth. Code Average
    Original Prompt None (10,295 tok) 39.7 38.7 26.5 67.0 37.8 54.2 44.0
    BM25 2,000 tokens 30.1 29.4 21.2 19.5 12.4 29.1 23.6
    SBERT 2,000 tokens 33.8 35.9 25.9 23.5 18.0 17.8 25.8
    OpenAI 2,000 tokens 34.3 36.3 24.7 32.4 26.3 24.8 29.8
    LongLLMLingua rkr_k 2,000 tokens 37.8 41.7 26.9 66.3 53.0 52.4 46.3
    Selective-Context 2,000 tokens 16.2 34.8 24.4 15.7 8.4 49.2 24.8
    LLMLingua 2,000 tokens 22.4 32.1 24.5 61.2 10.4 56.8 34.6
    LongLLMLingua 2,000 tokens 39.9 43.2 27.4 69.8 53.0 56.7 48.3
    LongLLMLingua 3,000 tokens 40.7 46.2 27.2 70.6 53.0 55.2 48.8
    Zero-shot 214 tokens 15.6 31.3 15.6 40.7 1.6 36.2 23.5

    At the 2,000-token constraint (compressed to 1,822 tokens, a 6×6\times reduction), LongLLMLingua attained an overall average score of 48.3, exceeding the original prompt (44.0) by 4.3 points while reducing average end-to-end latency from 15.6s to 6.1s (2.6×2.6\times speedup).

  9. Knowl 9 — Cost and Latency Reductions Across Long-Context Scenarios

    empirical result

    LongLLMLingua achieves substantial financial and system latency reductions when serving prompts to black-box LLMs such as GPT-3.5-Turbo:

    • Inference Cost Reductions (per 1,000 samples):

      • NaturalQuestions Multi-document QA: reduced from $4.6 to $1.3 (71.7% reduction).
      • LongBench: reduced from $31.5 to $3.0 (90.5% reduction).
      • ZeroSCROLLS: reduced from $30.6 to $3.2 (89.5% reduction).
      • MuSiQue (multi-hop QA): reduced from $3.8 to $1.8 (52.6% reduction).
      • LooGLE (long-dependency QA, ∼30k\sim 30\text{k} tokens): reduced from $93.6 to $5.6 (94.0% reduction).
    • Latency Speedups (measured on Tesla V100 32GB):

      • On LongBench, end-to-end inference latency decreased from 15.6s for the original prompt to 6.1s with LongLLMLingua (2.6×2.6\times speedup) at 6×6\times compression.
      • On NaturalQuestions at 4×4\times constraint, end-to-end latency decreased from 4.1s to 2.1s (2.0×2.0\times speedup).
      • Unlike embedding retrieval baselines that require multiple API roundtrips or entropy methods that perform sequential token scoring, LongLLMLingua achieves net inference acceleration despite the local small LM compression cost.
  10. Knowl 10 — Limitations of LongLLMLingua

    limitation

    The authors identify three primary limitations of LongLLMLingua:

    1. Inability to cache compressed context: Because coarse-grained ranking and token-level contrastive perplexity are conditioned on the question, context prompts must be re-compressed dynamically for each unique query, preventing static pre-compression caching.
    2. Increased compression computation: Incorporating question awareness doubles the computational workload on the small language model compared to question-agnostic LLMLingua.
    3. Sensitivity in complex multi-hop dependencies: On complex multi-hop question answering (such as MuSiQue), coarse-grained filtering conditioned solely on the final question can risk filtering out intermediate reasoning documents if their individual direct lexical or semantic alignment with the final question is low.

Coverage note — Specific per-task breakdowns for the MuSiQue and LooGLE benchmarks were integrated into the framework, cost, and general benchmark knowls rather than extracted as isolated tables to maintain high standalone significance.

References

  1. 1.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations.
  2. 2.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023. Longbench: A bilingual, multitask benchmark for long context understanding. ArXiv preprint, abs/2308.14508.
  3. 3.Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew R. Gormley. 2023. Unlimiformer: Long-range transformers with unlimited length input. In Thirty-seventh Conference on Neural Information Processing Systems.
  4. 4.Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. 2023. Token merging: Your vit but faster. In The Eleventh International Conference on Learning Representations.
  5. 5.Harrison Chase. 2022. LangChain.
  6. 6.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023. Extending context window of large language models via positional interpolation. ArXiv preprint, abs/2306.15595.
  7. 7.Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3829–3846, Singapore. Association for Computational Linguistics.
  8. 8.Kenneth Ward Church and Patrick Hanks. 1989. Word association norms, mutual information, and lexicography. In 27th Annual Meeting of the Association for Computational Linguistics, pages 76–83, Vancouver, British Columbia, Canada. Association for Computational Linguistics.
  9. 9.Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al. 2023. Language modeling is compression. ArXiv preprint, abs/2309.10668.
  10. 10.Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, and Furu Wei. 2023. Longnet: Scaling transformers to 1,000,000,000 tokens. ArXiv preprint, abs/2307.02486.
  11. 11.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2023. A survey for in-context learning. ArXiv preprint, abs/2301.00234.
  12. 12.Tao Ge, Hu Jing, Lei Wang, Xun Wang, Si-Qing Chen, and Furu Wei. 2024. In-context autoencoder for context compression in a large language model. In The Twelfth International Conference on Learning Representations.
  13. 13.Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020. Power-bert: Accelerating BERT inference via progressive word-vector elimination. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 3690–3699. PMLR.
  14. 14.Michael Günther, Jackmin Ong, Isabelle Mohr, Alaeddine Abdessalem, Tanguy Abel, Mohammad Kalim Akram, Susana Guzman, Georgios Mastrapas, Saba Sturua, Bo Wang, Maximilian Werk, Nan Wang, and Han Xiao. 2023. Jina embeddings 2: 8192-token general-purpose text embeddings for long documents. ArXiv preprint, abs/2310.19923.
  15. 15.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  16. 16.Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023a. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13358–13376. Association for Computational Linguistics.
  17. 17.Zhiying Jiang, Matthew Yang, Mikhail Tsirlin, Raphael Tang, Yiqin Dai, and Jimmy Lin. 2023b. “low-resource” text classification: A parameter-free classification method with compressors. In Findings of the Association for Computational Linguistics: ACL 2023, pages 6810–6828, Toronto, Canada. Association for Computational Linguistics.
  18. 18.Greg Kamradt. 2023. Needle In A Haystack - Pressure Testing LLMs.
  19. 19.Gyuwan Kim and Kyunghyun Cho. 2021. Length-adaptive transformer: Train once with length drop, use anytime with search. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6501–6511, Online. Association for Computational Linguistics.
  20. 20.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  21. 21.Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  22. 22.Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. 2023a. How long can open-source llms truly promise on context length?
  23. 23.Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang. 2023b. Loogle: Can long-context language models understand long contexts? ArXiv preprint, abs/2311.04939.
  24. 24.Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. 2023c. Compressing context to enhance inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6342–6353, Singapore. Association for Computational Linguistics.
  25. 25.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
  26. 26.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. MetaICL: Learning to learn in context. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2791–2809, Seattle, United States. Association for Computational Linguistics.
  27. 27.Ali Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. 2022. AdapLeR: Speeding up inference by adaptive length reduction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1–15, Dublin, Ireland. Association for Computational Linguistics.
  28. 28.Jesse Mu, Xiang Lisa Li, and Noah Goodman. 2023. Learning to compress prompts with gist tokens. In Thirty-seventh Conference on Neural Information Processing Systems.
  29. 29.Erik Nijkamp, Tian Xie, Hiroaki Hayashi, Bo Pang, Congying Xia, Chen Xing, Jesse Vig, Semih Yavuz, Philippe Laban, Ben Krause, Senthil Purushwalkam, Tong Niu, Wojciech Kryściński, Lidiya Murakhovs’ka, Prafulla Kumar Choubey, Alex Fabbri, Ye Liu, Rui Meng, Lifu Tu, Meghana Bhat, Chien-Sheng Wu, Silvio Savarese, Yingbo Zhou, Shafiq Joty, and Caiming Xiong. 2023. Xgen-7b technical report. ArXiv preprint, abs/2309.03450.
  30. 30.Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST ’23, New York, NY, USA. Association for Computing Machinery.
  31. 31.Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2024. YaRN: Efficient context window extension of large language models. In The Twelfth International Conference on Learning Representations.
  32. 32.Ofir Press, Noah A. Smith, and Mike Lewis. 2022. Train short, test long: Attention with linear biases enables input length extrapolation. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  33. 33.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  34. 34.Uri Shaham, Maor Ivgi, Avia Efrat, Jonathan Berant, and Omer Levy. 2023. ZeroSCROLLS: A zero-shot benchmark for long text understanding. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 7977–7989, Singapore. Association for Computational Linguistics.
  35. 35.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Advances in Neural Information Processing Systems, 36.
  36. 36.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pages 31210–31227. PMLR.
  37. 37.Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023. Retentive network: A successor to transformer for large language models. ArXiv preprint, abs/2307.08621.
  38. 38.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554.
  39. 39.Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłos. 2023. Focused transformer: Contrastive training for context scaling. In Thirty-seventh Conference on Neural Information Processing Systems.
  40. 40.Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. 2023a. Autogen: Enabling next-gen llm applications via multi-agent conversation framework. ArXiv preprint, abs/2308.08155.
  41. 41.Zhiyong Wu, Yaoxiang Wang, Jiacheng Ye, and Lingpeng Kong. 2023b. Self-adaptive in-context learning: An information compression perspective for in-context example selection and ordering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1423–1436, Toronto, Canada. Association for Computational Linguistics.
  42. 42.Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. C-pack: Packaged resources to advance general chinese embedding. ArXiv preprint, abs/2309.07597.
  43. 43.Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Retrieval meets long context large language models. In The Twelfth International Conference on Learning Representations.

Citation

MLA
Jiang, H., et al. “LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1658–77, https://doi.org/10.18653/v1/2024.acl-long.91.
APA
Jiang, H., Wu, Q., Luo, X., Li, D., Lin, C.-Y., Yang, Y., & Qiu, L. (2024). LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1658–1677. https://doi.org/10.18653/v1/2024.acl-long.91
Chicago
Jiang, H., Q. Wu, X. Luo, et al. 2024. “LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1658–77. https://doi.org/10.18653/v1/2024.acl-long.91.
Harvard
Jiang, H. et al. (2024) “LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1658–1677. Available at: https://doi.org/10.18653/v1/2024.acl-long.91.
Vancouver
1. Jiang H, Wu Q, Luo X, Li D, Lin C-Y, Yang Y, Qiu L (2024) LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1658–1677

BibTeX

@inproceedings{jiang-etal-2024-longllmlingua,
    title = "{L}ong{LLML}ingua: Accelerating and Enhancing {LLM}s in Long Context Scenarios via Prompt Compression",
    author = "Jiang, Huiqiang  and
      Wu, Qianhui  and
      Luo, Xufang  and
      Li, Dongsheng  and
      Lin, Chin-Yew  and
      Yang, Yuqing  and
      Qiu, Lili",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.91/",
    doi = "10.18653/v1/2024.acl-long.91",
    pages = "1658--1677"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/