LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
Huiqiang JiangQianhui WuXufang LuoDongsheng LiChin-Yew LinYuqing YangLili Qiu
Proposes a question-aware prompt compression method that improves large language model performance by up to 21.4% while reducing token counts by up to four times and mitigating position bias in long-context processing.
Modern artificial intelligence applications increasingly rely on long prompts containing thousands of words to support tasks such as document search, complex reasoning, and multi-turn interactions. However, feeding lengthy contexts into large language models introduces three major operational hurdles: high computational and financial costs, slower response times, and degraded accuracy. Models frequently get distracted by irrelevant details or suffer from position bias, also known as the "lost in the middle" phenomenon, where critical data placed in the middle of a prompt is overlooked.
The article demonstrates and evaluates LongLLMLingua, a prompt compression framework designed to improve how large language models perceive key information. The primary objective is to simultaneously reduce prompt token size, cut costs, speed up processing latency, and increase answering accuracy across various long-context tasks.
To achieve this, the authors developed a multi-stage approach using a smaller, cost-effective language model to filter and structure prompts before passing them to target models like GPT-3.5-Turbo and LongChat-13B. The technique applies coarse-to-fine compression by evaluating document relevance based on the specific question, reordering documents so key facts avoid the model's middle blind spot, dynamically allocating compression budgets, and using a post-processing algorithm to reconstruct any incomplete entities in the generated answer. This framework was tested across five standardized benchmarks covering multi-document question answering, code completion, multi-hop reasoning, and long-dependency analysis.
The experimental findings demonstrate substantial improvements across cost, speed, and accuracy metrics. LongLLMLingua compressed prompts by 2x to 6x while matching or exceeding the accuracy of uncompressed inputs. In multi-document question answering, the compressed prompts improved model accuracy by up to 21.4% with four times fewer tokens, overcoming the drop in performance typically caused by misplaced context. Economically, prompt compression reduced operational inference costs by 52.6% to 94.0% across the evaluated datasets, while accelerating end-to-end processing speeds by 1.4x to 2.6x.
These results show that smaller contexts with higher key-information density yield more reliable outputs than simply maximizing context length. For enterprise deployment, adopting prompt compression can lower API operating expenses, mitigate middle-context hallucination risks, and improve user turnaround times without sacrificing task performance.
Stakeholders and engineering teams deploying long-context workflows should consider integrating question-aware compression pipelines prior to querying primary language models. Next steps should focus on extending the technique from question-specific compression to broader task-aware frameworks, enabling prompt caching to lower computational overhead. Readers should note that because compression is tailored to individual questions, contexts cannot currently be pre-cached across different queries, and highly convoluted multi-hop dependencies may introduce marginal loss during initial document filtering.
- Paper: Lost in the Middle: How Language Models Use Long Contexts, Nelson F. Liu et al. (2024). This seminal work establishes the 'lost in the middle' phenomenon and positional bias in long-context language models, which directly motivates LongLLMLingua's document reordering and prompt compression techniques.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). LongBench introduces the standard multi-task, bilingual evaluation benchmark for long-context understanding that provides the primary empirical testbed used in LongLLMLingua.
- Paper: Learning by Distilling Context, Charlie Snell et al. (2022). Understanding context distillation methods provides foundational context for how lengthy prompts can be compressed into smaller token budgets.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). This paper analyzes the compositionality gap and multi-hop reasoning failures in language models that LongLLMLingua aims to mitigate when filtering and structuring multi-document prompts.
- Paper: SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Yuhang Wang et al. (2026). SWE-Pruner extends task-aware prompt and context pruning methods specifically to the domain of software engineering and multi-turn coding agents.
- Paper: ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning, Yanjun Zhao et al. (2026). ReContext builds upon prompt-level context utilization challenges by using recursive attention-based evidence replay as a training-free inference scaffold for long-context reasoning.
- Paper: Recursive Language Models, Alex L. Zhang et al. (2025). Recursive Language Models extend context handling beyond simple prompt compression by treating long prompts as external programmatic objects via REPL execution.
- Paper: REFRAG: Rethinking RAG based Decoding, Xiaoqiang Lin et al. (2025). REFRAG takes context compression further in retrieval-augmented pipelines by compressing retrieved chunks into compact latent embeddings during LLM decoding.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). FiCoCo generalizes training-free token reduction and context filtering principles from textual prompt compression to multimodal LLMs.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). This survey provides a comprehensive architectural perspective on how context compression and memory management strategies integrate into broader efficient agent workflows.
