Prompt Compression for Large Language Models: A Survey
Zongqian LiYinhong LiuYixuan SuNigel Collier
Categorizes prompt compression techniques into discrete hard prompt and continuous soft prompt approaches to help researchers reduce large language model computational costs while retaining critical context across diverse tasks.
As organizations deploy large language models for complex real-world tasks, inputs have grown increasingly lengthy to accommodate detailed instructions, reference documents, and contextual examples. These long-form prompts create substantial operational bottlenecks, notably high memory consumption, increased compute costs, and slower response latency. Prompt compression has emerged as a modular, input-focused strategy to alleviate these burdens without requiring fundamental alterations to the underlying language model parameters.
The article provides a systematic evaluation of contemporary prompt compression techniques. Its primary objective is to categorize these methods, analyze their core operational mechanisms across various downstream applications, identify existing performance trade-offs, and outline practical avenues for future system optimization.
The analysis reviews the landscape by categorizing prompt compression into two foundational paradigms: hard prompt methods and soft prompt methods. The evaluation synthesizes architectural characteristics, compression ratios, and computational overheads across representative frameworks, assessing their utility in standard tasks such as question answering, retrieval-augmented generation, and automated agent workflows.
The review establishes several core findings regarding the efficacy and mechanics of prompt compression. First, hard prompt methods operate through natural language filtering or paraphrasing to eliminate redundant words, achieving compression ratios up to 20x while remaining compatible with closed commercial application programming interfaces. Second, soft prompt methods use neural encoders to condense lengthy contexts into compact continuous vector representations, reaching extreme compression ratios ranging from 26x to as high as 480x while retaining 60% to 70% of original task performance. Third, current soft prompt techniques introduce major computational trade-offs: because their compression encoders are often as large as the primary language models, efficiency gains are primarily realized during output token generation rather than initial input processing. Fourth, compression encoders typically lack general transferability, requiring costly retraining whenever the primary model is updated.
These findings indicate that while prompt compression effectively lowers memory requirements for long-context tasks, its net operational savings depend heavily on task architecture. Workflows characterized by long inputs and brief generated answers may experience minimal latency benefits due to the initial compression overhead. Additionally, the risk of semantic information loss and degraded grammatical coherence necessitates careful validation before deploying compressed prompts in high-precision enterprise environments.
To balance performance and resource costs, stakeholders should match compression techniques to their deployment constraints, prioritizing hard prompt filtering for external commercial interfaces and soft prompt vectors for internal high-throughput pipelines. Developers should pursue hybrid architectures that combine natural language filtering with vector compression, while transitioning to smaller semantic encoders—such as models at least ten times smaller than the base system—to significantly accelerate encoding speeds.
The assessment's confidence is tempered by limitations in current benchmark evaluations, specifically the lack of direct empirical comparisons between prompt compression and established attention-optimization techniques, as well as potential data overlap in experimental benchmarks. Organizations should conduct targeted pilot tests on domain-specific workloads to verify true latency reductions and task accuracy before broad implementation.
- Paper: LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Huiqiang Jiang et al. (2023). Read this foundational hard-prompt compression method first to understand the token-pruning and budget-control approach that the survey uses to frame the field.
- Paper: Compressing Context to Enhance Inference Efficiency of Large Language Models, Yucheng Li et al. (2023). Its Selective Context method provides a key example of information-based text pruning, grounding the survey’s discussion of hard compression and its efficiency trade-offs.
- Paper: LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression, Huiqiang Jiang et al. (2024). This follow-up to LLMLingua adds question-aware, coarse-to-fine compression, making it useful for understanding the survey’s account of later hard-prompt methods.
- Paper: Adapting Language Models to Compress Contexts, Alexis Chevalier et al. (2023). Its AutoCompressors introduce learned summary vectors for long contexts, a foundational soft-compression approach needed to follow the survey’s contrast with text filtering.
No sufficiently relevant recommendations were found.
