Prompt Compression for Large Language Models: A Survey

Zongqian LiYinhong LiuYixuan SuNigel Collier

article2025NAACL86 citations

Categorizes prompt compression techniques into discrete hard prompt and continuous soft prompt approaches to help researchers reduce large language model computational costs while retaining critical context across diverse tasks.

Listen

As organizations deploy large language models for complex real-world tasks, inputs have grown increasingly lengthy to accommodate detailed instructions, reference documents, and contextual examples. These long-form prompts create substantial operational bottlenecks, notably high memory consumption, increased compute costs, and slower response latency. Prompt compression has emerged as a modular, input-focused strategy to alleviate these burdens without requiring fundamental alterations to the underlying language model parameters.

The article provides a systematic evaluation of contemporary prompt compression techniques. Its primary objective is to categorize these methods, analyze their core operational mechanisms across various downstream applications, identify existing performance trade-offs, and outline practical avenues for future system optimization.

The analysis reviews the landscape by categorizing prompt compression into two foundational paradigms: hard prompt methods and soft prompt methods. The evaluation synthesizes architectural characteristics, compression ratios, and computational overheads across representative frameworks, assessing their utility in standard tasks such as question answering, retrieval-augmented generation, and automated agent workflows.

The review establishes several core findings regarding the efficacy and mechanics of prompt compression. First, hard prompt methods operate through natural language filtering or paraphrasing to eliminate redundant words, achieving compression ratios up to 20x while remaining compatible with closed commercial application programming interfaces. Second, soft prompt methods use neural encoders to condense lengthy contexts into compact continuous vector representations, reaching extreme compression ratios ranging from 26x to as high as 480x while retaining 60% to 70% of original task performance. Third, current soft prompt techniques introduce major computational trade-offs: because their compression encoders are often as large as the primary language models, efficiency gains are primarily realized during output token generation rather than initial input processing. Fourth, compression encoders typically lack general transferability, requiring costly retraining whenever the primary model is updated.

These findings indicate that while prompt compression effectively lowers memory requirements for long-context tasks, its net operational savings depend heavily on task architecture. Workflows characterized by long inputs and brief generated answers may experience minimal latency benefits due to the initial compression overhead. Additionally, the risk of semantic information loss and degraded grammatical coherence necessitates careful validation before deploying compressed prompts in high-precision enterprise environments.

To balance performance and resource costs, stakeholders should match compression techniques to their deployment constraints, prioritizing hard prompt filtering for external commercial interfaces and soft prompt vectors for internal high-throughput pipelines. Developers should pursue hybrid architectures that combine natural language filtering with vector compression, while transitioning to smaller semantic encoders—such as models at least ten times smaller than the base system—to significantly accelerate encoding speeds.

The assessment's confidence is tempered by limitations in current benchmark evaluations, specifically the lack of direct empirical comparisons between prompt compression and established attention-optimization techniques, as well as potential data overlap in experimental benchmarks. Organizations should conduct targeted pilot tests on domain-specific workloads to verify true latency reductions and task accuracy before broad implementation.

Li et al (2025).pdf

No sufficiently relevant recommendations were found.

Cover for Prompt Compression for Large Language Models: A Survey

Abstract

Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing tasks, but the computational cost of processing long prompts remains a significant challenge. Prompt compression has emerged as a promising solution to reduce the input length while preserving essential information. In this survey, we provide a comprehensive overview of prompt compression techniques for LLMs. We categorize existing methods into two main groups: hard prompt compression, which selects or rewrites discrete tokens, and soft prompt compression, which encodes information into continuous embeddings. We review representative approaches in each category, discuss their underlying mechanisms, and compare their performance on various benchmarks. Furthermore, we analyze the trade-offs between compression ratio and task performance, and outline potential future research directions. Our survey aims to provide a systematic understanding of prompt compression and to facilitate the development of more efficient and effective LLM applications.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 3 Hard Prompt Methods
  • 4 Soft Prompt Methods
  • 4.1 Architectures
  • 4.2 Insights
  • 5 Downstream Adaptions
  • 6 Challenges and Future Work
  • 6.1 Current Challenges
  • 6.2 Future Directions
  • 7 Conclusions
  • Limitations
  • Ethics Statement
  • Availability Statement
  • References
  • A Appendix
  • Hard Prompt Methods
  • Soft Prompt Methods

Knowls

  1. Knowl 1 — Prompt compression reduces input length through discrete filtering or continuous encoding

    definition

    Prompt compression reduces the number of input tokens processed by a large language model (LLM), either by deleting or rewriting low-value natural-language content (hard prompt compression) or encoding the input as a shorter sequence of continuous vectors (soft prompt compression). The compression ratio is the number of original input tokens divided by the number of compressed prompt tokens. In the survey’s usage, prompt compression changes the input representation rather than modifying the key–value (KV) values of the original input tokens, distinguishing it from KV-cache compression. Methods are commonly assessed either at a fixed compression ratio, by comparing task performance, or at comparable performance, by comparing the compression ratios each method can achieve.

  2. Knowl 2 — Hard filtering ranks information and removes tokens while retaining natural-language inputs

    model/method

    Hard filtering produces a shorter prompt from discrete natural-language tokens. SelectiveContext estimates lexical-unit informativeness with self-information and uses SpaCy dependency parsing to group tokens into noun phrases, aiming to remove less informative phrases while maintaining coherence. It requires no external language model or additional learned parameters, but depends on accurate noun-phrase boundaries and does not provide a corresponding way to merge verb phrases. LLMLingua instead uses a smaller language model to estimate self-information: it first selects key demonstrations and then filters tokens throughout the prompt, including at subword level. It prioritizes retaining important items such as numbers and units, and the survey reports compression ratios up to 20×. Its compressor adds memory and may use a tokenizer different from the target LLM’s; selecting demonstrations is also a distinct operation from compressing prompt content. Other filtering variants include longer-window document reordering and subsequence recovery (LongLLMLingua), query-difficulty- and retrieval-quality-dependent document selection (AdaComp), distilled-data token classification (LLMLingua-2), reinforcement-learning-based token selection (PCRL and task-specific TACO-RL), and embedding-based relevance or semantic compression (CPC and TCRA-LLM).

  3. Knowl 3 — Hard paraphrasing compresses prompts by generating concise, fluent text

    model/method

    Nano-Capsulator compresses a prompt by having an independent, fine-tuned Vicuna-7B compressor produce a concise natural-language paraphrase that omits irrelevant information. Unlike ordinary summarization, its training includes a semantic-preservation loss to retain task-relevant meaning and a reward function that optimizes the compressed prompt’s utility to the downstream LLM. This targets both meaning retention and readable output, rather than merely deleting tokens. The trade-off is that the compressor has non-negligible memory requirements and must run as an additional generation step, making compression more computationally involved than simple token filtering.

  4. Knowl 4 — Decoder-only soft compression ranges from prompt-specific conditioning to reusable compressed contexts

    model/method

    Decoder-only soft prompt methods use continuous representations rather than a rewritten natural-language prompt. Contrastive Conditioning (CC) learns a short soft prompt to approximate the output distribution conditioned on a natural-language prompt, minimizing Kullback–Leibler divergence over generated token sequences; it estimates these distributions through repeated sampling. CC can steer attributes such as sentiment, but each soft prompt is trained for its particular natural-language prompt, so a new prompt requires training again. GIST appends learned compression tokens to an input: those tokens can attend to the input, while generated tokens attend only to the compression tokens. A fine-tuned LLM acts as both encoder and decoder, and GIST is reported to reach up to 26× compression on prompts within the length range represented in its training data. Its learned compression tokens are not directly usable with the original untuned LLM. AutoCompressor extends recursive compression to long contexts: it divides the input into sub-prompts, compresses each into continuous vectors, and passes those vectors along with the next sub-prompt for further compression. It can handle inputs up to 30,720 tokens, but its training is time-consuming and its compression tokens likewise require a fine-tuned LLM.

  5. Knowl 5 — ICAE and 500xCompressor trade compressed representation size against compression ratio

    model/method

    In-context Autoencoder (ICAE) compresses information-rich context while leaving the question uncompressed; a frozen LLM decoder answers using both. It compresses up to 512 input tokens into 32, 64, or 128 continuous vectors, corresponding to compression ratios from 16× down to 4×, and concatenating encoded groups allows inputs up to 5,120 tokens. Because the decoder is frozen, its compressed vectors can be used with the original LLM. The survey notes that ICAE’s training and evaluation on the Pile raise possible data-overlap and answer-retrieval concerns. 500xCompressor also uses a trainable encoder with a frozen LLM decoder, but supplies the decoder with the compression tokens’ KV pairs rather than the vectors themselves. It compresses up to 480 tokens into 1–16 tokens, with reported ratios from 6× to 480×, while retaining more than 60–70% of the uncompressed prompt capabilities. Its ArxivQA test set was generated from abstracts published January–April 2024, providing an evaluation on data described as strictly unseen. Although 500xCompressor uses KV pairs as its compressed representation, it does not modify those pairs and is therefore classified by the survey as prompt compression, not KV compression.

  6. Knowl 6 — Encoder–decoder designs use embedding models or projectors for specialized compression

    model/method

    Several soft compression methods train only a small component between an encoder and an LLM decoder. xRAG uses a frozen embedding model and a trainable adapter (projector) to encode retrieved text for question answering; it demonstrates that a single token can represent information from a document for this task. The survey identifies a deployment cost: its selected SFR-Embedding-Mistral encoder is itself LLM-based, so xRAG requires both that model and the decoder, unlike methods that use one LLM with LoRA parameters. UniICL compresses in-context demonstrations while leaving the query intact. Its encoder and decoder are the same frozen LLM, and only an intervening projector is trainable; its compressed vectors represent demonstrations and can avoid a separate embedding step during example selection. Related architectures include COCOM and LLoCO, which fine-tune encoder–decoder setups for retrieval-augmented generation (RAG), and QGC, which uses a frozen decoder.

  7. Knowl 7 — Soft prompts can be interpreted as attention compression, PEFT-like inputs, a modality, or an LLM-oriented language

    model/method

    The survey offers four complementary interpretations of soft prompt compression. First, it acts as attention optimization: compression tokens read the original input, then generated tokens attend to the shorter encoded representation rather than the full prompt. This differs from conventional sliding-window or sparse attention, which changes attention patterns while retaining the input tokens. Second, it resembles parameter-efficient fine-tuning (PEFT): ICAE-like methods generate prompt-tuning-style embeddings from each input, while 500xCompressor supplies KV pairs more analogous to prefix tuning. Third, continuous compressed text can be viewed as a new modality, analogous to image embeddings supplied to an LLM; the survey cautions that text is information-dense and therefore may be more sensitive to information loss than images. Fourth, the vectors can be viewed as a synthetic language for LLMs: they encode information, can be transmitted or saved, and allow the receiving model to condition its output on the encoded content. These are conceptual lenses proposed by the survey, not claims that the methods are equivalent to the corresponding established techniques.

  8. Knowl 8 — Prompt compression is applied to question answering, retrieval, demonstrations, agents, and specialized tasks

    model/method

    The survey groups applications of prompt compression across several task settings. In general question answering and instruction following, methods compress instructions or context before generating an answer. In RAG, compressed representations of retrieved documents can replace passing the complete retrieved text to the answering LLM; xRAG is an example. For in-context learning, UniICL compresses demonstrations while preserving the query, reducing the demonstration representation and supporting example use or selection. Agent applications include compressing API documentation to make tool use more efficient. The survey also identifies domain-specific and other applications, including specialized-domain tasks, role playing, and function-related tasks.

  9. Knowl 9 — Current methods face generalization, compressor-cost, and evaluation-comparison limitations

    limitation

    The survey identifies three broad limitations. First, soft compression can inherit fine-tuning problems such as forgetting, overfitting, and model drift; training robust encoders may require large and diverse datasets, and an encoder tied to one decoder may need retraining when that LLM is updated. Filtered hard prompts can also be less grammatical and unfamiliar to the target LLM. Second, compression itself can be costly: soft methods may add a large encoder and substantial processing time, while hard methods may require another model and the target LLM must still process the retained prompt. Thus, much of the efficiency benefit may accrue during output generation, and can be modest for tasks with short outputs. Third, the survey reports that prompt compression has not been adequately compared with attention-optimization methods such as sliding-window and sparse attention. Those alternatives avoid a separate encoder and use a consistent attention mechanism for input and generated tokens, so meaningful evaluation should compare compression and attention methods at equivalent computational efficiency.

  10. Knowl 10 — Future work should make encoders cheaper, combine compression types, and explore multimodal designs

    model/method

    The survey proposes three directions. Encoder optimization could reduce compression time by using smaller encoders that still capture useful semantics—for example, models such as BERT, described as at least ten times smaller than LLMs—or by investigating PEFT methods beyond LoRA, including QLoRA, DoRA, and MoRA. Combining hard filtering with soft encoding could increase compression because the two methods remove redundancy in different ways, although their sequential application would add compression latency. Finally, multimodal architectures may provide new designs: current soft prompt methods use self-attention to transfer input information to compression tokens, whereas cross-attention remains unexplored in this setting. The survey also suggests adapting image–text embedding-alignment approaches to train prompt compression models.

Coverage note — The survey’s catalog-only entries for individual RAG, agent, domain-specific, and other application methods are not each expanded into separate knowls because it provides limited technical detail about many of them; their application roles and the main architectural families are included.

References

  1. 1.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150.
  2. 2.Zhiwei Cao, Qian Cao, Yu Lu, Ningxin Peng, Luyang Huang, Shanbo Cheng, and Jinsong Su. 2024. Retaining key information under high compression ratios: Query-guided compressor for LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12685–12695, Bangkok, Thailand. Association for Computational Linguistics.
  3. 3.Kaiyan Chang, Songcheng Xu, Chenglong Wang, Yingfeng Luo, Tong Xiao, and Jingbo Zhu. 2024. Efficient prompting methods for large language models: A survey. Preprint, arXiv:2404.01077.
  4. 4.Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, and Dongyan Zhao. 2024. xrag: Extreme context compression for retrieval-augmented generation with one token. Preprint, arXiv:2405.13792.
  5. 5.Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. In The 2023 Conference on Empirical Methods in Natural Language Processing.
  6. 6.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509.
  7. 7.Yu-Neng Chuang, Tianwei Xing, Chia-Yuan Chang, Zirui Liu, Xun Chen, and Xia Hu. 2024. Learning to compress prompt in natural language formats. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 7756–7767, Mexico City, Mexico. Association for Computational Linguistics.
  8. 8.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. Gpt3.int8(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems, volume 35, pages 30318–30332. Curran Associates, Inc.
  9. 9.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. In Advances in Neural Information Processing Systems, volume 36, pages 10088–10115. Curran Associates, Inc.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, ..., and Zhiwei Zhao. 2024. The llama 3 herd of models. Preprint, arXiv:2407.21783.
  12. 12.Jun Gao, Ziqiang Cao, and Wenjie Li. 2024. Unifying demonstration selection and compression for in-context learning. Preprint, arXiv:2405.17062.
  13. 13.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020a. The pile: An 800gb dataset of diverse text for language modeling. Preprint, arXiv:2101.00027.
  14. 14.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020b. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  15. 15.Tao Ge, Hu Jing, Li Dong, Shaoguang Mao, Yan Xia, Xun Wang, Si-Qing Chen, and Furu Wei. 2023. Extensible prompts for language models on zero-shot language style customization. In Advances in Neural Information Processing Systems, volume 36, pages 35576–35591. Curran Associates, Inc.
  16. 16.Tao Ge, Hu Jing, Lei Wang, Xun Wang, Si-Qing Chen, and Furu Wei. 2024. In-context autoencoder for context compression in a large language model. In The Twelfth International Conference on Learning Representations.
  17. 17.Shangmin Guo, Yi Ren, Agnieszka Słowik, and Kory Mathewson. 2020. Inductive bias and language expressivity in emergent communication. 4th NeurIPS Workshop on Emergent Communication.
  18. 18.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  19. 19.John L Hennessy and David A Patterson. 2011. Computer architecture: a quantitative approach. Elsevier.
  20. 20.Siddharth Jha, Lutfi Eren Erdogan, Sehoon Kim, Kurt Keutzer, and Amir Gholami. 2024. Characterizing prompt compression methods for long context inference. In Workshop on Efficient Systems for Foundation Models II @ ICML2024.
  21. 21.Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023. LLMLingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13358–13376, Singapore. Association for Computational Linguistics.
  22. 22.Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2024a. LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1658–1677, Bangkok, Thailand. Association for Computational Linguistics.
  23. 23.Ting Jiang, Shaohan Huang, Shengyue Luo, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, et al. 2024b. Mora: High-rank updating for parameter-efficient fine-tuning. arXiv preprint arXiv:2405.12130.
  24. 24.Yichen Jiang, Marco Vecchio, Mohit Bansal, and Anders Johannsen. 2024c. Hierarchical and dynamic prompt compression for efficient zero-shot API usage. In Findings of the Association for Computational Linguistics: EACL 2024, pages 2162–2174, St. Julian’s, Malta. Association for Computational Linguistics.
  25. 25.Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, Yabiao Wang, Chengjie Wang, and Lizhuang Ma. 2024. Efficient multimodal large language models: A survey. Preprint, arXiv:2405.10739.
  26. 26.Dongwon Jung, Qin Liu, Tenghao Huang, Ben Zhou, and Muhao Chen. 2024. Familiarity-aware evidence compression for retrieval augmented generation. Preprint, arXiv:2409.12468.
  27. 27.Hoyoun Jung and Kyung-Joong Kim. 2024. Discrete prompt compression with reinforcement learning. IEEE Access, 12:72578–72587.
  28. 28.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. RACE: Large-scale ReAding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785–794, Copenhagen, Denmark. Association for Computational Linguistics.
  29. 29.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  30. 30.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474. Curran Associates, Inc.
  31. 31.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  32. 32.Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. 2023. Compressing context to enhance inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6342–6353, Singapore. Association for Computational Linguistics.
  33. 33.Zongqian Li, Yixuan Su, and Nigel Collier. 2024. 500xcompressor: Generalized prompt compression for large language models. Preprint, arXiv:2408.03094.
  34. 34.Barys Liskavets, Maxim Ushakov, Shuvendu Roy, Mark Klibanov, Ali Etemad, and Shane Luke. 2024. Prompt compression with context-aware sentence encoding for fast and improved llm inference. Preprint, arXiv:2409.01227.
  35. 35.Junyi Liu, Liangzhi Li, Tong Xiang, Bowen Wang, and Yiming Qian. 2023a. TCRA-LLM: Token compression retrieval augmented large language model for inference cost reduction. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9796–9810, Singapore. Association for Computational Linguistics.
  36. 36.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023b. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv., 55(9).
  37. 37.Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. Llm-pruner: On the structural pruning of large language models. In Advances in Neural Information Processing Systems, volume 36, pages 21702–21720. Curran Associates, Inc.
  38. 38.Paloma Ubeda Mansilla. 2004. Foundations of language (brain, meaning, grammar, evolution), by ray jackendoff. Ibérica: Revista de la Asociación Europea de Lenguas para Fines Específicos (AELFE), (7):150–152.
  39. 39.Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. 2024. Sfr-embedding-mistral:enhance text retrieval with transfer learning. Salesforce AI Research Blog.
  40. 40.Jesse Mu, Xiang Lisa Li, and Noah Goodman. 2024. Learning to compress prompts with gist tokens. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Curran Associates Inc.
  41. 41.Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. 2024. LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression. In Findings of the Association for Computational Linguistics ACL 2024, pages 963–981, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  42. 42.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  43. 43.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  44. 44.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  45. 45.David Rau, Shuai Wang, Hervé Déjean, and Stéphane Clinchant. 2024. Context embeddings for efficient answer generation in rag. Preprint, arXiv:2407.09252.
  46. 46.Hongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang, Jure Leskovec, Dale Schuurmans, and Bo Dai. 2021. Combiner: Full attention transformer with sparse computation cost. In Advances in Neural Information Processing Systems, volume 34, pages 22470–22482. Curran Associates, Inc.
  47. 47.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, CHI EA ’21, New York, NY, USA. Association for Computing Machinery.
  48. 48.Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, Pranav Sandeep Dulepet, Saurav Vidyadhara, Dayeon Ki, Sweta Agrawal, Chau Pham, Gerson Kroiz, Feileen Li, Hudson Tao, Ashay Srivastava, Hevander Da Costa, Saloni Gupta, Megan L. Rogers, Inna Goncearenco, Giuseppe Sarli, Igor Galynker, Denis Peskoff, Marine Carpuat, Jules White, Shyamal Anadkat, Alexander Hoyle, and Philip Resnik. 2024. The prompt report: A systematic survey of prompting techniques. Preprint, arXiv:2406.06608.
  49. 49.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  50. 50.Shivam Shandilya, Menglin Xia, Supriyo Ghosh, Huiqiang Jiang, Jue Zhang, Qianhui Wu, and Victor Rühle. 2024. Taco-rl: Task aware prompt compression optimization with reinforcement learning. Preprint, arXiv:2409.13035.
  51. 51.Junhong Shen, Neil Tenenholtz, James Brian Hall, David Alvarez-Melis, and Nicolo Fusi. 2024. Tag-llm: Repurposing general-purpose llms for specialized domains. arXiv preprint arXiv:2402.05140.
  52. 52.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  53. 53.Sijun Tan, Xiuyu Li, Shishir Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph E. Gonzalez, and Raluca Ada Popa. 2024. Lloco: Learning long contexts offline. Preprint, arXiv:2404.07979.
  54. 54.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html, 3(6):7.
  55. 55.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022. Efficient transformers: A survey. ACM Comput. Surv., 55(6).
  56. 56.Ryan Teehan, Brenden Lake, and Mengye Ren. 2024. CoLLEGe: Concept embedding generation for large language models. In First Conference on Language Modeling.
  57. 57.Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau. 2024. Function vectors in large language models. In The Twelfth International Conference on Learning Representations.
  58. 58.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  59. 59.Shubham Vatsal and Harsh Dubey. 2024. A survey of prompt engineering methods in large language models for different nlp tasks. Preprint, arXiv:2407.12994.
  60. 60.Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Jiachen Liu, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, and Mi Zhang. 2024. Efficient large language models: A survey. Transactions on Machine Learning Research. Survey Certification.
  61. 61.David Wingate, Mohammad Shoeybi, and Taylor Sorensen. 2022. Prompt compression and contrastive conditioning for controllability and toxicity reduction in language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5621–5634, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  62. 62.Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024a. RECOMP: Improving retrieval-augmented LMs with context compression and selective augmentation. In The Twelfth International Conference on Learning Representations.
  63. 63.Yang Xu, Yunlong Feng, Honglin Mu, Yutai Hou, Yitong Li, Xinghao Wang, Wanjun Zhong, Zhongyang Li, Dandan Tu, Qingfu Zhu, Min Zhang, and Wanxiang Che. 2024b. Concise and precise context compression for tool-using language models. Preprint, arXiv:2407.02043.
  64. 64.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  65. 65.Shih yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. 2024. DoRA: Weight-decomposed low-rank adaptation. In Forty-first International Conference on Machine Learning.
  66. 66.Chanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong, and Jaewoo Kang. 2024. Compact: Compressing retrieved documents actively for question answering. Preprint, arXiv:2407.09014.
  67. 67.Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024a. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8):5625–5644.
  68. 68.Qianchi Zhang, Hainan Zhang, Liang Pang, Hongwei Zheng, and Zhiming Zheng. 2024b. Adacomp: Extractive context compression with adaptive predictor for retrieval-augmented large language models. Preprint, arXiv:2409.01579.
  69. 69.Wenbo Zhao, Arpit Gupta, Tagyoung Chung, and Jing Huang. 2023. SPC: Soft prompt construction for cross domain generalization. In Proceedings of the 8th Workshop on Representation Learning for NLP (RepL4NLP 2023), pages 118–130, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Li, Z., et al. “Prompt Compression for Large Language Models: A Survey”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 7182–95, https://doi.org/10.18653/v1/2025.naacl-long.368.
APA
Li, Z., Liu, Y., Su, Y., & Collier, N. (2025). Prompt Compression for Large Language Models: A Survey. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7182–7195. https://doi.org/10.18653/v1/2025.naacl-long.368
Chicago
Li, Z., Y. Liu, Y. Su, and N. Collier. 2025. “Prompt Compression for Large Language Models: A Survey”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7182–95. https://doi.org/10.18653/v1/2025.naacl-long.368.
Harvard
Li, Z. et al. (2025) “Prompt Compression for Large Language Models: A Survey”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7182–7195. Available at: https://doi.org/10.18653/v1/2025.naacl-long.368.
Vancouver
1. Li Z, Liu Y, Su Y, Collier N (2025) Prompt Compression for Large Language Models: A Survey. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 7182–7195

BibTeX

@inproceedings{li-etal-2025-prompt,
    title = "Prompt Compression for Large Language Models: A Survey",
    author = "Li, Zongqian  and
      Liu, Yinhong  and
      Su, Yixuan  and
      Collier, Nigel",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.368/",
    doi = "10.18653/v1/2025.naacl-long.368",
    pages = "7182--7195",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/