keyword
prompt compression
Prompt compression is a technique in natural language processing that reduces the token count or length of input prompts given to large language models while preserving the essential semantic meaning and task performance. By filtering out redundant details, selecting high-value tokens, or summarizing context, prompt compression lowers computational resource demands, accelerates inference speed, and reduces the financial costs of querying models. Approaches to prompt compression generally fall into hard compression, which modifies discrete natural language text through selective pruning or rewriting, and soft compression, which encodes input text into dense, continuous embedding representations. This process enables models to handle extensive background information and long-context reasoning more efficiently while mitigating performance issues associated with long input sequences.
5 items

Learning to Compress Prompt in Natural Language Formats
Yu-Neng Chuang, Tianwei Xing, Chia-Yuan Chang, Zirui Liu, Xun Chen, Xia Ben Hu
Why you should read this
Introduces the Nano-Capsulator framework to compress long prompts into concise natural language text using semantics-preserving loss and reward-guided length constraints, cutting prompt length by over 80% while ensuring direct transferability across diverse black-box language models.
Large language models (LLMs) are excel at processing multiple natural language processing tasks, but their abilities are constrained by inferior performance with long context, slow inference speed, and the high cost of computing the results. Deploying LLMs with precise and informative context helps users process large-scale datasets more effectively and cost-efficiently. Existing works rely on compressing long prompt contexts into soft prompts. However, soft prompt compression encounters limitations in transferability across different LLMs, especially API-based LLMs. To this end, this work aims to compress lengthy prompts in the form of natural language with LLM transferability. This poses two challenges: (i) Natural Language (NL) prompts are incompatible with back-propagation, and (ii) NL prompts lack flexibility in imposing length constraints. In this work, we propose a Natural Language Prompt Encapsulation (Nano-Capsulator) framework compressing original prompts into NL formatted Capsule Prompt while maintaining the prompt utility and transferability. Specifically, to tackle the first challenge, the Nano-Capsulator is optimized by a reward function that interacts with the proposed semantics preserving loss. To address the second question, Nano-Capsulator is optimized by a reward function featuring length constraints. Experimental results demonstrate that the Capsule Prompt can reduce 81.4% of the original length, decrease inference latency up to 4.5×, and save 80.1% of budget overheads while providing transferability across diverse LLMs and different datasets.
Added
2026-10-03

CompAct: Compressing Retrieved Documents Actively for Question Answering
Chanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong, Jaewoo Kang
Why you should read this
Presents CompAct, an active context compression framework that dynamically integrates multi-hop evidence across retrieved documents and applies early termination, achieving up to 47x compression while boosting reader accuracy on complex question-answering benchmarks.
Retrieval-augmented generation supports language models to strengthen their factual groundings by providing external contexts. However, language models often face challenges when given extensive information, diminishing their effectiveness in solving questions. Context compression tackles this issue by filtering out irrelevant information, but current methods still struggle in realistic scenarios where crucial information cannot be captured with a single-step approach. To overcome this limitation, we introduce CompAct, a novel framework that employs an active strategy to condense extensive documents without losing key information. Our experiments demonstrate that CompAct brings significant improvements in both performance and compression rate on multi-hop question-answering benchmarks. CompAct flexibly operates as a cost-efficient plug-in module with various off-the-shelf retrievers or readers, achieving exceptionally high compression rates (47x).
Added
2026-10-03

Prompt Compression for Large Language Models: A Survey
Zongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier
Why you should read this
Categorizes prompt compression techniques into discrete hard prompt and continuous soft prompt approaches to help researchers reduce large language model computational costs while retaining critical context across diverse tasks.
Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing tasks, but the computational cost of processing long prompts remains a significant challenge. Prompt compression has emerged as a promising solution to reduce the input length while preserving essential information. In this survey, we provide a comprehensive overview of prompt compression techniques for LLMs. We categorize existing methods into two main groups: hard prompt compression, which selects or rewrites discrete tokens, and soft prompt compression, which encodes information into continuous embeddings. We review representative approaches in each category, discuss their underlying mechanisms, and compare their performance on various benchmarks. Furthermore, we analyze the trade-offs between compression ratio and task performance, and outline potential future research directions. Our survey aims to provide a systematic understanding of prompt compression and to facilitate the development of more efficient and effective LLM applications.
Added
2026-10-02

TokenSkip: Controllable Chain-of-Thought Compression in LLMs
Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, Wenjie Li
Why you should read this
Introduces TokenSkip, a method that fine-tunes language models on pruned reasoning trajectories to selectively skip low-importance tokens, cutting chain-of-thought length by up to 40% with negligible drops in accuracy.
Chain-of-Thought (CoT) has been proven effective in enhancing the reasoning capabilities of large language models (LLMs). Recent advancements, such as OpenAI’s o1 and DeepSeek-R1, suggest that scaling up the length of CoT sequences during inference could further boost LLM reasoning performance. However, due to the autoregressive nature of LLM decoding, longer CoT outputs lead to a linear increase in inference latency, adversely affecting user experience, particularly when the CoT exceeds 10,000 tokens. To address this limitation, we analyze the semantic importance of tokens within CoT outputs and reveal that their contributions to reasoning vary. Building on this insight, we propose TokenSkip, a simple yet effective approach that enables LLMs to selectively skip less important tokens, allowing for controllable CoT compression. Extensive experiments across various models and tasks demonstrate the effectiveness of TokenSkip in reducing CoT token usage while preserving strong reasoning performance. Notably, when applied to Qwen2.5-14B-Instruct, TokenSkip reduces reasoning tokens by 40% (from 313 to 181) on GSM8K, with less than a 0.4% performance drop. We release our code and checkpoints in https://github.com/hemingkx/TokenSkip.
Added
2026-09-28

LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu
Why you should read this
Proposes a question-aware prompt compression method that improves large language model performance by up to 21.4% while reducing token counts by up to four times and mitigating position bias in long-context processing.
In long context scenarios, large language models (LLMs) face three main challenges: higher computational cost, performance reduction, and position bias. Research indicates that LLM performance hinges on the density and position of key information in the input prompt. Inspired by these findings, we propose LongLLMLingua for prompt compression towards improving LLMs' perception of the key information to simultaneously address the three challenges. Our extensive evaluation across various long context scenarios demonstrates that LongLLMLingua not only enhances performance but also significantly reduces costs and latency. For instance, in the NaturalQuestions benchmark, LongLLMLingua boosts performance by up to 21.4% with around 4x fewer tokens in GPT-3.5-Turbo, leading to substantial cost savings. It achieves a 94.0% cost reduction in the LooGLE benchmark. Moreover, when compressing prompts of about 10k tokens at ratios of 2x-6x, LongLLMLingua can accelerate end-to-end latency by 1.4x-2.6x.
Added
2026-09-26
