Built independently by an author, for readers. Read the story and support ChapterPal

keyword

token-level prompt compression

Token-level prompt compression is a natural language processing technique that reduces the length of an input prompt for large language models by identifying and removing redundant or low-information individual tokens while preserving the original semantic meaning. Unlike coarse-grained approaches that eliminate entire sentences or generate abstractive summaries, this method operates at the granularity of individual words or sub-word tokens, evaluating their relative importance using metrics such as perplexity, information entropy, or attention scores. By filtering out non-essential tokens before the prompt is processed by a target model, token-level prompt compression lowers computational overhead, accelerates inference speed, reduces operational costs, and facilitates the handling of long contexts with minimal loss in task performance.

1 item

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, Lili Qiu

OrganizationsMicrosoft

Why you should read this

Proposes LLMLingua, a coarse-to-fine prompt compression framework that uses smaller language models to reduce prompt length by up to 20x while preserving key reasoning and context capabilities across black-box large language models.

Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens. To accelerate model inference and reduce cost, this paper presents LLMLingua, a coarse-to-fine prompt compression method that involves a budget controller to maintain semantic integrity under high compression ratios, a token-level iterative compression algorithm to better model the interdependence between compressed contents, and an instruction tuning based method for distribution alignment between language models. We conduct experiments and analysis over four datasets from different scenarios, i.e., GSM8K, BBH, ShareGPT, and Arxiv-March23; showing that the proposed approach yields state-of-the-art performance and allows for up to 20x compression with little performance loss.¹

Added

2026-09-28