keyword
token-level prompt compression
Token-level prompt compression is a natural language processing technique that reduces the length of an input prompt for large language models by identifying and removing redundant or low-information individual tokens while preserving the original semantic meaning. Unlike coarse-grained approaches that eliminate entire sentences or generate abstractive summaries, this method operates at the granularity of individual words or sub-word tokens, evaluating their relative importance using metrics such as perplexity, information entropy, or attention scores. By filtering out non-essential tokens before the prompt is processed by a target model, token-level prompt compression lowers computational overhead, accelerates inference speed, reduces operational costs, and facilitates the handling of long contexts with minimal loss in task performance.
1 item

