Learning to Compress Prompt in Natural Language Formats
Yu-Neng ChuangTianwei XingChia-Yuan ChangZirui LiuXun ChenXia Ben Hu
Introduces the Nano-Capsulator framework to compress long prompts into concise natural language text using semantics-preserving loss and reward-guided length constraints, cutting prompt length by over 80% while ensuring direct transferability across diverse black-box language models.
Large language models face severe computational bottlenecks when handling long input contexts, resulting in slow inference speeds, high financial costs, and memory constraints during deployment. While prior methods address these challenges by condensing prompts into specialized internal vector formats known as soft prompts, these vector solutions cannot transfer across different model architectures and cannot be used with commercial, closed application programming interface (API) models such as Claude or PaLM. The article addresses this operational bottleneck by investigating whether lengthy prompts can be compressed directly into concise, plain natural language while retaining their core utility and broad transferability across different language models.
To accomplish this, the article introduces and evaluates a framework called the Natural Language Prompt Encapsulation (Nano-Capsulator). The system trains an open-source model (Vicuna-7B) to summarize input prompts into compact natural language texts, referred to as Capsule Prompts. The approach uses an unsupervised semantic loss to preserve core logical meaning alongside a reward mechanism that penalizes outputs exceeding length limits or degrading downstream task accuracy. Credibility was established through evaluations across two main tasks—reasoning using few-shot chain-of-thought demonstrations and reading comprehension over lengthy passages—utilizing four standard benchmark datasets (CSQA, GSM8K, MultiRC, and TriviaQA-Long) tested across diverse models including Vicuna-13B, PaLM, and Claude2.
The findings show substantial operational gains with minimal impact on accuracy. The framework achieved up to an 81.4% reduction in prompt token length, which decreased cloud API operational costs by up to 80.1% across multiple benchmarks. Furthermore, the condensed prompts decreased inference latency by up to 4.5 times and allowed for larger processing batch sizes without running out of GPU memory. The resulting natural language prompts maintained nearly identical accuracy to original uncompressed prompts across downstream models and transferred effectively to unseen datasets within similar task domains without requiring model retraining. The framework also consistently outperformed zero-shot summarization, generic language model summarization, and direct word-dropping techniques.
These results demonstrate that natural language prompt compression can drastically cut inference expenses and infrastructure latency while maintaining high accuracy. Because the compressed outputs remain standard text, organizations can apply one lightweight compression step upstream and flexibly route requests across diverse commercial API providers and open-source models without vendor lock-in. When implementing this technology, organizations should consider adopting natural language encapsulation for high-volume, cost-sensitive text processing workflows, though engineering teams must tune task-specific target length constraints, as overly aggressive truncation or excessive length can introduce noise or degrade reasoning logic.
Confidence in these findings is supported by consistent performance across diverse reasoning benchmarks and distinct language models. However, the evaluation was primarily conducted within reasoning and reading comprehension tasks using prompt lengths capped around 2,000 tokens. Further validation and pilot testing are recommended for broader enterprise domains and extremely long-document applications before deploying the framework into mission-critical production systems.
- Paper: LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models, Huiqiang Jiang et al. (2023). LLMLingua establishes an earlier prompt-compression approach for reducing inference cost, providing the direct baseline and context for the source’s natural-language alternative.
- Paper: Compressing Context to Enhance Inference Efficiency of Large Language Models, Yucheng Li et al. (2023). Selective Context shows how pruning low-information text can preserve utility while reducing context length, clarifying the source’s shift from deletion-based compression to compact natural-language capsules.
- Paper: Prompt Compression for Large Language Models: A Survey, Zongqian Li et al. (2025). This later survey organizes prompt-compression methods, including natural-language approaches, into a broader taxonomy that situates the source’s findings within the field’s subsequent development.
