Learning to Compress Prompt in Natural Language Formats

Yu-Neng ChuangTianwei XingChia-Yuan ChangZirui LiuXun ChenXia Ben Hu

article2024NAACL50 citations

Introduces the Nano-Capsulator framework to compress long prompts into concise natural language text using semantics-preserving loss and reward-guided length constraints, cutting prompt length by over 80% while ensuring direct transferability across diverse black-box language models.

Listen

Large language models face severe computational bottlenecks when handling long input contexts, resulting in slow inference speeds, high financial costs, and memory constraints during deployment. While prior methods address these challenges by condensing prompts into specialized internal vector formats known as soft prompts, these vector solutions cannot transfer across different model architectures and cannot be used with commercial, closed application programming interface (API) models such as Claude or PaLM. The article addresses this operational bottleneck by investigating whether lengthy prompts can be compressed directly into concise, plain natural language while retaining their core utility and broad transferability across different language models.

To accomplish this, the article introduces and evaluates a framework called the Natural Language Prompt Encapsulation (Nano-Capsulator). The system trains an open-source model (Vicuna-7B) to summarize input prompts into compact natural language texts, referred to as Capsule Prompts. The approach uses an unsupervised semantic loss to preserve core logical meaning alongside a reward mechanism that penalizes outputs exceeding length limits or degrading downstream task accuracy. Credibility was established through evaluations across two main tasks—reasoning using few-shot chain-of-thought demonstrations and reading comprehension over lengthy passages—utilizing four standard benchmark datasets (CSQA, GSM8K, MultiRC, and TriviaQA-Long) tested across diverse models including Vicuna-13B, PaLM, and Claude2.

The findings show substantial operational gains with minimal impact on accuracy. The framework achieved up to an 81.4% reduction in prompt token length, which decreased cloud API operational costs by up to 80.1% across multiple benchmarks. Furthermore, the condensed prompts decreased inference latency by up to 4.5 times and allowed for larger processing batch sizes without running out of GPU memory. The resulting natural language prompts maintained nearly identical accuracy to original uncompressed prompts across downstream models and transferred effectively to unseen datasets within similar task domains without requiring model retraining. The framework also consistently outperformed zero-shot summarization, generic language model summarization, and direct word-dropping techniques.

These results demonstrate that natural language prompt compression can drastically cut inference expenses and infrastructure latency while maintaining high accuracy. Because the compressed outputs remain standard text, organizations can apply one lightweight compression step upstream and flexibly route requests across diverse commercial API providers and open-source models without vendor lock-in. When implementing this technology, organizations should consider adopting natural language encapsulation for high-volume, cost-sensitive text processing workflows, though engineering teams must tune task-specific target length constraints, as overly aggressive truncation or excessive length can introduce noise or degrade reasoning logic.

Confidence in these findings is supported by consistent performance across diverse reasoning benchmarks and distinct language models. However, the evaluation was primarily conducted within reasoning and reading comprehension tasks using prompt lengths capped around 2,000 tokens. Further validation and pilot testing are recommended for broader enterprise domains and extremely long-document applications before deploying the framework into mission-critical production systems.

arXiv: 2402.18700
  • Paper: Prompt Compression for Large Language Models: A Survey, Zongqian Li et al. (2025). This later survey organizes prompt-compression methods, including natural-language approaches, into a broader taxonomy that situates the source’s findings within the field’s subsequent development.
Cover for Learning to Compress Prompt in Natural Language Formats

Abstract

Large language models (LLMs) are excel at processing multiple natural language processing tasks, but their abilities are constrained by inferior performance with long context, slow inference speed, and the high cost of computing the results. Deploying LLMs with precise and informative context helps users process large-scale datasets more effectively and cost-efficiently. Existing works rely on compressing long prompt contexts into soft prompts. However, soft prompt compression encounters limitations in transferability across different LLMs, especially API-based LLMs. To this end, this work aims to compress lengthy prompts in the form of natural language with LLM transferability. This poses two challenges: (i) Natural Language (NL) prompts are incompatible with back-propagation, and (ii) NL prompts lack flexibility in imposing length constraints. In this work, we propose a Natural Language Prompt Encapsulation (Nano-Capsulator) framework compressing original prompts into NL formatted Capsule Prompt while maintaining the prompt utility and transferability. Specifically, to tackle the first challenge, the Nano-Capsulator is optimized by a reward function that interacts with the proposed semantics preserving loss. To address the second question, Nano-Capsulator is optimized by a reward function featuring length constraints. Experimental results demonstrate that the Capsule Prompt can reduce 81.4% of the original length, decrease inference latency up to 4.5×, and save 80.1% of budget overheads while providing transferability across diverse LLMs and different datasets.

Table of Contents

  • 2 Related Work
  • 2.1 Soft Prompt Compression
  • 2.2 Context Distillation for Compression
  • 3 Long Prompt Encapsulation
  • 3.1 Prompt Encapsulation
  • 3.1.1 NL-formatted Prompt Compression
  • 3.1.2 Prompt Utility Preservation
  • 3.1.3 Compression with Reward
  • 3.2 Algorithm of Nano-Capsulator
  • 4 Experiments
  • 4.1 Dataset
  • 4.2 Experiment Settings
  • 4.3 Main Results (RQ1)
  • 4.4 Contributions of Utility Preservation (RQ2)
  • 4.5 Exploration of Impact Factors (RQ3)
  • 4.6 Latency of Nano-Capsulator (RQ3)
  • 5 Conclusion
  • References
  • Appendix
  • A Computation Infrastructure
  • B Additional Experiments of Comparison to Soft Prompt Baselines
  • C API Cost of PaLM
  • D Training Costs of Nano-Capsulator
  • E The Case Studies of Capsule Prompt
  • F Instruction Usage in Inference LLMs

Knowls

  1. Knowl 1 — Nano-Capsulator training and inference procedure

    algorithm

    Nano-Capsulator learns a compressor that turns a long natural-language prompt into a shorter natural-language Capsule Prompt. Let KK be an input prompt, CC its compressed version, F(⋅∣θC)F(\cdot\mid\theta_C) the trainable compressor, and G∗(⋅)G^*(\cdot) a frozen downstream language model. Training uses a replication instruction TRepT_{Rep}, a summarization instruction TSummT_{Summ}, and sampled downstream questions QQ.

    Input: Long prompts K, instructions TRep and TSumm, frozen downstream LLM G*, sampled downstream questions Q, and a token-length limit.
    Output: Natural-language Capsule Prompts C.
    Initialize the compressor F with parameters θC and keep G* frozen.
    Repeat until the compressor training converges:
        Generate C from F(K | TRep, TSumm, θC).
        Sample downstream questions Q.
        Compute semantic-preservation loss between representations of K and C.
        Truncate C to the token-length limit and score downstream utility against K using G* and Q.
        Combine semantic loss and utility reward into the Nano-Capsulator objective.
        Update θC to minimize the combined objective.
    Return the trained compressor F.

    At inference, the trained compressor generates a Capsule Prompt from the supplied long prompt in a forward pass; the downstream model can then receive the natural-language compressed prompt without requiring a model-specific soft prompt.

  2. Knowl 2 — Semantic-preservation objective for natural-language compression

    equation

    For an original prompt K={k1,…,kn}K=\{k_1,\ldots,k_n\} of nn tokens and a generated Capsule Prompt C={c1,…,cm}C=\{c_1,\ldots,c_m\} of mm tokens, with n≫mn\gg m, Nano-Capsulator encourages the two texts to have similar hidden representations under compressor F(⋅∣θC)F(\cdot\mid\theta_C). The compressor is instructed to replicate KK to obtain its representation and summarize it to obtain the representation of CC:

    eK∼PF(K∣θC,TRep),eC∼PF(C∣θC,TSumm),eK,eC∈Rd,e_K\sim P_F(K\mid\theta_C,T_{Rep}),\qquad e_C\sim P_F(C\mid\theta_C,T_{Summ}),\qquad e_K,e_C\in\mathbb{R}^d, LComp=EC[Ddist(eK∥eC)].\mathcal{L}_{Comp}=\mathbb{E}_{C}\left[D_{dist}(e_K\|e_C)\right].

    Here, TRepT_{Rep} and TSummT_{Summ} are replication and summarization instructions, dd is the hidden-state dimension, and DdistD_{dist} is a distance between representations; the experiments use mean squared error. The semantic content to preserve is task-dependent: the logical reasoning structure in few-shot chain-of-thought demonstrations or useful information in reading passages.

  3. Knowl 3 — Downstream-utility reward and strict length constraint

    equation

    Nano-Capsulator scores a compressed prompt by comparing downstream behavior when a frozen language model receives the original prompt versus the length-limited Capsule Prompt. Let G∗G^* be the frozen downstream model, QQ a set of sampled task questions, KiK_i and CiC_i the original and compressed prompts associated with question QiQ_i, Φ(Ci)\Phi(C_i) the compressed prompt truncated to the predetermined token limit, and ⊕\oplus concatenation. The reward is:

    Rcap=EQ[I(G∗(Φ(Ci)⊕Qi) ∥ G∗(Ki⊕Qi))].R_{cap}=\mathbb{E}_{Q}\left[\mathcal{I}\left(G^*(\Phi(C_i)\oplus Q_i)\,\middle\|\,G^*(K_i\oplus Q_i)\right)\right].

    The comparison function I\mathcal{I} supplies a downstream-utility score. In the reported experiments it uses mean squared error between hidden-state embeddings from G∗G^*; the authors also identify accuracy and GPT-4 evaluation scores as possible alternatives. Truncation is applied before scoring, so text beyond the prescribed limit cannot contribute to the evaluated Capsule Prompt.

  4. Knowl 4 — Joint objective for semantic and utility preservation

    equation

    Nano-Capsulator combines semantic-preservation loss with the downstream-utility reward in a single training objective:

    LNano=LComp(⋅∣θC)  Rcap(⋅∣θ∗).\mathcal{L}_{Nano}=\mathcal{L}_{Comp}(\cdot\mid\theta_C)\;R_{cap}(\cdot\mid\theta^*).

    Here, θC\theta_C denotes trainable compressor parameters, and θ∗\theta^* denotes the frozen parameters of downstream model G∗G^*. The intended training behavior is to penalize compressed prompts that lose downstream utility: the paper states that a low reward causes a larger penalty on semantic loss. The two components are optimized together so compression targets both retained meaning and task performance under a length limit.

  5. Knowl 5 — Experimental tasks, transfer tests, and training configuration

    experimental setup

    The experiments assess two kinds of prompts: few-shot chain-of-thought (CoT) demonstrations for CommonsenseQA (CSQA) and GSM8K, and reading-comprehension passages for MultiRC and TriviaQA-Long. The CoT prompts are built from seven CSQA examples and eight GSM8K examples; 1,000 CoT samples train the compressor, while inference also evaluates manually constructed CoT demonstrations not used in training. Reading-comprehension training uses 2,000 MultiRC question-answer-passage examples and all TriviaQA-Long training data. TriviaQA passages longer than 2,000 tokens are excluded, leaving an average passage length of about 900 tokens. Accuracy is the evaluation metric.

    For primary experiments, the token limits are 150 for CSQA and MultiRC, 350 for GSM8K, and 500 for TriviaQA-Long. The compressor is initialized from Vicuna-7B, and Vicuna-7B is also used as the frozen reward model during training; evaluation tests transfer to Vicuna-13B, PaLM, and Claude2. Model-transfer tests use a 70% training/validation and 30% test split. The reported data-transfer test trains on MultiRC and evaluates on BoolQ without further training. Training uses LoRA, Adam with learning rate 5×10−65\times10^{-6}, gradient clipping at 0.8, and two 48-GB NVIDIA A40 GPUs. Approximate training time is eight hours for few-shot CoT compression and four hours for reading-comprehension compression.

  6. Knowl 6 — Accuracy, compression, and cross-model results

    data/table

    The table reports accuracy (%) for manual few-shot CoT, zero-shot CoT, and Capsule Prompt on CSQA and GSM8K, and for original versus Capsule Prompt passages on MultiRC and TriviaQA-Long. It also gives the prompt lengths and compression rates reported by the authors. Capsule Prompt largely retains the original-prompt performance across the three evaluation LLMs; results vary by task and model, with some gains and some modest drops.

    CSQA GSM8K MultiRC TriviaQA-Long
    Model Manual Zero-shot Ours Manual Zero-shot Ours Original Ours Original Ours
    Vicuna-13B 60.4 44.6 58.8 34.4 25.3 31.9 57.3 57.1 86.0 88.8
    PaLM 73.7 67.5 75.5 62.8 56.8 59.5 72.7 72.2 78.9 78.8
    Claude2 76.6 69.4 74.6 85.6 52.7 84.9 59.4 58.2 95.0 92.3
    Length (tokens) 831 – 154 751 – 231 378.39 95.66 915.7 422.6
    Compression rate – – 81.4% – – 69.3% – 74.71% – 53.84%

    The comparison uses manual and zero-shot CoT baselines for the reasoning tasks, and uncompressed passages as the reading-comprehension baseline. The reported reductions in length are substantial, while the corresponding accuracy results are generally close to the uncompressed or manual-prompt results rather than uniformly identical.

  7. Knowl 7 — Transfer to an unseen reading-comprehension dataset

    empirical result

    For a data-transfer test, Nano-Capsulator is trained on MultiRC and then applied to the unseen BoolQ reading-comprehension dataset without further training. The paper reports that Capsule Prompt gives competitive accuracy under both Vicuna-13B and Claude2, with only a slight drop relative to a Capsule Prompt trained for the test data; training on the target data performs better. This experiment supports transfer to an unseen dataset with a similar downstream task, but does not establish transfer to unrelated task domains.

  8. Knowl 8 — Ablations and comparisons with alternative compression methods

    empirical result

    The paper evaluates the roles of semantic preservation and the downstream reward, and compares Capsule Prompt with summarization and text-removal baselines. On GSM8K and MultiRC, Capsule Prompt generated by Nano-Capsulator outperforms in-context zero-shot summarization by Vicuna-7B under the tested Vicuna-13B and Claude2 evaluations. On CSQA and GSM8K, it also performs better than direct GPT-3.5-Turbo summarization. On TriviaQA-Long, removing the reward function degrades downstream performance under Vicuna-13B and Claude2, supporting the utility of task-based reward during training.

    On Claude2, Capsule Prompt also outperforms the evaluated text-dropping approaches: Selective Context on CSQA and MultiRC, and random demonstration elimination on CSQA, including comparisons at similar prompt lengths. The plotted ablation results are presented graphically rather than as exact numerical values in the text.

  9. Knowl 9 — Effect of the Capsule Prompt length

    empirical result

    In experiments on TriviaQA, the tested Capsule Prompt length affects accuracy, and the best tested length differs by downstream model: 150 tokens performs best among the shown settings for Claude2, whereas 200 tokens performs best for Vicuna-13B. The authors observe that adding more text can preserve additional useful information but can also introduce noise or misinformation, so longer compressed prompts do not guarantee better task performance. Zero-shot summaries show similar performance trends across the tested length settings.

  10. Knowl 10 — Inference latency and API cost savings

    empirical result

    Latency experiments use OPT-2.7B and Vicuna-13B with generated output length fixed at 200 tokens and varying batch sizes. Compared with original prompts, Capsule Prompt reduces inference latency by 2.1×–4.5× overall; the maximum reported speedup is 4.5× for OPT-2.7B and 4.1× for Vicuna-13B. On OPT-2.7B at batch size 16, the original prompt runs out of memory whereas Capsule Prompt fits. The paper reports that latency benefits increase with batch size.

    The following Claude2 API costs are in dollars. Savings are calculated relative to the original prompt; the greatest listed saving is 80.1% on TriviaQA-Long.

    Dataset Original cost ($) Capsule Prompt cost ($) Saved
    CSQA 15.03 3.30 77.9%
    GSM8K 5.22 1.88 63.9%
    MultiRC 45.91 13.01 71.6%
    TriviaQA-Long 2.14 0.42 80.1%

    The paper also reports up to 80.1% API-cost savings for PaLM. These latency and cost results accompany the reduction in input length and are measured for the evaluated models, tasks, and batch settings.

Coverage note — The paper’s illustrative prompt case studies are not separate knowls because they exemplify the already captured semantic-preservation findings rather than adding independent results.

References

  1. 1.Anthropic. 2023. Claude.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. arXiv preprint arXiv:2305.14788.
  4. 4.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113.
  6. 6.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  8. 8.Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. Eraser: A benchmark to evaluate rationalized nlp models.
  9. 9.Tao Ge, Jing Hu, Xun Wang, Si-Qing Chen, and Furu Wei. 2023. In-context autoencoder for context compression in a large language model. arXiv preprint arXiv:2307.06945.
  10. 10.Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023. Llmlingua: Compressing prompts for accelerated inference of large language models. arXiv preprint arXiv:2310.05736.
  11. 11.Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Zirui Liu, Chia-Yuan Chang, Huiyuan Chen, and Xia Hu. 2024. Llm maybe longlm: Self-extend llm context window without tuning. arXiv preprint arXiv:2401.01325.
  12. 12.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. arXiv e-prints, page arXiv:1705.03551.
  13. 13.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface:a challenge set for reading comprehension over multiple sentences. In Proceedings of North American Chapter of the Association for Computational Linguistics (NAACL).
  14. 14.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  15. 15.Yucheng Li, Bo Dong, Chenghua Lin, and Frank Guerin. 2023. Compressing context to enhance inference efficiency of large language models. arXiv preprint arXiv:2310.06201.
  16. 16.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment, may 2023. arXiv preprint arXiv:2303.16634, 6.
  17. 17.Jesse Mu, Xiang Lisa Li, and Noah Goodman. 2023. Learning to compress prompts with gist tokens. arXiv preprint arXiv:2304.08467.
  18. 18.Siyu Ren, Qi Jia, and Kenny Q Zhu. 2023. Context compression for auto-regressive transformers with sentinel tokens. arXiv preprint arXiv:2310.08152.
  19. 19.Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  20. 20.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  21. 21.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  22. 22.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  23. 23.David Wingate, Mohammad Shoeybi, and Taylor Sorensen. 2022. Prompt compression and contrastive conditioning for controllability and toxicity reduction in language models. arXiv preprint arXiv:2210.03162.
  24. 24.Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu. 2023. Harnessing the power of llms in practice: A survey on chatgpt and beyond. arXiv preprint arXiv:2304.13712.
  25. 25.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pretrained transformer language models.

Citation

MLA
Chuang, Y.-N., et al. “Learning to Compress Prompt in Natural Language Formats”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 7756–67, https://doi.org/10.18653/v1/2024.naacl-long.429.
APA
Chuang, Y.-N., Xing, T., Chang, C.-Y., Liu, Z., Chen, X., & Hu, X. (2024). Learning to Compress Prompt in Natural Language Formats. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7756–7767. https://doi.org/10.18653/v1/2024.naacl-long.429
Chicago
Chuang, Y.-N., T. Xing, C.-Y. Chang, Z. Liu, X. Chen, and X. Hu. 2024. “Learning to Compress Prompt in Natural Language Formats”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7756–67. https://doi.org/10.18653/v1/2024.naacl-long.429.
Harvard
Chuang, Y.-N. et al. (2024) “Learning to Compress Prompt in Natural Language Formats”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7756–7767. Available at: https://doi.org/10.18653/v1/2024.naacl-long.429.
Vancouver
1. Chuang Y-N, Xing T, Chang C-Y, Liu Z, Chen X, Hu X (2024) Learning to Compress Prompt in Natural Language Formats. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 7756–7767

BibTeX

@inproceedings{chuang-etal-2024-learning,
    title = "Learning to Compress Prompt in Natural Language Formats",
    author = "Chuang, Yu-Neng  and
      Xing, Tianwei  and
      Chang, Chia-Yuan  and
      Liu, Zirui  and
      Chen, Xun  and
      Hu, Xia",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.429/",
    doi = "10.18653/v1/2024.naacl-long.429",
    pages = "7756--7767"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/