Adapting Language Models to Compress Contexts

Alexis ChevalierAlexander WettigAnirudh AjithDanqi Chen

article2023EMNLP395 citations

Introduces AutoCompressors, an unsupervised method for fine-tuning pre-trained language models to recursively compress long contexts into compact summary vectors that serve as soft prompts, significantly extending effective context windows up to 30,720 tokens while lowering inference costs for in-context learning and retrieval tasks.

Listen

Transformer-based language models are central to modern artificial intelligence applications, but they face severe practical constraints due to finite context windows and steep computational costs when processing long texts. Standard attention mechanisms scale quadratically with sequence length, making the ingestion of long documents prohibitively expensive in memory and compute. The article addresses this bottleneck by evaluating whether pre-trained language models can be adapted into "AutoCompressors" that recursively compress long context sequences into compact "summary vectors" (short soft prompts) to extend context capacity and accelerate inference.

The authors develop an unsupervised fine-tuning approach that equips models like OPT (1.3B and 2.7B parameters) and Llama-2 (7B parameters) with special summary tokens. Documents are divided into segments, compressed into summary vectors, and accumulated across segments as soft prompts for future text. To enable efficient training on a single 80GB GPU, the framework incorporates summary accumulation, randomized segment lengths, and backpropagation through time with stopped gradients. The study evaluates long-range language modeling on sequences up to 30,720 tokens, in-context learning across 11 benchmark tasks, retrieval-augmented generation on multi-billion token corpora, and unsupervised document re-ranking.

The findings show that AutoCompressors effectively compress long contexts while retaining critical factual and semantic information. First, in long-context evaluations, AutoCompressors successfully leveraged contexts up to 28,000 tokens to consistently reduce perplexity, outperforming baseline Recurrent Memory Transformers. Second, for in-context learning, substituting plain-text examples with summary vectors yielded higher accuracy than 150 plain tokens on 8 out of 11 tasks, and outperformed 750 plain tokens on 8 tasks while substantially lowering token processing requirements. Third, when applied to retrieval-augmented modeling, fusing pre-computed summary vectors achieved 1.5 times the perplexity gain of plain-text passages and delivered a 1.7-fold throughput increase over traditional multi-passage ensembling. Finally, in passage re-ranking, caching summary vectors established a Pareto-optimal trade-off between retrieval recall and computational throughput.

These results demonstrate that pre-computing and caching summary vectors offers a scalable, cost-effective method to expand context windows and speed up high-volume inference workflows. Organizations running large-scale retrieval or few-shot classification systems can lower runtime latency and operational costs without training massive architectures from scratch. Next steps supported by the article include evaluating the approach on larger foundation models, refining training to better capture granular information that full attention retains, and optimizing methods for aggregating high volumes of summary vectors.

Readers should note certain limitations: testing was restricted to models up to 7B parameters, and summary vectors exhibited a slight performance gap compared to full attention over very long spans due to information loss during compression. Nevertheless, the empirical findings provide high confidence that context compression via summary vectors is a viable and efficient enhancement for production language model pipelines.

arXiv: 2305.14788
Cover for Adapting Language Models to Compress Contexts

Abstract

Transformer-based language models (LMs) are powerful and widely-applicable tools, but their usefulness is constrained by a finite context window and the expensive computational cost of processing long text documents. We propose to adapt pre-trained LMs into AutoCompressors. These language models are capable of compressing long contexts into compact summary vectors, which are then accessible to the model as soft prompts. Summary vectors are trained with an unsupervised objective, whereby long documents are processed in segments, and summary vectors from all previous segments are used in language modeling. We fine-tune OPT and Llama-2 models on sequences of up to 30,720 tokens and show that AutoCompressors can utilize long contexts to improve perplexity. We evaluate AutoCompressors on in-context learning by compressing task demonstrations and find that summary vectors are good substitutes for plain-text demonstrations, increasing accuracy while reducing inference costs. Finally, we explore the benefits of pre-computing summary vectors for large corpora by applying summary vectors to retrieval-augmented language modeling and a passage re-ranking task. Overall, AutoCompressors emerge as a simple and inexpensive solution to extend the context window of LMs while speeding up inference over long contexts.^1

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Training Summary Vectors
  • 4 Language Modeling Evaluation
  • 4.1 Experiments on 8K-Token Sequences
  • 4.2 Experiments on 30K-Token Sequences
  • 4.3 Scaling Up AutoCompressors to Llama-2
  • 4.4 Analysis
  • 5 Compressing Demonstrations for In-Context Learning
  • 6 Compressing Retrieval Corpora for Efficient Inference
  • 6.1 Retrieval-augmented Language Modeling
  • 6.2 Unsupervised Passage Re-ranking
  • 7 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Models and Data
  • A.1 OPT Experiments on 8K Tokens
  • A.2 OPT Experiments on 30K Tokens
  • A.3 Llama-2 Experiments on 8K Tokens
  • B No-context Language Modeling
  • C AutoCompressor Ablations
  • D Token-level AutoCompressor Analysis
  • E In-Context Learning Details
  • F Fused Retrieval-augmented Language Modeling

Knowls

  1. Knowl 1 — AutoCompressor Architecture and Summary Accumulation

    model/method

    An AutoCompressor adapts a pre-trained Transformer language model to compress long sequences into compact continuous representations called summary vectors, which are passed to subsequent segments as soft prompts.

    Let the model vocabulary be expanded by κ\kappa special summary tokens ⟨Sum⟩1,…,⟨Sum⟩κ\langle\text{Sum}\rangle_1, \dots, \langle\text{Sum}\rangle_\kappa, each associated with an initialized input embedding. A long context document is partitioned into sequential segments S1,S2,…,SnS_1, S_2, \dots, S_n. When the token sequence ⟨Sum⟩1…⟨Sum⟩κ\langle\text{Sum}\rangle_1 \dots \langle\text{Sum}\rangle_\kappa is appended to the input of segment SiS_i, the Transformer's output hidden states at these positions form κ\kappa summary vectors σi∈Rκ×d\sigma_i \in \mathbb{R}^{\kappa \times d}, where dd is the hidden dimension.

    To allow direct information flow from all preceding text to the current segment, AutoCompressors employ summary accumulation: summary vectors from all previous segments are concatenated to form σ<i=Concat(σ1,…,σi−1)∈R(i−1)κ×d\sigma_{<i} = \text{Concat}(\sigma_1, \dots, \sigma_{i-1}) \in \mathbb{R}^{(i-1)\kappa \times d} and prepended as soft prompts to the token embeddings of segment SiS_i. The prepended prompt length (i−1)κ(i-1)\kappa grows linearly with the number of processed segments.

    Handling of positional embeddings depends on the base model architecture:

    • For models with learned absolute positional embeddings (such as OPT), no positional embeddings are added to the summary tokens or summary vectors. This preserves the pre-trained positional embeddings entirely for context tokens and allows scaling to an arbitrary number of compression steps without exceeding the position vocabulary.
    • For models using relative positional encodings (such as Rotary Position Embeddings / RoPE in Llama-2), standard positional embeddings are applied directly to the summary tokens and vectors without modification.
  2. Knowl 2 — AutoCompressor Unsupervised Training Objective and Optimization

    equation

    AutoCompressors are trained using an unsupervised next-token prediction objective across multi-segment documents without requiring a teacher model for knowledge distillation.

    Let a document be partitioned into nn segments S1,…,SnS_1, \dots, S_n, where segment Si=(x1i,…,xmii)S_i = (x_1^i, \dots, x_{m_i}^i) contains mim_i tokens, and let N=∑i=1nmiN = \sum_{i=1}^n m_i be the total token count. Conditioning on the accumulated summary vectors σ<i=Concat(σ1,…,σi−1)\sigma_{<i} = \text{Concat}(\sigma_1, \dots, \sigma_{i-1}) (with σ<1=∅\sigma_{<1} = \emptyset), the model minimizes the total cross-entropy loss:

    L=−1N∑i=1n∑t=1milog⁡p(xti∣x1i,…,xt−1i,σ<i)\mathcal{L} = -\frac{1}{N} \sum_{i=1}^n \sum_{t=1}^{m_i} \log p(x_t^i \mid x_1^i, \dots, x_{t-1}^i, \sigma_{<i})

    where p(xti∣x1i,…,xt−1i,σ<i)p(x_t^i \mid x_1^i, \dots, x_{t-1}^i, \sigma_{<i}) is obtained by projecting the Transformer hidden state at position t−1t-1 through the language modeling head.

    Training relies on two core optimization strategies:

    1. Randomized Segmenting: The segment lengths mim_i are varied randomly during training (e.g., uniformly between 1,024 and 2,048 tokens), which enables the model to compress variable-length context sequences and generalize to segment lengths unseen during training.
    2. Backpropagation Through Time (BPTT) with Stop-Gradients: Summary vectors are cached, and their gradient computation is truncated after two compression steps (ii and i+1i+1). This bounds the computational graph and GPU memory requirements while maintaining sufficient training signal for compression.
  3. Knowl 3 — Language Modeling Perplexity on 8K-Token Sequences for OPT-2.7B

    data/table

    Evaluating the OPT-2.7B AutoCompressor fine-tuned on 2B tokens from The Pile demonstrates that summary vectors reduce perplexity over long contexts while adding minimal prompt overhead. The AutoCompressor uses κ=50\kappa = 50 summary tokens per segment (a 40:1 compression ratio for 2,048-token segments). Perplexity is evaluated on the final 2,048 tokens of 8,192-token documents across 4 in-domain subdomains (Books3, FreeLaw, GitHub, Wikipedia) and 4 out-of-domain subdomains (ArXiv, Gutenberg, HackerNews, YouTubeSubtitles).

    In-domain Out-of-domain
    Context tokens 128 512 2048 4096 6144 128 512 2048 4096 6144
    Extended Full Attention 6.33 6.15 5.94 – – 8.57 8.28 7.93 – –
    RMT 6.42 6.19 6.02 6.02 6.01 8.76 8.44 8.21 8.20 8.20
    AutoCompressor 6.14 6.04 5.98 5.94 5.93 8.39 8.26 8.17 8.12 8.10

    The zero-context fine-tuned OPT-2.7B baseline achieves a perplexity of 6.28 in-domain and 8.53 out-of-domain.

    Key findings:

    • The AutoCompressor consistently outperforms the Recurrent Memory Transformer (RMT), which does not accumulate summary vectors and stagnates beyond 2,048 context tokens.
    • The AutoCompressor benefits from contexts as short as 128 tokens, outperforming full attention on short contexts due to randomized segmenting during training.
    • Over 6,144 context tokens, the AutoCompressor achieves a 5.6% in-domain and 5.0% out-of-domain perplexity gain using only 3×50=1503 \times 50 = 150 soft prompt tokens, whereas Extended Full Attention cannot scale past 2,048 additional tokens due to GPU memory constraints.
  4. Knowl 4 — Scaling AutoCompressors to 30K-Token Sequences and Memory Efficiency

    empirical result

    AutoCompressors can scale to sequences of 30,720 tokens with 20 compression steps during fine-tuning on a single 80GB NVIDIA A100 GPU.

    When fine-tuned on Books3 data from The Pile with κ=50\kappa = 50 summary tokens:

    • OPT-1.3B AutoCompressor reduces held-out perplexity on the final 2,048 tokens from 13.21 (0 context tokens) to 12.49 (14,336 context tokens) and 12.47 (28,672 context tokens). In contrast, the RMT-1.3B baseline plateaus at 12.50 perplexity at both 14,336 and 28,672 context tokens.
    • OPT-2.7B AutoCompressor reduces perplexity from 11.86 (0 context tokens) to 11.21 (14,336 context tokens) and 11.18 (28,672 context tokens).

    Memory Efficiency: Because AutoCompressors stop gradients after two compression steps, fine-tuning OPT-1.3B requires 38GB CUDA memory (compared to 54GB for RMT-1.3B), and fine-tuning OPT-2.7B on 30,720 tokens requires 75GB CUDA memory. In contrast, RMT-2.7B encounters an Out-Of-Memory (OOM) error under identical hardware and gradient checkpointing conditions.

  5. Knowl 5 — Scaling AutoCompressors to Llama-2-7B via LoRA

    experimental setup

    To adapt larger language models without full parameter fine-tuning, AutoCompressors can be applied to Llama-2-7B on a single GPU by freezing the base model and training only the newly initialized summary token embeddings and attention projection weights using Low-Rank Adaptation (LoRA).

    Configuration parameters:

    • Base Model: Llama-2-7B (native 4,096-token context window, RoPE positional embeddings).
    • LoRA Parameters: Rank r=16r = 16 applied to Query, Key, Value, and Output projection matrices in all attention heads.
    • Training Dataset: 15B tokens from RedPajama (10B CommonCrawl, 1B each from ArXiv, Books, C4, GitHub; 800M Wikipedia, 70M StackExchange).
    • Sequence Length: 6,144 tokens split into 4 segments with randomized segmenting.
    • Summary Tokens: κ=50\kappa = 50.
    • Optimization: Adam optimizer, learning rate 4×10−44\times 10^{-4}, batch size 200K tokens, 5,000 warmup steps, stop-gradients after 2 compression steps.

    On 8,192-token evaluation documents, compressing a 4,096-token context into 100 summary vectors yields a perplexity of 5.08 on the final 2,048 tokens (matching Extended Full Attention with 512 plain tokens at 5.06), and compressing 6,144 tokens into 150 summary vectors achieves 5.07 perplexity (down from the 5.40 zero-context baseline).

  6. Knowl 6 — In-Context Learning via Compressed Demonstration Vectors

    empirical result

    In-context learning (ICL) demonstrations can be compressed into summary vectors using an AutoCompressor, matching or exceeding the performance of plain-text demonstrations while reducing prompt length and inference cost.

    When evaluating the Llama-2-7B AutoCompressor across 11 NLP classification and multiple-choice benchmarks (AG News, SST-2, BoolQ, WiC, WSC, RTE, CB, COPA, MultiRC, MR, Subj):

    • Summary vectors from 1 to 3 compressed segments of demonstrations (each segment containing up to 750 tokens of demonstrations, compressed into 50, 100, or 150 summary vectors) consistently outperform zero-shot baselines on all 11 tasks.
    • Prompting with summary vectors (50 to 150 vectors) outperforms standard few-shot ICL using 150 plain-text demonstration tokens on 8 out of 11 tasks.
    • Summary vectors also outperform standard ICL using 750 plain-text demonstration tokens on 8 out of 11 tasks (AG News, SST-2, BoolQ, WiC, WSC, CB, COPA, MultiRC).
    • On OPT-2.7B, the AutoCompressor with summary accumulation benefits from scaling up to 3 compression steps, whereas the RMT baseline fails to benefit from multiple compression steps on 7 of 11 tasks.
  7. Knowl 7 — Fused Summaries for Retrieval-Augmented Language Modeling

    model/method

    Fused Summaries is a retrieval-augmented language modeling framework that incorporates multiple external passages into a single forward pass by fusing pre-computed summary vectors.

    Given an external corpus C\mathcal{C}, summary vectors σd∈Rκ×d\sigma_d \in \mathbb{R}^{\kappa \times d} are pre-computed offline and cached for all passages d∈Cd \in \mathcal{C}. Given a query context xx, an off-the-shelf dense retriever (such as Contriever) retrieves the top-kk relevant passages D={d1,d2,…,dk}D = \{d_1, d_2, \dots, d_k\}.

    The retrieved summary vectors are concatenated in order of increasing relevance (least to most relevant) to form the fused summary prefix: σD=Concat(σdk,…,σd1)\sigma_D = \text{Concat}(\sigma_{d_k}, \dots, \sigma_{d_1})

    To improve output calibration, next-segment token predictions are computed by smoothing the summary-conditioned probability with the unconditioned probability: p(y∣x,D)=p(y∣Concat(σD,x))+p(y∣x)2p(y \mid x, D) = \frac{p(y \mid \text{Concat}(\sigma_D, x)) + p(y \mid x)}{2}

    Additionally, re-ordering the candidate passages by the smallest ℓ2\ell_2 distance between their summary vectors σdi\sigma_{d_i} and the prompt context's summary vector σx\sigma_x (computed in the same forward pass as p(y∣x)p(y \mid x)) further enhances language modeling perplexity compared to raw retriever scores.

  8. Knowl 8 — Perplexity and Throughput of Retrieval-Augmented Fused Summaries

    data/table

    Evaluating retrieval-augmented language modeling on The Pile (with OPT-2.7B and Contriever retrieving top-kk passages from a 10B-token corpus per domain) shows that Fused Summaries achieves an advantageous trade-off between perplexity gain and inference throughput.

    Perplexity Gain (%) Throughput (examples/s)
    Passage Format Method top-1 top-2 top-5 top-10 top-1 top-2 top-5 top-10
    50 tokens REPLUG -0.64 0.58 1.68 2.35 51 38 16 9
    50 tokens Fused Passages 0.71 1.01 1.70 2.60 28 27 23 17
    512 tokens →\rightarrow 50 sum. vecs Fused Summaries 1.04 1.67 2.63 3.74 28 27 23 17
    512 tokens REPLUG -1.47 2.24 5.25 8.30 18 10 6 3

    Key results:

    • Fused Summaries (compressing 512-token passages into 50 summary vectors) outperforms REPLUG with 50-token passages at all top-kk levels and yields 1.5×\times the perplexity gain of Fused Passages at top-10 (3.74% vs 2.60%).
    • Fused Summaries at top-10 outperforms REPLUG top-2 with 512-token passages (3.74% vs 2.24% gain) while achieving 1.7×1.7\times higher throughput (17 vs 10 examples/s on an A100 GPU).
    • Storage footprint for 10B tokens compressed into summary vectors is 5TB in half-precision format, compared to 51TB for token-level representations and 3,276TB for full attention key-value caches.
  9. Knowl 9 — Unsupervised Passage Re-ranking with Cached Summary Vectors

    empirical result

    AutoCompressor summary vectors can be used for zero-shot passage re-ranking without fine-tuning, placing the system on the Pareto frontier of Recall@20 versus query throughput.

    In this setup, BM25 first retrieves candidate passages from 21M Wikipedia passages for queries from the Natural Questions (NQ) test set. The language model re-ranks candidate passages pip_i by computing the likelihood of query qq conditioned on the prompt template: "Passage: {pi}. Please write a question based on this passage."\text{"Passage: } \{p_i\}\text{. Please write a question based on this passage."}

    When substituting the text of passage pip_i with pre-computed summary vectors (using κ=20\kappa = 20 or κ=50\kappa = 50 vectors per passage, requiring 2.1TB and 5.4TB disk storage across 21M passages in half-precision):

    • The evaluation throughput on an unbatched NVIDIA A100 GPU increases substantially compared to feeding full plain-text passages into standard OPT models (125M to 2.7B parameters).
    • Re-ranking with summary vectors retains high Recall@20, outperforming smaller full-text models at equivalent or higher query processing speeds and establishing a Pareto-optimal compute-performance trade-off.
  10. Knowl 10 — Ablation of AutoCompressor Design Components

    empirical result

    Ablation experiments on OPT-2.7B identify the performance contributions of each architectural and training choice:

    1. Summary Accumulation: Removing accumulation (reverting to single-segment recurrent passing as in RMT) causes perplexity gains to plateau after 2,048 context tokens, whereas accumulation continues to reduce perplexity across 4,096 and 6,144 tokens.
    2. Randomized Segmenting: Training on randomly sized segments between 1,024 and 2,048 tokens enables the model to compress short contexts (e.g., 128 and 512 tokens) effectively during inference, while models trained on fixed-length segments fail to compress short contexts effectively.
    3. Stop-Gradients: Truncating backpropagation after two compression steps yields identical held-out perplexity curves across context lengths up to 6,144 tokens while drastically reducing peak VRAM usage from out-of-memory to runnable on a single 80GB GPU.
    4. Summary Vector Count κ\kappa: Comparing κ∈{20,50,70,100}\kappa \in \{20, 50, 70, 100\} on 8,192-token documents reveals that κ=50\kappa = 50 achieves the lowest perplexity (6.93 at 6,144 context tokens vs 7.00 for κ=20\kappa=20, 6.95 for κ=70\kappa=70, and 7.00 for κ=100\kappa=100). Increasing soft prompt length beyond 50 vectors does not improve performance without scaling the training budget.
  11. Knowl 11 — Limitations of AutoCompressors

    limitation

    The AutoCompressor approach exhibits three main limitations:

    1. Information Loss Compared to Full Attention: Compressed summary vectors do not retain all context information accessible via uncompressed full attention over long sequences, leading to an empirical performance gap relative to extended full-attention baselines when sequence length is within the baseline's context budget.
    2. Quadratic Complexity of Summary Accumulation: Although the growth rate is reduced by a factor of m/κm/\kappa (e.g., 40×40\times), concatenating summary vectors across nn segments still results in an attention sequence length of nκn\kappa, which retains quadratic computational complexity with respect to the total number of segments.
    3. Saturation with Prompt Length: Language modeling performance does not improve when increasing the number of summary vectors per segment beyond κ=50\kappa = 50, likely because the supervision signal for learning summary vectors is bounded by the model's ability to predict local tokens from uncompressed current-segment inputs.

Coverage note — Qualitative token-level inspection examples from Appendices D/E and dataset-specific prompt templates from Appendix Table 11 were omitted as they provide individual illustrative examples rather than core reusable findings.

References

  1. 1.Joshua Ainslie, Tao Lei, Michiel de Jong, Santiago Ontañón, Siddhartha Brahma, Yury Zemlyanskiy, David Uthus, Mandy Guo, James Lee-Thorp, Yi Tay, Yun-Hsuan Sung, and Sumit Sanghai. 2023. CoLT5: Faster long-range transformers with conditional computation.
  2. 2.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861.
  3. 3.Vidhisha Balachandran, Bhuwan Dhingra, Haitian Sun, Michael Collins, and William Cohen. 2021. Investigating the effect of background knowledge on natural questions. In Proceedings of Deep Learning Inside Out (DeeLIO): The 2nd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 25–30, Online. Association for Computational Linguistics.
  4. 4.Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009. The fifth PASCAL recognizing textual entailment challenge. In TAC.
  5. 5.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  7. 7.Aydar Bulatov, Yuri Kuratov, and Mikhail Burtsev. 2022. Recurrent memory transformer. In Advances in Neural Information Processing Systems.
  8. 8.Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174.
  9. 9.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509.
  10. 10.Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. 2021. Rethinking attention with Performers. In International Conference on Learning Representations.
  11. 11.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2924–2936, Minneapolis, Minnesota. Association for Computational Linguistics.
  12. 12.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The PASCAL recognising textual entailment challenge. In the First International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, and Recognizing Textual Entailment.
  13. 13.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019. Transformer-XL: Attentive language models beyond a fixed-length context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2978–2988, Florence, Italy. Association for Computational Linguistics.
  14. 14.Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Re. 2022. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems.
  15. 15.Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The CommitmentBank: Investigating projection in naturally occurring discourse. Proceedings of Sinn und Bedeutung, 23(2):107–124.
  16. 16.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The Pile: An 800GB dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  17. 17.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 3929–3938. PMLR.
  18. 18.R Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006. The second pascal recognising textual entailment challenge. In Proceedings of the Second PASCAL Challenges Workshop on Recognising Textual Entailment, volume 7, pages 785–794.
  19. 19.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  20. 20.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  21. 21.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  22. 22.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  23. 23.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations.
  24. 24.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 252–262, New Orleans, Louisiana. Association for Computational Linguistics.
  25. 25.Diederik Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), San Diega, CA, USA.
  26. 26.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  27. 27.Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012. The winograd schema challenge. In 13th International Conference on the Principles of Knowledge Representation and Reasoning, KR 2012, Proceedings of the International Conference on Knowledge Representation and Reasoning, pages 552–561. Institute of Electrical and Electronics Engineers Inc. 13th International Conference on the Principles of Knowledge Representation and Reasoning, KR 2012 ; Conference date: 10-06-2012 Through 14-06-2012.
  28. 28.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  29. 29.Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. 2023. Scaling down to scale up: A guide to parameter-efficient fine-tuning. arXiv preprint arXiv:2303.15647.
  30. 30.Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 61–68, Dublin, Ireland. Association for Computational Linguistics.
  31. 31.Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. 2022. Mega: moving average equipped gated attention. arXiv preprint arXiv:2209.10655.
  32. 32.Jesse Mu, Xiang Lisa Li, and Noah Goodman. 2023. Learning to compress prompts with gist tokens. arXiv preprint arXiv:2304.08467.
  33. 33.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity. In Proceedings of ACL, pages 271–278.
  34. 34.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of ACL, pages 115–124.
  35. 35.Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. WiC: the word-in-context dataset for evaluating context-sensitive meaning representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1267–1273, Minneapolis, Minnesota. Association for Computational Linguistics.
  36. 36.Ofir Press, Noah Smith, and Mike Lewis. 2022. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations.
  37. 37.Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020. Compressive transformers for long-range sequence modelling. In International Conference on Learning Representations.
  38. 38.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI Spring Symposium: Logical Formalizations of Commonsense Reasoning.
  39. 39.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023. Code llama: Open foundation models for code.
  40. 40.Devendra Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022. Improving passage retrieval with zero-shot question generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3781–3797, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  41. 41.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. REPLUG: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652.
  42. 42.Charlie Snell, Dan Klein, and Ruiqi Zhong. 2022. Learning by distilling context. arXiv preprint arXiv:2209.15189.
  43. 43.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  44. 44.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2022. Roformer: Enhanced transformer with rotary position embedding.
  45. 45.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022. Efficient transformers: A survey. ACM Comput. Surv., 55(6).
  46. 46.TogetherAI. 2023. RedPajama: An open source recipe to reproduce llama training dataset.
  47. 47.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  49. 49.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  50. 50.David Wingate, Mohammad Shoeybi, and Taylor Sorensen. 2022. Prompt compression and contrastive conditioning for controllability and toxicity reduction in language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5621–5634, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  51. 51.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. OPT: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  52. 52.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.
  53. 53.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.
  54. 54.Lin Zheng, Chong Wang, and Lingpeng Kong. 2022. Linear complexity randomized self-attention mechanism. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 27011–27041. PMLR.
  55. 55.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [MASK]: Learning vs. learning to recall. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5017–5033, Online. Association for Computational Linguistics.
  56. 56.Zexuan Zhong, Tao Lei, and Danqi Chen. 2022. Training language models with memory augmentation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5657–5673, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.

Citation

MLA
Chevalier, A., et al. “Adapting Language Models to Compress Contexts”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 3829–46, https://doi.org/10.18653/v1/2023.emnlp-main.232.
APA
Chevalier, A., Wettig, A., Ajith, A., & Chen, D. (2023). Adapting Language Models to Compress Contexts. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3829–3846. https://doi.org/10.18653/v1/2023.emnlp-main.232
Chicago
Chevalier, A., A. Wettig, A. Ajith, and D. Chen. 2023. “Adapting Language Models to Compress Contexts”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 3829–46. https://doi.org/10.18653/v1/2023.emnlp-main.232.
Harvard
Chevalier, A. et al. (2023) “Adapting Language Models to Compress Contexts”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3829–3846. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.232.
Vancouver
1. Chevalier A, Wettig A, Ajith A, Chen D (2023) Adapting Language Models to Compress Contexts. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3829–3846

BibTeX

@inproceedings{chevalier-etal-2023-adapting,
    title = "Adapting Language Models to Compress Contexts",
    author = "Chevalier, Alexis  and
      Wettig, Alexander  and
      Ajith, Anirudh  and
      Chen, Danqi",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.232/",
    doi = "10.18653/v1/2023.emnlp-main.232",
    pages = "3829--3846"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/