Training Transformers for KV Cache Compressibility

Yoav GelbergYam EitanMichael BronsteinYarin GalHaggai Maron

article2026arXiv9 citations

Introduces KV-Compression Aware Training (KV-CAT), a pretraining method that masks key-value slots during training to produce representations that substantially improve the performance of downstream KV cache compression algorithms on long-context tasks.

Listen

Deploying autoregressive transformer language models for long-horizon tasks—such as code repository analysis, document synthesis, and autonomous agent operations—is severely bottlenecked by the Key-Value (KV) cache. During inference, models must store key and value states for every token across all attention heads and layers. Memory requirements and decoding latency scale linearly with sequence length, making long context windows a major operational constraint. While post-hoc context and cache compression methods have emerged to mitigate this burden, they operate on fixed, pretrained models whose internal representations may be inherently resistant to compression.

The article establishes both theoretically and empirically that KV cache compressibility is a learned structural property of the model itself rather than a trait of the input data alone. Its primary objective is to demonstrate that transformers can be guided during training to learn internal representations that are substantially more amenable to downstream compression, thereby improving inference efficiency without degrading standard model accuracy.

To achieve this, the authors propose KV-Compression Aware Training (KV-CAT), a continued pretraining framework. Under KV-CAT, the transformer undergoes dual forward passes: a dense pass and a masked pass where lightweight, learned linear-attention routers enforce a train-time KV sparsification policy by masking a set fraction of KV slots. The training objective combines a self-distillation loss that forces the masked representations to mirror the dense output, an anchoring next-token prediction loss to preserve base capabilities, and a budget constraint loss. The authors evaluated the approach on Qwen 2.5 models (0.5B and 1.5B parameters) across standard benchmarks, long-context question answering, needle-in-a-haystack retrieval, and suffix-continuation perplexity under post-hoc compression techniques.

Empirical findings confirm the effectiveness of this approach across several key operational dimensions. First, KV-CAT models fully preserved baseline capabilities, matching standard short-context multiple-choice benchmark accuracy within 0.5 to 0.7 percentage points without compression applied. Second, when downstream compression methods were applied, KV-CAT checkpoints demonstrated up to a 3.21-fold improvement in suffix perplexity retention and reached equivalent performance levels up to 5 times faster during gradient-based cache optimization. Third, in retrieval tasks from compressed contexts, KV-CAT raised mean retrieval accuracy by 5.2 to 6.4 percentage points overall, with improvements reaching 11 to 19 percentage points at moderate retention budgets (30% to 50% keep ratios). Finally, on long-context question answering tasks from LongBench v2, KV-CAT delivered an average accuracy improvement of up to 39% across multiple domains.

These results demonstrate that the bottleneck of post-hoc cache compression can be addressed at the training stage. For engineering and deployment leaders, this offers a practical pathway to substantially lower the operational compute and memory footprint of long-context applications, potentially reducing serving costs and hardware constraints for high-throughput language model pipelines.

Organizations developing or hosting long-context language models should consider incorporating compression-aware objectives, such as KV-CAT, during domain adaptation or continued pretraining phases. When adopting this method, engineering teams can choose between fixed heuristic sparsification policies and adaptive learned routers; the evidence indicates that learned routers offer the best trade-off by preserving uncompressed baseline accuracy while maximizing downstream compressibility.

Decision-makers should note several operational boundaries and uncertainties. KV-CAT introduces non-trivial training overhead and implementation complexity due to the dual forward passes and auxiliary routing modules. Furthermore, empirical evaluations were conducted on smaller open-weight models (up to 1.5B parameters) using specific continued pretraining corpora. While confidence in the reported experimental improvements is high, broader deployment across larger-scale frontier models and specialized domains will require targeted pilot validation.

arXiv: 2605.05971

No sufficiently relevant recommendations were found.

Cover for Training Transformers for KV Cache Compressibility

Abstract

Long-context language modeling is increasingly constrained by the Key-Value (KV) cache, whose memory and decode-time access costs scale linearly with the prefix length. This bottleneck has motivated a range of context-compression methods, from token-level summarization to recent optimization-based KV compression methods. These post-hoc methods operate on the KV cache of a fixed pretrained model, so their effectiveness is fundamentally limited by how well the model's internal representations can be compressed. In this work, we formalize the notion of KV compressibility and show that it is a property of the learned representations, rather than of the context alone. We prove that almost any sequence-to-vector function admits both highly compressible and inherently non-compressible transformer implementations, highlighting the need to guide transformers toward compressible representations during training. Motivated by this, we propose KV-Compression Aware Training (KV-CAT), a continued pretraining procedure that incentivizes the emergence of compressible representations. We introduce a train-time KV sparsification policy that masks KV slots during training. This forces the model to use fewer KV slots and encourages it to learn representations amenable to post-hoc compression. Empirically, we show that KV-CAT improves the quality-budget tradeoff of downstream compression methods across retrieval, long-context question answering, and perplexity-based evaluation of compressed-prefix continuation.

Table of Contents

  • 1 Introduction
  • 2 KV Cache Compression: Problem Formulation
  • 3 Motivation and Theoretical Results
  • 4 Training for KV Cache Compressibility
  • 5 Empirical Evaluation
  • 6 Related Work
  • 7 Conclusion
  • 8 Acknowledgements
  • References
  • A Theory
  • A.1 Transformers and KV Cache Compression
  • A.2 Motivating example: histogram computation
  • A.3 Compressible Transformers for General Functions
  • B KV-CAT implementation details
  • C Extended Experimental Details
  • C.1 Continued pretraining runs
  • C.2 No compression QA evaluation
  • C.3 Suffix perplexity under prefix KV cache compression
  • C.4 Needle-in-a-haystack retrieval under KV cache compression
  • C.5 Long-form question answering under KV cache compression
  • D Additional Experimental Results
  • D.1 Train-time KV sparsification policy ablation
  • D.2 Cross-domain attention matching evaluation
  • D.3 Additional KV keep-ratio curves
  • E Limitations
  • F Broader Impact

Knowls

  1. Knowl 1 — A sequence-to-vector function can have both highly compressible and non-compressible transformer implementations

    theoretical result

    Let A\mathcal A be a finite alphabet, and let ff map sequences of length at most NN over A\mathcal A to vectors in Rdout\mathbb R^{d_{\mathrm{out}}}. Suppose there is a prefix aa and two suffixes b1,b2b_1,b_2 of the same length, with ∣a∣+∣bi∣≤N|a|+|b_i|\le N, such that f([a,b1])≠f([a,b2])f([a,b_1])\ne f([a,b_2]). For every approximation tolerance ε>0\varepsilon>0 and every error threshold C>0C>0, there are transformer implementations of the same architecture that both approximate ff on sequences of length at most NN, but have sharply different KV compressibility: one admits a compression policy retaining just one KV pair per compressed prefix while keeping output error below ε\varepsilon; another has output error greater than CC on some prefix–suffix input for every compression policy that reduces any prefix of length nn to fewer than nn KV pairs. Thus, for functions meeting the stated condition, compressibility is not determined by the function or input alone; it can depend on the transformer’s learned representation.

  2. Knowl 2 — Formal definition of transformer KV compressibility

    definition

    Fix a maximum combined sequence length NN, an output-error tolerance ε>0\varepsilon>0, and a budget function rr that maps a prefix length nn to a number of retained KV pairs no greater than nn. A KV compression policy C=(c1,…,cL)C=(c_1,\ldots,c_L) has one compressor per transformer layer; each cℓc_\ell maps that layer’s sequence of prefix key–value pairs to a sequence of r(n)r(n) pairs. For a prefix token sequence aa and suffix token sequence bb, let M([a,b])M([a,b]) be the original transformer’s output vector for their concatenation, and let MC,a(b)M_{C,a}(b) be the output when processing bb with the prefix KV cache compressed by CC. The transformer is (N,ε,r)(N,\varepsilon,r)-compressible if there exists one policy CC such that, for every such pair with ∣a∣+∣b∣≤N|a|+|b|\le N, ∥M([a,b])−MC,a(b)∥<ε\|M([a,b])-M_{C,a}(b)\|<\varepsilon. The comparison is between output vectors; compression changes the prefix KV pairs used during suffix processing.

  3. Knowl 3 — KV-CAT trains with masked self-distillation and a dense behavior anchor

    model/method

    KV-Compression Aware Training (KV-CAT) is continued pretraining of a pretrained decoder-only transformer with both masked and unmasked forward passes. For a token sequence a=(a1,…,an)a=(a_1,\ldots,a_n), let pθmask(⋅∣a<i)p^{\mathrm{mask}}_\theta(\cdot\mid a_{<i}) and pθdense(⋅∣a<i)p^{\mathrm{dense}}_\theta(\cdot\mid a_{<i}) denote the next-token distributions from the masked and dense passes, respectively. The dense distribution is stop-gradient in the self-distillation term. The objective is

    L(θ)=λmaskLmask+λbudgetLbudget+λanchorLanchor,\mathcal L(\theta)=\lambda_{\mathrm{mask}}\mathcal L_{\mathrm{mask}}+\lambda_{\mathrm{budget}}\mathcal L_{\mathrm{budget}}+\lambda_{\mathrm{anchor}}\mathcal L_{\mathrm{anchor}}, Lmask=1n∑i=1nDKL ⁣(sg⁡[pθdense(⋅∣a<i)] ∥ pθmask(⋅∣a<i)),Lanchor=−1n∑i=1nlog⁡pθdense(ai∣a<i).\mathcal L_{\mathrm{mask}}=\frac{1}{n}\sum_{i=1}^n D_{\mathrm{KL}}\!\left(\operatorname{sg}[p^{\mathrm{dense}}_\theta(\cdot\mid a_{<i})]\,\middle\|\,p^{\mathrm{mask}}_\theta(\cdot\mid a_{<i})\right),\qquad \mathcal L_{\mathrm{anchor}}=-\frac{1}{n}\sum_{i=1}^n\log p^{\mathrm{dense}}_\theta(a_i\mid a_{<i}).

    Here θ\theta denotes transformer parameters, DKLD_{\mathrm{KL}} is Kullback–Leibler divergence, and sg⁡\operatorname{sg} stops gradients through its argument. The budget term regulates the fraction of active KV slots: if FF is the mean binary active-mask value, GG is the mean router keep score, and ρ∈(0,1)\rho\in(0,1) is the target retention rate, the implementation uses Lbudget=ρ−1FG+(1−ρ)−1(1−F)(1−G)\mathcal L_{\mathrm{budget}}=\rho^{-1}FG+(1-\rho)^{-1}(1-F)(1-G). The masked pass learns to match the same model’s dense distribution, while the dense next-token-prediction anchor preserves unmasked behavior and supplies the distillation target. Transformer and router parameters are updated jointly. At evaluation, masking is disabled and ordinary post-hoc KV compressors are applied to the resulting standard transformer.

  4. Knowl 4 — Histogram computation illustrates representation-dependent compressibility

    theoretical result

    For an alphabet [m]={1,…,m}[m]=\{1,\ldots,m\}, the histogram of a length-nn sequence is its vector of empirical symbol frequencies. A simple two-layer transformer computes this by mapping each token to its one-hot symbol vector and using uniform attention to average the vectors. This implementation is not robust to nontrivial prefix compression: for any maximum length NN and budget with r(N−1)<N−1r(N-1)<N-1, there is a positive output-error lower bound for some prefix and suffix, regardless of the compression policy. In contrast, a modified implementation with the same architecture can preserve positional and length information and use a single compressed KV pair whose value stores the prefix’s unnormalized histogram and length. Its final feed-forward computation corrects the scaling introduced when the compressed prefix summary is averaged with suffix tokens. For every NN and every ε>0\varepsilon>0, this structured implementation approximates the histogram with a one-pair prefix budget and error below ε\varepsilon. The example exhibits both forms of compressibility for the same aggregate-statistics function.

  5. Knowl 5 — Learned causal routers implement KV sparsification during training

    model/method

    KV-CAT’s reported sparsification policy uses lightweight learned routers shared across groups of consecutive transformer layers. At a routing boundary, a router receives hidden states hth_t for tokens tt and uses causal linear attention to form a summary ata_t. In the paper’s parameterization, it layer-normalizes hth_t, applies learned query, key, and value projections followed by ϕ(x)=ELU⁡(x)+1\phi(x)=\operatorname{ELU}(x)+1, and accumulates St=∑j≤tkjrj⊤S_t=\sum_{j\le t}k_jr_j^\top and zt=∑j≤tkjz_t=\sum_{j\le t}k_j; the summary is at=WO(qt⊤St)/(qt⊤zt+ϵ0)a_t=W_O(q_t^\top S_t)/(q_t^\top z_t+\epsilon_0), where ϵ0>0\epsilon_0>0 stabilizes the denominator. A keep score is computed as pt=(1−⟨ut,wt⟩)/2p_t=(1-\langle u_t,w_t\rangle)/2, with ut=WPht/∥WPht∥2u_t=W_Ph_t/\|W_Ph_t\|_2 and wt=(ht+αat)/∥ht+αat∥2w_t=(h_t+\alpha a_t)/\|h_t+\alpha a_t\|_2. The binary decision is mt=1{pt>τ}m_t=\mathbf 1\{p_t>\tau\}; the main runs use τ=0.5\tau=0.5 and a straight-through estimator for gradients. Routers start with WP=−IW_P=-I and α=0\alpha=0, giving all-active masks. Masks restrict which past KV slots are visible to attention, are independent at different routing boundaries, and are disabled for evaluation.

    The authors continued-pretrained QWEN2.5-0.5B and QWEN2.5-1.5B on FineWeb-Edu for 5.24×1095.24\times10^9 tokens per model, using sequences of 1,024 tokens and a target keep rate of 50%. Four router boundaries were used: layers 0, 6, 12, and 18 for the 0.5B model, and layers 0, 7, 14, and 21 for the 1.5B model; router feature dimension was 64. Both runs used λmask=1\lambda_{\mathrm{mask}}=1, λanchor=1\lambda_{\mathrm{anchor}}=1, and λbudget=0.1\lambda_{\mathrm{budget}}=0.1, with AdamW for 40,000 steps, peak learning rate 10−410^{-4}, 600 warmup steps, minimum learning rate 5×10−65\times10^{-6}, weight decay 0.01, and gradient clipping at norm 1.0. In the policy ablation, the learned router performed best on the reported compression metrics and retained higher uncompressed average accuracy than the random and attention-based fixed policies.

  6. Knowl 6 — Attention Matching finds lower-distortion caches at matched budgets

    data/table

    The comparison below reports prefix-compression degradation on held-out FineWeb examples: each example has a 768-token prefix and a 256-token suffix, and only the prefix KV cache is compressed. Attention Matching uses suffix query states to construct compact caches; metrics are averaged over 128 examples per keep ratio. Δ\DeltaPPL is the increase in suffix perplexity relative to the model’s dense-prefix pass (lower is better), KL is divergence from the dense-prefix token distribution (lower is better), and Top-1 is the percentage of suffix tokens whose most likely token agrees with the dense-prefix pass (higher is better). KV-CAT improves all three metrics for every reported model and keep ratio. Its largest proportional reduction in Δ\DeltaPPL is 3.21×3.21\times, for QWEN2.5-1.5B at 10% keep (1.0271.027 versus 0.3200.320).

    Model Keep Variant Δ\DeltaPPL (↓\downarrow) KL (↓\downarrow) Top-1 (↑\uparrow)
    QWEN2.5-0.5B 5% Base 0.780 0.0562 88.5%
    KV-CAT 0.461 0.0385 91.1%
    10% Base 0.662 0.0483 89.3%
    KV-CAT 0.422 0.0321 91.4%
    20% Base 0.777 0.0549 88.7%
    KV-CAT 0.438 0.0374 90.8%
    40% Base 0.592 0.0431 90.3%
    KV-CAT 0.360 0.0272 92.5%
    QWEN2.5-1.5B 5% Base 0.573 0.0570 89.2%
    KV-CAT 0.292 0.0361 90.9%
    10% Base 1.027 0.0796 88.3%
    KV-CAT 0.320 0.0366 90.7%
    20% Base 1.093 0.0876 87.7%
    KV-CAT 0.536 0.0491 90.2%
    40% Base 0.742 0.0653 90.0%
    KV-CAT 0.259 0.0326 92.4%

    On the same prefix–suffix compression problem, a separate gradient-based method directly optimizes compact KV tensors against dense-model suffix logits. Across keep ratios, KV-CAT reaches comparable suffix-perplexity degradation in fewer optimization steps, with up to a 5× reduction in steps relative to the base model. The paper also reports transfer of the Attention Matching gains to WikiText-103, PG-19, and arXiv: KL and Top-1 improve at every tested ratio, and Δ\DeltaPPL improves in 11 of 12 domain–ratio settings.

  7. Knowl 7 — KV-CAT improves passkey retrieval from optimized compressed contexts

    data/table

    This needle-in-a-haystack evaluation measures exact-match retrieval after compressing a 1,024-token prompt prefix. For each keep ratio, the compact cache is optimized for 100 steps using a reconstruction sequence that repeats the haystack and applies loss to the repeated passage; the final query remains uncompressed. Accuracy is reported over 100 examples per keep ratio. KV-CAT raises mean exact-match accuracy by 6.4 percentage points for QWEN2.5-0.5B and 5.2 points for QWEN2.5-1.5B. Gains are especially visible at moderate budgets: at 50% keep, accuracy rises from 28% to 47% and from 49% to 67%, respectively.

    Keep QWEN2.5-0.5B QWEN2.5-1.5B
    Base KV-CAT Base KV-CAT
    5% 18 17 20 20
    10% 12 15 20 20
    15% 15 16 21 21
    20% 15 13 22 23
    25% 20 22 24 28
    30% 23 34 41 44
    35% 21 32 46 54
    40% 24 38 42 55
    50% 28 47 49 67
    Mean 19.6 26.0 31.7 36.9
  8. Knowl 8 — KV-CAT raises LongBench v2 accuracy after context compression

    data/table

    The reported long-context QA comparison covers seven LongBench v2 subdomains and a total of 221 examples. For each context, a compact cache is optimized with model weights frozen; the question and answer choices are then processed without compression. The table gives accuracy for the base and KV-CAT checkpoints at 10%, 20%, and 50% KV retention. KV-CAT has a higher total score at all three budgets, with the largest relative improvement at 20% keep: 30.8 versus 22.2, approximately a 39% increase.

    10% keep 20% keep 50% keep
    Subdomain Base KV-CAT Base KV-CAT Base KV-CAT
    Academic 18.1 26.6 19.1 27.7 21.3 26.6
    Agent history QA 20.0 30.0 25.0 30.0 30.0 30.0
    Knowledge graph reasoning 33.3 33.3 26.7 26.7 20.0 26.7
    Legal 27.3 48.5 27.3 45.5 27.3 42.4
    Many-shot learning 28.6 19.0 28.6 28.6 38.1 33.3
    New language translation 30.0 30.0 15.0 35.0 15.0 35.0
    Table QA 11.1 27.8 22.2 22.2 38.9 33.3
    Total 22.2 30.3 22.2 30.8 25.3 31.2
  9. Knowl 9 — Dense benchmark accuracy remains close to the base checkpoints

    empirical result

    With masking and post-hoc compression disabled, the KV-CAT checkpoints were evaluated on six multiple-choice benchmarks using 1,000 examples per task. Scores are normalized accuracy percentages. The average changes by +0.7 points for QWEN2.5-0.5B and −0.5 points for QWEN2.5-1.5B, indicating that the reported compression gains do not require a large loss of ordinary uncompressed benchmark performance.

    Model Variant HellaSwag WinoGrande PIQA OpenBookQA ARC-E ARC-C Avg.
    QWEN2.5-0.5B Base 54.6 52.9 70.4 38.2 60.8 32.2 51.5
    KV-CAT 53.6 54.1 72.1 37.4 63.6 32.4 52.2
    QWEN2.5-1.5B Base 68.6 60.5 76.2 40.0 73.3 45.2 60.6
    KV-CAT 66.4 60.5 75.5 40.6 73.9 43.5 60.1
  10. Knowl 10 — KV-CAT adds training cost and has limited evaluation coverage

    limitation

    The method requires continued pretraining with additional routing modules, which introduces nontrivial training overhead and may limit use in low-resource settings. The routing mechanism and objective also add implementation complexity that may make production integration harder. Evaluation covers only a small set of model sizes and tasks; how well the approach generalizes to larger models and other domains remains unresolved.

Coverage note — The broader-impact discussion is omitted because it does not add a technical method, result, or limitation; no substantial contributed theory, methods, empirical findings, or stated limitations were deliberately omitted.

References

  1. 1.Simran Arora and Christopher Ré. Can foundation models help us achieve perfect secrecy? arXiv preprint arXiv:2205.13722, 2022.
  2. 2.Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, et al. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3639–3664, 2025.
  3. 3.Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  4. 4.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. PIQA: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7432–7439, 2020. doi: 10.1609/aaai.v34i05.6239. URL https://ojs.aaai.org/index.php/AAAI/article/view/6239.
  5. 5.Aydar Bulatov, Yuri Kuratov, and Mikhail S. Burtsev. Recurrent memory transformer, 2022.
  6. 6.Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Yucheng Li, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Junjie Hu, and Wen Xiao. PyramidKV: Dynamic kv cache compression based on pyramidal information funneling, 2025.
  7. 7.Rujikorn Charakorn, Edoardo Cetin, Shinnosuke Uesaka, and Robert Tjarko Lange. Doc-to-lora: Learning to instantly internalize contexts. arXiv preprint arXiv:2602.15902, 2026.
  8. 8.Qiguang Chen, Libo Qin, Jinhao Liu, Dengyun Peng, Jiannan Guan, Peng Wang, Mengkang Hu, Yuhang Zhou, Te Gao, and Wanxiang Che. Towards reasoning era: A survey of long chain-of-thought for reasoning large language models. arXiv preprint arXiv:2503.09567, 2025.
  9. 9.Tong Chen, Hao Fang, Patrick Xia, Xiaodong Liu, Benjamin Van Durme, Luke Zettlemoyer, Jianfeng Gao, and Hao Cheng. Generative adapter: Contextualizing language models in parameters with a single forward pass. arXiv preprint arXiv:2411.05877, 2024.
  10. 10.Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. Adapting language models to compress contexts, 2023.
  11. 11.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020.
  12. 12.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018. URL https://arxiv.org/abs/1803.05457.
  13. 13.Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 615–621, New Orleans, Louisiana, 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-2097. URL https://aclanthology.org/N18-2097/.
  14. 14.George Cybenko. Approximation by superpositions of a sigmoidal function. Mathematics of control, signals and systems, 2(4):303–314, 1989.
  15. 15.Yam Eitan. The centered convex body whose marginals have the heaviest tails. arXiv preprint arXiv:2110.14382, 2021.
  16. 16.Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily Liu, Will Tennien, Atri Rudra, James Zou, Azalia Mirhoseini, and Christopher Ré. Cartridges: Lightweight and general-purpose long context representations via self-study, 2025.
  17. 17.Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023.
  18. 18.Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021.
  19. 19.Nathan Habib, Clémentine Fourrier, Hynek Kydlícek, Thomas Wolf, and Lewis Tunstall. Lighteval: A lightweight framework for llm evaluation. https://github.com/huggingface/lighteval, 2023. GitHub repository.
  20. 20.Alex Horn, Ali Kheradmand, and Mukul Prasad. Delta-net: Real-time network verification using atoms. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17), pages 735–749, 2017.
  21. 21.Kurt Hornik. Approximation capabilities of multilayer feedforward networks. Neural networks, 4(2):251–257, 1991.
  22. 22.Sukjun Hwang, Brandon Wang, and Albert Gu. Dynamic chunking for end-to-end hierarchical sequence modeling. arXiv preprint arXiv:2507.07955, 2025.
  23. 23.Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen. Financebench: A new benchmark for financial question answering. arXiv preprint arXiv:2311.11944, 2023.
  24. 24.Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LLMLingua: Compressing prompts for accelerated inference of large language models, 2023.
  25. 25.Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. LongLLMLingua: Accelerating and enhancing llms in long context scenarios via prompt compression, 2024.
  26. 26.Samuel Karlin and William J Studden. Optimal experimental designs. The Annals of Mathematical Statistics, 37(4):783–815, 1966.
  27. 27.Samuel Karlin and William J Studden. Tchebycheff systems: With applications in analysis and statistics. (No Title), 1966.
  28. 28.Samuel Karlin and Zvi Ziegler. Chebyshevian spline functions. Siam Journal on Numerical Analysis, 3(3):514–543, 1966.
  29. 29.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International conference on machine learning, pages 5156–5165. PMLR, 2020.
  30. 30.Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W Lee, Sangdoo Yun, and Hyun Oh Song. Kvzip: Query-agnostic kv cache compression with context reconstruction. arXiv preprint arXiv:2505.23416, 2025.
  31. 31.Junhyuck Kim, Jongho Park, Jaewoong Cho, and Dimitris Papailiopoulos. Lexico: Extreme KV cache compression via sparse coding over universal dictionaries. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 30672–30687, 2025.
  32. 32.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474, 2020.
  33. 33.Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. Compressing context to enhance inference efficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6342–6353, 2023. doi: 10.18653/v1/2023.emnlp-main.391.
  34. 34.Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. SnapKV: Llm knows what you are looking for before generation, 2024.
  35. 35.Zhuoling Li, Xiaogang Xu, Zhenhua Xu, SerNam Lim, and Hengshuang Zhao. Larm: Large auto-regressive model for long-horizon embodied intelligence. arXiv preprint arXiv:2405.17424, 2024.
  36. 36.Yewei Liu, Xiyuan Wang, Yansheng Mao, Yoav Gelbery, Haggai Maron, and Muhan Zhang. Shine: A scalable in-context hypernetwork for mapping context to lora in a single pass. arXiv preprint arXiv:2602.06358, 2026.
  37. 37.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models, 2016.
  38. 38.CA Micchelli and Allan Pinkus. Moment theory for weak chebyshev systems with applications to monosplines, quadrature formulae and best one-sided lˆ1-approximation by spline functions with fixed knots. SIAM Journal on Mathematical Analysis, 8(2):206–230, 1977.
  39. 39.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium, 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1260. URL https://aclanthology.org/D18-1260/.
  40. 40.Jesse Mu, Xiang Lisa Li, and Noah Goodman. Learning to compress prompts with gist tokens, 2023.
  41. 41.Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. Using an llm to help with code understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, pages 1–13, 2024.
  42. 42.Emre Okcular. Context Engineering - Short-Term Memory Management with Sessions from OpenAI Agents SDK, September 2025.
  43. 43.Matanel Oren, Michael Hassid, Nir Yarden, Yossi Adi, and Roy Schwartz. Transformers are multi-state RNNs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 18724–18741, 2024. doi: 10.18653/v1/2024.emnlp-main.1043.
  44. 44.Guilherme Penedo, Hynek Kydlícek, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, Thomas Wolf, et al. The fineweb datasets: Decanting the web for the finest text data at scale. Advances in Neural Information Processing Systems, 37:30811–30849, 2024.
  45. 45.Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint, 2019. URL https://arxiv.org/abs/1911.05507.
  46. 46.Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield. Effective context engineering for ai agents, September 2025.
  47. 47.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An adversarial winograd schema challenge at scale. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732–8740, 2020. doi: 10.1609/aaai.v34i05.6399. URL https://ojs.aaai.org/index.php/AAAI/article/view/6399.
  48. 48.Clayton Sanford, Daniel J Hsu, and Matus Telgarsky. Representational strengths and limitations of transformers. Advances in Neural Information Processing Systems, 36:36677–36707, 2023.
  49. 49.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. Social IQa: Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4463–4473, Hong Kong, China, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1454. URL https://aclanthology.org/D19-1454/.
  50. 50.Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. QUEST: Query-aware sparsity for efficient long-context LLM inference. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 47901–47911, 2024.
  51. 51.Qwen Team. Qwen2.5: A party of foundation models, September 2024. URL https://qwenlm.github.io/blog/qwen2.5/.
  52. 52.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks, 2024.
  53. 53.Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han. Duoattention: Efficient long-context LLM inference with retrieval and streaming heads. In International Conference on Learning Representations, 2025. URL https://arxiv.org/abs/2410.10819.
  54. 54.Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving mamba2 with delta rule. arXiv preprint arXiv:2412.06464, 2024.
  55. 55.Gilad Yehudai, Haim Kaplan, Guy Dar, Royi Rassin, Asma Ghandeharioun, Mor Geva, and Amir Globerson. When can transformers count to n? arXiv preprint arXiv:2407.15160, 2024.
  56. 56.Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola. Deep sets. Advances in neural information processing systems, 30, 2017.
  57. 57.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. Big bird: Transformers for longer sequences. Advances in neural information processing systems, 33:17283–17297, 2020.
  58. 58.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy, 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1472. URL https://aclanthology.org/P19-1472/.
  59. 59.Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen. H2O: Heavy-hitter oracle for efficient generative inference of large language models, 2023.
  60. 60.Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu, and Qianli Ma. Lifelong learning of large language model based agents: A roadmap. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026.
  61. 61.Adam Zweiger, Xinghong Fu, Han Guo, and Yoon Kim. Fast kv compaction via attention matching, 2026.

Citation

MLA
Gelberg, Y., et al. “Training Transformers for KV Cache Compressibility”. arXiv, 2026, http://arxiv.org/abs/2605.05971v2.
APA
Gelberg, Y., Eitan, Y., Bronstein, M., Gal, Y., & Maron, H. (2026). Training Transformers for KV Cache Compressibility. arXiv. http://arxiv.org/abs/2605.05971v2
Chicago
Gelberg, Y., Y. Eitan, M. Bronstein, Y. Gal, and H. Maron. 2026. “Training Transformers for KV Cache Compressibility”. arXiv. http://arxiv.org/abs/2605.05971v2.
Harvard
Gelberg, Y. et al. (2026) “Training Transformers for KV Cache Compressibility”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.05971v2.
Vancouver
1. Gelberg Y, Eitan Y, Bronstein M, Gal Y, Maron H (2026) Training Transformers for KV Cache Compressibility. arXiv

BibTeX

@article{gelberg2026training,
  title = {Training Transformers for KV Cache Compressibility},
  author = {Gelberg, Yoav and Eitan, Yam and Bronstein, Michael and Gal, Yarin and Maron, Haggai},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.05971v2},
  eprint = {2605.05971}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/