Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Piotr NawrotAdrian LancuckiMarcin ChochowskiDavid TarjanEdoardo M. Ponti

article2024ICML103 citations

Introduces Dynamic Memory Compression, an approach for retrofitting pre-trained large language models to adaptively compress key-value caches across heads and layers, increasing inference throughput up to 3.7x without requiring extra parameters or degrading downstream accuracy.

Listen

Deploying large language models for real-world generative AI applications is severely limited by inference inefficiency. In modern Transformer architectures, generating responses requires storing intermediate key-value representations in memory for past tokens. This memory footprint scales linearly with sequence length and batch size, rapidly exhausting GPU hardware capacity during long-context processing or high-volume concurrent serving. Existing mitigation methods, such as token-dropping heuristics or fixed grouping approaches, often suffer from severe degradation in downstream task accuracy.

The article introduces and evaluates Dynamic Memory Compression, a method designed to compress key-value memory on the fly during inference. The main objective is to retrofit existing pre-trained models into compressed variants that reduce memory usage and accelerate inference speeds without degrading generation quality or adding extra model parameters.

The researchers retrofitted open-source language models across multiple scales (7-billion, 13-billion, and 70-billion parameters) through continued pre-training on a minimal fraction of data—amounting to roughly 2% to 4% of original training volumes. Using gradient descent with continuous relaxations, the system learns at each decoding step whether to append key-value states to memory or accumulate them with the preceding state via a weighted average. The evaluation measured downstream performance on standard benchmarks for factual knowledge, common-sense reasoning, and code generation, alongside physical hardware latency and throughput metrics on high-performance GPUs.

The findings show that Dynamic Memory Compression successfully preserves baseline model accuracy at up to 4-fold memory compression, occasionally outperforming original baselines due to light continued training. It systematically outperforms common eviction baselines and Grouped Query Attention, demonstrating superior sample efficiency during training. Compounded gains are also achievable: applying a 2-fold compression to a 70-billion parameter model already utilizing 8-fold Grouped Query Attention yielded a combined 16-fold memory compression without performance loss. In hardware testing, this compression enabled larger batch sizes and delivered between 3.4-fold and 3.7-fold increases in serving throughput for 7-billion and 13-billion models on advanced GPUs. Analysis of learned compression behavior revealed that models naturally prefer higher compression ratios in deeper transformer layers.

These results demonstrate that Dynamic Memory Compression can substantially reduce hardware operational costs and carbon footprint while accelerating response times and supporting longer contexts. Because it uses existing parameter dimensions, it serves as an efficient drop-in upgrade for pre-trained architectures. Organizations serving large language models should consider adopting this compression technique to improve GPU utilization and throughput. For optimal results, implementations should leverage memory management tools like PagedAttention that accommodate variable-length cache allocations across attention heads.

Confidence in these findings is high for retrofitting scenarios, though practitioners should note key limitations. The technique currently applies to continuing the training of pre-existing base models; preliminary attempts to train models from scratch with dynamic compression yielded negative results due to instability between representation learning and token boundary segmentation. Additional research into staged training schedules is required before this method can be reliably used for initial pre-training from scratch.

Cover for Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Abstract

Transformers have emerged as the backbone of large language models (LLMs). However, generation remains inefficient due to the need to store in memory a cache of key–value representations for past tokens, whose size scales linearly with the input sequence length and batch size. As a solution, we propose Dynamic Memory Compression (DMC), a method for on-line key–value cache compression at inference time. Most importantly, the model learns to apply different compression ratios in different heads and layers. We retrofit pre-trained LLMs such as Llama 2 (7B, 13B and 70B) into DMC Transformers, achieving up to ~3.7× throughput increase during auto-regressive inference on an NVIDIA H100 GPU. DMC is applied via continued pre-training on a negligible percentage of the original data without adding any extra parameters. We find that DMC preserves the original downstream performance with up to 4× cache compression, outperforming up-trained grouped-query attention (GQA) and key–value eviction policies (H2O, TOVA). GQA and DMC can be even combined to obtain compounded gains. As a result DMC fits longer contexts and larger batches within any given memory budget. We release the DMC code and models at https://github.com/NVIDIA/Megatron-LM/tree/DMC.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Multi-Head Self-Attention
  • 2.2. KV Caching During Inference
  • 3. Method: Dynamic Memory Compression
  • 3.1. Inference
  • 3.2. Training
  • 3.3. Practical Considerations
  • 4. Experimental Setup
  • 5. Results
  • 5.1. Main Results
  • 5.2. Throughput and Latency Measurements
  • 5.3. Per-Head Learned Compression Ratios
  • 6. Related Work
  • 7. Conclusions and Future Work
  • Acknowledgements
  • Impact Statement
  • References
  • Appendix
  • A. Memory-Bound Operations in Transformers
  • B. Replicating the Original Results
  • C. Retrofitting Data
  • D. Analysis of the Compression Schema Learned by DMC
  • E. Similar Per-layer Compression Rates
  • F. Training Ablations
  • G. DMC Ablations
  • H. Masking Implementation Details
  • I. Limitations

Knowls

  1. Knowl 1 — Dynamic Memory Compression cache update

    algorithm

    Dynamic Memory Compression (DMC) adaptively shortens each attention head’s key–value cache during autoregressive inference. For a head, let KK and VV be the current key and value caches of length ll, and let qt,kt,vtq_t,k_t,v_t be the query, key, and value vectors for the newly generated token at time tt. DMC reuses the first scalar coordinate of ktk_t and qtq_t to predict a binary segmentation decision and a continuous importance weight, so it introduces no parameters.

    Input: Current key cache K of length l, value cache V of length l, and current vectors qt, kt, vt
    Output: Updated K and V, with the first coordinates of qt and kt removed from attention
    αt ← round(sigmoid(kt[0]))
    ωt ← sigmoid(qt[0])
    Set kt[0] and qt[0] to zero
    if αt = 1 then
        Accumulate the current token into the last cache slot
        zt ← zt−1 + ωt
        Kl ← (Kl zt−1 + kt ωt) / zt
        Vl ← (Vl zt−1 + vt ωt) / zt
    else
        Append a new cache slot
        zt ← ωt
        K ← [K, kt]
        V ← [V, vt]
    end if
    return K, V, qt, kt

    Here ztz_t is the running sum of importance weights since the most recent append decision. Thus, \alpha_t=1 merges the new key and value into the last cache item, whereas \alpha_t=0 starts a new segment and appends a cache item. For a sequence of nn tokens, the resulting cache length is l=∑t=1n(1−αt)l=\sum_{t=1}^{n}(1-\alpha_t), and the compression ratio is CR=n/l\mathrm{CR}=n/l. The procedure is applied independently to every layer and head, so different heads may retain different numbers of cache items.

  2. Knowl 2 — Differentiable training objective for DMC

    model/method

    DMC is trained end-to-end by replacing its discrete append-or-accumulate decisions with continuous stochastic relaxations. For each token, the decision is sampled as αt∼GumbelSigmoid⁡(kt[0]−c,τ)∈[0,1]\alpha_t\sim\operatorname{GumbelSigmoid}(k_t[0]-c,\tau)\in[0,1], where kt[0]k_t[0] is the first coordinate of the current key, cc is a bias that initially suppresses compression, and τ\tau is the temperature. The importance score is ωt=sigmoid⁡(qt[0]+c)\omega_t=\operatorname{sigmoid}(q_t[0]+c), where qt[0]q_t[0] is the first query coordinate. Low τ\tau makes the relaxed decisions nearly binary.

    For a head, the partially accumulated state for token ii is defined from the current key kik_i, value viv_i, decision αi\alpha_i, and importance ωi\omega_i by

    z0=ω0,kˉ0=k0,vˉ0=v0,z_0=\omega_0,\qquad \bar{k}_0=k_0,\qquad \bar{v}_0=v_0, zi=αizi−1+ωi,kˉi=αizi−1kˉi−1+ωikizi,vˉi=αizi−1vˉi−1+ωivizi.z_i=\alpha_i z_{i-1}+\omega_i,\qquad \bar{k}_i=\frac{\alpha_i z_{i-1}\bar{k}_{i-1}+\omega_i k_i}{z_i},\qquad \bar{v}_i=\frac{\alpha_i z_{i-1}\bar{v}_{i-1}+\omega_i v_i}{z_i}.

    When αi=0\alpha_i=0, the state is the current token’s key and value; when αi=1\alpha_i=1, it is a weighted continuation of the preceding segment. During parallel training, all intermediate states are retained, and an additive attention mask prevents queries from using intermediate states that would be discarded at inference. For a query position ii and key position jj, the mask adds log⁡(1−αj+1)\log(1-\alpha_{j+1}) to the normalized attention score, in addition to the usual causal restriction. Consequently, a nearly binary αj+1=1\alpha_{j+1}=1 blocks that intermediate state with a value approaching −∞-\infty, while αj+1=0\alpha_{j+1}=0 leaves it unmasked.

    Let LL be the number of Transformer layers, HH the number of heads per layer, NN the number of training tokens, and CRtarget\mathrm{CR}_{\mathrm{target}} the desired global compression ratio. DMC uses the one-sided global compression penalty

    ℓCR=1LHNmax⁡(0,∑ℓ=1L∑h=1H∑t=1N(1−αℓht)−LHNCRtarget),\ell_{\mathrm{CR}}=\frac{1}{LHN}\max\left(0,\sum_{\ell=1}^{L}\sum_{h=1}^{H}\sum_{t=1}^{N}(1-\alpha_{\ell ht})-\frac{LHN}{\mathrm{CR}_{\mathrm{target}}}\right),

    where αℓht\alpha_{\ell ht} is the relaxed decision for layer ℓ\ell, head hh, and token tt. The model parameters θ\theta minimize ℓLM+ℓCR\ell_{\mathrm{LM}}+\ell_{\mathrm{CR}}, combining the language-modeling loss with a penalty for retaining more tokens than the target cache budget.

  3. Knowl 3 — Retrofitting schedule and training requirements

    experimental setup

    The authors retrofit pretrained Llama 2 models rather than training DMC from scratch. DMC reuses the first dimensions of the query and key vectors as control signals, so it adds no model parameters. Because zeroing those dimensions initially disrupts attention, an initialization phase first trains the model to disregard them: for training step ss out of S=250S=250, the first coordinates are annealed according to

    qs[0]←qs[0](1−sS),ks[0]←ks[0](1−sS).q_s[0]\leftarrow q_s[0]\left(1-\frac{s}{S}\right),\qquad k_s[0]\leftarrow k_s[0]\left(1-\frac{s}{S}\right).

    This phase processes 1 billion tokens. The main retrofitting phase then linearly increases the global compression ratio from 1×1\times to the target ratio, followed by an 8-billion-token solidification phase at fixed compression. During solidification, the learning rate follows a cosine schedule down to 10% of its initial value. The standard continued-pretraining schedules use 24B, 48B, and 72B tokens for target ratios of 2×2\times, 3×3\times, and 4×4\times, respectively; Llama 2 7B and 13B are trained up to 4×4\times, while Llama 2 70B is trained only up to an additional 2×2\times because it already has 8×8\times GQA.

    The training uses AdamW with β1=0.9\beta_1=0.9, β2=0.95\beta_2=0.95, ϵ=10−5\epsilon=10^{-5}, weight decay 0.10.1, gradient clipping at 1.01.0, batch size 10241024, and sequence length 40964096. The constant learning rate is 3×10−53\times10^{-5} for 7B and 13B and 1.5×10−51.5\times10^{-5} for 70B. The Gumbel-sigmoid bias is c=5c=5, giving initial values approximately αt=0.0067\alpha_t=0.0067 and ωt=0.9933\omega_t=0.9933; the temperature is fixed at 0.10.1. A window size of 1212 is used for the partial-accumulation approximation.

    Retrofitting uses approximately 2% of the original pretraining data for 2×2\times compression and approximately 4% for 4×4\times compression. Ablations show that gradually increasing the target ratio is important: setting the target compression immediately causes an early perplexity spike and lowers downstream accuracy even when perplexity later recovers. The gradual schedule also permits usable checkpoints at intermediate compression ratios.

  4. Knowl 4 — Downstream performance under cache compression

    data/table

    DMC preserves the downstream quality of Llama 2 at 2×2\times and 4×4\times cache compression, whereas GQA and cache-eviction baselines generally lose accuracy as compression increases. MMLU is accuracy on a 5-shot factuality evaluation, CS-QA is accuracy averaged over six 0-shot commonsense question-answering tasks, and HumanEval is Python-generation pass@1. The 70B model already has 8×8\times GQA; its DMC result adds 2×2\times compression for a total 16×16\times reduction relative to a model with neither method.

    Scale Method CR MMLU CS-QA HumanEval
    7B Original 1×1\times 44.6 70.5 14.0
    7B GQA 2×2\times 39.8 68.9 12.8
    7B H2O 2×2\times 45.2 67.5 9.8
    7B TOVA 2×2\times 44.9 70.0 6.1
    7B DMC 2×2\times 45.2 70.8 15.2
    7B GQA 4×4\times 34.7 68.3 14.0
    7B H2O 4×4\times 41.1 56.8 4.9
    7B TOVA 4×4\times 43.4 64.2 1.8
    7B DMC 4×4\times 43.9 70.2 16.5
    13B Original 1×1\times 54.5 73.5 17.5
    13B GQA 2×2\times 50.2 72.7 15.9
    13B H2O 2×2\times 54.1 70.3 16.5
    13B TOVA 2×2\times 54.4 72.8 12.2
    13B DMC 2×2\times 54.8 74.2 20.7
    13B GQA 4×4\times 48.6 72.2 16.5
    13B H2O 4×4\times 50.7 60.0 8.5
    13B TOVA 4×4\times 53.3 67.8 2.4
    13B DMC 4×4\times 54.2 73.2 22.0
    70B Original + GQA 8×8\times 68.8 78.0 29.6
    70B H2O 16×16\times 68.7 74.1 18.3
    70B TOVA 16×16\times 68.1 77.6 29.9
    70B DMC + GQA 16×16\times 68.8 77.9 29.9

    At 2×2\times, DMC sometimes improves over the original model, including MMLU and CS-QA for 7B and 13B. At 4×4\times, its largest reported degradations relative to the corresponding original model are only 0.70.7 MMLU points and 0.30.3 CS-QA points for 7B, and 0.30.3 points on both metrics for 13B; HumanEval improves for both model sizes.

  5. Knowl 5 — Inference throughput and latency gains

    empirical result

    DMC’s smaller key–value caches produce practical inference gains, not just lower memory usage. The measurements use Megatron-LM in bfloat16, with Llama 2 7B and 13B on one NVIDIA A100 80GB SXM or H100 SXM GPU and Llama 2 70B on two same-type GPUs with tensor parallelism. Each run processes 2,000 prompt tokens and generates 2,000 additional tokens; throughput is averaged over the final 1,000 generated tokens, with batch size increased to the largest value fitting in GPU memory.

    Relative to the uncompressed model, 2×2\times DMC increases throughput by more than 1.8×1.8\times for 7B, 13B, and 70B on both tested GPU types. 4×4\times DMC increases throughput by 3.4×3.4\times–3.7×3.7\times for the 7B and 13B models. These gains are close to the theoretical cache-size benefit because autoregressive attention is memory-bound, and the saved memory also permits larger batches.

    For next-token latency, when each model is run at its own maximum batch size, the cost begins scaling approximately linearly with context length after about 2,200 generated tokens as reading the cache from high-bandwidth memory becomes dominant. If DMC 4× uses the same batch size as the original model instead of a larger maximum batch, its reduced cache footprint substantially lowers latency for longer contexts.

  6. Knowl 6 — Advantages over GQA and cache eviction

    empirical result

    At equal compression ratios, DMC outperforms up-trained grouped-query attention (GQA) on both the 7B and 13B Llama 2 models. On MMLU, DMC’s advantage over GQA is +5.4+5.4 points at 2×2\times and +9.2+9.2 points at 4×4\times for 7B, and +4.6+4.6 points at 2×2\times and +5.6+5.6 points at 4×4\times for 13B. DMC also shows comparable gains over GQA on CS-QA and HumanEval. In sample-efficiency experiments, DMC reaches its reported 2×2\times performance after about 8,000 fine-tuning steps, whereas GQA still does not match it after more than 17,000 steps.

    H2O and TOVA perform token eviction using post-softmax attention scores. This can require materializing an n×nn\times n attention-score tensor for sequence length nn, making them incompatible with efficient FlashAttention-style execution and potentially slowing inference despite reducing cache size. Their downstream performance also falls sharply at high compression, especially on HumanEval.

    DMC composes with GQA rather than replacing it. Llama 2 70B, which has an 8×8\times GQA cache reduction, retains essentially unchanged downstream performance after an additional 2×2\times DMC reduction, yielding a total 16×16\times cache compression.

  7. Knowl 7 — Learned compression is heterogeneous across layers and heads

    empirical result

    DMC does not impose a predefined allocation of compression across layers or heads. Instead, each attention head learns its own compression ratio from the global compression objective. Across Llama 2 7B, 13B, and 70B, the dominant pattern is to compress deeper layers: layers above approximately 16 in 7B, 22 in 13B, and 44 in 70B are most heavily compressed. The final few layers are compressed somewhat less than the preceding deep layers. At 4×4\times global compression, some heads also learn unusually high ratios in the first few layers, although this is counterproductive because early token representations are not yet sufficiently contextualized.

    The layerwise allocation evolves during training: compression first concentrates in deeper layers, then spreads to earlier and sometimes noncontiguous intermediate ranges as the global target increases. For a fixed trained model, the achieved compression ratio increases approximately logarithmically with total sequence length rather than remaining uniform across lengths. Average decision values are approximately independent of absolute token position, indicating that DMC does not simply follow a fixed positional compression pattern.

    Some learned groupings align with linguistic structure. For example, a Llama 2 13B head at compression ratio 4×4\times merges subwords back into words and occasionally groups semantic phrases such as “1 9 th century,” “5 0 percent,” and “a week back later.” Many other heads are not linguistically interpretable, while higher layers generally merge longer token sequences.

  8. Knowl 8 — Efficient storage and windowed accumulation

    model/method

    Because DMC learns compression independently for every layer and head, the resulting key–value sequences have variable lengths. The authors store these sequences with PagedAttention, allocating memory pages on demand separately for each head instead of padding every head to a common length. This permits the learned heterogeneous compression pattern with little storage overhead and supports efficient FlashAttention-compatible inference.

    For a sequence of length nn, computing the exact partially accumulated state sequentially has an O(n)O(n) dependency span. To reduce this cost, DMC uses a windowed approximation: at each position, the accumulation recurrence is evaluated over only the most recent ww tokens, with w=12w=12 in the experiments. With at least nn parallel execution threads, the dependency span becomes O(w)O(w) rather than O(n)O(n). The same approximation accelerates prompt processing during inference. Its limitation is that heads whose segments accumulate more than ww tokens must retain the sliding window in cache.

  9. Knowl 9 — Constrained per-layer compression trades accuracy for simpler storage

    empirical result

    The constrained variant DMC-C adds a penalty that encourages all heads within a layer to have similar compression ratios. This makes the cache easier to store as padded tensors, but it removes some of DMC’s ability to assign compression selectively across heads. The following results report MMLU, CS-QA, and HumanEval in that order:

    • For 7B at 2×2\times, DMC scores 45.245.2, 70.870.8, and 15.215.2, whereas DMC-C scores 45.545.5, 70.670.6, and 14.614.6; at 4×4\times, DMC scores 43.943.9, 70.270.2, and 16.516.5, whereas DMC-C scores 38.238.2, 69.669.6, and 14.614.6.
    • For 13B at 2×2\times, DMC scores 54.854.8, 74.274.2, and 20.720.7, whereas DMC-C scores 54.854.8, 73.973.9, and 18.318.3; at 4×4\times, DMC scores 54.254.2, 73.273.2, and 22.022.0, whereas DMC-C scores 52.452.4, 72.972.9, and 18.318.3.
    • For 70B with total compression 16×16\times, DMC scores 68.868.8, 77.977.9, and 29.929.9, whereas DMC-C scores 67.467.4, 78.278.2, and 31.131.1.

    The largest loss occurs for 7B at 4×4\times, where DMC-C is 6.46.4 MMLU points below the original model, while unconstrained DMC remains close to the original. Thus, when custom variable-length attention storage is available, the unconstrained head-specific DMC is preferred.

  10. Knowl 10 — Limitation of training DMC from scratch

    limitation

    The paper’s successful results concern retrofitting pretrained LLMs. Preliminary experiments that trained LLMs with DMC from scratch performed worse than the corresponding GQA training curve. The authors attribute this result tentatively to a mutual dependency: early, low-quality token representations make append-or-accumulate boundary decisions unreliable, while incorrect boundaries in turn produce poor token representations. They suggest that an alternating procedure, such as an expectation-maximization-style alternation between representation modeling and segmentation, may be needed to break this feedback loop.

Coverage note — Detailed secondary ablations on short versus long compression schedules, shared hard decisions, uniform accumulation weights, and fixed pooling were omitted because they support the main design choices already captured by the training and constrained-compression results.

References

  1. 1.Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S. GQA: training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023. URL https://doi.org/10.18653/v1/2023.emnlp-main.298.
  2. 2.Anagnostidis, S., Pavllo, D., Biggio, L., Noci, L., Lucchi, A., and Hofmann, T. Dynamic context pruning for efficient and interpretable autoregressive transformers. In Advances in Neural Information Processing Systems 36, 2023.
  3. 3.Bahdanau, D., Cho, K., and Bengio, Y. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, 2015. URL http://arxiv.org/abs/1409.0473.
  4. 4.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer. ArXiv, abs/2004.05150, 2020. URL https://arxiv.org/abs/2004.05150.
  5. 5.Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020. URL https://doi.org/10.1609/aaai.v34i05.6239.
  6. 6.Bolya, D., Fu, C., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J. Token merging: Your vit but faster. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/pdf?id=JroZRaRw7Eu.
  7. 7.Chen, M., Tworek, J., Jun, H., Yuan, Q., Ponde, H., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D. W., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Babuschkin, I., Balaji, S., Jain, S., Carr, A., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M. M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. ArXiv, abs/2107.03374, 2021.
  8. 8.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. ArXiv, abs/1904.10509, 2019.
  9. 9.Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlós, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A. R. Rethinking attention with performers. In 9th International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Ua6zuk0WRH.
  10. 10.Clark, C., Lee, K., Chang, M., Kwiatkowski, T., Collins, M., and Toutanova, K. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, 2019. URL https://doi.org/10.18653/v1/n19-1300.
  11. 11.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try ARC, the AI2 reasoning challenge. ArXiv, abs/1803.05457, 2018.
  12. 12.Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems, volume 35, 2022.
  13. 13.DeepSeek-AI. DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model, 2024.
  14. 14.Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Re, C. Hungry Hungry Hippos: Towards language modeling with state space models. In The International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=COZDy0WYGg.
  15. 15.Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J. Model tells you what to discard: Adaptive kv cache compression for llms. ArXiv, abs/2310.01801, 2023.
  16. 16.Gu, A. and Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. ArXiv, abs/2312.00752, 2023.
  17. 17.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ.
  18. 18.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In 8th International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rygGQyrFvH.
  19. 19.Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A. KVQuant: Towards 10 million context length llm inference with kv cache quantization. ArXiv, abs/2401.18079, 2024.
  20. 20.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7B. ArXiv, abs/2310.06825, 2023.
  21. 21.Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, 2023. URL https://doi.org/10.1145/3600006.3613165.
  22. 22.Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N. Generating wikipedia by summarizing long sequences. In 6th International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=Hyg0vbWC-.
  23. 23.Liu, Z., Desai, A., Liao, F., Wang, W., Xie, V., Xu, Z., Kyrillidis, A., and Shrivastava, A. Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time. In Advances in Neural Information Processing Systems 36, 2023a.
  24. 24.Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V. LLM-QAT: Data-free quantization aware training for large language models. ArXiv, abs/2305.17888, 2023b.
  25. 25.Liu, Z., Yuan, J., Jin, H., Zhong, S., Xu, Z., Braverman, V., Chen, B., and Hu, X. KIVI: A tuning-free asymmetric 2bit quantization for KV cache. ArXiv, abs/2402.02750, 2024.
  26. 26.Mu, J., Li, X., and Goodman, N. D. Learning to compress prompts with gist tokens. In Advances in Neural Information Processing Systems 36, 2023.
  27. 27.Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V. A., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., and Zaharia, M. A. Efficient large-scale language model training on gpu clusters using megatron-lm. SC21: International Conference for High Performance Computing, Networking, Storage and Analysis, 2021. URL https://doi.org/10.1145/3458817.3476209.
  28. 28.Nawrot, P., Chorowski, J., Łancucki, A., and Ponti, E. M. Efficient transformers with dynamic token pooling. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 2023. URL https://aclanthology.org/2023.acl-long.353.
  29. 29.Oren, M., Hassid, M., Adi, Y., and Schwartz, R. Transformers are multi-state RNNs. ArXiv, abs/2401.06104, 2024.
  30. 30.Patterson, D. A., Gonzalez, J., Le, Q. V., Liang, C., Munguía, L.-M., Rothchild, D., So, D. R., Texier, M., and Dean, J. Carbon emissions and large neural network training. ArXiv, abs/2104.10350, 2021.
  31. 31.Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J. Efficiently scaling transformer inference. ArXiv, abs/2211.05102, 2022. URL https://doi.org/10.48550/arXiv.2211.05102.
  32. 32.Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P. Compressive transformers for long-range sequence modelling. In 8th International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SylKikSYDH.
  33. 33.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020. URL https://doi.org/10.1609/aaai.v34i05.6399.
  34. 34.Shazeer, N. M. Fast transformer decoding: One write-head is all you need. ArXiv, abs/1911.02150, 2019. URL http://arxiv.org/abs/1911.02150.
  35. 35.Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Fu, D. Y., Xie, Z., Chen, B., Barrett, C. W., Gonzalez, J., Liang, P., Ré, C., Stoica, I., and Zhang, C. High-throughput generative inference of large language models with a single gpu. In International Conference on Machine Learning, 2023. URL https://proceedings.mlr.press/v202/sheng23a.html.
  36. 36.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, R., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models. ArXiv, abs/2307.09288, 2023. URL https://arxiv.org/abs/2307.09288.
  37. 37.Treviso, M. V., Ji, T., Lee, J.-U., van Aken, B., Cao, Q., Ciosici, M. R., Hassid, M., Heafield, K., Hooker, S., Martins, P. H., Martins, A. F. T., Milder, P., Raffel, C., Simpson, E., Slonim, N., Balasubramanian, N., Derczynski, L., and Schwartz, R. Efficient methods for natural language processing: A survey. Transactions of the Association for Computational Linguistics, 11, 2022.
  38. 38.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  39. 39.Wang, H., Zhang, Z., and Han, S. Spatten: Efficient sparse attention architecture with cascade token and head pruning. 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2020. URL https://doi.org/10.1109/HPCA51647.2021.00018.
  40. 40.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. URL https://aclanthology.org/P19-1472.
  41. 41.Zhang, B., Xiong, D., and Su, J. Accelerating neural transformer via an average attention network. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018. URL https://aclanthology.org/P18-1166.
  42. 42.Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., Wang, Z. A., and Chen, B. H2o: Heavy-hitter oracle for efficient generative inference of large language models. In Advances in Neural Information Processing Systems 36, 2023.

Citation

MLA
Nawrot, P., et al. “Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference”. Proceedings of the 41st International Conference on Machine Learning (2024) 37396-37412, 2024, http://arxiv.org/abs/2403.09636v2.
APA
Nawrot, P., Łańcucki, A., Chochowski, M., Tarjan, D., & Ponti, E. M. (2024). Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference. Proceedings of the 41st International Conference on Machine Learning (2024) 37396-37412. http://arxiv.org/abs/2403.09636v2
Chicago
Nawrot, P., A. Łańcucki, M. Chochowski, D. Tarjan, and E. M. Ponti. 2024. “Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference”. Proceedings of the 41st International Conference on Machine Learning (2024) 37396-37412. http://arxiv.org/abs/2403.09636v2.
Harvard
Nawrot, P. et al. (2024) “Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference”, Proceedings of the 41st International Conference on Machine Learning (2024) 37396-37412 [Preprint]. Available at: http://arxiv.org/abs/2403.09636v2.
Vancouver
1. Nawrot P, Łańcucki A, Chochowski M, Tarjan D, Ponti EM (2024) Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference. Proceedings of the 41st International Conference on Machine Learning (2024) 37396-37412

BibTeX

@article{nawrot2024dynamic,
  title = {Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference},
  author = {Nawrot, Piotr and Łańcucki, Adrian and Chochowski, Marcin and Tarjan, David and Ponti, Edoardo M.},
  year = {2024},
  journal = {Proceedings of the 41st International Conference on Machine Learning (2024) 37396-37412},
  url = {http://arxiv.org/abs/2403.09636v2},
  eprint = {2403.09636}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/