Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference
Piotr NawrotAdrian LancuckiMarcin ChochowskiDavid TarjanEdoardo M. Ponti
Introduces Dynamic Memory Compression, an approach for retrofitting pre-trained large language models to adaptively compress key-value caches across heads and layers, increasing inference throughput up to 3.7x without requiring extra parameters or degrading downstream accuracy.
Deploying large language models for real-world generative AI applications is severely limited by inference inefficiency. In modern Transformer architectures, generating responses requires storing intermediate key-value representations in memory for past tokens. This memory footprint scales linearly with sequence length and batch size, rapidly exhausting GPU hardware capacity during long-context processing or high-volume concurrent serving. Existing mitigation methods, such as token-dropping heuristics or fixed grouping approaches, often suffer from severe degradation in downstream task accuracy.
The article introduces and evaluates Dynamic Memory Compression, a method designed to compress key-value memory on the fly during inference. The main objective is to retrofit existing pre-trained models into compressed variants that reduce memory usage and accelerate inference speeds without degrading generation quality or adding extra model parameters.
The researchers retrofitted open-source language models across multiple scales (7-billion, 13-billion, and 70-billion parameters) through continued pre-training on a minimal fraction of data—amounting to roughly 2% to 4% of original training volumes. Using gradient descent with continuous relaxations, the system learns at each decoding step whether to append key-value states to memory or accumulate them with the preceding state via a weighted average. The evaluation measured downstream performance on standard benchmarks for factual knowledge, common-sense reasoning, and code generation, alongside physical hardware latency and throughput metrics on high-performance GPUs.
The findings show that Dynamic Memory Compression successfully preserves baseline model accuracy at up to 4-fold memory compression, occasionally outperforming original baselines due to light continued training. It systematically outperforms common eviction baselines and Grouped Query Attention, demonstrating superior sample efficiency during training. Compounded gains are also achievable: applying a 2-fold compression to a 70-billion parameter model already utilizing 8-fold Grouped Query Attention yielded a combined 16-fold memory compression without performance loss. In hardware testing, this compression enabled larger batch sizes and delivered between 3.4-fold and 3.7-fold increases in serving throughput for 7-billion and 13-billion models on advanced GPUs. Analysis of learned compression behavior revealed that models naturally prefer higher compression ratios in deeper transformer layers.
These results demonstrate that Dynamic Memory Compression can substantially reduce hardware operational costs and carbon footprint while accelerating response times and supporting longer contexts. Because it uses existing parameter dimensions, it serves as an efficient drop-in upgrade for pre-trained architectures. Organizations serving large language models should consider adopting this compression technique to improve GPU utilization and throughput. For optimal results, implementations should leverage memory management tools like PagedAttention that accommodate variable-length cache allocations across attention heads.
Confidence in these findings is high for retrofitting scenarios, though practitioners should note key limitations. The technique currently applies to continuing the training of pre-existing base models; preliminary attempts to train models from scratch with dynamic compression yielded negative results due to instability between representation learning and token boundary segmentation. Additional research into staged training schedules is required before this method can be reliably used for initial pre-training from scratch.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). This work introduces PagedAttention and efficient KV cache management systems that Dynamic Memory Compression explicitly relies on to support variable-length cache allocations during serving.
- Paper: GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints, Joshua Ainslie et al. (2023). It establishes Grouped Query Attention (GQA) and checkpoint adaptation methods, serving as the primary structural baseline and compounding foundation evaluated in Dynamic Memory Compression.
- Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). This seminal paper introduces Multi-Query Attention to resolve key-value memory bandwidth bottlenecks during autoregressive decoding, providing the foundational architectural context for KV compression.
- Paper: Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers, Sotiris Anagnostidis et al. (2023). It demonstrates dynamic context pruning to drop uninformative tokens from the KV cache during autoregressive generation, introducing the learnable cache reduction paradigm that Dynamic Memory Compression improves upon.
- Paper: PolarQuant: Quantizing KV Caches with Polar Transformation, Insu Han et al. (2025). It builds on KV cache footprint reduction for long-context LLMs by introducing a polar transformation framework for low-bit cache quantization.
- Paper: ThinK: Thinner Key Cache by Query-Driven Pruning, Yuhui Xu et al. (2025). It extends KV cache compression by demonstrating query-driven channel pruning across key dimensions to complement sequence-level and token-merging methods.
- Paper: QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead, Amir Zandieh et al. (2025). It continues the advancement of online KV cache compression by proposing zero-overhead 1-bit quantization via randomized projections.
- Paper: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, DeepSeek-AI (2026). It advances extreme KV cache compression techniques to massive scale using cross-layer reuse and low-precision state storage in long-horizon LLMs.
- Paper: Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning, Heng Wang et al. (2026). It critically evaluates learned and heuristic KV cache reduction policies against simple baseline strategies during long-trace LLM reasoning.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). This survey contextualizes Dynamic Memory Compression within the broader landscape of modern efficient sequence modeling and KV cache optimization architectures.
