Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time
Zichang LiuJue WangTri DaoTianyi ZhouBinhang YuanZhao SongAnshumali ShrivastavaCe ZhangYuandong TianChristopher Ré
Introduces DejaVu, a hardware-aware system that dynamically predicts input-dependent contextual sparsity to cut the inference latency of 175-billion-parameter language models by more than half without retraining or sacrificing accuracy.
Large language models deliver state-of-the-art performance across diverse tasks, but their massive parameter scale creates substantial computational and latency bottlenecks during inference. In latency-sensitive generation settings, loading billions of parameters from memory dominates runtime. Prior model compression techniques—such as static weight pruning, task-specific pruning, or zero-shot unstructured pruning—often necessitate expensive retraining, degrade in-context learning capabilities, or fail to produce actual wall-clock speedups on modern hardware.
The article evaluates whether pre-trained large language models exhibit dynamic, input-dependent contextual sparsity that can be accurately predicted and exploited to accelerate inference without modifying model weights or degrading accuracy. To leverage this, the authors develop DEJAVU, an inference acceleration framework featuring lightweight sparsity predictors, asynchronous cross-layer execution, and hardware-aware graphics processing unit implementations.
The investigation combines empirical evaluations across multiple open-source models—primarily OPT models ranging up to 175 billion parameters, as well as the BLOOM architecture—with theoretical analyses of residual connections and self-attention clustering dynamics. The researchers evaluated model accuracy on standard language modeling benchmarks (WikiText and C4) and seven zero-shot and few-shot downstream reasoning tasks using evaluation suites on eight NVIDIA A100 graphics processing units.
The article demonstrates that contextual sparsity naturally exists at high rates in dense models: on average, individual inputs activate only about 20% of attention heads and 5% of multilayer perceptron parameters, yielding an overall structured parameter reduction of approximately 85%. Because token representations evolve slowly across layers due to strong residual connections, small neural network predictors can asynchronously forecast which parameters are required for upcoming layers, eliminating sequential prediction overhead. In end-to-end evaluations on the 175-billion-parameter OPT model, DEJAVU achieved over a 2× latency speedup compared to NVIDIA FasterTransformer and up to a 6× speedup compared to Hugging Face implementations at batch size 1, maintaining full baseline accuracy up to 75% sparsity while demonstrating strong compatibility with 4-bit weight quantization.
These findings indicate that large models do not require full dense activation during inference, providing a viable path to substantially lower serving costs and improve response times for interactive artificial intelligence services. By achieving wall-clock acceleration without retraining or losing in-context reasoning abilities, this approach overcomes the primary practical hurdles that previously limited dynamic model pruning on modern accelerator hardware.
Organizations serving large language models should evaluate the integration of contextual sparsity systems into their inference pipelines to reduce operational costs and latency. Systems teams should adopt fused GPU memory kernels and asynchronous prediction architectures to capitalize on hardware memory hierarchies. Future efforts should focus on deploying contextual sparsity in high-throughput, large-batch environments and exploring model depth optimizations such as dynamic layer skipping.
The primary limitations include lower sparsity savings under large batch sizes—where the union of activated parameters across disparate queries increases—and the reliance on custom hardware kernels optimized for memory bandwidth. Nevertheless, the findings offer high empirical and theoretical confidence for small-batch, latency-critical inference workloads.
- Paper: MoEfication: Transformer Feed-forward Layers are Mixtures of Experts, Zhengyan Zhang et al. (2022). It demonstrates that transformer feed-forward layers exhibit naturally sparse, input-dependent activation patterns that can be routed dynamically, providing the empirical foundation for DejaVu's contextual sparsity.
- Paper: DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale, Reza Yazdani Aminabadi et al. (2022). It establishes key high-performance system and kernel optimization principles for large-scale transformer inference that DejaVu directly compares against and builds upon.
- Paper: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, William Fedus et al. (2022). It introduces foundational concepts of dynamic sparse routing per token to reduce active computation during transformer execution.
- Paper: Are Sixteen Heads Really Better than One?, Paul Michel et al. (2019). It provides foundational empirical evidence that individual attention heads can be dynamically bypassed at inference time with minimal performance loss.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). It introduces IO-aware hardware execution principles that underpin modern low-latency transformer inference and block-level sparsity optimizations.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). It explores post-training sparsity at the scale of 175B parameter models, framing the challenge of achieving wall-clock speedup without retraining that DejaVu tackles.
- Paper: Mixture-of-Depths: Dynamically allocating compute in transformer-based language models, David Raposo et al. (2024). It extends dynamic conditional computation from layer-internal MLP/attention subsets to dynamic layer skipping and depth routing per token.
- Paper: Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference, Dhruv Deshmukh et al. (2025). It builds on dynamic inference-time sparsity by predicting and reusing anchor sparse attention patterns across transformer layers without retraining.
- Paper: Language Models Can Control Their Own Attention, Namgyu Ho et al. (2026). It advances dynamic inference sparsity by letting models natively control their attention masks and KV-cache loads during decoding.
- Paper: LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing, Wen Zan et al. (2026). It extends hardware-aligned dynamic sparse attention to massive long-context models using hierarchical cross-layer indexing.
- Paper: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention, Jingyang Yuan et al. (2025). It explores natively trainable, hardware-aligned sparse attention to achieve high inference efficiency across extended contexts.
- Paper: Sparser, Faster, Lighter Transformer Language Models, Edoardo Cetin et al. (2026). It builds upon high activation sparsity in feed-forward layers with custom GPU CUDA kernels to maximize wall-clock inference speedups.
