keyword
cache reuse
Cache reuse is the practice of storing and repurposing previously computed intermediate representations, particularly key-value states in transformer-based models, across different inference requests, prompt segments, or generation steps. Rather than recalculating attention keys and values for repeated contexts, shared system prompts, or retrieved document chunks during the prefill phase, inference engines retrieve the saved states directly from memory or storage. By bypassing redundant computation over overlapping or pre-existing text, cache reuse significantly reduces processing latency, lowers the time required to produce the first generated token, and optimizes memory bandwidth and resource utilization in large language model serving systems.
2 items

TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, Yaohua Tang
Why you should read this
Proposes a hybrid offline-online framework that precomputes chunk-level key-value caches and stitches them at inference using independent attention and reordered rotary position embeddings, accelerating time-to-first-token in retrieval-augmented generation by up to 9.4x without sacrificing accuracy or requiring architectural modifications.
Current Retrieval-Augmented Generation (RAG) systems concatenate and process numerous retrieved document chunks for prefill which requires a large volume of online computation, therefore leading to significant latency in time-to-first-token (TTFT). To reduce the computation overhead as well as TTFT, we introduce TurboRAG, a hybrid offline-online paradigm that (i) pre-computes chunk-level key-value (KV) caches, (ii) stitches them together at inference time using independent-attention and reordered-RoPE techniques, and (iii) preserves answer quality without changing the model architecture. Our approach is applicable to most existing large language models and their applications without any requirement in modification of models and inference systems. Experimental results across a suite of RAG benchmarks demonstrate that TurboRAG reduces TTFT by up to 9.4x compared to the conventional RAG systems (on an average of 8.6x), but reserving comparable performance to the standard RAG systems.
Added
2026-10-04

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-AI
Why you should read this
Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model that achieves significant reductions in KV cache footprint and prefill costs for long-context agentic workloads.
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens.
Added
2026-09-10
