Built independently by an author, for readers. Read the story and support ChapterPal

keyword

cache reuse

Cache reuse is the practice of storing and repurposing previously computed intermediate representations, particularly key-value states in transformer-based models, across different inference requests, prompt segments, or generation steps. Rather than recalculating attention keys and values for repeated contexts, shared system prompts, or retrieved document chunks during the prefill phase, inference engines retrieve the saved states directly from memory or storage. By bypassing redundant computation over overlapping or pre-existing text, cache reuse significantly reduces processing latency, lowers the time required to produce the first generated token, and optimizes memory bandwidth and resource utilization in large language model serving systems.

2 items

TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text

TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text

Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, Yaohua Tang

OrganizationsMThreads, Inc.

Why you should read this

Proposes a hybrid offline-online framework that precomputes chunk-level key-value caches and stitches them at inference using independent attention and reordered rotary position embeddings, accelerating time-to-first-token in retrieval-augmented generation by up to 9.4x without sacrificing accuracy or requiring architectural modifications.

Current Retrieval-Augmented Generation (RAG) systems concatenate and process numerous retrieved document chunks for prefill which requires a large volume of online computation, therefore leading to significant latency in time-to-first-token (TTFT). To reduce the computation overhead as well as TTFT, we introduce TurboRAG, a hybrid offline-online paradigm that (i) pre-computes chunk-level key-value (KV) caches, (ii) stitches them together at inference time using independent-attention and reordered-RoPE techniques, and (iii) preserves answer quality without changing the model architecture. Our approach is applicable to most existing large language models and their applications without any requirement in modification of models and inference systems. Experimental results across a suite of RAG benchmarks demonstrate that TurboRAG reduces TTFT by up to 9.4x compared to the conventional RAG systems (on an average of 8.6x), but reserving comparable performance to the standard RAG systems.

Added

2026-10-04