TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
Songshuo LuHua WangYutian RongZhi ChenYaohua Tang
Proposes a hybrid offline-online framework that precomputes chunk-level key-value caches and stitches them at inference using independent attention and reordered rotary position embeddings, accelerating time-to-first-token in retrieval-augmented generation by up to 9.4x without sacrificing accuracy or requiring architectural modifications.
Deploying retrieval-augmented generation systems in real-world applications is often constrained by high latency during the initial response phase, known as time-to-first-token. Under the standard approach, systems concatenate retrieved text passages and compute their internal representations—known as key-value caches—online on every user query. This process leads to redundant computation, quadratic processing delays as document length expands, and heavy memory usage that limits system throughput.
The article introduces and evaluates TurboRAG, a hybrid inference framework designed to eliminate online document processing overhead by precomputing chunk-level caches offline and stitching them dynamically during queries without altering the underlying model architecture.
To establish this framework, the authors leveraged empirical observations showing that attention between separate retrieved documents is inherently sparse and that modern positional encodings depend primarily on relative offsets. The approach precomputes and stores key-value caches for isolated text passages offline. During online inference, the system retrieves the precomputed caches and stitches them using an independent-attention mask and reordered position indices. The authors evaluated the approach primarily using fine-tuned open-source language models (including 7B to 72B parameter configurations) across multiple standard question-answering benchmarks and general reasoning evaluation suites.
The evaluation yielded several key findings. First, TurboRAG achieved an average 8.6-fold reduction in time-to-first-token on multi-document question-answering benchmarks, reaching peak speedups of up to 9.4-fold on long contexts and over 2.4-fold on shorter documents. Second, online computational resource consumption dropped by approximately 98.5% compared to standard systems, allowing larger batch sizes and higher throughput on constrained hardware. Third, response accuracy remained comparable to standard systems, with the reordered-position configuration maintaining performance within 1% of standard baselines after fine-tuning. Finally, general regression testing showed no degradation across reasoning, coding, or dialogue capabilities.
These findings indicate that large language models do not require full online cross-attention between reference documents to generate accurate answers. Organizations can significantly reduce infrastructure operating costs and response latency for customer-facing applications and edge deployments by shifting computation to one-time offline precomputation. The authors note that the financial cost of extra disk storage is substantially lower than the compute resources typically needed to achieve sub-second response times.
Organizations operating latency-sensitive knowledge bases should consider piloting offline cache reuse strategies. For immediate deployment, the reordered-position scheme is recommended due to its minimal accuracy loss even without fine-tuning, though incorporating lightweight supervised fine-tuning provides optimal performance. Engineering teams must also plan for secondary technical requirements, including high-throughput disk storage and cache invalidation strategies for updating dynamic content.
While the results demonstrate robust performance across multiple benchmarks and model sizes, two main constraints remain. Storing precomputed caches introduces disk storage and memory transfer overhead, making compression techniques an important area for future integration. Additionally, while the approach is broadly applicable to rotary-position-based models, achieving optimal accuracy currently relies on lightweight model fine-tuning.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Lewis et al. establish the retrieve-and-generate pipeline TurboRAG accelerates, making its offline cache reuse strategy easier to place in the original RAG setting.
No sufficiently relevant recommendations were found.
