TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text

Songshuo LuHua WangYutian RongZhi ChenYaohua Tang

article2025EMNLP80 citations

Proposes a hybrid offline-online framework that precomputes chunk-level key-value caches and stitches them at inference using independent attention and reordered rotary position embeddings, accelerating time-to-first-token in retrieval-augmented generation by up to 9.4x without sacrificing accuracy or requiring architectural modifications.

Listen

Deploying retrieval-augmented generation systems in real-world applications is often constrained by high latency during the initial response phase, known as time-to-first-token. Under the standard approach, systems concatenate retrieved text passages and compute their internal representations—known as key-value caches—online on every user query. This process leads to redundant computation, quadratic processing delays as document length expands, and heavy memory usage that limits system throughput.

The article introduces and evaluates TurboRAG, a hybrid inference framework designed to eliminate online document processing overhead by precomputing chunk-level caches offline and stitching them dynamically during queries without altering the underlying model architecture.

To establish this framework, the authors leveraged empirical observations showing that attention between separate retrieved documents is inherently sparse and that modern positional encodings depend primarily on relative offsets. The approach precomputes and stores key-value caches for isolated text passages offline. During online inference, the system retrieves the precomputed caches and stitches them using an independent-attention mask and reordered position indices. The authors evaluated the approach primarily using fine-tuned open-source language models (including 7B to 72B parameter configurations) across multiple standard question-answering benchmarks and general reasoning evaluation suites.

The evaluation yielded several key findings. First, TurboRAG achieved an average 8.6-fold reduction in time-to-first-token on multi-document question-answering benchmarks, reaching peak speedups of up to 9.4-fold on long contexts and over 2.4-fold on shorter documents. Second, online computational resource consumption dropped by approximately 98.5% compared to standard systems, allowing larger batch sizes and higher throughput on constrained hardware. Third, response accuracy remained comparable to standard systems, with the reordered-position configuration maintaining performance within 1% of standard baselines after fine-tuning. Finally, general regression testing showed no degradation across reasoning, coding, or dialogue capabilities.

These findings indicate that large language models do not require full online cross-attention between reference documents to generate accurate answers. Organizations can significantly reduce infrastructure operating costs and response latency for customer-facing applications and edge deployments by shifting computation to one-time offline precomputation. The authors note that the financial cost of extra disk storage is substantially lower than the compute resources typically needed to achieve sub-second response times.

Organizations operating latency-sensitive knowledge bases should consider piloting offline cache reuse strategies. For immediate deployment, the reordered-position scheme is recommended due to its minimal accuracy loss even without fine-tuning, though incorporating lightweight supervised fine-tuning provides optimal performance. Engineering teams must also plan for secondary technical requirements, including high-throughput disk storage and cache invalidation strategies for updating dynamic content.

While the results demonstrate robust performance across multiple benchmarks and model sizes, two main constraints remain. Storing precomputed caches introduces disk storage and memory transfer overhead, making compression techniques an important area for future integration. Additionally, while the approach is broadly applicable to rotary-position-based models, achieving optimal accuracy currently relies on lightweight model fine-tuning.

No sufficiently relevant recommendations were found.

Cover for TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text

Abstract

Current Retrieval-Augmented Generation (RAG) systems concatenate and process numerous retrieved document chunks for prefill which requires a large volume of online computation, therefore leading to significant latency in time-to-first-token (TTFT). To reduce the computation overhead as well as TTFT, we introduce TurboRAG, a hybrid offline-online paradigm that (i) pre-computes chunk-level key-value (KV) caches, (ii) stitches them together at inference time using independent-attention and reordered-RoPE techniques, and (iii) preserves answer quality without changing the model architecture. Our approach is applicable to most existing large language models and their applications without any requirement in modification of models and inference systems. Experimental results across a suite of RAG benchmarks demonstrate that TurboRAG reduces TTFT by up to 9.4x compared to the conventional RAG systems (on an average of 8.6x), but reserving comparable performance to the standard RAG systems.

Table of Contents

  • 1 Introduction
  • 2 RELATED WORK
  • 3 Methodology
  • 3.1 PROBLEM FORMALIZATION
  • 3.2 Position ID Rearrangement
  • 3.3 Adapting LLMs for Precomputed Cache Concatenation
  • 3.4 The TurboRAG Pipeline
  • 4 Experiments
  • 4.1 Experiment Setup
  • 4.2 Document QA Accuracy
  • 4.3 General Capability Regression
  • 4.4 TTFT Performance
  • 4.5 Batch Scaling
  • 5 CONCLUSION AND DISCUSSION
  • References
  • A Document Q&A Example
  • B Data proportions
  • C Computational Load Calculation

Knowls

  1. Knowl 1 — TurboRAG shifts retrieved-document prefill offline

    model/method

    TurboRAG changes retrieval-augmented generation (RAG) from processing all retrieved text afresh for every query to reusing per-chunk key-value (KV) caches. Offline, a retriever’s document chunks are embedded and indexed, and a language model pre-fills each chunk independently and stores its KV cache. Online, the query is embedded, the top-ranked chunks are retrieved, their cached keys and values are assembled, and the language model pre-fills the query against that context before generating an answer. The page-1 workflow graphic contrasts this cache-reuse path with standard RAG, where retrieved chunks are concatenated and prefilled online. TurboRAG targets the document-prefill computation; it does not require changing the language model architecture.

  2. Knowl 2 — Independent attention and reordered RoPE positions enable cache stitching

    model/method

    TurboRAG stitches independently precomputed chunk caches using two changes to the attention layout. First, document tokens are prevented from attending to tokens in other document chunks, while the query and generated answer tokens can attend to all retrieved documents and to the preceding query or answer tokens under causal decoding. Second, document-token positions are made globally monotone across the stitched context: for chunks of length ll, chunk jj (indexed from zero) receives positions jl,…,(j+1)l−1jl,\ldots,(j+1)l-1. In a RoPE-based model, the cached keys are rotated using these reordered positions when the chunks are joined; values are not position-rotated. This restores the document-to-query relative offsets that would be lost if every chunk retained positions 0,…,l−10,\ldots,l-1. The alternative, called composite positions, leaves each chunk’s positions starting at zero and therefore does not preserve those relative offsets. In the page-2 attention-map visualization and analysis, cross-chunk attention is reported to be sparse in the examined RAG example, and masking it does not substantially change the query’s attention distribution over documents: attention remains concentrated on relevant material.

  3. Knowl 3 — Supervised fine-tuning adapts models to the stitched-cache layout

    model/method

    Because the independent-attention mask and reordered positions differ from standard causal prefill, TurboRAG uses supervised fine-tuning (SFT) so the model learns to answer with the cache layout it will encounter at inference. The authors train both composite-position and reordered-position variants, using their corresponding nonstandard attention and position layouts during training. The Qwen2-7B SFT mixture contains 50% document question answering, 25% general dialogue, 10% reasoning, 10% code, and 5% other tasks. Training used 32 NVIDIA A100 80-GB GPUs, a batch size of 256 samples, learning rate 10−510^{-5}, and AdamW; the reported run took approximately 888 GPU-hours. The Naïve RAG and TurboRAG models were trained with the same proportions of data for comparison.

  4. Knowl 4 — Reordered positions preserve RGB document-QA accuracy better than composite positions

    empirical result

    On the bilingual RGB document-QA benchmark, each query was evaluated with five retrieved documents and noise ratios from 0.2 to 0.8; the reported values are accuracy scores. The tuned reordered-position model is close to Naïve RAG on the aggregate results, while the untuned reordered variant is more robust than the untuned composite variant, especially at high noise. The full reported scores are shown below.

    Chinese English
    Model 0.2 0.4 0.6 0.8 Avg. 0.2 0.4 0.6 0.8 Avg.
    GPT-4o-2024-08-06 98.3 98.0 96.6 87.7 95.2 99.0 99.3 98.3 96.3 98.2
    Naïve RAG 99.0 98.0 96.7 87.3 95.3 99.7 99.3 99.3 94.3 98.2
    TurboRAG-composite, no fine-tuning 98.3 96.3 93.7 79.0 91.8 98.0 96.3 91.3 75.0 90.2
    TurboRAG-reordered, no fine-tuning 98.0 96.7 93.3 81.3 92.3 98.0 97.3 90.7 85.7 92.9
    TurboRAG-composite 99.0 97.3 96.0 86.7 94.8 99.3 98.0 96.7 92.7 96.7
    TurboRAG-reordered 98.7 97.3 96.0 90.7 95.7 99.0 98.3 96.0 93.7 96.8

    The same reordered-position approach was evaluated with LLaMA-3.1-8B on RGB. Its average scores were 95.0 versus 97.1 for Naïve RAG in Chinese, and 97.4 versus 98.4 in English. Thus, the reported results extend beyond the Qwen2-7B experiments, although the LLaMA scores also show a remaining accuracy gap.

  5. Knowl 5 — TurboRAG matches Naïve RAG on RGB information-integration tasks after fine-tuning

    empirical result

    RGB’s information-integration task tests whether an answer can be synthesized from multiple retrieved documents; its reported scores are given below by language and noise ratio. Fine-tuning substantially improves both TurboRAG variants over their untuned forms. The tuned reordered model averages 44 in Chinese and 47 in English, compared with 42 and 47 for Naïve RAG, respectively; at noise ratio 0.6, all three tuned systems score 32–34 in Chinese and 34 in English.

    Chinese English
    Model 0.2 0.4 0.6 Avg. 0.2 0.4 0.6 Avg.
    Naïve RAG 50 46 29 42 57 48 36 47
    TurboRAG-composite, no fine-tuning 35 27 18 27 40 27 27 31
    TurboRAG-reordered, no fine-tuning 30 21 20 24 31 23 19 24
    TurboRAG-composite 53 41 32 42 58 48 34 47
    TurboRAG-reordered 56 44 32 44 57 51 34 47
  6. Knowl 6 — LongBench multi-document QA shows an average 8.6× TTFT speedup

    empirical result

    On LongBench multi-document QA, TurboRAG-reordered substantially reduces time-to-first-token (TTFT) while retaining competitive answer scores. Evaluation used Hugging Face Transformers with FlashAttention 2 on one NVIDIA A100 80-GB GPU. The table gives the reported QA scores and TTFT measurements; context and query lengths are token counts, and TTFT is in milliseconds.

    Task (metric) Context tokens Query tokens Naïve score Composite score Reordered score Naïve TTFT (ms) Reordered TTFT (ms) Speedup
    MuSiQue (F1) 16349 18.8 22.12 23.64 27.37 1610 171 9.4x
    2WikimQA (F1) 7553 17.0 35.02 34.28 39.51 709 101 7.0x
    DuReader (Rouge-L) 10642 6.0 34.57 33.37 33.03 1007 116 8.7x
    HotpotQA (F1) 13453 20.1 40.21 35.78 45.28 1333 147 9.1x
    Average 11999 15.5 32.99 31.76 36.29 1165 134 8.6x

    The largest reported TTFT gain is 9.4× on MuSiQue; the smallest is 7.0× on 2WikimQA. A separate RGB measurement with 743 context tokens reports 87 ms for Naïve RAG and 36 ms for TurboRAG, a 2.42× speedup.

  7. Knowl 7 — Cache reuse cuts online computation by about 98.5% in a fixed-context batch test

    empirical result

    A batch-scaling experiment compared Naïve RAG with TurboRAG for a fixed total recalled-text length of 8192 tokens and query length of 128 tokens. Cache transfer used PCIe Gen4; the TurboRAG no-H2D condition assumes the cache is already on the GPU. TTFT is in milliseconds and computational load is in TFLOPs. The reported TFLOPs are the same for TurboRAG with and without transfer; the paper reports an approximately 98.46% reduction relative to Naïve RAG.

    Batch size Naïve TTFT TurboRAG TTFT Speedup TurboRAG no-H2D TTFT No-H2D speedup
    1 711 175 4.1x 44 16.1x
    2 1408 325 4.3x 56 25.1x
    4 2842 666 4.3x 97 29.3x
    6 4373 928 4.7x 134 32.6x
    8 5812 1429 4.1x 177 32.8x
    TFLOPs by batch size: Naïve RAG = 136.36, 272.72, 545.46, 818.20, 1090.93; TurboRAG = 2.09, 4.19, 8.39, 12.58, 16.78.

    Even with host-to-device cache transfer, TurboRAG is about four times faster in this setup. With the cache already on the GPU, the measured speedup grows from 16.1× at batch size 1 to 32.8× at batch size 8.

  8. Knowl 8 — TTFT gains persist across Qwen2 model sizes

    empirical result

    The authors report TTFT measurements for Qwen2 models from 1.5B to 72B parameters, with batch sizes from 1 to 8; the 32B and 72B tests use two and four GPUs, respectively. The endpoint measurements below show that TurboRAG remains faster across model sizes and batch sizes. TTFT values are in milliseconds.

    Model Batch size 1: Naïve / Turbo / speedup Batch size 8: Naïve / Turbo / speedup
    1.5B 197 / 85 / 2.31x 1479 / 292 / 5.06x
    3B 363 / 96 / 3.78x 2855 / 491 / 5.81x
    14B 1413 / 272 / 5.19x 11888 / 2452 / 4.84x
    32B (2 GPUs) 2923 / 383 / 7.63x 23205 / 2884 / 8.04x
    72B (4 GPUs) 6157 / 653 / 9.42x 50595 / 5600 / 9.03x

    The reported speedup ranges from 2.31× for the 1.5B model at batch size 1 to 11.23× for the 72B model at batch size 2. These results support scalability of the latency benefit across the tested model family, but do not establish performance for model architectures outside those evaluated.

  9. Knowl 9 — General-capability regression is small on the reported OpenCompass tasks

    empirical result

    On the reported OpenCompass regression tests, TurboRAG-reordered’s scores remain close to those of Naïve RAG. The final column is the TurboRAG-reordered score minus the Naïve RAG score, in the benchmark’s reported units.

    Metric Naïve RAG TurboRAG-reordered Difference
    MMLU 69.57 70.73 +1.16
    TriviaQA 56.90 56.47 -0.43
    GSM-8K 79.12 79.45 +0.33
    MATH 39.54 40.58 +1.04
    HumanEval 58.26 57.32 -0.94
    AlpacaEval2 7.83 8.32 +0.49

    The differences range from -0.94 on HumanEval to +1.16 on MMLU. These results indicate no large regression on the tested tasks; they do not measure general capabilities beyond the listed benchmarks.

  10. Knowl 10 — Storage requirements and fine-tuning constrain deployment

    limitation

    TurboRAG trades storage and cache-loading demands for lower online prefill latency, and the evaluated pipeline requires model fine-tuning. For Qwen2-7B, the paper estimates that one 512-token chunk requires 28 MB for an FP16 KV cache; one million such chunks therefore require approximately 28 TB. It estimates that a collection of 10410^4–10510^5 passages would occupy 0.3–3 TB, or 40–400 GB after 8-bit KV quantization. Loading caches from disk also places pressure on memory. The authors identify the fine-tuning requirement as a limitation for direct adoption on newly emerging models and leave reducing or eliminating that requirement as future work.

Coverage note — The short-context TTFT measurements and the appendix’s detailed FLOPs-accounting formula are omitted as supplementary measurements and calculation details rather than distinct contributions.

References

  1. 1.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508.
  2. 2.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR.
  3. 3.Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17754–17762.
  4. 4.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595.
  5. 5.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. 2020. Rethinking attention with performers. arXiv preprint arXiv:2009.14794.
  6. 6.Tri Dao. 2023. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. Preprint, arXiv:1810.04805.
  8. 8.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  9. 9.Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282.
  10. 10.Wenqi Jiang, Shuai Zhang, Boran Han, Jie Wang, Bernie Wang, and Tim Kraska. 2024. Piperag: Fast retrieval-augmented generation via algorithm-system co-design. arXiv preprint arXiv:2403.05676.
  11. 11.Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Xin Liu, Xuanzhe Liu, and Xin Jin. 2024. Ragcache: Efficient knowledge caching for retrieval-augmented generation. arXiv preprint arXiv:2404.12457.
  12. 12.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451.
  13. 13.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  14. 14.Xiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li, Bo Li, Xuming Hu, and Xiaowen Chu. 2025. Chunkkv: Semantic-preserving kv cache compression for efficient long-context llm inference. arXiv preprint arXiv:2502.00299.
  15. 15.Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, et al. 2024. Cachegen: Kv cache compression and streaming for fast large language model serving. In Proceedings of the ACM SIGCOMM 2024 Conference, pages 38–56.
  16. 16.I Loshchilov. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  17. 17.Dongyang Ma, Yan Wang, and Lan Tian. 2024. Block-attention for efficient prefilling. arXiv preprint arXiv:2409.15355.
  18. 18.Rui Pan, Zhuang Wang, Zhen Jia, Can Karakus, Luca Zancato, Tri Dao, Yida Wang, and Ravi Netravali. 2024. Marconi: Prefix caching for the era of hybrid llms. arXiv preprint arXiv:2411.19379.
  19. 19.Guofeng Quan, Wenfeng Feng, Chuzhan Hao, Guochao Jiang, Yuewei Zhang, and Hao Wang. 2025. Rasd: Retrieval-augmented speculative decoding. arXiv preprint arXiv:2503.03434.
  20. 20.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. Transactions of the Association for Computational Linguistics, 11:1316–1331.
  21. 21.Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2022. Parallel context windows for large language models. arXiv preprint arXiv:2212.10947.
  22. 22.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063.
  23. 23.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509.
  24. 24.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768.
  25. 25.Zheng Wang, Boxiao Jin, Zhongzhi Yu, and Minjia Zhang. 2024a. Model tells you where to merge: Adaptive kv cache merging for llms on long-context tasks. arXiv preprint arXiv:2407.08454.
  26. 26.Zhibin Wang, Rui Ning, Chao Fang, Zhonghui Zhang, Xi Lin, Shaobo Ma, Mo Zhou, Xue Li, Zhongfeng Wang, Chengying Huan, et al. 2025. Flashforge: Ultra-efficient prefix-aware attention for llm decoding. arXiv preprint arXiv:2505.17694.
  27. 27.Zilong Wang, Zifeng Wang, Long Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang, Anush Mattapalli, Ankur Taly, Jingbo Shang, et al. 2024b. Speculative rag: Enhancing retrieval augmented generation through drafting. arXiv preprint arXiv:2407.08223.
  28. 28.An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671.
  29. 29.Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. 2025. Cacheblend: Fast large language model serving for rag with cached knowledge fusion. In Proceedings of the Twentieth European Conference on Computer Systems, pages 94–109.
  30. 30.Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al. 2024. H2o: Heavy-hitter oracle for efficient generative inference of large language models. Advances in Neural Information Processing Systems, 36.
  31. 31.Yun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko, Yinxiao Liu, Chu-Cheng Lin, Lei Shu, Liangchen Luo, Lei Meng, Bang Liu, et al. 2024. Accelerating inference of retrieval-augmented generation via sparse context selection. arXiv preprint arXiv:2405.16178.

Citation

MLA
Lu, S., et al. “TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 6588–601, https://doi.org/10.18653/v1/2025.emnlp-main.334.
APA
Lu, S., Wang, H., Rong, Y., Chen, Z., & Tang, Y. (2025). TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 6588–6601. https://doi.org/10.18653/v1/2025.emnlp-main.334
Chicago
Lu, S., H. Wang, Y. Rong, Z. Chen, and Y. Tang. 2025. “TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 6588–6601. https://doi.org/10.18653/v1/2025.emnlp-main.334.
Harvard
Lu, S. et al. (2025) “TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text”, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6588–6601. Available at: https://doi.org/10.18653/v1/2025.emnlp-main.334.
Vancouver
1. Lu S, Wang H, Rong Y, Chen Z, Tang Y (2025) TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6588–6601

BibTeX

@inproceedings{lu-etal-2025-turborag,
    title = "{T}urbo{RAG}: Accelerating Retrieval-Augmented Generation with Precomputed {KV} Caches for Chunked Text",
    author = "Lu, Songshuo  and
      Wang, Hua  and
      Rong, Yutian  and
      Chen, Zhi  and
      Tang, Yaohua",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.334/",
    doi = "10.18653/v1/2025.emnlp-main.334",
    pages = "6588--6601",
    ISBN = "979-8-89176-332-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/