RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs

Yue YuWei PingZihan LiuBoxin WangJiaxuan YouChao ZhangMohammad ShoeybiBryan Catanzaro

article2024NeurIPS315 citations

Proposes an instruction fine-tuning framework that trains a single language model to both rerank retrieved contexts and generate answers, outperforming GPT-4 across multiple knowledge-intensive retrieval-augmented generation benchmarks.

Listen

Retrieval-augmented generation (RAG) is essential for grounding large language models in external knowledge, keeping information up to date, and reducing hallucinations without altering underlying model weights. However, conventional pipelines face a critical trade-off: initial search retrievers often have limited capacity to accurately identify the best documents, yet feeding too many retrieved contexts into a language model introduces noise and degrades answer accuracy. While deploying dedicated, external ranking models can help filter these contexts, such models often generalize poorly across diverse tasks and add pipeline complexity.

The article demonstrates that a single large language model can be instruction-tuned to simultaneously perform high-accuracy context reranking and final answer generation, eliminating the need for a separate specialized ranking model. The authors evaluate this unified framework, named RankRAG, across multiple open-domain, conversational, and specialized biomedical question-answering benchmarks against leading models such as GPT-4 and ChatQA-1.5.

The approach uses a two-stage instruction-tuning strategy. Following initial general supervised fine-tuning, the second stage trains the model on a unified blend of context-rich question answering, retrieval-augmented generation with hard negatives, and relevance ranking framed as simple instruction-based tasks. At inference time, the model executes a "retrieve-rerank-generate" process: a standard search retriever gathers candidate passages, the model calculates relevance scores to filter down to the top few passages, and it then generates the final response using only those refined contexts. Experiments were conducted using open-weight models across nine general benchmarks and five biomedical benchmarks.

The key findings show significant performance and efficiency gains. First, the 8-billion and 70-billion parameter RankRAG models consistently outperform established baselines like ChatQA-1.5 across nine general benchmarks, with the 8-billion model outperforming systems with five to eight times more parameters. Second, the performance advantages are largest on challenging tasks, achieving over 10% improvements on long-tail and multi-hop reasoning datasets where initial retrieval quality is poor. Third, the framework exhibits high data efficiency: adding a modest amount of ranking data (around 50,000 pairs, or about 10% of standard ranking datasets) allowed the model to outperform dedicated ranking systems trained on up to ten times more data. Fourth, on specialized biomedical benchmarks, the model achieved competitive performance—reaching over 98% of GPT-4's accuracy—without any domain-specific fine-tuning.

These results demonstrate that context ranking and generation capabilities mutually reinforce one another within a single model. For organizations deploying generative systems, this unified approach lowers operational complexity, reduces dependency on multi-model pipelines, and improves factual accuracy on difficult tasks while requiring only small amounts of ranking training data. The findings indicate that the primary bottleneck in current retrieval systems often lies in context selection rather than generator capacity.

Organizations seeking to improve the factual reliability of knowledge-intensive generative applications should consider adopting a unified reranking and generation framework. When deploying this architecture, teams can manage processing latency by tuning the initial candidate pool size, as reranking 20 to 30 documents captures most accuracy gains with minimal execution overhead. Future efforts should explore combining this approach with multi-step or iterative retrieval workflows.

Confidence in these findings is supported by consistent gains across multiple model architectures, parameter sizes, and diverse benchmark suites. However, decision-makers should note that the reranking stage introduces modest latency overhead compared to single-step generation pipelines. Furthermore, the evaluations focus on single-turn retrieval settings, so additional validation is warranted before deploying in highly interactive, multi-turn conversational environments with dynamic query updates.

arXiv: 2407.02485
Cover for RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs

Abstract

Large language models (LLMs) typically utilize the top-k contexts from a retriever in retrieval-augmented generation (RAG). In this work, we propose a novel instruction fine-tuning framework RankRAG, which instruction-tunes a single LLM for the dual purpose of context ranking and answer generation in RAG. In particular, the instruction-tuned LLMs work surprisingly well by adding a small fraction of ranking data into the training blend, and outperform existing expert ranking models, including the same LLM exclusively fine-tuned on a large amount of ranking data. For generation, we compare our model with many strong baselines, including GPT-4-0613, GPT-4-turbo-2024-0409, and ChatQA-1.5, an open-sourced model with the state-of-the-art performance on RAG benchmarks. Specifically, our Llama3-RankRAG significantly outperforms Llama3-ChatQA-1.5 and GPT-4 models on nine knowledge-intensive benchmarks. In addition, it also performs comparably to GPT-4 on five RAG benchmarks in the biomedical domain without instruction fine-tuning on biomedical data, demonstrating its superb capability for generalization to new domains.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 3.1 Problem Setup
  • 3.2 Limitation of Current RAG Pipelines
  • 4 RankRAG
  • 4.1 Stage-I: Supervised Fine-Tuning (SFT)
  • 4.2 Stage-II: Unified Instruction-Tuning for Ranking and Generation
  • 4.3 RankRAG Inference: Retrieve-Rerank-Generate Pipeline
  • 5 Experiments
  • 5.1 Experiment Setup
  • 5.2 Main Experiments
  • 5.3 Ablation Studies
  • 5.4 Experiment on Domain-specific RAG Benchmarks
  • 5.5 A Closer Look at the Ranking Module
  • 5.6 Case Study
  • 6 Conclusion
  • References
  • A Dataset Description
  • A.1 Main Experiments
  • A.2 Biomedical Benchmarks
  • B Data Blending Details for Ranking-enhanced Instruction Finetuning
  • C Prompt Formats of Instruction Tuning
  • C.1 Stage I: Supervised Fine-tuning
  • C.2 Stage-II: Unified Instruction-Tuning for Ranking and Generation
  • D Prompt Formats of Target Tasks
  • D.1 Context Ranking
  • E Additional Experiment Results
  • E.1 Ranking Performance Using DPR and Contriever as Retrievers
  • E.2 RAG Performance with Different k
  • F Performance of NQ and Trivia QA on DPR Splits
  • G Additional Case Studies
  • NeurIPS Paper Checklist
  • 1. Claims
  • 2. Limitations
  • 3. Theory Assumptions and Proofs
  • 4. Experimental Result Reproducibility
  • 6. Experimental Setting/Details
  • 7. Experiment Statistical Significance
  • 8. Experiments Compute Resources
  • 9. Code Of Ethics
  • 10. Broader Impacts
  • 11. Safeguards
  • 12. Licenses for existing assets
  • 13. New Assets
  • 14. Crowdsourcing and Research with Human Subjects
  • 15. Institutional Review Board (IRB) Approvals or Equivalent for Research with Human Subjects

Knowls

  1. Knowl 1 — RankRAG Two-Stage Instruction Tuning Framework

    model/method

    RankRAG is an instruction fine-tuning framework that trains a single autoregressive large language model (LLM) to perform both candidate context reranking and final answer generation for retrieval-augmented generation (RAG). The framework uses a two-stage training strategy:

    1. Stage-I: Supervised Fine-Tuning (SFT): The base LLM is fine-tuned on a collection of 128k general instruction-following examples formatted as multi-turn conversations. Training datasets include OpenAssistant, Dolly, SODA, ELI5, Self-Instruct, Unnatural Instructions, FLAN, and Chain-of-Thought datasets. Optimization is performed by calculating autoregressive cross-entropy loss exclusively on the final assistant response tokens.

    2. Stage-II: Unified Instruction-Tuning for Ranking and Generation: The model from Stage-I is fine-tuned on a unified mixture consisting of five task categories, all mapped to a standardized (x,c,y)(x, c, y) instruction format (where xx is the question/instruction, cc is the context passage or set of passages, and yy is the target answer):

    • Stage-I SFT data (blending weight: 0.20) to maintain general instruction-following capability.
    • Context-rich QA data (blending weight: ~0.415 total) from DROP (0.069), NarrativeQA (0.09), Quoref (0.026), ROPES (0.026), NewsQA (0.09), TAT-QA (0.15 arithmetic, 0.08 others), and conversational QA datasets HumanAnnotatedConvQA and SyntheticConvQA (0.20).
    • Retrieval-augmented QA data (blending weight: 0.18 total) from SQuAD (0.09) and WebQuestions (0.09), where the gold passage is mixed with BM25-retrieved hard negative contexts to form a 5-passage context, training the model to answer questions while ignoring irrelevant distractors.
    • Context ranking data (blending weight: 0.18 total) from MS MARCO passage ranking (0.15) and synthetic conversational QA pairs (0.03), where the model receives a question and a single candidate passage and must predict True if relevant or False if irrelevant.
    • Retrieval-augmented ranking data (blending weight: 0.04 total) from SQuAD (0.02) and WebQuestions (0.02), where 5 shuffled candidate passages are presented together and the LLM is instructed to output the passage indexes containing the answer.
  2. Knowl 2 — Retrieve-Rerank-Generate Inference Algorithm

    algorithm

    At inference time, RankRAG executes a three-step retrieve-rerank-generate process using the underlying retriever R\mathcal{R} and the dual-purpose RankRAG LLM:

    Input: Query question qq, Document corpus D\mathcal{D}, Dense retriever R\mathcal{R}, RankRAG LLM M\mathcal{M}, Candidate pool size NN, Final context budget kk (where k≪Nk \ll N)
    Output: Generated answer text $y^*
    1. Retrieve Candidate Set:
       Retrieve top-NN candidate passages CN={c1,c2,…,cN}⊂D\mathcal{C}_N = \{c_1, c_2, \dots, c_N\} \subset \mathcal{D} using retriever R(q,D)\mathcal{R}(q, \mathcal{D}).
    2. Relevance Scoring and Reranking:
       for each passage ci∈CNc_i \in \mathcal{C}_N do
           Format context ranking prompt: xi←"For the question "+q+", assess whether the passage is relevant to the question. Return True if relevant, otherwise False."x_i \leftarrow \text{"For the question "} + q + \text{", assess whether the passage is relevant to the question. Return True if relevant, otherwise False."}
           Compute relevance score: si=PM("True"∣xi,ci)s_i = P_{\mathcal{M}}(\text{"True"} \mid x_i, c_i)
       end for
       Sort passages in CN\mathcal{C}_N descending by sis_i and select the top-kk passages Ck={c(1),c(2),…,c(k)}\mathcal{C}_k = \{c_{(1)}, c_{(2)}, \dots, c_{(k)}\}.
    3. Retrieval-Augmented Generation:
       Construct input context: C="Passage 1: "+c(1)+⋯+"Passage k: "+c(k)C = \text{"Passage 1: "} + c_{(1)} + \dots + \text{"Passage } k\text{: "} + c_{(k)}
       Generate answer: y∗=arg⁡max⁡yPM(y∣q,C)y^* = \arg\max_y P_{\mathcal{M}}(y \mid q, C)
       return y∗y^*

    In standard setups, N=100N=100 (for 8B models) or N=30N=30 (for 70B models), and k=5k=5. Scoring each candidate context cic_i requires generating only a single token (True vs False), significantly reducing inference latency relative to multi-token decoding.

  3. Knowl 3 — Unified Instruction Formats and Pseudo-Relevance Mining

    model/method

    In RankRAG Stage-II instruction tuning, all tasks are structured under a unified format (x,c,y)(x, c, y), where xx represents the prompt instruction/question, cc is the passage context, and yy is the ground-truth target output:

    • Context-rich QA: xx is Answer the following question from context. {question}; cc is Passage: {Passage} (1 passage); yy is a target phrase or sentence.
    • Retrieval-augmented QA: xx is Answer the following question from context. {question}; cc contains concatenated shuffled passages Passage 1: {Passage 1} ... Passage 5: {Passage 5}; yy is the gold answer.
    • Context ranking: xx is For the question {question}, assess whether the passage is relevant to the question. Return True if relevant, otherwise False.; cc is Passage: {Passage}; yy is True for relevant pairs and False for hard negatives.
    • Retrieval-augmented ranking: xx is For the question {question}, access whether the above passages are relevant to the question. Return all the relevant passage id.; cc contains 5 candidate passages Passage 1: {Passage 1} ... Passage 5: {Passage 5}; yy is the list of relevant passage indexes.

    Conversational Pseudo-Relevance Mining: To generate single-passage ranking pairs from conversational QA datasets where questions span multi-turn dialogue over long documents dd without passage-level labels, document dd is segmented into 150-word chunks (c1,c2,…,cn)(c_1, c_2, \dots, c_n). For each chunk cic_i and ground-truth answer aa, the 4-gram recall score R4(ci,a)R_4(c_i, a) is computed:

    • Chunk cic_i is labeled relevant (y=Truey = \text{True}) if R4(ci,a)>0.5R_4(c_i, a) > 0.5.
    • Chunk cic_i is labeled irrelevant (y=Falsey = \text{False}) if R4(ci,a)<0.1R_4(c_i, a) < 0.1.
    • Chunks with 0.1≤R4(ci,a)≤0.50.1 \le R_4(c_i, a) \le 0.5 are discarded.
  4. Knowl 4 — Zero-Shot RAG Performance Across 9 General-Domain Benchmarks

    data/table

    RankRAG was evaluated under zero-shot conditions across 9 knowledge-intensive benchmarks comprising single-hop OpenQA (Natural Questions [NQ], TriviaQA, PopQA), multi-hop OpenQA (HotpotQA, 2WikimQA), fact verification (FEVER), and conversational QA (Doc2Dial, TopiOCQA, INSCIT). Dragon retriever was used by default (N=100N=100 for 8B models, N=30N=30 for 70B models, k=5k=5).

    Model NQ TriviaQA PopQA HotpotQA 2WikimQA FEVER Doc2Dial TopiOCQA INSCIT Avg.
    EM EM / Acc. EM / Acc. EM / F1 EM / F1 Acc. F1 F1 F1 –
    Without Retrieval
    GPT-3.5-turbo-1106 38.6 82.9 / 91.7 28.4 / 32.2 29.9 / 42.0 23.9 / 30.4 82.7 20.1 28.5 27.2 38.5
    GPT-4-0613 40.3 84.8 / 94.5 31.3 / 34.8 34.5 / 46.9 29.8 / 36.6 87.7 27.6 30.1 27.0 42.0
    GPT-4-turbo-2024-0409 41.5 80.0 / 94.3 25.0 / 33.5 26.6 / 43.8 24.1 / 35.5 87.0 27.6 26.4 24.4 38.6
    With Retrieval
    Atlas 11B 26.7 56.9 / – – / – 34.7 / – – / – 77.0 – – – –
    RA-DIT 65B 35.2 75.4 / – – / – 39.7 / – – / – 80.7 – – – –
    InstructRetro 43B 38.9 78.3 / – – / – – / – – / – – 36.0 – – –
    Llama3-Instruct 8B 30.9 70.7 / 80.4 34.9 / 55.8 26.0 / 35.8 9.6 / 25.2 88.9 33.6 44.9 32.6 40.8
    Llama3-Instruct 70B 42.7 82.4 / 89.3 45.3 / 56.4 35.5 / 43.3 13.5 / 27.9 91.4 37.9 49.7 36.2 47.1
    Llama3-ChatQA-1.5 8B 42.4 81.0 / 87.6 52.6 / 59.8 33.4 / 44.6 26.8 / 31.9 90.9 39.3 49.9 30.1 49.6
    Llama3-ChatQA-1.5 70B 47.0 85.6 / 91.4 50.9 / 58.3 42.2 / 54.4 34.9 / 37.4 92.7 41.3 55.6 32.3 53.6
    Llama3-RankRAG 8B 50.6 82.9 / 89.5 57.6 / 64.1 35.3 / 46.7 31.4 / 36.9 92.0 40.4 50.4 33.3 52.6
    Llama3-RankRAG 70B 54.2 86.5 / 92.3 59.9 / 65.4 42.7 / 55.4 38.2 / 43.9 93.8 41.5 52.8 35.2 56.1

    Llama3-RankRAG-8B achieves an average score of 52.6, outperforming Llama3-ChatQA-1.5-8B (49.6), larger baselines such as RA-DIT 65B and InstructRetro 43B, and proprietary models without RAG. Llama3-RankRAG-70B achieves the highest overall average score of 56.1. Gains are especially pronounced on challenging datasets requiring precise passage filtering, such as long-tail QA (PopQA: +5.0 EM over ChatQA-1.5-8B) and multi-hop QA (2WikimQA: +4.6 EM over ChatQA-1.5-8B).

  5. Knowl 5 — Context Ranking Recall and Data Efficiency

    data/table

    The context ranking performance of RankRAG was evaluated against dedicated ranking models and off-the-shelf LLMs by reranking the top-100 retrieved passages from Dragon. Recall@kk (R@kR@k, for k∈{5,10,20}k \in \{5, 10, 20\}) measures whether passages containing the answer are retained in top-kk.

    Model # Rank Data NQ TriviaQA HotpotQA INSCIT
    R@5 R@10 R@20 R@5 R@10 R@20 R@5 R@10 R@20 R@5
    Dragon (Retriever) – 74.9 80.3 84.3 89.0 92.9 95.3 47.5 52.4 60.1 43.4
    Finetuned Ranking Models
    RankBERT 110M ∼\sim503k 73.5 79.3 84.0 88.4 92.0 95.5 54.6 59.8 63.7 45.6
    monoT5 3B ∼\sim503k 75.6 80.9 84.9 90.7 93.6 95.9 54.8 60.2 63.3 48.6
    BGE-Rerank-v2-m3 568M ∼\sim1.6M 78.0 82.8 85.6 91.6 94.5 97.1 58.5 61.8 65.0 51.3
    RankLLaMA 8B ∼\sim503k 77.8 83.1 86.0 91.2 93.1 96.4 57.1 62.1 64.8 57.8
    ChatQA-1.5 8B N/A 68.2 75.7 82.0 85.4 91.1 94.0 37.4 45.0 53.6 32.3
    Off-the-shelf LLMs
    GPT-3.5 (top 100) Unk. 77.8 82.5 85.7 91.1 94.4 96.7 52.1 56.6 62.4 50.2
    GPT-4 (top 30) Unk. 79.3 83.2 85.1 92.8 95.5 96.8 53.2 57.0 61.0 52.3
    RankRAG
    RankRAG 8B (top 100) ∼\sim50k 80.3 84.0 86.3 93.2 95.4 97.3 57.6 61.8 65.2 60.9
    RankRAG 70B (top 30) ∼\sim50k 80.6 84.0 85.4 93.6 95.9 97.1 56.3 59.7 62.2 61.3

    Despite using only ∼\sim50k ranking pairs (a ≈10×\approx 10\times to 30×30\times reduction compared to dedicated rerankers like RankLLaMA and BGE-Reranker), RankRAG 8B attains superior Recall@kk across most benchmarks (e.g., NQ R@5 reaches 80.3% vs. RankLLaMA's 77.8% and BGE's 78.0%). Fine-tuning with as few as 5k ranking pairs already yields strong recall improvements.

  6. Knowl 6 — Zero-Shot Biomedical Domain Generalization on MIRAGE Benchmark

    data/table

    RankRAG models trained exclusively on general-domain data were evaluated zero-shot on the MIRAGE biomedical RAG benchmark using MedCPT as the retriever and MedCorp as the knowledge corpus.

    Model MMLU-med PubmedQA BioASQ MedQA MedMCQA Avg.
    GPT-4-0613 87.24 70.60 92.56 82.80 66.65 79.97
    GPT-3.5 75.48 67.40 90.29 66.61 58.04 71.56
    Mixtral 8x7B 75.85 67.60 87.54 60.02 56.42 69.49
    Llama2 70B 54.55 50.40 73.95 44.93 43.08 53.38
    Meditron 70B 65.38 56.40 76.86 49.57 52.67 60.18
    PMC-LLaMA 13B 52.53 42.58 48.29 56.00 65.21 52.92
    Llama3-ChatQA-1.5 8B 61.40 66.40 82.69 42.36 46.97 59.96
    Llama3-ChatQA-1.5 70B 80.51 74.80 83.17 68.89 62.54 73.98
    Llama3-RankRAG 8B 64.55 65.00 84.44 48.86 56.90 63.95
    Llama3-RankRAG 70B 81.44 79.80 90.76 69.21 69.11 78.06

    Without any fine-tuning on biomedical corpora, Llama3-RankRAG-8B (63.95 avg.) outperforms specialized medical LLMs including Meditron-70B (60.18) and PMC-LLaMA-13B (52.92). Llama3-RankRAG-70B achieves an average score of 78.06, outperforming ChatQA-1.5-70B (73.98) and reaching over 97.6% of GPT-4's performance (79.97).

  7. Knowl 7 — Ablation of RankRAG Training Data and Inference Reranking

    data/table

    An ablation study evaluated the individual contributions of RankRAG components using Llama3-8B across 9 general-domain benchmarks (NQ, TriviaQA, PopQA, HotpotQA, 2WikimQA, FEVER, Doc2Dial, TopiOCQA, INSCIT).

    Configuration NQ TriviaQA PopQA HotpotQA 2WikimQA FEVER Doc2Dial TopiOCQA INSCIT Avg.
    EM EM / Acc. EM / Acc. EM / F1 EM / F1 Acc. F1 F1 F1 –
    RankRAG 8B 50.6 82.9 / 89.5 57.6 / 64.1 35.3 / 46.7 31.4 / 36.9 92.0 40.4 50.4 33.3 52.6
    w/o reranking 48.0 80.3 / 86.8 49.3 / 59.0 31.3 / 41.6 26.4 / 30.5 91.1 39.7 49.4 30.9 49.8
    w/o RQA 49.4 82.0 / 88.9 55.1 / 62.9 35.6 / 45.9 31.8 / 37.5 92.1 39.4 46.8 32.4 51.6
    w/o RAR 48.6 82.2 / 89.1 56.0 / 62.6 35.1 / 45.2 31.2 / 35.7 91.4 39.6 48.6 33.5 51.8
    w/ RAFT 43.3 80.8 / 87.6 48.9 / 56.3 30.5 / 41.8 25.2 / 29.6 91.2 36.8 46.4 30.1 48.1
    w/ Stage-I SFT Only 38.3 63.7 / 76.6 49.8 / 54.6 26.5 / 40.3 18.0 / 25.9 85.7 33.3 33.7 30.5 42.2

    Key takeaways:

    • Removing the test-time reranking step (w/o reranking) drops the average performance by 2.8 points (from 52.6 to 49.8), with sharp decreases on PopQA (-8.3 EM) and 2WikimQA (-5.0 EM).
    • Removing Retrieval-augmented QA (w/o RQA) or Retrieval-augmented Ranking (w/o RAR) degrades performance across multiple datasets, demonstrating the value of exposing the model to multi-passage contexts containing distractors during training.
    • Formatting data using RAFT (which processes retrieved contexts separately rather than together) achieves an average of only 48.1, showing that RankRAG's unified multi-context formulation is superior.
  8. Knowl 8 — RankRAG Generality Across LLM Backbones and Dense Retrievers

    empirical result

    RankRAG consistently outperforms baseline RAG methods across different model families, model sizes, and retriever architectures:

    1. Model Backbone Generalization: When implemented on Llama-2 foundation models, Llama2-RankRAG consistently outperforms Llama2-ChatQA-1.0 across all scales:
    • 7B scale: 50.3 average (Llama2-RankRAG-7B) vs. 46.6 (ChatQA-1.0-7B), a +7.8% relative gain.
    • 13B scale: 52.5 average (Llama2-RankRAG-13B) vs. 49.6 (ChatQA-1.0-13B), a +6.4% relative gain.
    • 70B scale: 55.0 average (Llama2-RankRAG-70B) vs. 51.8 (ChatQA-1.0-70B), a +6.3% relative gain.
    1. Retriever Architecture Generalization: Evaluating RankRAG-8B paired with different retrievers demonstrates robust improvements over ChatQA-1.5 across benchmarks:
    • With Dense Passage Retriever (DPR), RankRAG increases Exact Match by over 10% on average over ChatQA-1.5 across NQ, TriviaQA, and PopQA.
    • With Contriever-MS MARCO, RankRAG similarly improves answer recall (R@5R@5 increases from 67.60% to 75.32% on NQ, 81.95% to 88.71% on TriviaQA, and 60.61% to 65.11% on PopQA) and downstream generation accuracy over ChatQA-1.5.
  9. Knowl 9 — Context Budget Trade-Off and Optimal Context Size $k$

    empirical result

    In standard retrieval-augmented generation pipelines without context reranking (such as ChatQA-1.5), selecting context size kk introduces a fundamental trade-off: small kk (e.g., k=5k=5) limits answer recall due to imperfect retrieval, while large kk (e.g., k=20k=20) incorporates irrelevant or noisy distractor passages that degrade generation accuracy. Consequently, baseline models plateau or decline in accuracy as kk increases beyond 10.

    In contrast, RankRAG achieves peak or near-peak generation performance at k=5k=5 across evaluated tasks (NQ, TriviaQA, PopQA, FEVER). Because the LLM reranks a larger candidate pool (N=100N=100 or N=30N=30) and concentrates the most relevant evidence into the top positions, setting k=5k=5 captures high answer recall while avoiding distracting noise, reducing the input context length required during the generation phase.

  10. Knowl 10 — Inference Latency Overhead and Ranking Budget Trade-Off

    model/method

    In the retrieve-rerank-generate pipeline of RankRAG, reranking introduces an additional computational overhead. Let t1t_1 denote the dense index retrieval time per query, t2t_2 the inference time for the LLM to score a single candidate passage via single-token generation probability P("True"∣x,ci)P(\text{"True"} \mid x, c_i), and t3t_3 the final multi-token answer generation time given top-kk passages. The relative time overhead added by reranking NN candidate passages is:

    Overhead Ratio=N⋅t2t1+t3\text{Overhead Ratio} = \frac{N \cdot t_2}{t_1 + t_3}

    Because relevance scoring only evaluates the probability of a single token (True) over short inputs without autoregressive token-by-token decoding, t2≪t3t_2 \ll t_3. Empirical evaluation shows that reranking N=20N=20 to N=100N=100 candidates improves exact match accuracy by 5.9%5.9\% to 9.1%9.1\% across NQ, TriviaQA, and HotpotQA while increasing overall query latency by 0.9×0.9\times to 6.0×6.0\times, far below the theoretical upper bound of an NN-fold latency expansion.

Coverage note — Case study qualitative examples and dataset licensing checklist responses were omitted as standalone knowls because their quantitative conclusions and method specifications are fully captured in the main empirical, algorithmic, and data table knowls.

References

  1. 1.Adlakha, V., Dhuliawala, S., Suleman, K., de Vries, H., and Reddy, S. Topiocqa: Open-domain conversational question answering with topic switching. TACL, 2022.
  2. 2.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  3. 3.Anthropic. Model card and evaluations for claude models. 2023.
  4. 4.Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. In ICLR, 2024a.
  5. 5.Asai, A., Zhong, Z., Chen, D., Koh, P. W., Zettlemoyer, L., Hajishirzi, H., and Yih, W.-t. Reliable, adaptable, and attributable language models with retrieval. arXiv preprint arXiv:2403.03187, 2024b.
  6. 6.Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., Nguyen, T., et al. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268, 2016.
  7. 7.Berant, J., Chou, A., Frostig, R., and Liang, P. Semantic parsing on freebase from question-answer pairs. In EMNLP, 2013.
  8. 8.Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al. Improving language models by retrieving from trillions of tokens. In ICML. PMLR, 2022.
  9. 9.Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation, 2023a.
  10. 10.Chen, Z., Cano, A. H., Romanou, A., Bonnet, A., Matoba, K., Salvi, F., Pagliardini, M., Fan, S., Köpf, A., Mohtashami, A., et al. Meditron-70b: Scaling medical pretraining for large language models. arXiv preprint arXiv:2311.16079, 2023b.
  11. 11.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. JMLR, 25(70), 2024.
  12. 12.Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R. Free Dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.
  13. 13.Dasigi, P., Liu, N. F., Marasovic, A., Smith, N. A., and Gardner, M. Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. In EMNLP, 2019.
  14. 14.DeepSeek. Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.
  15. 15.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. Glam: Efficient scaling of language models with mixture-of-experts. In ICML, 2022.
  16. 16.Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In NAACL, 2019.
  17. 17.Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M. Eli5: Long form question answering. In ACL, 2019.
  18. 18.Feng, S., Wan, H., Gunasekara, C., Patel, S., Joshi, S., and Lastras, L. doc2dial: A goal-oriented document-grounded dialogue dataset. In EMNLP, 2020.
  19. 19.Glass, M., Rossiello, G., Chowdhury, M. F. M., Naik, A., Cai, P., and Gliozzo, A. Re2G: Retrieve, rerank, generate. In NAACL, 2022.
  20. 20.Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M. Retrieval augmented language model pre-training. In ICML, 2020.
  21. 21.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. In ICLR, 2021.
  22. 22.Ho, X., Nguyen, A.-K. D., Sugawara, S., and Aizawa, A. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. In COLING, 2020.
  23. 23.Honovich, O., Scialom, T., Levy, O., and Schick, T. Unnatural instructions: Tuning language models with (almost) no human labor. In ACL, 2023.
  24. 24.Huang, J., Ping, W., Xu, P., Shoeybi, M., Chang, K. C.-C., and Catanzaro, B. Raven: In-context learning with retrieval augmented encoder-decoder language models. arXiv preprint arXiv:2308.07922, 2023.
  25. 25.Izacard, G. and Grave, E. Leveraging passage retrieval with generative models for open domain question answering. In EACL, 2021.
  26. 26.Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E. Unsupervised dense information retrieval with contrastive learning. TMLR, 2022.
  27. 27.Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E. Atlas: Few-shot learning with retrieval augmented language models. JMLR, 24(251):1–43, 2023.
  28. 28.Jeong, S., Baek, J., Cho, S., Hwang, S. J., and Park, J. C. Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity. In NAACL, 2024.
  29. 29.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  30. 30.Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G. Active retrieval augmented generation. In EMNLP, 2023.
  31. 31.Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421, 2021.
  32. 32.Jin, Q., Dhingra, B., Liu, Z., Cohen, W., and Lu, X. Pubmedqa: A dataset for biomedical research question answering. In EMNLP, 2019.
  33. 33.Jin, Q., Kim, W., Chen, Q., Comeau, D. C., Yeganova, L., Wilbur, W. J., and Lu, Z. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics, 39(11), 2023.
  34. 34.Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In ACL, 2017.
  35. 35.Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. Dense passage retrieval for open-domain question answering. In EMNLP, 2020.
  36. 36.Kasai, J., Sakaguchi, K., yoichi takahashi, Bras, R. L., Asai, A., Yu, X. V., Radev, D., Smith, N. A., Choi, Y., and Inui, K. Realtime QA: What’s the answer right now? In NeurIPS, 2023.
  37. 37.Khalifa, M., Logeswaran, L., Lee, M., Lee, H., and Wang, L. Few-shot reranking for multi-hop QA via language model prompting. In ACL, 2023.
  38. 38.Khattab, O., Santhanam, K., Li, X. L., Hall, D., Liang, P., Potts, C., and Zaharia, M. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp. arXiv preprint arXiv:2212.14024, 2022.
  39. 39.Kim, H., Hessel, J., Jiang, L., Lu, X., Yu, Y., Zhou, P., Bras, R. L., Alikhani, M., Kim, G., Sap, M., et al. Soda: Million-scale dialogue distillation with social commonsense contextualization. In EMNLP, 2023.
  40. 40.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  41. 41.Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E. The narrativeqa reading comprehension challenge. TACL, 2018.
  42. 42.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., et al. Natural questions: a benchmark for question answering research. TACL, 2019.
  43. 43.Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., ES, S., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., and Mattick, A. Openassistant conversations - democratizing large language model alignment. arXiv preprint arXiv: 2304.07327, 2023.
  44. 44.Lazaridou, A., Gribovskaya, E., Stokowiec, W., and Grigorev, N. Internet-augmented language models through few-shot prompting for open-domain question answering. arXiv preprint arXiv:2203.05115, 2022.
  45. 45.Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W. Nv-embed: Improved techniques for training llms as generalist embedding models. arXiv preprint arXiv:2405.17428, 2024.
  46. 46.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. NeurIPS, 33, 2020.
  47. 47.Lin, K., Tafjord, O., Clark, P., and Gardner, M. Reasoning over paragraph effects in situations. In Workshop on Machine Reading for Question Answering, 2019.
  48. 48.Lin, S.-C., Asai, A., Li, M., Oguz, B., Lin, J., Mehdad, Y., Yih, W.-t., and Chen, X. How to train your dragon: Diverse augmentation towards generalizable dense retrieval. In Findings of EMNLP, 2023.
  49. 49.Lin, X. V., Chen, X., Chen, M., Shi, W., Lomeli, M., James, R., Rodriguez, P., Kahn, J., Szilvasy, G., Lewis, M., Zettlemoyer, L., and tau Yih, W. RA-DIT: Retrieval-augmented dual instruction tuning. In ICLR, 2024.
  50. 50.Liu, Z., Ping, W., Roy, R., Xu, P., Shoeybi, M., and Catanzaro, B. Chatqa: Surpassing gpt-4 on conversational qa and rag. arXiv preprint arXiv:2401.10225, 2024.
  51. 51.Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., et al. The flan collection: Designing data and methods for effective instruction tuning. In ICML, 2023.
  52. 52.Luan, Y., Eisenstein, J., Toutanova, K., and Collins, M. Sparse, dense, and attentional representations for text retrieval. TACL, 2021.
  53. 53.Luo, H., Chuang, Y.-S., Gong, Y., Zhang, T., Kim, Y., Wu, X., Fox, D., Meng, H., and Glass, J. Sail: Search-augmented instruction learning. arXiv preprint arXiv:2305.15225, 2023.
  54. 54.Ma, X., Wang, L., Yang, N., Wei, F., and Lin, J. Fine-tuning llama for multi-stage text retrieval. arXiv preprint arXiv:2310.08319, 2023.
  55. 55.Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., and Hajishirzi, H. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In ACL, 2023.
  56. 56.Menon, A., Jayasumana, S., Rawat, A. S., Kim, S., Reddi, S., and Kumar, S. In defense of dual-encoders for neural ranking. In ICML, 2022.
  57. 57.Meta-AI. Llama 3 model card. 2024.
  58. 58.Mistral. Mixtral 8x22b. 2024. URL https://mistral.ai/news/mixtral-8x22b/.
  59. 59.Mitra, B., Craswell, N., et al. An introduction to neural information retrieval. Foundations and Trends® in Information Retrieval, 2018.
  60. 60.Muennighoff, N., Su, H., Wang, L., Yang, N., Wei, F., Yu, T., Singh, A., and Kiela, D. Generative representational instruction tuning. arXiv preprint arXiv:2402.09906, 2024.
  61. 61.Nogueira, R., Jiang, Z., Pradeep, R., and Lin, J. Document ranking with a pretrained sequence-to-sequence model. In Findings of EMNLP, 2020.
  62. 62.OpenAI. Introducing ChatGPT, 2022.
  63. 63.OpenAI. GPT-4, 2023.
  64. 64.Oren, Y., Meister, N., Chatterji, N. S., Ladhak, F., and Hashimoto, T. Proving test set contamination in black-box language models. In ICLR, 2024.
  65. 65.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. NeurIPS, 35, 2022.
  66. 66.Pal, A., Umapathi, L. K., and Sankarasubbu, M. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In CHIL, 2022.
  67. 67.Petroni, F., Piktus, A., Fan, A., Lewis, P., Yazdani, M., De Cao, N., Thorne, J., Jernite, Y., Karpukhin, V., Maillard, J., Plachouras, V., Rocktäschel, T., and Riedel, S. KILT: a benchmark for knowledge intensive language tasks. In NAACL, 2021.
  68. 68.Qin, Z., Jagerman, R., Hui, K., Zhuang, H., Wu, J., Shen, J., Liu, T., Liu, J., Metzler, D., Wang, X., et al. Large language models are effective text rankers with pairwise ranking prompting. In Findings of NAACL, 2024.
  69. 69.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100,000+ questions for machine comprehension of text. In EMNLP, 2016.
  70. 70.Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y. In-context retrieval-augmented language models. TACL, 2023.
  71. 71.Robertson, S., Zaragoza, H., and Taylor, M. Simple bm25 extension to multiple weighted fields. In CIKM, 2004.
  72. 72.Sachan, D. S., Reddy, S., Hamilton, W. L., Dyer, C., and Yogatama, D. End-to-end training of multi-document reader and retriever for open-domain question answering. In NeurIPS, 2021.
  73. 73.Shao, Z., Gong, Y., Shen, Y., Huang, M., Duan, N., and Chen, W. Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy. In Findings of EMNLP, 2023.
  74. 74.Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t. Replug: Retrieval-augmented black-box language models. In NAACL, 2024.
  75. 75.Sun, W., Yan, L., Ma, X., Wang, S., Ren, P., Chen, Z., Yin, D., and Ren, Z. Is ChatGPT good at search? investigating large language models as re-ranking agents. In EMNLP, 2023.
  76. 76.Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., and Gurevych, I. Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In NeurIPS, 2021.
  77. 77.Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A. Fever: A large-scale dataset for fact extraction and verification. In NAACL, 2018.
  78. 78.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  79. 79.Trischler, A., Wang, T., Yuan, X., Harris, J., Sordoni, A., Bachman, P., and Suleman, K. Newsqa: A machine comprehension dataset. In RepL4NLP Workshop at ACL, 2017.
  80. 80.Trivedi, H., Balasubramanian, N., Khot, T., and Sabharwal, A. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. In ACL, 2023.
  81. 81.Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers, M. R., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., et al. An overview of the bioasq large-scale biomedical semantic indexing and question answering competition. BMC bioinformatics, 2015.
  82. 82.Wang, B., Ping, W., McAfee, L., Xu, P., Li, B., Shoeybi, M., and Catanzaro, B. Instructretro: Instruction tuning post retrieval-augmented pretraining. In ICML, 2024.
  83. 83.Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., and Wei, F. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533, 2022.
  84. 84.Wang, L., Yang, N., Huang, X., Yang, L., Majumder, R., and Wei, F. Improving text embeddings with large language models. arXiv preprint arXiv:2401.00368, 2023a.
  85. 85.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language models with self-generated instructions. In ACL, 2023b.
  86. 86.Wang, Z., Araki, J., Jiang, Z., Parvez, M. R., and Neubig, G. Learning to filter context for retrieval-augmented generation. arXiv preprint arXiv:2311.08377, 2023c.
  87. 87.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. In ICLR, 2022.
  88. 88.Wu, C., Lin, W., Zhang, X., Zhang, Y., Xie, W., and Wang, Y. Pmc-llama: toward building open-source language models for medicine. JAMIA, 2024.
  89. 89.Wu, Z., Parish, R., Cheng, H., Min, S., Ammanabrolu, P., Ostendorf, M., and Hajishirzi, H. Inscit: Information-seeking conversations with mixed-initiative interactions. TACL, 2023.
  90. 90.Xiong, G., Jin, Q., Lu, Z., and Zhang, A. Benchmarking retrieval-augmented generation for medicine. arXiv preprint arXiv:2402.13178, 2024.
  91. 91.Xu, F., Shi, W., and Choi, E. RECOMP: Improving retrieval-augmented LMs with context compression and selective augmentation. In ICLR, 2024a.
  92. 92.Xu, P., Ping, W., Wu, X., McAfee, L., Zhu, C., Liu, Z., Subramanian, S., Bakhturina, E., Shoeybi, M., and Catanzaro, B. Retrieval meets long context large language models. In ICLR, 2024b.
  93. 93.Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In EMNLP, 2018.
  94. 94.Yoran, O., Wolfson, T., Ram, O., and Berant, J. Making retrieval-augmented language models robust to irrelevant context. In ICLR, 2024.
  95. 95.Yu, W., Iter, D., Wang, S., Xu, Y., Ju, M., Sanyal, S., Zhu, C., Zeng, M., and Jiang, M. Generate rather than retrieve: Large language models are strong context generators. In ICLR, 2023a.
  96. 96.Yu, W., Zhang, H., Pan, X., Ma, K., Wang, H., and Yu, D. Chain-of-note: Enhancing robustness in retrieval-augmented language models. arXiv preprint arXiv:2311.09210, 2023b.
  97. 97.Yu, W., Zhang, Z., Liang, Z., Jiang, M., and Sabharwal, A. Improving language models via plug-and-play retrieval feedback, 2024.
  98. 98.Yu, Y., Xiong, C., Sun, S., Zhang, C., and Overwijk, A. Coco-dr: Combating distribution shift in zero-shot dense retrieval with contrastive and distributionally robust learning. In EMNLP, 2022.
  99. 99.Zhang, T., Patil, S. G., Jain, N., Shen, S., Zaharia, M., Stoica, I., and Gonzalez, J. E. Raft: Adapting language model to domain specific rag. arXiv preprint arXiv:2403.10131, 2024.
  100. 100.Zhu, F., Lei, W., Huang, Y., Wang, C., Zhang, S., Lv, J., Feng, F., and Chua, T.-S. Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance. In ACL, 2021.
  101. 101.Zhu, Y., Zhang, P., Zhang, C., Chen, Y., Xie, B., Dou, Z., Liu, Z., and Wen, J.-R. Inters: Unlocking the power of large language models in search with instruction tuning. In ACL, 2024.

Citation

MLA
Yu, Y., et al. “RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs”. arXiv, 2024, http://arxiv.org/abs/2407.02485v1.
APA
Yu, Y., Ping, W., Liu, Z., Wang, B., You, J., Zhang, C., Shoeybi, M., & Catanzaro, B. (2024). RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs. arXiv. http://arxiv.org/abs/2407.02485v1
Chicago
Yu, Y., W. Ping, Z. Liu, et al. 2024. “RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs”. arXiv. http://arxiv.org/abs/2407.02485v1.
Harvard
Yu, Y. et al. (2024) “RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2407.02485v1.
Vancouver
1. Yu Y, Ping W, Liu Z, Wang B, You J, Zhang C, Shoeybi M, Catanzaro B (2024) RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs. arXiv

BibTeX

@article{yu2024rankrag,
  title = {RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs},
  author = {Yu, Yue and Ping, Wei and Liu, Zihan and Wang, Boxin and You, Jiaxuan and Zhang, Chao and Shoeybi, Mohammad and Catanzaro, Bryan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2407.02485v1},
  eprint = {2407.02485}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors