Re2G: Retrieve, Rerank, Generate
Michael R. GlassGaetano RossielloMd. Faisal Mahbub ChowdhuryAnkita NaikPengshan CaiAlfio Gliozzo
Proposes an end-to-end architecture that integrates neural and keyword passage retrieval with a reranking stage into sequence-to-sequence generation, achieving substantial performance gains across multiple knowledge-intensive NLP tasks on the KILT benchmark using a novel knowledge distillation training scheme.
Large language models require vast amounts of factual knowledge to handle information-heavy tasks such as answering questions, verifying facts, and conducting informed dialogue. However, continuously expanding the parameter size of neural networks is computationally expensive and memory-intensive. Incorporating external knowledge retrieval allows models to scale their access to facts without proportionate increases in hardware and computational costs. Existing retrieval-augmented generation systems remain limited by their initial search accuracy and their difficulty in combining diverse search mechanisms.
The article demonstrates the effectiveness of a novel framework called Retrieve, Rerank, Generate (Re2G), which integrates neural initial retrieval, keyword search, and neural reranking into an end-to-end conditional text generation system. The primary objective is to evaluate whether adding an intermediate reranking stage and ensembling search methods improves both passage retrieval and final generation quality across diverse knowledge-intensive benchmarks.
To evaluate this framework, the authors conducted experiments using the standardized KILT benchmark, testing across four distinct tasks: slot filling, question answering, fact checking, and dialogue. The Re2G architecture combines initial candidate passages retrieved via both neural dense retrieval and traditional keyword-based search (BM25). A neural interaction reranker evaluates these candidates jointly with the query, passing the top five passages to a sequence-to-sequence generator based on BART. To train the entire pipeline end-to-end, the authors introduced an online knowledge distillation method where the reranker acts as a teacher to guide the initial neural retrieval system using target output data.
The evaluation revealed substantial performance gains across multiple tasks. First, Re2G established new state-of-the-art results on headline metrics across five diverse benchmark datasets, achieving relative improvements of 34% on TriviaQA, 31% on Natural Questions, 22% on FEVER fact checking, 10% on Wizard of Wikipedia dialogue, and 9% on T-REx slot filling. Second, reranking significantly boosted retrieval precision across all tasks compared to single-stage retrieval. Third, ablation analyses confirmed that combining keyword search with dense retrieval improved downstream results in four out of five datasets, demonstrating the value of merging disparate retrieval scoring systems. Finally, the analysis showed that between 28% and 68% of output improvements directly resulted from better passage retrieval, while the remainder stemmed from generative models learning better reasoning when trained alongside superior retrieval components.
These findings indicate that generative AI performance can be substantially improved through better document retrieval and reranking architectures rather than simply increasing the size of language models. For organizations deploying knowledge-driven systems, this approach reduces computational overhead while improving factual reliability and provenance tracking. It also demonstrates that classic keyword search remains highly valuable when paired with modern neural rerankers.
Based on these results, engineering teams developing knowledge-intensive generative applications should adopt multi-stage retrieval pipelines that combine dense and sparse search alongside neural rerankers. Organizations can immediately leverage the authors' open-source codebase to evaluate domain-specific implementations. Further work should focus on testing the domain adaptation of this architecture on proprietary enterprise data and evaluating performance on broader, non-standardized knowledge sources.
Confidence in these findings is supported by consistent gains across multiple standardized benchmarks and rigorous ablation testing. However, decision-makers should note certain limitations: performance in dialogue tasks is less pronounced due to ambiguous and noisy reference datasets, and error analyses revealed that many apparent model failures stemmed from incomplete ground-truth annotations rather than generation errors. In addition, the system requires specialized hardware with substantial memory (e.g., 128 GB) to host and index large-scale passage collections.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Re2G builds directly on RAG’s retrieval-conditioned generation setup, so this paper establishes the model and pipeline it extends.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). Re2G develops the neural-retrieval approach represented by REALM, making REALM’s learned external-memory framework useful groundwork.
- Paper: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, Mike Lewis et al. (2020). Re2G uses BART as its sequence-to-sequence generator, so BART explains the pretrained architecture at the heart of its system.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). RankRAG carries Re2G’s retrieve-rerank-generate idea forward by training one language model to rank retrieved contexts and generate answers.
