Re2G: Retrieve, Rerank, Generate

Michael R. GlassGaetano RossielloMd. Faisal Mahbub ChowdhuryAnkita NaikPengshan CaiAlfio Gliozzo

article2022NAACL180 citations

Proposes an end-to-end architecture that integrates neural and keyword passage retrieval with a reranking stage into sequence-to-sequence generation, achieving substantial performance gains across multiple knowledge-intensive NLP tasks on the KILT benchmark using a novel knowledge distillation training scheme.

Listen

Large language models require vast amounts of factual knowledge to handle information-heavy tasks such as answering questions, verifying facts, and conducting informed dialogue. However, continuously expanding the parameter size of neural networks is computationally expensive and memory-intensive. Incorporating external knowledge retrieval allows models to scale their access to facts without proportionate increases in hardware and computational costs. Existing retrieval-augmented generation systems remain limited by their initial search accuracy and their difficulty in combining diverse search mechanisms.

The article demonstrates the effectiveness of a novel framework called Retrieve, Rerank, Generate (Re2G), which integrates neural initial retrieval, keyword search, and neural reranking into an end-to-end conditional text generation system. The primary objective is to evaluate whether adding an intermediate reranking stage and ensembling search methods improves both passage retrieval and final generation quality across diverse knowledge-intensive benchmarks.

To evaluate this framework, the authors conducted experiments using the standardized KILT benchmark, testing across four distinct tasks: slot filling, question answering, fact checking, and dialogue. The Re2G architecture combines initial candidate passages retrieved via both neural dense retrieval and traditional keyword-based search (BM25). A neural interaction reranker evaluates these candidates jointly with the query, passing the top five passages to a sequence-to-sequence generator based on BART. To train the entire pipeline end-to-end, the authors introduced an online knowledge distillation method where the reranker acts as a teacher to guide the initial neural retrieval system using target output data.

The evaluation revealed substantial performance gains across multiple tasks. First, Re2G established new state-of-the-art results on headline metrics across five diverse benchmark datasets, achieving relative improvements of 34% on TriviaQA, 31% on Natural Questions, 22% on FEVER fact checking, 10% on Wizard of Wikipedia dialogue, and 9% on T-REx slot filling. Second, reranking significantly boosted retrieval precision across all tasks compared to single-stage retrieval. Third, ablation analyses confirmed that combining keyword search with dense retrieval improved downstream results in four out of five datasets, demonstrating the value of merging disparate retrieval scoring systems. Finally, the analysis showed that between 28% and 68% of output improvements directly resulted from better passage retrieval, while the remainder stemmed from generative models learning better reasoning when trained alongside superior retrieval components.

These findings indicate that generative AI performance can be substantially improved through better document retrieval and reranking architectures rather than simply increasing the size of language models. For organizations deploying knowledge-driven systems, this approach reduces computational overhead while improving factual reliability and provenance tracking. It also demonstrates that classic keyword search remains highly valuable when paired with modern neural rerankers.

Based on these results, engineering teams developing knowledge-intensive generative applications should adopt multi-stage retrieval pipelines that combine dense and sparse search alongside neural rerankers. Organizations can immediately leverage the authors' open-source codebase to evaluate domain-specific implementations. Further work should focus on testing the domain adaptation of this architecture on proprietary enterprise data and evaluating performance on broader, non-standardized knowledge sources.

Confidence in these findings is supported by consistent gains across multiple standardized benchmarks and rigorous ablation testing. However, decision-makers should note certain limitations: performance in dialogue tasks is less pronounced due to ambiguous and noisy reference datasets, and error analyses revealed that many apparent model failures stemmed from incomplete ground-truth annotations rather than generation errors. In addition, the system requires specialized hardware with substantial memory (e.g., 128 GB) to host and index large-scale passage collections.

Cover for Re2G: Retrieve, Rerank, Generate

Abstract

As demonstrated by GPT-3 and T5, transformers grow in capability as parameter spaces become larger and larger. However, for tasks that require a large amount of knowledge, non-parametric memory allows models to grow dramatically with a sub-linear increase in computational cost and GPU memory requirements. Recent models such as RAG and REALM have introduced retrieval into conditional generation. These models incorporate neural initial retrieval from a corpus of passages. We build on this line of research, proposing Re²G, which combines both neural initial retrieval and reranking into a BART-based sequence-to-sequence generation. Our reranking approach also permits merging retrieval results from sources with incomparable scores, enabling an ensemble of BM25 and neural initial retrieval. To train our system end-to-end, we introduce a novel variation of knowledge distillation to train the initial retrieval, reranker and generation using only ground truth on the target sequence output. We find large gains in four diverse tasks: zero-shot slot filling, question answering, fact checking and dialog, with relative gains of 9% to 34% over the previous state-of-the-art on the KILT leaderboard. We make our code available as open source¹.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Reranker
  • 3.2 Training
  • 3.3 Reranking Training
  • 3.4 End-to-End Training
  • 3.5 Inference
  • 4 Experiments
  • 4.1 Retrieval
  • 4.2 Ablations
  • 5 Analysis
  • 5.1 Slot filling error analysis
  • 6 Conclusions
  • References
  • Appendix
  • A Hyperparameters
  • B Software Details
  • C Model Details
  • D.1 Generation Quality
  • D Generation Analysis

Knowls

  1. Knowl 1 — Re2G reranks heterogeneous retrieval results before generation

    model/method

    Re2G (Retrieve, Rerank, Generate) is a retrieval-augmented sequence-to-sequence system that places a learned reranker between retrieval and generation. It collects candidate passages from neural dense retrieval and, optionally, a keyword retriever such as BM25; the reranker scores the candidates together, so passages from retrieval sources with incomparable score scales can be combined without directly comparing those source scores. The highest-ranked passages are paired with the query and supplied to BART. During generation, the reranker scores weight the contributions of the passage-conditioned outputs. This design uses scalable initial retrieval to form candidates and a more computationally intensive reranker to select evidence for generation.

  2. Knowl 2 — Re2G combines independent dense encoders with a cross-attention reranker

    model/method

    The initial DPR retriever encodes a query and each passage independently with separate BERT encoders, then scores a query–passage pair by the inner product of their representation vectors. Because passage vectors can be precomputed and stored in an approximate-nearest-neighbor index, this representation model can search a large corpus efficiently. Re2G's reranker is instead an interaction model: it jointly encodes the query and passage with BERT, allowing cross-attention between their tokens, and outputs a relevance score for the pair. The reranker was initialized from a BERT model trained on MS MARCO by NBoost. Re2G uses three BERT-base models—the query encoder, passage encoder, and reranker, each with 110 million parameters—and a 400-million-parameter BART-large generator, for 730 million parameters in total.

  3. Knowl 3 — Re2G trains retrieval, reranking, and generation in four phases

    model/method

    A KILT training instance consists of a query qq, target sequence tt, and provenance set Prov\mathrm{Prov} of passages supporting the target. Re2G trains in this order: (1) DPR training uses triples (q,p+,p−)(q,p^+,p^-), where p+∈Provp^+\in\mathrm{Prov} and p−∈BM25(q)∖Provp^-\in\mathrm{BM25}(q)\setminus\mathrm{Prov}; (2) generation training trains BART and the query encoder using the target sequence; (3) reranker training scores the merged DPR and BM25 candidate set PP using provenance labels; and (4) end-to-end training uses query–target pairs (q,t)(q,t) rather than provenance labels. DPR Stage 1 uses batches of 128; BM25 hard negatives and positive and negative passages from other batch instances provide negatives, and the loss is the negative log-likelihood of the positive passage. During reranker training, multiple provenance passages can be positive, so the loss sums their negative log probabilities:

    Lrank=−∑i∈Provlog⁡(softmax⁡(z)i).\mathcal{L}_{\mathrm{rank}}=-\sum_{i\in\mathrm{Prov}}\log\left(\operatorname{softmax}(z)_i\right).

    Here zz is the vector of reranker logits over candidate passages, and ii indexes a candidate that is in the ground-truth provenance set. The DPR and generation phases follow the KGI0 training procedure; the reranker and final end-to-end phases extend it.

  4. Knowl 4 — Online distillation trains DPR when generation gradients do not reach it

    model/method

    In Re2G end-to-end training, the reranker scores—not the DPR similarity scores—weight passage-conditioned generation. Consequently, the target-output loss can train the reranker and generator but supplies no gradient to the DPR query encoder through passage marginalization. Re2G addresses this with online knowledge distillation: while the reranker is trained, it acts as a teacher that supplies soft passage labels to the DPR student across the retrieved candidates. With student logits zsz_s, teacher logits ztz_t, and temperature TT, the paper gives the distillation loss as

    LKD=DKL(softmax⁡(zs/T)  ∥  softmax⁡(zt/T))T2.\mathcal{L}_{\mathrm{KD}}=D_{\mathrm{KL}}\left(\operatorname{softmax}(z_s/T)\;\middle\|\;\operatorname{softmax}(z_t/T)\right)T^2.

    The temperature softens the distributions; the reported setting is T=10T=10 with a knowledge-distillation learning-rate scaling of 1.01.0. The soft labels convey degrees of relevance rather than only positive/negative status, and supervise more retrieved passages than the five used for generation. The authors found that adding DPR and reranker scores together instead made DPR learn a complementary signal that could rank relevant passages poorly, reducing retrieval and overall performance. Freezing the query encoder was another option and worked best for Wizard of Wikipedia, but distillation was the proposed way to continue training DPR.

  5. Knowl 5 — Inference uses two top-12 candidate lists and five passage-conditioned outputs

    algorithm

    For each input query, Re2G retrieves the top 12 passages from its DPR HNSW index and the top 12 from Anserini BM25. It scores the resulting candidates with the reranker, selects the top five, and pairs each selected passage with the query as input to BART-large. BART uses beam size 6, length penalty 1.0, minimum output length 2, and maximum output length 64. The five passage-conditioned output sequences are combined using weights from the softmax of their reranker scores. Thus, retrieval supplies up to 24 candidates, reranking selects the five generation inputs, and reranker probabilities determine their relative influence on the final output.

  6. Knowl 6 — Evaluation covers five KILT datasets across four task types

    experimental setup

    Re2G was evaluated on T-REx slot filling, Natural Questions and TriviaQA open-domain question answering, FEVER fact checking, and Wizard of Wikipedia dialogue, using the KILT Wikipedia knowledge source and its provenance evaluation. T-REx has 2.3 million instances, but the authors trained on a downsampled 370,000; development and test each contain 5,000. Natural Questions has 87,000 training, 3,000 development, and 1,400 test questions; TriviaQA has 62,000, 5,000, and 6,500. FEVER has about 100,000 training and 10,000 each for development and test; the KILT version omits the NOTENOUGHINFO class. Wizard of Wikipedia has about 64,000 training and 3,000 each for development and test. Provenance is assessed by R-Precision and Recall@5; generated output is assessed by accuracy and token-level F1, except dialogue uses ROUGE-L instead of accuracy. KILT metrics count output quality only when provenance is correct. The reported training settings use learning rates of 5×10−55\times10^{-5} for DPR and 3×10−53\times10^{-5} for reranking and generation; batch sizes are 128, 32, and 128, respectively, with 2 DPR epochs and 1 reranker and generation epoch. The main results use a single run with random seed 42.

  7. Knowl 7 — Re2G reports leading KILT scores on four tasks

    empirical result

    The headline KILT metrics in the paper's leaderboard comparison show Re2G with the following scores. The table also gives the strongest comparison score shown for each task in that leaderboard table; the Wizard of Wikipedia comparison is Hindsight, which was reported after Re2G's submission.

    Task Headline metric Re2G Comparison system Comparison score
    T-REx KILT-F1 77.05 KGI1 70.58
    Natural Questions KILT-F1 49.80 SEAL 44.40
    TriviaQA KILT-F1 61.78 SEAL 54.99
    FEVER KILT-Accuracy 78.53 SEAL 71.28
    Wizard of Wikipedia KILT-ROUGE-L 11.39 Hindsight 11.92

    The authors report historical relative gains of 9%, 31%, 34%, 22%, and 10% for T-REx, Natural Questions, TriviaQA, FEVER, and Wizard of Wikipedia, respectively, over the previous state of the art used for those claims. In the leaderboard comparison reproduced above, Re2G leads the listed systems on the first four tasks; Hindsight's 11.92 exceeds Re2G's 11.39 on Wizard of Wikipedia.

  8. Knowl 8 — Reranking substantially improves retrieval, while end-to-end effects vary by task

    empirical result

    On the development sets, provenance retrieval was measured by R-Precision and Recall@5. The table compares KGI0 DPR, a reranker trained on provenance labels before end-to-end training (Reranker Stage 1), and the final Re2G reranker after end-to-end training. Reranker Stage 1 improves both metrics over KGI0 DPR on all five datasets. Subsequent end-to-end training has mixed retrieval effects: it raises both metrics on FEVER, raises R-Precision but slightly lowers Recall@5 on Wizard of Wikipedia, is approximately flat on T-REx and Natural Questions, and sharply lowers TriviaQA retrieval scores even though the paper reports improved answer accuracy and F1 there.

    Dataset Retrieval system R-Precision Recall@5
    T-REx KGI0 DPR 65.02 75.52
    T-REx Reranker Stage 1 81.22 87.00
    T-REx Re2G reranker 81.24 88.58
    Natural Questions KGI0 DPR 64.65 69.60
    Natural Questions Reranker Stage 1 70.78 73.05
    Natural Questions Re2G reranker 70.92 74.79
    TriviaQA KGI0 DPR 60.55 63.65
    TriviaQA Reranker Stage 1 71.80 71.98
    TriviaQA Re2G reranker 60.37 70.61
    FEVER KGI0 DPR 80.34 86.53
    FEVER Reranker Stage 1 87.71 92.43
    FEVER Re2G reranker 90.06 92.91
    Wizard of Wikipedia KGI0 DPR 48.04 71.02
    Wizard of Wikipedia Reranker Stage 1 55.50 74.98
    Wizard of Wikipedia Re2G reranker 57.89 74.62

    The authors suggest that the TriviaQA retrieval decline may reflect incomplete provenance ground truth: retrieved passages can help answer generation even when their ranking is not rewarded by the provenance metric.

  9. Knowl 9 — Ablations show benefits from reranking, distillation, and BM25 on most tasks

    empirical result

    Development-set ablations report point estimates with 95% confidence intervals. Re2G-KD removes online knowledge distillation and freezes the query encoder during end-to-end training. Re2G-BM25 removes BM25 candidates and instead retrieves 24 passages from DPR before reranking. KGI0 has no reranker, BM25 candidates, or online distillation. The table reports the relevant KILT headline metric for each task. Relative to KGI0, full Re2G improves the headline score on every dataset. Removing distillation lowers the score on Natural Questions and FEVER, is close to full Re2G on T-REx and TriviaQA, and improves Wizard of Wikipedia. Removing BM25 lowers scores on four datasets; Natural Questions is the exception, where Re2G-BM25 is essentially tied with full Re2G.

    Dataset KILT metric Re2G Re2G-KD Re2G-BM25 KGI0
    T-REx KILT-F1 77.08±\pm1.15 77.00±\pm1.15 67.93±\pm1.28 61.38±\pm1.34
    Natural Questions KILT-F1 50.90±\pm1.76 49.93±\pm1.76 50.91±\pm1.76 42.87±\pm1.75
    TriviaQA KILT-F1 60.91±\pm1.27 60.84±\pm1.28 58.37±\pm1.29 47.35±\pm1.31
    FEVER KILT-Accuracy 80.56±\pm0.76 80.14±\pm0.77 78.74±\pm0.78 70.06±\pm0.88
    Wizard of Wikipedia KILT-ROUGE-L 11.37±\pm0.58 11.61±\pm0.58 11.13±\pm0.57 9.48±\pm0.53

    These ablations support the authors' broader finding that both BM25 ensembling and online distillation helped on four of the five datasets, but neither helped uniformly.

  10. Knowl 10 — Output gains are not always explained by better passage ranking, and labels are imperfect

    limitation

    The authors examined cases where Re2G generated better output than KGI0 and measured how often Re2G also ranked the first correct passage higher. This coincidence occurred in 67.73% of output improvements on T-REx, 61.08% on Natural Questions, and 66.86% on FEVER, but only 36.86% on TriviaQA and 27.74% on Wizard of Wikipedia. Thus, improved ranking accompanied many output gains, especially on three tasks, but did not account for all of them. In a manual sample of 50 T-REx development errors with zero accuracy and F1, 33 were attributed to incomplete ground truth: 19 involved ambiguous head entities and 16 involved relations with multiple valid fillers. The authors also note that TriviaQA retrieval scores can fall despite better answer generation, suggesting that provenance labels may be incomplete. In a separate Wizard of Wikipedia inspection, 5 of 20 ground-truth target texts were judged inconsistent; Re2G and KGI0 had similar counts judged GOOD/OK/INCONSISTENT (8/2/10 and 9/2/9, respectively). These observations limit how directly benchmark provenance and text-match metrics can be interpreted as measures of evidence quality or conversational coherence.

Coverage note — No substantial contributed method or result was deliberately omitted; software-version details and the code release are excluded as reproducibility metadata rather than standalone scientific findings.

References

  1. 1.Michele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Wen tau Yih, Sebastian Riedel, and Fabio Petroni. Autoregressive search engines: Generating substrings as document identifiers. ArXiv, abs/2204.10628, 2022.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In NeurIPS, 2020.
  3. 3.Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. Autoregressive entity retrieval. In International Conference on Learning Representations. OpenReview.net, 2021. URL https://openreview.net/forum?id=5k8F6UU39V.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  5. 5.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. Wizard of wikipedia: Knowledge-powered conversational agents. In International Conference on Learning Representations, 2018.
  6. 6.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. Wizard of wikipedia: Knowledge-powered conversational agents. In ICLR (Poster). OpenReview.net, 2019.
  7. 7.Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. T-REx: A large scale alignment of natural language with knowledge base triples. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan, May 2018. European Language Resources Association (ELRA). URL https://aclanthology.org/L18-1544.
  8. 8.P. Ferragina and G. Manzini. Opportunistic data structures with applications. In Proceedings 41st Annual Symposium on Foundations of Computer Science, pages 390–398, 2000. doi: 10.1109/SFCS.2000.892127.
  9. 9.Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, and Alfio Gliozzo. Robust retrieval augmented generation for zero-shot slot filling. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1939–1949, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.148. URL https://aclanthology.org/2021.emnlp-main.148.
  10. 10.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909, 2020.
  11. 11.Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. In NIPS Deep Learning and Representation Learning Workshop, 2015. URL http://arxiv.org/abs/1503.02531.
  12. 12.Gautier Izacard and Edouard Grave. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.74. URL https://aclanthology.org/2021.eacl-main.74.
  13. 13.Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with gpus. arXiv preprint arXiv:1702.08734, 2017.
  14. 14.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1147. URL https://aclanthology.org/P17-1147.
  15. 15.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. SpanBERT: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77, 2020. doi: 10.1162/tacl_a_00300. URL https://aclanthology.org/2020.tacl-1.5.
  16. 16.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.550. URL https://aclanthology.org/2020.emnlp-main.550.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466, March 2019. doi: 10.1162/tacl_a_00276. URL https://aclanthology.org/Q19-1026.
  18. 18.Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, and Danqi Chen. Learning dense representations of phrases at scale. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6634–6647, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.518. URL https://aclanthology.org/2021.acl-long.518.
  19. 19.Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. Zero-shot relation extraction via reading comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 333–342, Vancouver, Canada, August 2017. Association for Computational Linguistics. doi: 10.18653/v1/K17-1034. URL https://aclanthology.org/K17-1034.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online, July 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.703. URL https://aclanthology.org/2020.acl-main.703.
  21. 21.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive nlp tasks. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 9459–9474. Curran Associates, Inc., 2020b.
  22. 22.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004.
  23. 23.Tie-Yan Liu. Learning to rank for information retrieval. Information Retrieval, 3(3):225–331, 2009.
  24. 24.Jean Maillard, Vladimir Karpukhin, Fabio Petroni, Wen-tau Yih, Barlas Oguz, Veselin Stoyanov, and Gargi Ghosh. Multi-task retrieval for knowledge-intensive tasks. In ACL/IJCNLP (1), pages 1098–1111. Association for Computational Linguistics, 2021.
  25. 25.Yu A Malkov and Dmitry A Yashunin. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE transactions on pattern analysis and machine intelligence, 42(4):824–836, 2018.
  26. 26.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. Ms marco: A human generated machine reading comprehension dataset. In CoCo@ NIPS, 2016.
  27. 27.Rodrigo Nogueira and Kyunghyun Cho. Passage reranking with bert. arXiv preprint arXiv:1901.04085, 2019.
  28. 28.Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 708–718, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.63. URL https://aclanthology.org/2020.findings-emnlp.63.
  29. 29.Ashwin Paranjape, Omar Khattab, Christopher Potts, Matei Zaharia, and Christopher D Manning. Hindsight: Posterior-guided training of retrievers for improved open-ended generation. arXiv preprint arXiv:2110.07752, 2021.
  30. 30.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2523–2544, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.200. URL https://aclanthology.org/2021.naacl-main.200.
  31. 31.Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Dmytro Okhonko, Samuel Broscheit, Gautier Izacard, Patrick Lewis, Barlas Oguz, Edouard Grave, Wen-tau Yih, et al. The web is your oyster–knowledge-intensive nlp against a very large web corpus. arXiv preprint arXiv:2112.09924, 2021.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, 2020.
  33. 33.Stephen Robertson and Hugo Zaragoza. The probabilistic relevance framework: Bm25 and beyond. Found. Trends Inf. Retr., 3(4):333–389, April 2009. ISSN 1554-0669. doi: 10.1561/1500000019. URL http://dx.doi.org/10.1561/1500000019.
  34. 34.Cole Thienes and Jack Pertschuk. Nboost: Neural boosting search results. https://github.com/koursaros-ai/nboost, 2019.
  35. 35.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. FEVER: a large-scale dataset for fact extraction and verification. In NAACL-HLT, pages 809–819. Association for Computational Linguistics, 2018a.
  36. 36.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. Fever: a large-scale dataset for fact extraction and verification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, 2018b.
  37. 37.James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal. The fact extraction and verification (FEVER) shared task. CoRR, abs/1811.10971, 2018c.
  38. 38.James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal. The fever2. 0 shared task. In Proceedings of the Second Workshop on Fact Extraction and VERification (FEVER), pages 1–6, 2019.
  39. 39.Lidan Wang, Jimmy Lin, and Donald Metzler. A cascade ranking model for efficient ranked retrieval. In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval, pages 105–114, 2011.
  40. 40.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4003–4012, Marseille, France, May 2020. European Language Resources Association. ISBN 979-10-95546-34-4. URL https://aclanthology.org/2020.lrec-1.494.
  41. 41.Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel, and Luke Zettlemoyer. Scalable zero-shot entity linking with dense entity retrieval. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6397–6407, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.519. URL https://aclanthology.org/2020.emnlp-main.519.

Citation

MLA
Glass, M., et al. “Re2G: Retrieve, Rerank, Generate”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 2701–15, https://doi.org/10.18653/v1/2022.naacl-main.194.
APA
Glass, M., Rossiello, G., Chowdhury, M. F. M., Naik, A. R., Cai, P., & Gliozzo, A. (2022). Re2G: Retrieve, Rerank, Generate. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2701–2715. https://doi.org/10.18653/v1/2022.naacl-main.194
Chicago
Glass, M., G. Rossiello, M. F. M. Chowdhury, A. R. Naik, P. Cai, and A. Gliozzo. 2022. “Re2G: Retrieve, Rerank, Generate”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2701–15. https://doi.org/10.18653/v1/2022.naacl-main.194.
Harvard
Glass, M. et al. (2022) “Re2G: Retrieve, Rerank, Generate”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 2701–2715. Available at: https://doi.org/10.18653/v1/2022.naacl-main.194.
Vancouver
1. Glass M, Rossiello G, Chowdhury MFM, Naik AR, Cai P, Gliozzo A (2022) Re2G: Retrieve, Rerank, Generate. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 2701–2715

BibTeX

@inproceedings{glass-etal-2022-re2g,
    title = "{R}e2{G}: Retrieve, Rerank, Generate",
    author = "Glass, Michael  and
      Rossiello, Gaetano  and
      Chowdhury, Md Faisal Mahbub  and
      Naik, Ankita Rajaram  and
      Cai, Pengshan  and
      Gliozzo, Alfio",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.194/",
    doi = "10.18653/v1/2022.naacl-main.194",
    pages = "2701--2715"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/