Improving Passage Retrieval with Zero-Shot Question Generation

Devendra Singh SachanMike LewisMandar JoshiArmen AghajanyanWen-tau YihJoelle PineauLuke Zettlemoyer

article2022EMNLP268 citations

Proposes a simple zero-shot passage re-ranking method that scores retrieved texts by their likelihood of generating the input question using off-the-shelf language models, consistently outperforming supervised retrieval pipelines across multiple open-domain question answering benchmarks without any fine-tuning.

Listen

Modern information systems and open question-answering engines must accurately find relevant text passages across millions of documents before generating answers. Traditional retrieval systems rely on keyword matching or supervised neural models that require large quantities of expensive, human-annotated training data. These supervised search systems often struggle when transferred to new domains or question formats, creating a performance bottleneck for knowledge-intensive enterprise applications.

The article evaluates whether an unsupervised, zero-shot question generation method can improve text passage retrieval across diverse benchmarks without requiring task-specific training data or fine-tuning. The main objective is to demonstrate that an off-the-shelf pre-trained language model can effectively re-score and re-order candidate passages by calculating the likelihood of generating the query given a retrieved passage.

To assess this concept, the researchers developed the Unsupervised Passage Re-ranker (UPR) and tested it on standard open-domain question answering datasets (such as Natural Questions, TriviaQA, SQuAD-Open, and WebQuestions), entity-heavy benchmarks, and the multi-domain BEIR suite. The framework uses a two-stage approach: a standard retriever first gathers a broad pool of the top 1,000 candidate passages, and a 3-billion-parameter language model then re-ranks them based on average token likelihood using a simple prompt instruction. The evaluation compares this pipeline against leading keyword-based, unsupervised dense, and supervised neural retrievers, as well as downstream question-answering readers.

The experimental findings show substantial accuracy improvements across the board. First, re-ranking candidate passages with UPR increased top-20 passage retrieval accuracy by 6 to 18 percentage points for unsupervised retrievers and up to 12 percentage points for strong supervised baselines. Second, combining an unsupervised dense retriever with UPR outperformed strong supervised retrieval models like Dense Passage Retriever (DPR) by an average of 7 percentage points in top-20 accuracy, establishing that fully unsupervised retrieval pipelines can surpass supervised alternatives. Third, instruction-tuned language models (such as T0) proved most effective at re-scoring candidate passages compared to standard pre-trained architectures. Finally, applying the re-ranked passages to complete open-domain question-answering systems improved answer accuracy by 1 to 3 exact-match points, establishing new state-of-the-art results without retraining the downstream reader models.

These findings indicate that organizations can significantly improve search and question-answering accuracy without the financial and operational overhead of collecting annotated datasets or conducting resource-heavy joint training. Because the system relies entirely on general-purpose language models, it reduces maintenance complexity and adapts more reliably across domain shifts. However, leaders should weigh retrieval accuracy against system latency, as running cross-attention across hundreds of candidate passages per query increases computational time.

For engineering and product teams building search or question-answering pipelines, the article supports adopting zero-shot re-ranking on top of existing first-stage retrieval systems. To manage latency trade-offs, practitioners should benchmark candidate pool sizes (such as evaluating 100 versus 1,000 passages) and consider efficiency optimizations like model quantization or caching before broad deployment. Furthermore, because re-ranking based on question likelihood underperformed on claim-verification tasks where queries are assertions rather than questions, teams should adapt instruction prompts to match specific query formats.

While confidence in the core retrieval gains is high across standard question-answering benchmarks, several boundary conditions apply. Re-ranking performance remains bounded by the quality of the initial candidate pool retrieved in the first stage. Additionally, the approach may produce lower gains or performance drops when applied to non-question queries or out-of-domain technical areas unless appropriate prompt instructions and specialized language models are used.

Cover for Improving Passage Retrieval with Zero-Shot Question Generation

Abstract

We propose a simple and effective re-ranking method for improving passage retrieval in open question answering. The re-ranker re-scores retrieved passages with a zero-shot question generation model, which uses a pre-trained language model to compute the probability of the input question conditioned on a retrieved passage. This approach can be applied on top of any retrieval method (e.g. neural or keyword-based), does not require any domain- or task-specific training (and therefore is expected to generalize better to data distribution shifts), and provides rich cross-attention between query and passage (i.e. it must explain every token in the question). When evaluated on a number of open-domain retrieval datasets, our re-ranker improves strong unsupervised retrieval models by 6%-18% absolute and strong supervised models by up to 12% in terms of top-20 passage retrieval accuracy. We also obtain new state-of-the-art results on full open-domain question answering by simply adding the new re-ranker to existing models with no further changes.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Retriever
  • 2.2 Unsupervised Passage Re-ranking (UPR)
  • 3 Experimental Setup
  • 3.1 Open-Domain QA Datasets
  • 3.2 Keyword-centric Datasets
  • 3.3 Retrievers
  • 3.3.1 Unsupervised Retrievers
  • 3.3.2 Supervised Retrievers
  • 3.4 Pre-Trained Language Models (PLMs)
  • 3.5 Implementation Details
  • 4 Experiments: Passage Retrieval
  • 4.1 Main Task
  • 4.2 Ablation Studies
  • 4.2.1 Importance of Question Generation
  • 4.2.2 Impact of Pre-trained Language Models
  • 4.2.3 Passage Candidate Size vs Latency
  • 4.3 Zero-Shot Supervised Transfer
  • 4.4 Evaluation on Keyword-centric Datasets
  • 4.4.1 Entity Questions
  • 4.4.2 BEIR Benchmark
  • 5 Experiments: Question Answering
  • 5.1 Method
  • 5.2 Results
  • 6 Related Work
  • Generative Pre-training and Instruction Tuning
  • Document Ranking based on Query Likelihood
  • 7 Conclusions and Future Work
  • Acknowledgements
  • Limitations
  • Ethics Statement
  • References
  • A Appendix
  • A.1 Training Hyperparameters
  • A.2 Instruction Prompt Selection
  • A.3 Analysis
  • A.4 BEIR Benchmark Results
  • B Reproducibility Checklist
  • B.1 For all reported experimental results
  • B.2 For all results involving multiple experiments, such as hyperparameter search
  • B.3 For all datasets used

Knowls

  1. Knowl 1 — Unsupervised Passage Re-ranking Formulation and Scoring

    model/method

    Unsupervised Passage Re-ranking (UPR) is a zero-shot passage re-ranking approach for open-domain retrieval that re-scores candidate passages by estimating the conditional likelihood of generating the input question given a retrieved passage, using an off-the-shelf pre-trained language model (PLM) without domain- or task-specific fine-tuning.

    Given an input question qq and a set of KK candidate passages Z={z1,z2,…,zK}\mathcal{Z} = \{z_1, z_2, \dots, z_K\} retrieved from a text corpus D\mathcal{D} by a first-stage retriever (such as BM25, DPR, or Contriever), the passage relevance is modeled via Bayes' rule:

    log⁡p(zi∣q)=log⁡p(q∣zi)+log⁡p(zi)+c\log p(z_i \mid q) = \log p(q \mid z_i) + \log p(z_i) + c

    where p(zi)p(z_i) is the prior probability of passage ziz_i, and cc is a normalization constant independent of ziz_i. Under the assumption of a uniform prior p(zi)=constantp(z_i) = \text{constant}, ranking by posterior probability reduces to ranking by the conditional question generation likelihood:

    log⁡p(zi∣q)∝log⁡p(q∣zi)\log p(z_i \mid q) \propto \log p(q \mid z_i)

    The conditional log-likelihood log⁡p(q∣zi)\log p(q \mid z_i) is computed as the average log-likelihood over the question tokens q=(q1,q2,…,q∣q∣)q = (q_1, q_2, \dots, q_{|q|}):

    log⁡p(q∣zi)=1∣q∣∑t=1∣q∣log⁡p(qt∣q<t,zi;Θ)\log p(q \mid z_i) = \frac{1}{|q|} \sum_{t=1}^{|q|} \log p(q_t \mid q_{<t}, z_i; \Theta)

    where Θ\Theta denotes the parameters of the PLM and ∣q∣|q| is the number of tokens in question qq.

    In practice, the context fed to the language model is formed by appending a natural language prompt to the passage tokens: Passage: {z_i}. Please write a question based on this passage. Candidate passages are then sorted in descending order of their length-normalized log-likelihood score log⁡p(q∣zi)\log p(q \mid z_i).

  2. Knowl 2 — Passage Retrieval Accuracy Gains with UPR Across Retrievers and Datasets

    data/table

    Evaluating UPR with the T0-3B pre-trained language model to re-rank the top-1000 retrieved passages across unsupervised retrievers (MSS, BM25, Contriever) and supervised dense retrievers (DPR, MSS-DPR) demonstrates consistent gains on the SQuAD-Open, TriviaQA, Natural Questions (NQ), and WebQuestions (WebQ) open-domain QA test benchmarks. Retrieval performance is evaluated by top-KK accuracy, defined as the percentage of questions for which at least one of the top-KK retrieved passages contains an exact match string for the human-annotated answer.

    Retriever SQuAD-Open TriviaQA NQ WebQ Average
    Top-20 Top-100 Top-20 Top-100 Top-20 Top-100 Top-20 Top-100 Top-20 Top-100
    Unsupervised Retrievers
    MSS 51.3 68.4 67.2 79.1 60.0 75.6 49.2 68.4 56.9 72.9
    MSS + UPR 75.7 80.8 81.3 85.0 77.3 81.5 71.8 80.4 76.5 81.9
    BM25 71.1 81.8 76.4 83.2 62.9 78.3 62.4 75.5 68.2 79.7
    BM25 + UPR 83.6 87.4 83.0 86.4 78.6 85.2 72.9 81.4 79.5 85.1
    Contriever 63.4 78.2 73.9 82.9 67.9 80.6 65.7 80.1 70.0 80.5
    Contriever + UPR 81.3 85.6 82.8 86.4 80.4 87.0 75.7 83.5 80.1 85.6
    Supervised Retrievers
    DPR 59.4 74.5 79.8 85.1 79.2 85.7 74.6 81.6 73.3 81.7
    DPR + UPR 80.7 85.4 84.3 87.2 83.4 88.6 77.7 84.1 81.5 86.3
    MSS-DPR 73.1 84.5 81.9 86.6 81.4 88.1 76.9 84.6 78.3 86.0
    MSS-DPR + UPR 85.2 89.4 84.8 88.0 83.9 89.4 77.2 85.2 82.8 88.0
    E2E Supervised - - 84.1 87.8 84.8 89.8 79.1 85.2 - -

    UPR delivers absolute improvements of 6% to 18% in top-20 accuracy across all unsupervised retrievers and up to 12% across supervised models. Re-ranking the unsupervised Contriever model achieves 80.1% average top-20 accuracy and 85.6% average top-100 accuracy, outperforming the supervised DPR retriever (73.3% top-20, 81.7% top-100).

  3. Knowl 3 — Open-Domain Question Answering with Fusion-in-Decoder and UPR Passages

    empirical result

    In open-domain question answering pipelines using a Fusion-in-Decoder (FiD) reader model, replacing initial retrieved passages with passages re-ranked by UPR directly increases Exact Match (EM) scores without modifying or retraining the reader.

    In FiD, each retrieved passage is concatenated with the query and encoded independently with a T5 encoder; the concatenated encoder representations are then processed via joint cross-attention by the T5 decoder to generate an answer autoregressively. FiD models are trained on the top-100 retrieved passages from each baseline retriever. At test time, inference is performed using the top-100 passages re-ranked from the top-1000 candidates via UPR (using T0-3B).

    Model / Retriever Setup SQuAD-Open TriviaQA NQ
    Dev Test Dev Test Dev Test
    FiD-base (MSS retriever, T5 reader) 36.2 39.6 60.9 60.3 43.7 44.5
    + Inference with UPR re-ranked passages 43.7 50.1 68.5 68.9 45.8 47.3
    FiD-base (DPR retriever, T5 reader) 48.8 45.8 67.9 68.5 49.4 50.8
    + Inference with UPR re-ranked passages 51.5 54.0 70.1 71.2 49.8 51.3
    FiD-base (MSS-DPR retriever, T5 reader) 50.1 52.2 69.9 70.2 49.7 50.8
    + Inference with UPR re-ranked passages 51.9 55.6 71.5 71.8 49.9 51.5
    FiD-large (MSS-DPR retriever, T5 reader) 51.9 54.4 71.5 71.6 51.8 53.6
    + Inference with UPR re-ranked passages 53.1 58.1 72.7 73.2 51.5 54.5

    Using UPR-re-ranked passages with a fixed FiD-large model improves test exact match scores by 3.7 points on SQuAD-Open (54.4 to 58.1), 1.6 points on TriviaQA (71.6 to 73.2), and 0.9 points on NQ (53.6 to 54.5), outperforming joint end-to-end trained retrieval-reader architectures without retraining.

  4. Knowl 4 — Asymmetry of Question Generation versus Passage Generation in Zero-Shot Re-ranking

    empirical result

    Passage re-ranking based on question generation likelihood p(q∣z)p(q \mid z) outperforms passage generation likelihood p(z∣q)p(z \mid q), which substantially degrades retrieval accuracy below baseline retrievers.

    In passage generation scoring, the teacher-forced average log-likelihood of passage tokens z=(z1,…,z∣z∣)z = (z_1, \dots, z_{|z|}) conditioned on query qq is computed as:

    log⁡p(z∣q;Θ)=1∣z∣∑t=1∣z∣log⁡p(zt∣z<t,q;Θ)\log p(z \mid q; \Theta) = \frac{1}{|z|} \sum_{t=1}^{|z|} \log p(z_t \mid z_{<t}, q; \Theta)

    where ∣z∣|z| is the token length of passage zz and Θ\Theta are the language model parameters.

    When evaluated on the Natural Questions development set across the union of top-1000 passages from BM25 and MSS:

    • Scoring with question generation p(q∣z)p(q \mid z) using T0-3B and GPT-neo-2.7B increases top-KK accuracy monotonically, yielding ~79% top-20 accuracy (compared to 62.3% for BM25 and 57.4% for MSS).
    • Scoring with passage generation p(z∣q)p(z \mid q) using the same language models leads to a severe drop in performance across all top-KK cutoffs, falling well below both the initial BM25 and MSS curves (top-1 accuracy drops below 10%).

    This discrepancy arises because question generation requires the language model's cross-attention to explain every token in the concise query conditioned on the candidate passage. In contrast, average passage token likelihood conditioned on a short query penalizes relevant passages containing valid unmentioned background details and rewards generic passages.

  5. Knowl 5 — Impact of Pre-trained Language Model Architecture, Scale, and Instruction Tuning on UPR

    empirical result

    Re-ranking performance in UPR is sensitive to the pre-training objective, size, and fine-tuning history of the language model backbone. Evaluating different models on the union of BM25 and MSS top-1000 retrieved passages on the Natural Questions (NQ) development set shows:

    Retriever / Re-Ranker Top-1 Top-5 Top-20 Top-100
    BM25 Baseline 22.3 43.8 62.3 76.0
    MSS Baseline 17.7 38.6 57.4 72.4
    T5 (3B) 22.0 50.5 71.4 84.0
    GPT-neo (2.7B) 27.2 55.0 73.9 84.2
    GPT-j (6B) 29.8 59.5 76.8 85.6
    T5-lm-adapt (250M) 23.9 51.4 70.7 83.1
    T5-lm-adapt (800M) 29.1 57.5 75.1 84.8
    T5-lm-adapt (3B) 29.7 59.9 76.9 85.6
    T5-lm-adapt (11B) 32.1 62.3 78.5 85.8
    T0-3B 36.7 64.9 79.1 86.1
    T0-11B 37.4 64.9 79.1 86.0

    Key behavioral patterns include:

    1. Span-denoising pre-training (standard T5-3B) yields poor top-1 accuracy (22.0%, trailing BM25 at 22.3%) because it is not pre-trained for left-to-right generation.
    2. Autoregressive and language-model-adapted architectures (GPT-neo, GPT-j, T5-lm-adapt) improve top-1 and top-5 accuracy substantially over standard T5.
    3. Model scaling monotonically improves accuracy within the same model family (T5-lm-adapt top-1 accuracy increases from 23.9% at 250M to 32.1% at 11B).
    4. Instruction-tuned models (T0-3B and T0-11B) achieve the highest scores (36.7% and 37.4% top-1, 79.1% top-20), showing that multi-task prompted pre-training significantly transfers to zero-shot likelihood re-ranking.
  6. Knowl 6 — Zero-Shot UPR versus Supervised monoT5 Re-ranking

    empirical result

    Comparing zero-shot UPR against monoT5 re-rankers (which are T5 models supervised on MS MARCO passage ranking pairs using binary classification cross-entropy) on BM25 top-1000 retrieved passages from the Natural Questions development set reveals precision differences at narrow vs. broad cutoff levels:

    Retriever / Re-Ranker Top-1 Top-5 Top-20 Top-100
    BM25 (Initial) 22.3 43.8 62.3 76.0
    UPR (T0-3B, zero-shot) 36.1 62.8 76.8 83.1
    monoT5 (250M, supervised) 39.1 62.4 75.6 82.6
    monoT5 (800M, supervised) 43.5 66.1 77.5 83.3
    monoT5 (3B, supervised) 44.2 68.3 78.7 83.7

    Supervised monoT5 achieves higher top-1 accuracy (44.2% vs. 36.1% for UPR-T0-3B) and top-5 accuracy (68.3% vs. 62.8%). However, for top-20 and top-100 metrics—which supply candidate pools for downstream QA reader models—UPR approaches or matches supervised models (76.8% vs. 78.7% for top-20; 83.1% vs. 83.7% for top-100) without requiring annotated ranking data or supervised fine-tuning.

  7. Knowl 7 — UPR Performance on Entity Questions and BEIR Benchmark

    empirical result

    Evaluating UPR on entity-centric and domain-diverse retrieval benchmarks demonstrates broad robustness:

    1. Entity Questions Benchmark (22K entity-centric questions where dense retrievers typically underperform sparse models): Re-ranking the top-1000 passages retrieved from Wikipedia using UPR (T0-3B) improves all retriever baselines:
    • MSS: Top-20 increases from 51.2% to 71.3%; Top-100 increases from 66.3% to 76.7%.
    • DPR: Top-20 increases from 51.1% to 65.4%; Top-100 increases from 63.8% to 72.0%.
    • MSS-DPR: Top-20 increases from 60.6% to 73.9%; Top-100 increases from 73.7% to 80.1%.
    • Contriever: Top-20 increases from 63.0% to 76.0%; Top-100 increases from 75.1% to 81.6%.
    • BM25: Top-20 increases from 71.2% to 79.3%; Top-100 increases from 79.8% to 83.9%.
    • Union of BM25 and Contriever candidates re-ranked with UPR reaches 80.2% Top-20 and 85.4% Top-100 accuracy, outperforming the prior state-of-the-art SPAR retriever (74.0% Top-20, 82.0% Top-100).
    1. BEIR Benchmark (15 heterogeneous zero-shot retrieval tasks): Re-ranking the top-1000 documents with UPR (T0-3B) increases the 15-dataset macro-average metrics:
    • Contriever: nDCG@10 increases from 36.0 to 44.6 (+8.6 absolute); Recall@100 increases from 60.1 to 66.3 (+6.2 absolute).
    • BM25: nDCG@10 increases from 41.6 to 44.9 (+3.3 absolute); Recall@100 increases from 63.6 to 68.0 (+4.4 absolute). Re-ranking yields improvements on 13/15 datasets for Contriever and 12/15 for BM25, with largest gains observed on question-based queries (e.g., NQ, MS-MARCO, HotpotQA, FIQA-2018).
  8. Knowl 8 — Instruction Prompt Phrasing Sensitivity in UPR

    empirical result

    The wording of the natural language instruction prompt prepended to candidate passages affects the likelihood assigned by the language model during UPR re-ranking. Evaluating various prompt templates on the BM25 top-1000 retrieved passages on the Natural Questions development set using T0-3B yields:

    Instruction Prompt Top-1 Top-5 Top-20 Top-100
    BM25 baseline (no re-ranking) 22.3 43.8 62.3 76.0
    No prompt (empty string: ) 28.3 56.1 73.2 82.4
    SScore the following question based on this passage." 35.3 62.6 76.4 83.0
    Ä possible question based on this passage is." 33.8 61.6 76.2 83.1
    "This is a relevant document for the following question." 33.7 61.8 76.0 83.0
    "Please write a question based on this passage." 36.1 62.8 76.8 83.1

    Prompting with natural language instructions consistently outperforms the unprompted setting (36.1% vs. 28.3% Top-1 accuracy), and direct generation directives ("Please write a question based on this passage." ) produce the strongest retrieval performance across all Top-KK metrics.

  9. Knowl 9 — Computational Latency, Candidate Pool Saturation, and Domain Limitations of UPR

    limitation

    UPR has three primary operational and structural limitations:

    1. Computational Latency: Re-ranking requires executing full cross-attention between each passage ziz_i and query qq across all layers LL of a large pre-trained language model, incurring O(L⋅∣q∣⋅(∣zi∣+∣q∣))\mathcal{O}(L \cdot |q| \cdot (|z_i| + |q|)) per candidate. On an NVIDIA A100 (40GB) GPU with T0-3B re-ranking BM25 passages on NQ dev, query latency scales linearly with candidate count:

      • 100 candidates: 1.0 s/query (Top-20 accuracy: 71.6%)
      • 250 candidates: 2.7 s/query (Top-20 accuracy: 74.2%)
      • 500 candidates: 5.5 s/query (Top-20 accuracy: 75.6%)
      • 750 candidates: 8.5 s/query (Top-20 accuracy: 76.4%)
      • 900 candidates: 10.5 s/query (Top-20 accuracy: 76.6%)
      • 1000 candidates: 11.6 s/query (Top-20 accuracy: 76.7%) Marginal accuracy gains plateau beyond 500 candidate passages despite linear latency increases.
    2. First-Stage Recall Ceiling: As a second-stage re-ranker, the upper bound on retrieval accuracy at top-KK after re-ranking NN passages is strictly constrained by the initial retriever's top-NN recall (e.g., re-ranking 1000 passages cannot recover relevant documents not present in the top-1000 pool).

    3. Query Format Mismatch on Claim Verification: The default question generation prompt causes performance degradation on datasets where queries are declarative claims rather than information-seeking questions. On the BEIR benchmark fact-verification datasets, UPR re-ranking with T0-3B degrades performance compared to the original retriever:

      • FEVER: BM25 nDCG@10 drops from 75.3 to 59.1; Contriever nDCG@10 drops from 68.2 to 57.3.
      • Climate-FEVER: BM25 nDCG@10 drops from 21.3 to 11.7; Contriever nDCG@10 drops from 15.5 to 9.5.

Coverage note — None omitted; all core contributions, formulations, experimental results on open-domain QA and BEIR, reader experiments, model/prompt ablations, and stated limitations are included.

References

  1. 1.Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268.
  2. 2.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on Freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing.
  3. 3.Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33.
  5. 5.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  6. 6.Xilun Chen, Kushal Lakhotia, Barlas Oğuz, Anchit Gupta, Patrick Lewis, Stan Peshterliev, Yashar Mehdad, Sonal Gupta, and Wen-tau Yih. 2021. Salient phrase aware dense retrieval: Can a dense retriever imitate a sparse one? arXiv preprint arXiv:2110.06918.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Oliveira Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. ArXiv, abs/2204.02311.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers).
  9. 9.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  10. 10.Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.
  11. 11.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning.
  12. 12.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  13. 13.Gautier Izacard and Edouard Grave. 2021a. Distilling knowledge from reader to retriever for question answering. In International Conference on Learning Representations.
  14. 14.Gautier Izacard and Edouard Grave. 2021b. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume.
  15. 15.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  16. 16.Jia-Huei Ju, Jheng-Hong Yang, and Chuan-Ju Wang. 2021. Text-to-text multi-view learning for passage re-ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval.
  17. 17.Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  18. 18.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In The 2015 International Conference for Learning Representations.
  19. 19.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics.
  20. 20.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  21. 21.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  22. 22.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems.
  23. 23.Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python toolkit for reproducible information retrieval research with sparse and dense representations. In Proceedings of the 44th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2021).
  24. 24.Xueguang Ma, Kai Sun, Ronak Pradeep, and Jimmy Lin. 2021. A replication study of dense passage retriever. arXiv preprint arXiv:2104.05740.
  25. 25.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. MetaICL: Learning to learn in context. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  26. 26.Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Document ranking with a pre-trained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020.
  27. 27.Cicero Nogueira dos Santos, Xiaofei Ma, Ramesh Nallapati, Zhiheng Huang, and Bing Xiang. 2020. Beyond [CLS] through ranking by generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  28. 28.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems.
  29. 29.Jay M. Ponte and W. Bruce Croft. 1998. A language modeling approach to information retrieval. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval.
  30. 30.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  31. 31.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John F. J. Mellor, Irina Higgins, Antonia Creswell, Nathan McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, L. Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, N. K. Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Tobias Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew G. Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem W. Ayoub, Jeff Stanway, L. L. Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. ArXiv, abs/2112.11446.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  33. 33.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing.
  34. 34.Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends in Information Retrieval.
  35. 35.Devendra Singh Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant, Wei Ping, William L Hamilton, and Bryan Catanzaro. 2021a. End-to-end training of neural retrievers for open-domain question answering. In Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP).
  36. 36.Devendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer, and Dani Yogatama. 2021b. End-to-end training of multi-document reader and retriever for open-domain question answering. In Advances in Neural Information Processing Systems.
  37. 37.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  38. 38.Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021. Simple entity-centric questions challenge dense retrievers. In Empirical Methods in Natural Language Processing (EMNLP).
  39. 39.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  40. 40.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  41. 41.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  42. 42.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.
  43. 43.Chengxiang Zhai and John Lafferty. 2001. A study of smoothing methods for language models applied to ad hoc information retrieval. In Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval.
  44. 44.Giulio Zhou and Jacob Devlin. 2021. Multi-vector attention models for deep re-ranking. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.

Citation

MLA
Sachan, D., et al. “Improving Passage Retrieval with Zero-Shot Question Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3781–97, https://doi.org/10.18653/v1/2022.emnlp-main.249.
APA
Sachan, D., Lewis, M., Joshi, M., Aghajanyan, A., Yih, W.-. tau ., Pineau, J., & Zettlemoyer, L. (2022). Improving Passage Retrieval with Zero-Shot Question Generation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3781–3797. https://doi.org/10.18653/v1/2022.emnlp-main.249
Chicago
Sachan, D., M. Lewis, M. Joshi, et al. 2022. “Improving Passage Retrieval with Zero-Shot Question Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3781–97. https://doi.org/10.18653/v1/2022.emnlp-main.249.
Harvard
Sachan, D. et al. (2022) “Improving Passage Retrieval with Zero-Shot Question Generation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3781–3797. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.249.
Vancouver
1. Sachan D, Lewis M, Joshi M, Aghajanyan A, Yih W-tau, Pineau J, Zettlemoyer L (2022) Improving Passage Retrieval with Zero-Shot Question Generation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3781–3797

BibTeX

@inproceedings{sachan-etal-2022-improving,
    title = "Improving Passage Retrieval with Zero-Shot Question Generation",
    author = "Sachan, Devendra  and
      Lewis, Mike  and
      Joshi, Mandar  and
      Aghajanyan, Armen  and
      Yih, Wen-tau  and
      Pineau, Joelle  and
      Zettlemoyer, Luke",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.249/",
    doi = "10.18653/v1/2022.emnlp-main.249",
    pages = "3781--3797"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/