Improving Passage Retrieval with Zero-Shot Question Generation
Devendra Singh SachanMike LewisMandar JoshiArmen AghajanyanWen-tau YihJoelle PineauLuke Zettlemoyer
Proposes a simple zero-shot passage re-ranking method that scores retrieved texts by their likelihood of generating the input question using off-the-shelf language models, consistently outperforming supervised retrieval pipelines across multiple open-domain question answering benchmarks without any fine-tuning.
Modern information systems and open question-answering engines must accurately find relevant text passages across millions of documents before generating answers. Traditional retrieval systems rely on keyword matching or supervised neural models that require large quantities of expensive, human-annotated training data. These supervised search systems often struggle when transferred to new domains or question formats, creating a performance bottleneck for knowledge-intensive enterprise applications.
The article evaluates whether an unsupervised, zero-shot question generation method can improve text passage retrieval across diverse benchmarks without requiring task-specific training data or fine-tuning. The main objective is to demonstrate that an off-the-shelf pre-trained language model can effectively re-score and re-order candidate passages by calculating the likelihood of generating the query given a retrieved passage.
To assess this concept, the researchers developed the Unsupervised Passage Re-ranker (UPR) and tested it on standard open-domain question answering datasets (such as Natural Questions, TriviaQA, SQuAD-Open, and WebQuestions), entity-heavy benchmarks, and the multi-domain BEIR suite. The framework uses a two-stage approach: a standard retriever first gathers a broad pool of the top 1,000 candidate passages, and a 3-billion-parameter language model then re-ranks them based on average token likelihood using a simple prompt instruction. The evaluation compares this pipeline against leading keyword-based, unsupervised dense, and supervised neural retrievers, as well as downstream question-answering readers.
The experimental findings show substantial accuracy improvements across the board. First, re-ranking candidate passages with UPR increased top-20 passage retrieval accuracy by 6 to 18 percentage points for unsupervised retrievers and up to 12 percentage points for strong supervised baselines. Second, combining an unsupervised dense retriever with UPR outperformed strong supervised retrieval models like Dense Passage Retriever (DPR) by an average of 7 percentage points in top-20 accuracy, establishing that fully unsupervised retrieval pipelines can surpass supervised alternatives. Third, instruction-tuned language models (such as T0) proved most effective at re-scoring candidate passages compared to standard pre-trained architectures. Finally, applying the re-ranked passages to complete open-domain question-answering systems improved answer accuracy by 1 to 3 exact-match points, establishing new state-of-the-art results without retraining the downstream reader models.
These findings indicate that organizations can significantly improve search and question-answering accuracy without the financial and operational overhead of collecting annotated datasets or conducting resource-heavy joint training. Because the system relies entirely on general-purpose language models, it reduces maintenance complexity and adapts more reliably across domain shifts. However, leaders should weigh retrieval accuracy against system latency, as running cross-attention across hundreds of candidate passages per query increases computational time.
For engineering and product teams building search or question-answering pipelines, the article supports adopting zero-shot re-ranking on top of existing first-stage retrieval systems. To manage latency trade-offs, practitioners should benchmark candidate pool sizes (such as evaluating 100 versus 1,000 passages) and consider efficiency optimizations like model quantization or caching before broad deployment. Furthermore, because re-ranking based on question likelihood underperformed on claim-verification tasks where queries are assertions rather than questions, teams should adapt instruction prompts to match specific query formats.
While confidence in the core retrieval gains is high across standard question-answering benchmarks, several boundary conditions apply. Re-ranking performance remains bounded by the quality of the initial candidate pool retrieved in the first stage. Additionally, the approach may produce lower gains or performance drops when applied to non-question queries or out-of-domain technical areas unless appropriate prompt instructions and specialized language models are used.
- Paper: Dense Passage Retrieval for Open-Domain Question Answering, Vladimir Karpukhin et al. (2020). Dense Passage Retrieval establishes the standard dense dual-encoder retrieval baseline and open-domain QA evaluation benchmark that this work directly builds upon and re-ranks.
- Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). This seminal work introduced cross-attention neural passage re-ranking using pre-trained language models, providing the conceptual foundation that zero-shot question generation scoring seeks to generalize.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). Fusion-in-Decoder established the core retrieve-and-read open-domain QA architecture that the zero-shot question-generation re-ranker is evaluated on to achieve state-of-the-art reader performance.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational paper formulated end-to-end Retrieval-Augmented Generation for knowledge-intensive NLP tasks, setting up the passage-conditioned generation paradigm exploited by this paper's scoring technique.
- Paper: Reading Wikipedia to Answer Open-Domain Questions, Danqi Chen et al. (2017). DrQA formulated the standard modern pipeline of passage retrieval followed by machine reading over Wikipedia for open-domain question answering.
- Paper: Generative Language Models for Paragraph-Level Question Generation, Asahi Ushio et al. (2022). This study establishes the standard methodology and sequence-to-sequence model behavior for generating questions conditioned on context paragraphs.
- Paper: Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents, Weiwei Sun et al. (2023). This work extends zero-shot and few-shot language model passage re-ranking by investigating instruction-driven permutation generation prompts with modern LLMs like ChatGPT and GPT-4.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). RankRAG advances beyond separate zero-shot re-ranking heuristics by unifying context ranking and response generation within a single instruction-tuned large language model.
- Paper: Learning to Rank in Generative Retrieval, Yongqi Li et al. (2024). This paper continues the exploration of generative models in retrieval by integrating classical learning-to-rank loss functions directly into generative string-decoding retrieval systems.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Self-RAG builds upon passage-relevance evaluation by training models to dynamically critique, retrieve, and score context passages through learned reflection tokens.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). This comprehensive survey categorizes the evolution of retrieval, re-ranking, and post-retrieval optimizations across naive, advanced, and modular RAG systems.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). In-Context RALM demonstrates how lightweight retrieval and re-ranking pipelines can be prepended directly to general-purpose language model inputs without additional fine-tuning.
