How Does Generative Retrieval Scale to Millions of Passages?
Ronak PradeepKai HuiJai GuptaÁdám D. LelkesHonglei ZhuangJimmy LinDonald MetzlerVinh Q. Tran
Presents the first comprehensive empirical evaluation of generative retrieval scaled up to 8.8 million passages and 11 billion parameters, showing that while synthetic queries are vital for indexing, current architectures struggle to match standard dual encoders as corpus size grows.
Modern information retrieval is shifting toward generative retrieval, an emerging paradigm where a single sequence-to-sequence language model replaces traditional external index structures by directly mapping search queries to document identifiers stored in model memory. While early studies showed generative retrieval matching or exceeding dense retrieval models on small collections of roughly 100,000 documents, its viability on realistic, large-scale systems remained unproven. The article presents the first comprehensive empirical evaluation of generative retrieval techniques scaled to a corpus of 8.8 million passages, determining which design choices hold up as data volume and model size grow.
To conduct this evaluation, the researchers tested various document representations, identifier structures, and specialized decoding architectures across standard benchmarks, including Natural Questions, TriviaQA, and scaled versions of the MS MARCO passage dataset. They evaluated models ranging from 220 million to 11 billion parameters. Across these experiments, synthetic query generation emerged as the single most critical factor for performance, providing a two- to threefold accuracy improvement over conventional document text indexing. Exposing the model to diverse, predicted questions bridged severe coverage gaps on large corpora, where fewer than 6% of documents had human-labeled training queries. Conversely, complex architectural additions like prefix-aware weight-adaptive decoders and constrained decoding offered marginal or no benefit once compute budgets were equalized.
Critically, the findings reveal that current generative retrieval approaches degrade substantially at scale and fail to outperform conventional dense retrievers on large collections. On the full 8.8-million passage benchmark, the best generative model configuration achieved a Mean Reciprocal Rank of 26.7, trailing the dense retriever baseline of 34.8. Furthermore, naively expanding the model from 3 billion to 11 billion parameters caused retrieval accuracy to deteriorate to 24.3, contradicting the common assumption that increasing model capacity alone solves retrieval scaling bottlenecks. While atomic document identifiers yielded low inference latency, they required billions of additional parameters that scaled poorly with corpus size compared to standard string identifiers.
These results indicate that organizations should exercise caution before deploying generative retrieval systems for large-scale enterprise or production search workloads. Because current generative architectures incur massive computational training costs without matching dense retrieval quality at scale, traditional dense and hybrid search architectures remain the recommended production standard. Future research must develop better scaling principles, investigate why oversized models experience performance degradation on memorization-heavy tasks, and formulate novel identifier structures that balance inference efficiency with manageable parameter growth.
- Paper: Transformer Memory as a Differentiable Search Index, Yi Tay et al. (2022). This seminal work introduced the Differentiable Search Index (DSI) paradigm where language models directly map queries to document identifiers, establishing the foundational generative retrieval framework that the source systematically stress-tests at scale.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). This paper establishes standard zero-shot retrieval benchmarks across diverse datasets, providing the evaluation principles and baseline methodology used to test generative and dense retrieval models.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This study demonstrates how parametric model memory degrades on tail knowledge and benefits from external retrieval, motivating the source's investigation into whether pure model memory can scale to millions of passages.
- Paper: ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT, Omar Khattab et al. (2020). This paper introduces late-interaction neural retrieval on the multi-million MS MARCO passage benchmark, providing the key dense retrieval baseline against which the source compares generative architectures.
- Paper: Dense Passage Retrieval for Open-Domain Question Answering, Vladimir Karpukhin et al. (2020). This foundational paper presents Dense Passage Retrieval (DPR) for open-domain question answering, supplying the dual-encoder baseline concepts and benchmark setups utilized throughout the source.
- Paper: Unsupervised Dense Information Retrieval with Contrastive Learning, Gautier Izacard et al. (2021). This work develops Contriever for contrastive dense passage retrieval, defining the standard dense retriever performance standards that the source empirically tests generative indexing against.
- Paper: SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking, Thibault Formal et al. (2021). This work introduces neural sparse lexical expansion for first-stage ranking, establishing the trade-offs between inverted-index search efficiency and neural representation that generative indexing aims to replace.
- Paper: Learning to Rank in Generative Retrieval, Yongqi Li et al. (2024). This work tackles the objective mismatch and ranking bottlenecks identified in generative retrieval systems by integrating margin-based learning-to-rank objectives into string identifier decoding.
- Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Siddharth Gollapudi et al. (2026). This study explores an alternative paradigm to generative weight memorization by evaluating whether language models can perform in-context retrieval directly over million-token scale corpora.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). Following the source's findings on the limitations of pure generative retrieval at scale, this work unifies context reranking and retrieval-augmented generation within a single LLM.
- Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). This theoretical analysis reveals mathematical capacity bounds on embedding dimensions across large corpora, providing formal foundations for scaling bottlenecks in retrieval.
