How Does Generative Retrieval Scale to Millions of Passages?

Ronak PradeepKai HuiJai GuptaÁdám D. LelkesHonglei ZhuangJimmy LinDonald MetzlerVinh Q. Tran

article2023EMNLP90 citations

Presents the first comprehensive empirical evaluation of generative retrieval scaled up to 8.8 million passages and 11 billion parameters, showing that while synthetic queries are vital for indexing, current architectures struggle to match standard dual encoders as corpus size grows.

Listen

Modern information retrieval is shifting toward generative retrieval, an emerging paradigm where a single sequence-to-sequence language model replaces traditional external index structures by directly mapping search queries to document identifiers stored in model memory. While early studies showed generative retrieval matching or exceeding dense retrieval models on small collections of roughly 100,000 documents, its viability on realistic, large-scale systems remained unproven. The article presents the first comprehensive empirical evaluation of generative retrieval techniques scaled to a corpus of 8.8 million passages, determining which design choices hold up as data volume and model size grow.

To conduct this evaluation, the researchers tested various document representations, identifier structures, and specialized decoding architectures across standard benchmarks, including Natural Questions, TriviaQA, and scaled versions of the MS MARCO passage dataset. They evaluated models ranging from 220 million to 11 billion parameters. Across these experiments, synthetic query generation emerged as the single most critical factor for performance, providing a two- to threefold accuracy improvement over conventional document text indexing. Exposing the model to diverse, predicted questions bridged severe coverage gaps on large corpora, where fewer than 6% of documents had human-labeled training queries. Conversely, complex architectural additions like prefix-aware weight-adaptive decoders and constrained decoding offered marginal or no benefit once compute budgets were equalized.

Critically, the findings reveal that current generative retrieval approaches degrade substantially at scale and fail to outperform conventional dense retrievers on large collections. On the full 8.8-million passage benchmark, the best generative model configuration achieved a Mean Reciprocal Rank of 26.7, trailing the dense retriever baseline of 34.8. Furthermore, naively expanding the model from 3 billion to 11 billion parameters caused retrieval accuracy to deteriorate to 24.3, contradicting the common assumption that increasing model capacity alone solves retrieval scaling bottlenecks. While atomic document identifiers yielded low inference latency, they required billions of additional parameters that scaled poorly with corpus size compared to standard string identifiers.

These results indicate that organizations should exercise caution before deploying generative retrieval systems for large-scale enterprise or production search workloads. Because current generative architectures incur massive computational training costs without matching dense retrieval quality at scale, traditional dense and hybrid search architectures remain the recommended production standard. Future research must develop better scaling principles, investigate why oversized models experience performance degradation on memorization-heavy tasks, and formulate novel identifier structures that balance inference efficiency with manageable parameter growth.

arXiv: 2305.11841
Cover for How Does Generative Retrieval Scale to Millions of Passages?

Abstract

The emerging paradigm of generative retrieval re-frames the classic information retrieval problem into a sequence-to-sequence modeling task, forgoing external indices and encoding an entire document corpus within a single Transformer. Although many different approaches have been proposed to improve the effectiveness of generative retrieval, they have only been evaluated on document corpora on the order of 100K in size. We conduct the first empirical study of generative retrieval techniques across various corpus scales, ultimately scaling up to the entire MS MARCO passage ranking task with a corpus of 8.8M passages and evaluating model sizes up to 11B parameters. We uncover several findings about scaling generative retrieval to millions of passages; notably, the central importance of using synthetic queries as document representations during indexing, the ineffectiveness of existing proposed architecture modifications when accounting for compute cost, and the limits of naively scaling model parameters with respect to retrieval performance. While we find that generative retrieval is competitive with state-of-the-art dual encoders on small corpora, scaling to millions of passages remains an important and unsolved challenge. We believe these findings will be valuable for the community to clarify the current state of generative retrieval, highlight the unique challenges, and inspire new research directions.

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Background
  • 2.2 Inputs and Targets
  • 2.2.1 Document Representations
  • 2.2.2 Synthetic Query Generation
  • 2.2.3 Document Identifiers
  • 2.3 Model Variants
  • 3 Experimental Setting
  • 3.1 Corpus and Training Data
  • 3.2 Synthetic Query Generation
  • 3.3 Evaluation Dataset and Metrics
  • 3.4 Model Variants
  • 4 Experimental Results
  • 4.1 Ablations over Small Corpora
  • 4.2 Scaling Corpus Size
  • 4.3 Scaling Model Size
  • 5 Analysis & Discussion
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Appendix
  • A.1 Related Work
  • A.2 Implementation Details
  • A.3 Discussion
  • A.3.1 Why are synthetic queries effective?
  • A.3.2 Which model scaling approach is best?
  • A.4 Future Directions
  • A.5 Additional Results

Knowls

  1. Knowl 1 — Generative Retrieval Effectiveness Across Document Corpus Scales from 100K to 8.8M Passages

    empirical result

    When evaluating generative retrieval across different corpus scales on the MS MARCO passage ranking task using a T5-Base backbone, retrieval effectiveness (measured by Mean Reciprocal Rank at top 10, MRR@10) drops drastically as the corpus size increases from 100K passages (MSMarco100K) to 1M passages (MSMarco1M) and 8.8M passages (MSMarcoFULL):

    MSMarco100K MSMarco1M MSMarcoFULL
    Training Configuration At. Nv. Sm. At. Nv. Sm. At. Nv. Sm.
    Labeled Queries (No Indexing) 0.0 1.1 0.0 0.0 0.5 0.0 0.0 0.0 0.0
    FirstP/DaQ + Labeled Queries (DSI) 0.0 23.9 19.2 2.1 12.4 7.4 0.0 7.5 3.1
    FirstP/DaQ + D2Q + Labeled Queries 79.2 77.7 76.8 53.3 48.2 47.1 14.2 13.2 6.4
    Synthetic Queries Only (D2Q only) 80.3 78.7 78.5 55.8 55.4 54.0 24.2 13.3 11.8
    D2Q only + PAWA (2D Semantic DocIDs) - - 78.2 - - 54.1 - - 17.3
    D2Q only + Constrained Decoding - - 78.6 - - 54.0 - - 12.0

    Here, At. denotes Atomic DocIDs, Nv. denotes Naive DocIDs, and Sm. denotes Semantic DocIDs. D2Q denotes docT5query synthetic queries, FirstP denotes document prefix representations, and DaQ denotes chunked document-as-query representations. Architectural modifications (PAWA decoder, Constrained Decoding, Consistency Loss) fail to provide substantial benefits over simpler target formulations.

  2. Knowl 2 — Parameter Scaling Behavior and Performance Peak of Generative Retrieval on Large Corpora

    empirical result

    When scaling model parameter capacity for generative retrieval on the 8.8M MS MARCO passage corpus (MSMarcoFULL) using synthetic query training (D2Q only), retrieval effectiveness improves up to 3B parameters (T5-XL) but degrades when scaling further to 11B parameters (T5-XXL):

    T5 Scale Configuration Parameters Inference FLOPs MRR@10
    Base D2Q Only + Atomic DocID 7.0B 0.9×10120.9 \times 10^{12} 24.2
    Base D2Q Only + Naive DocID 220M 1.4×10121.4 \times 10^{12} 13.3
    Base D2Q Only + PAWA (2D Semantic) 761M 6.8×10126.8 \times 10^{12} 17.3
    Large D2Q Only + Naive DocID 783M 3.5×10123.5 \times 10^{12} 21.4
    Large D2Q Only + PAWA (2D Semantic) 2.1B 1.1×10131.1 \times 10^{13} 19.8
    XL D2Q Only + Naive DocID 2.8B 9.3×10129.3 \times 10^{12} 26.7
    XXL D2Q Only + Naive DocID 11B 4.3×10134.3 \times 10^{13} 24.3

    Under identical hyperparameter and optimization settings, T5-XXL converges faster during training but achieves lower MRR@10 (24.3) compared to T5-XL (26.7). On TREC Deep Learning evaluation sets, T5-XL also outperforms T5-XXL (nDCG@10 of 55.0 vs. 52.0 on TREC DL 19; 52.2 vs. 49.0 on TREC DL 20). Furthermore, naively scaling standard Transformer parameters with Naive DocIDs outperforms specialized architectures like the Prefix-Aware Weight-Adaptive (PAWA) decoder at equivalent parameter counts (21.4 vs 19.8 MRR@10 at ~800M/2B parameters).

  3. Knowl 3 — Performance Gap Between Scaled Generative Retrieval Models and Dense Dual Encoders on Large Corpora

    limitation

    On the full MS MARCO 8.8M passage ranking benchmark (MSMarcoFULL), the highest-performing generative retrieval model (T5-XL with Naive DocIDs, 2.8B parameters) achieves 26.7 MRR@10, 84.7 Recall@100, and 33.2 nDCG@10. In contrast, standard first-stage retrieval baselines achieve higher retrieval effectiveness with significantly fewer parameters or computation:

    • GTR-Base (a dual encoder with 110M parameters) achieves 34.8 MRR@10, 89.8 Recall@100, and 42.0 nDCG@10.
    • BM25 expanded with docT5query achieves 27.2 MRR@10, 81.9 Recall@100, and 33.8 nDCG@10.
    • Unexpanded BM25 achieves 18.4 MRR@10, 65.8 Recall@100, and 22.8 nDCG@10.

    While generative retrieval performs competitively with dual encoders on small corpora (≤100K \le 100\text{K} passages), scaling to millions of documents remains an unsolved challenge where generative retrieval trails top-tier dense retrieval models.

  4. Knowl 4 — Synthetic Query Generation as the Essential Training Component for Generative Retrieval at Scale

    empirical result

    Synthetic query generation (e.g., via docT5query) is the single most critical component for training generative retrieval models on large corpora. On the 8.8M passage MS MARCO corpus, training a T5-Base generative retrieval model solely on document text representations (FirstP/DaQ) and labeled dataset queries yields 7.5 MRR@10 with Naive DocIDs and 0.0 MRR@10 with Atomic DocIDs. Training solely on synthetic queries (D2Q only) boosts performance to 13.3 MRR@10 (an 80% relative gain) for Naive DocIDs and 24.2 MRR@10 for Atomic DocIDs. Using synthetic queries as the exclusive document representation during indexing is superior to combining synthetic queries with raw passage text spans.

  5. Knowl 5 — Inference Compute and Memory Trade-Offs Between Atomic and Sequential Document Identifiers

    model/method

    Generative retrieval models exhibit contrasting compute and memory trade-offs depending on whether document identifiers (DocIDs) are formulated as Atomic tokens or Sequential strings:

    1. Atomic DocIDs: Each document identifier in a corpus of size NN is treated as a single token in the decoder vocabulary. For N=8.8×106N = 8.8\times 10^6 passages and embedding dimension d=768d=768, the decoder output projection matrix requires an extra N×d≈6.8×109N \times d \approx 6.8\times 10^9 parameters, expanding a 220M T5-Base model to 7.0B parameters. However, ranking the entire corpus requires only a single decoder forward step followed by a top-kk sort over the output logits, consuming 0.9×10120.9 \times 10^{12} inference FLOPs on MS MARCO.
    2. Sequential (Naive) DocIDs: Document identifiers are tokenized as text strings (e.g., SentencePiece tokens) with zero additional parameter overhead. However, retrieving top-kk documents requires autoregressive beam search decoding across kk beams (k=40k=40) and sequence length dseqd_{\text{seq}}, requiring O(dseq⋅k)O(d_{\text{seq}} \cdot k) decoding steps. Consequently, T5-XL with Naive DocIDs requires 9.3×10129.3 \times 10^{12} inference FLOPs (over 10×10\times more than T5-Base with Atomic DocIDs) to achieve 26.7 MRR@10.
  6. Knowl 6 — Taxonomy of Target Document Identifier Representations in Generative Retrieval

    definition

    Generative retrieval models map input queries to document identifiers (DocIDs) using four distinct target formulations:

    1. Atomic DocIDs: Each document is treated as an individual, unique token in the decoder vocabulary. The output projection matrix size is ∣V∣=N|V| = N, where NN is the corpus size.
    2. Naive DocIDs: The predefined document ID (e.g., the string "42915") is directly tokenized into subword tokens using the base language model vocabulary (e.g., SentencePiece).
    3. Semantic DocIDs: Dense embeddings of documents (e.g., from SentenceT5-Base) are recursively clustered using hierarchical kk-means into kk clusters until each leaf cluster contains at most cc documents. The identifier is the sequence of cluster indices (from 00 to k−1k-1) followed by an offset index in the leaf cluster (from 00 to c−1c-1).
    4. 2D Semantic DocIDs: A variant of Semantic DocIDs where tokens at different hierarchy depths are assigned an explicit position/level coordinate, resolving token semantic ambiguity across tree levels.
  7. Knowl 7 — Prefix-Aware Weight-Adaptive Decoder for 2D Semantic Document Identifiers

    model/method

    The Prefix-Aware Weight-Adaptive (PAWA) decoder replaces the static vocabulary projection matrix W∈Rd×∣V∣W \in \mathbb{R}^{d \times |V|} of a standard Transformer decoder with an adaptive, position-specific projection tensor Wpawa∈Rd×l×∣V∣W_{\text{pawa}} \in \mathbb{R}^{d \times l \times |V|}, where dd is the model hidden dimension, ll is the sequence length, and ∣V∣|V| is the target vocabulary size.

    At each generation timestep ii, WpawaW_{\text{pawa}} is generated by an auxiliary Transformer decoder conditioned on the input query qq and the prefix of previously decoded DocID tokens t<it_{<i}:

    Wpawa(i)=Decoderaux(q,t<i)W_{\text{pawa}}^{(i)} = \text{Decoder}_{\text{aux}}(q, t_{<i})

    The final hidden state hi∈Rdh_i \in \mathbb{R}^d from the primary decoder is then projected using Wpawa(i)W_{\text{pawa}}^{(i)} to obtain logits over the vocabulary at step ii. This parameterization allows the model to alter the projection mapping based on clustering depth and path context in 2D Semantic DocID trees.

  8. Knowl 8 — Mechanisms of Synthetic Query Efficacy: Mitigating Coverage and Query Distribution Gaps

    empirical result

    Synthetic query generation improves generative retrieval through two primary mechanisms:

    1. Document Coverage Gap: In large datasets, only a small fraction of corpus documents have associated queries in the supervised training set (92.9% coverage in MSMarco100K, 51.6% in MSMarco1M, and 5.8% in MSMarcoFULL). Adding synthetic queries provides positive retrieval targets for 100% of the corpus passages, boosting retrieval MRR@10 by 3.3×3.3\times on MSMarco100K and 3.9×3.9\times on MSMarco1M over raw text indexing.
    2. Query Distribution Gap: Synthetic queries align the indexing input format (short question-like queries) with inference inputs. For MS MARCO evaluation queries, higher Jaccard lexical similarity between synthetic training queries and ground-truth validation queries correlates monotonically with higher MRR@10 across all similarity deciles.
  9. Knowl 9 — Impact of Synthetic Query Sample Volume and Diversity on Generative Retrieval Effectiveness

    empirical result

    On the MSMarco100K dataset with T5-Base and Atomic DocIDs, retrieval effectiveness (MRR@10) increases monotonically as the number of synthetic queries generated per passage grows from 10 to 100:

    • 10 sampled queries: ~72.0 MRR@10
    • 20 sampled queries: ~78.0 MRR@10
    • 30 sampled queries: ~80.0 MRR@10
    • 40 sampled queries: 80.3 MRR@10
    • 100 sampled queries: 82.4 MRR@10

    Filtering or selecting the top-kk synthetic queries using a cross-encoder re-ranker (RankT5-XL) degrades retrieval MRR@10 compared to taking a uniform random sample of kk queries (e.g., ~77.5 MRR@10 for RankT5 top-20 vs. ~79.0 MRR@10 for random 20). High query diversity and sample volume provide superior regularization and distribution coverage compared to quality-filtered subsets.

  10. Knowl 10 — Component Ablation of Generative Retrieval on Small-Scale Corpora (NQ100K and TriviaQA)

    data/table

    On small-scale open-domain QA corpora (NQ100K with 110K documents and TriviaQA with 74K documents), evaluating T5-Base across document representations, DocID types, and architecture extensions yields the following performance:

    NQ100K (Recall@1) TriviaQA (Recall@5)
    Training / Architecture Setup At. Nv. Sm. At. Nv. Sm.
    (1a) Labeled Queries Only (No Indexing) 50.7 49.2 49.0 60.9 56.7 61.4
    (2a) FirstP + Labeled Queries (DSI) 60.0 58.4 58.7 71.6 75.2 78.9
    (2b) DaQ + Labeled Queries 61.4 60.4 60.0 81.0 80.4 77.6
    (3a) DaQ + D2Q + Labeled Queries 69.6 67.9 67.9 88.2 85.7 86.3
    (3b) FirstP + DaQ + D2Q + Labeled Queries 69.0 68.2 67.2 88.9 86.9 87.4
    (4a) 3b + PAWA (w/ 2D Semantic DocIDs) - - 66.3 - - 86.5
    (4b) 3b + Constrained Decoding - - 67.3 - - 87.3
    (5) 4b + Consistency Loss (NCI) - - 66.3 - - 86.6
    (7) 3b + In-Domain D2Q 70.7 69.7 69.5 90.0 88.0 89.2

    Here, At. is Atomic DocIDs, Nv. is Naive DocIDs, and Sm. is Semantic DocIDs. Using synthetic queries generated by an in-domain query generator (trained on DPR for NQ) paired with standard Transformer sequence-to-sequence training achieves state-of-the-art results without requiring specialized architectures like PAWA or constrained decoding.

  11. Knowl 11 — Training Instability and NaN Divergence of Consistency Regularization in Generative Retrieval

    limitation

    In generative retrieval, consistency regularization penalizes discrepancies between the output probability distributions of two forward passes generated with different dropout masks. Given token output distributions pi,1p_{i,1} and pi,2p_{i,2} over the vocabulary space at decoding step ii, the bidirectional Kullback-Leibler (KL) divergence loss is defined as:

    Lreg=12[DKL(pi,1∥pi,2)+DKL(pi,2∥pi,1)]L_{\text{reg}} = \frac{1}{2} \left[ D_{\text{KL}}(p_{i,1} \parallel p_{i,2}) + D_{\text{KL}}(p_{i,2} \parallel p_{i,1}) \right]

    When optimizing sequence-to-sequence generative retrieval models on large corpora, incorporating LregL_{\text{reg}} causes severe numerical instability during training, frequently leading the loss function to diverge into NaN\text{NaN} values.

Coverage note — None was omitted; all key empirical scaling experiments, target DocID formulations, architectural ablations, synthetic query analyses, compute trade-offs, and negative scaling results across MS MARCO and NQ/TriviaQA are covered.

References

  1. 1.Akari Asai, Jungo Kasai, Jonathan Clark, Kenton Lee, Eunsol Choi, and Hannaneh Hajishirzi. 2021. XOR QA: Cross-lingual open-retrieval question answer-ing. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 547–564. Association for Computational Linguistics.
  2. 2.Michele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Wen-tau Yih, Sebastian Riedel, and Fabio Petroni. 2022. Autoregressive search engines: Generating substrings as document identifiers. arXiv:2204.10628.
  3. 3.Xiaoyang Chen, Yanjiang Liu, Ben He, Le Sun, and Yingfei Sun. 2023. Understanding differential search index for text retrieval. arXiv:2305.02073.
  4. 4.Xuanang Chen, Jian Luo, Ben He, Le Sun, and Yingfei Sun. 2022. Towards robust dense retrieval via local ranking alignment. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI, pages 1980–1986.
  5. 5.Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021. Overview of the TREC 2020 deep learning track. In Text REtrieval Conference (TREC).
  6. 6.Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Jimmy Lin. 2022. Overview of the TREC 2022 deep learning track. In Text REtrieval Conference (TREC).
  7. 7.Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M Voorhees. 2020. Overview of the TREC 2019 deep learning track. In Text REtrieval Conference (TREC).
  8. 8.Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2020. Autoregressive entity retrieval. arXiv:2010.00904.
  9. 9.Mostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer, and Ashish Vaswani. 2022. The efficiency misnomer. In International Conference on Learning Representations.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1).
  11. 11.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910. Association for Computational Linguistics.
  12. 12.Daniel Gillick, Alessandro Presta, and Gaurav Singh Tomar. 2018. End-to-end retrieval in continuous space. arXiv:1811.08008.
  13. 13.Mitko Gospodinov, Sean MacAvaney, and Craig Macdonald. 2023. Doc2Query–: When less is more. arXiv:2301.03266.
  14. 14.Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006. Dimensionality reduction by learning an invariant mapping. 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06), 2:1735–1742.
  15. 15.Kai Hui, Honglei Zhuang, Tao Chen, Zhen Qin, Jing Lu, Dara Bahri, Ji Ma, Jai Gupta, Cicero Nogueira dos Santos, Yi Tay, and Donald Metzler. 2022. ED2LM: Encoder-decoder to language model for faster document re-ranking inference. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3747–3758. Association for Computational Linguistics.
  16. 16.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2021. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 7(3):535–547.
  17. 17.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. arXiv:1705.03551.
  18. 18.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781.
  19. 19.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  20. 20.Oleg Lesota, Navid Rekabsaz, Daniel Cohen, Klaus Antonius Grasserbauer, Carsten Eickhoff, and Markus Schedl. 2021. A modern perspective on query likelihood with deep generative retrieval models. Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval.
  21. 21.Xueguang Ma, Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2022. Document expansion baselines and learned sparse lexical representations for ms marco v1 and v2. Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval.
  22. 22.Sanket Vaibhav Mehta, Jai Gupta, Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Jinfeng Rao, Marc Najork, Emma Strubell, and Donald Metzler. 2022. DSI++: Updating transformer memory with new documents. arXiv:2212.09744.
  23. 23.Donald Metzler, Yi Tay, Dara Bahri, and Marc Najork. 2021. Rethinking search: Making domain experts out of dilettantes. SIGIR Forum, 55(1).
  24. 24.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human generated machine reading comprehension dataset. In CoCo@ NIPS.
  25. 25.Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith B. Hall, Daniel Cer, and Yinfei Yang. 2022a. Sentence-T5: Scaling up sentence encoder from pre-trained text-to-text transfer transformer. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1864–1874. Association for Computational Linguistics.
  26. 26.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. 2022b. Large dual encoders are generalizable retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9844–9855. Association for Computational Linguistics.
  27. 27.Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with BERT. arXiv:1901.04085.
  28. 28.Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 708–718. Association for Computational Linguistics.
  29. 29.Rodrigo Nogueira and Jimmy Lin. 2019. From doc2query to docTTTTTquery. Online preprint.
  30. 30.Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019. Document expansion by query prediction. arXiv:1904.08375.
  31. 31.Ronak Pradeep, Yilin Li, Yuetong Wang, and Jimmy Lin. 2022. Neural query synthesis and domain-specific ranking templates for multi-stage clinical trial matching. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '22, page 2325–2330. Association for Computing Machinery.
  32. 32.Ronak Pradeep, Xueguang Ma, Rodrigo Nogueira, and Jimmy Lin. 2021a. Vera: Prediction techniques for reducing harmful misinformation in consumer health search. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '21, page 2066–2070. Association for Computing Machinery.
  33. 33.Ronak Pradeep, Rodrigo Nogueira, and Jimmy J. Lin. 2021b. The Expando-Mono-Duo design pattern for text ranking with pretrained sequence-to-sequence models. arXiv:2101.05667.
  34. 34.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  35. 35.Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan H. Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q. Tran, Jonah Samost, Maciej Kula, Ed H. Chi, and Maheswaran Sathiamoorthy. 2023. Recommender systems with generative retrieval. arXiv:2305.05065.
  36. 36.Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo. 2022. Scaling up models and data with t5x and seqio. arXiv:2203.17189.
  37. 37.Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Now Publishers Inc.
  38. 38.Weiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang, Haichao Zhu, Pengjie Ren, Zhumin Chen, Dawei Yin, Maarten de Rijke, and Zhaochun Ren. 2023. Learning to tokenize for generative retrieval. arXiv:2304.04171.
  39. 39.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. arXiv:1409.3215.
  40. 40.Yi Tay, Vinh Q. Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, Tal Schuster, William W. Cohen, and Donald Metzler. 2022. Transformer memory as a differentiable search index. arXiv:2202.06991.
  41. 41.Dan Vanderkam, Robert B Schonberger, H. Rowley, and Sanjiv Kumar. 2013. Nearest neighbor search in Google Correlate. Online preprint.
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS.
  43. 43.Yujing Wang, Ying Hou, Hong Wang, Ziming Miao, Shibin Wu, Hao Sun, Qi Chen, Yuqing Xia, Chengmin Chi, Guoshuai Zhao, Zheng Liu, Xing Xie, Hao Sun, Weiwei Deng, Qi Zhang, and Mao Yang. 2022. A neural corpus indexer for document retrieval. arXiv:2206.02743.
  44. 44.Yidan Zhang, Ting Zhang, Dong Chen, Yujing Wang, Qi Chen, Xing Xie, Hao Sun, Weiwei Deng, Qi Zhang, Fan Yang, Mao Yang, Qingmin Liao, and Baining Guo. 2023. IRGen: Generative modeling for image retrieval. arXiv:2303.10126.
  45. 45.Yujia Zhou, Jing Yao, Zhicheng Dou, Ledell Wu, Peitian Zhang, and Ji-Rong Wen. 2022. Ultron: An ultimate retriever on corpus with a model-based indexer. arXiv:2208.09257.
  46. 46.Honglei Zhuang, Zhen Qin, Rolf Jagerman, Kai Hui, Ji Ma, Jing Lu, Jianmo Ni, Xuanhui Wang, and Michael Bendersky. 2022a. RankT5: Fine-tuning T5 for text ranking with ranking losses. arXiv:2210.10634.
  47. 47.Shengyao Zhuang, Houxing Ren, Linjun Shou, Jian Pei, Ming Gong, Guido Zuccon, and Daxin Jiang. 2022b. Bridging the gap between indexing and retrieval for differentiable search index with query generation. arXiv:2206.10128.
  48. 48.Noah Ziems, Wenhao Yu, Zhihan Zhang, and Meng Jiang. 2023. Large language models are built-in autoregressive search engines. arXiv:2305.09612.

Citation

MLA
Pradeep, R., et al. “How Does Generative Retrieval Scale to Millions of Passages?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1305–21, https://doi.org/10.18653/v1/2023.emnlp-main.83.
APA
Pradeep, R., Hui, K., Gupta, J., Lelkes, A., Zhuang, H., Lin, J., Metzler, D., & Tran, V. (2023). How Does Generative Retrieval Scale to Millions of Passages?. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1305–1321. https://doi.org/10.18653/v1/2023.emnlp-main.83
Chicago
Pradeep, R., K. Hui, J. Gupta, et al. 2023. “How Does Generative Retrieval Scale to Millions of Passages?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1305–21. https://doi.org/10.18653/v1/2023.emnlp-main.83.
Harvard
Pradeep, R. et al. (2023) “How Does Generative Retrieval Scale to Millions of Passages?”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1305–1321. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.83.
Vancouver
1. Pradeep R, Hui K, Gupta J, Lelkes A, Zhuang H, Lin J, Metzler D, Tran V (2023) How Does Generative Retrieval Scale to Millions of Passages?. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1305–1321

BibTeX

@inproceedings{pradeep-etal-2023-generative,
    title = "How Does Generative Retrieval Scale to Millions of Passages?",
    author = "Pradeep, Ronak  and
      Hui, Kai  and
      Gupta, Jai  and
      Lelkes, Adam  and
      Zhuang, Honglei  and
      Lin, Jimmy  and
      Metzler, Donald  and
      Tran, Vinh",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.83/",
    doi = "10.18653/v1/2023.emnlp-main.83",
    pages = "1305--1321"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/