Learning to Rank in Generative Retrieval

Yongqi LiNan YangLiang WangFuru WeiWenjie Li

article2024AAAI80 citations

Proposes LTRGR, a framework that incorporates ranking losses into autoregressive models to bridge the gap between identifier generation and final passage ranking without adding computational overhead during inference.

Listen

Modern information retrieval is increasingly exploring generative retrieval, where an autoregressive language model generates text identifier strings—such as titles, substrings, or pseudo-queries—rather than matching indexed dense vectors. However, existing generative retrieval methods suffer from an objective mismatch: they optimize the model purely to generate valid text identifiers and rely on an external heuristic function to convert these generated strings into a final ranked list of documents. This disconnect between string-generation training and document-ranking goals creates a substantial performance bottleneck.

The article demonstrates that integrating classical learning-to-rank techniques directly into generative retrieval resolves this disconnect. The authors introduce LTRGR, a framework that augments existing generative models with an additional training phase driven by a margin-based passage rank loss alongside standard generation loss.

The approach evaluates full corpus benchmarks using two established question-answering datasets (Natural Questions and TriviaQA) and a large-scale web search dataset (MSMARCO). The model builds upon the MINDER architecture using a BART-Large backbone. It first generates identifier candidates via constrained decoding, maps these identifiers to passage rank scores, and directly propagates rank loss gradients through the model logits without altering the inference procedure.

The findings establish new state-of-the-art results across generative retrieval benchmarks. On the MSMARCO web search task, LTRGR improved top-10 Mean Reciprocal Rank by 28.8% (achieving 25.5 compared to the prior best generative benchmark of 18.6) and Recall@5 by 36.3%, outperforming even larger baseline models. On Natural Questions, LTRGR increased Hits@5 by 4.56% and became the first generative system to outperform Dense Passage Retrieval across Hits@5, Hits@20, and Hits@100 on the full corpus. Analysis confirms that these performance gains concentrate primarily at top-ranked positions and generalize across other generative architectures like SEAL.

These results demonstrate that generative retrieval can match or exceed classical dense retrieval when properly aligned with ranking objectives, requiring no extra compute during live query inference. Organizations building search or question-answering systems can adopt this secondary ranking phase to achieve substantial retrieval accuracy without adding operational latency. Future development should explore normalized list-wise loss formulations and enhanced negative sample mining strategies to close remaining gaps on complex benchmarks like TriviaQA.

  • Paper: How Does Generative Retrieval Scale to Millions of Passages?, Ronak Pradeep et al. (2023). This paper examines how generative retrieval scales across millions of passages in MS MARCO, establishing the core sequence-to-sequence document retrieval problem that LTRGR directly improves upon.
  • Paper: From RankNet to LambdaRank to LambdaMART: An Overview, Christopher J. C. Burges (2010). This work details foundational learning-to-rank algorithms and gradient-based ranking objectives that directly inform LTRGR's passage margin ranking losses.
  • Paper: Learning to rank using gradient descent, Chris Burges et al. (2005). This seminal text introduces pairwise ranking loss optimization via gradient descent, providing the mathematical basis for the ranking objectives integrated into the generative retrieval framework.
  • Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). This foundational paper establishes standard passage re-ranking methodologies on MS MARCO that classical and generative retrieval models aim to replicate or replace.
  • Paper: Learning to rank: from pairwise approach to listwise approach, Zhe Cao et al. (2007). This work outlines listwise learning-to-rank formulations, providing the foundational principles for moving beyond naive generation scoring toward direct ranking optimization.
  • Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This paper introduces the standard dense retrieval and generative framework for knowledge-intensive benchmarks like Natural Questions and TriviaQA evaluated in the source.
  • Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). This study analyzes metric mismatch and sequence-level training objectives in autoregressive generation, which motivates resolving the objective mismatch in generative retrieval.
  • Paper: IR evaluation methods for retrieving highly relevant documents, Kalervo Järvelin et al. (2000). This foundational work formalizes Discounted Cumulative Gain and evaluation metrics designed to prioritize top-ranked documents in information retrieval.
Cover for Learning to Rank in Generative Retrieval

Abstract

Generative retrieval stands out as a promising new paradigm in text retrieval that aims to generate identifier strings of relevant passages as the retrieval target. This generative paradigm taps into powerful generative language models, distinct from traditional sparse or dense retrieval methods. However, only learning to generate is insufficient for generative retrieval. Generative retrieval learns to generate identifiers of relevant passages as an intermediate goal and then converts predicted identifiers into the final passage rank list. The disconnect between the learning objective of autoregressive models and the desired passage ranking target leads to a learning gap. To bridge this gap, we propose a learning-to-rank framework for generative retrieval, dubbed LTRGR. LTRGR enables generative retrieval to learn to rank passages directly, optimizing the autoregressive model toward the final passage ranking target via a rank loss. This framework only requires an additional learning-to-rank training phase to enhance current generative retrieval systems and does not add any burden to the inference stage. We conducted experiments on three public benchmarks, and the results demonstrate that LTRGR achieves state-of-the-art performance among generative retrieval methods. The code and checkpoints are released at https://github.com/liyongqi67/LTRGR.

Table of Contents

  • Introduction
  • Related Work
  • Generative Retrieval
  • Learning to Rank
  • Method
  • Learning to Generate
  • Learning to Rank
  • Experiments
  • Datasets
  • Baselines
  • Implementation Details
  • Retrieval Results on QA
  • Retrieval Results on Web Search
  • Ablation Study
  • In-depth Analysis
  • Effectiveness Analysis of Learning to Rank
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Two-Stage Learning-to-Rank Framework for Generative Retrieval

    model/method

    Generative retrieval treats document retrieval as sequence-to-sequence generation by training an autoregressive language model to map input queries to passage identifier strings (such as titles, substrings, or synthetic pseudo-queries), which are subsequently aggregated and mapped to candidate passages via a scoring function. Standard generative retrieval models optimize solely for token-level generation likelihood, creating a disconnect between the sequence-generation training objective and the target passage-ranking order.

    Learning to Rank in Generative Retrieval (LTRGR) resolves this gap by introducing a two-stage training paradigm:

    1. Learning-to-Generate Phase: The autoregressive language model is trained using standard sequence-to-sequence cross-entropy loss to predict valid multiview identifiers (titles, substrings, pseudo-queries) given an input query and identifier prefix.
    2. Learning-to-Rank Phase: Using the trained generative model, a candidate list of passages is retrieved and scored for each query in the training set. The autoregressive model parameters are then continuously optimized using a joint ranking loss defined over positive and negative candidate passages together with the generation loss.

    During inference, LTRGR uses constrained beam search over an FM-index to generate identifiers and scores candidate passages identically to standard generative retrieval, incurring zero additional computational or latency overhead during inference.

  2. Knowl 2 — Passage Relevance Scoring via Aggregated Identifier Logits

    equation

    In generative retrieval with non-unique text identifiers (such as multiview substrings, titles, and pseudo-queries), the relevance score of a passage p∈Cp \in \mathcal{C} from a corpus C\mathcal{C} with respect to a query qq is defined as the sum of the sequence scores of all generated identifiers that appear within pp:

    s(q,p)=∑ip∈Ipsips(q, p) = \sum_{i_p \in \mathcal{I}_p} s_{i_p}

    where I\mathcal{I} is the set of identifiers generated for query qq via constrained beam search, Ip⊆I\mathcal{I}_p \subseteq \mathcal{I} is the subset of predicted identifiers that match or belong to passage pp, and sips_{i_p} is the sequence generation logit score assigned to identifier ipi_p by the autoregressive model AM\text{AM}.

    Because sips_{i_p} corresponds directly to the unnormalized logits produced by the neural network, passage relevance score s(q,p)s(q, p) is differentiable with respect to the autoregressive model parameters θ\theta, allowing passage-level ranking gradients to propagate back into the token-level generative model.

  3. Knowl 3 — Multi-Task Passage Ranking Loss in LTRGR

    equation

    To directly align autoregressive identifier generation with passage ranking quality, the learning-to-rank stage of LTRGR optimizes a multi-task loss combining two margin-based ranking losses and an identifier generation loss:

    L=Lrank1+Lrank2+λLgen\mathcal{L} = \mathcal{L}_{\text{rank1}} + \mathcal{L}_{\text{rank2}} + \lambda \mathcal{L}_{\text{gen}}

    where λ∈R+\lambda \in \mathbb{R}^+ is a hyperparameter balancing the rank losses and the generation loss. Each margin ranking loss Lrank\mathcal{L}_{\text{rank}} enforces a score margin m>0m > 0 between a positive passage ppp_p (relevant to query qq) and a negative passage pnp_n (irrelevant to query qq) retrieved in the passage candidate list P\mathcal{P}:

    Lrank=max⁡(0,s(q,pn)−s(q,pp)+m)\mathcal{L}_{\text{rank}} = \max\left(0, s(q, p_n) - s(q, p_p) + m\right)

    The two ranking loss terms employ distinct sample mining strategies over the retrieved candidate list P={p1,p2,…,pn}\mathcal{P} = \{p_1, p_2, \dots, p_n\}:

    • Lrank1\mathcal{L}_{\text{rank1}} (Hard Mining): ppp_p and pnp_n are selected as the positive passage and negative passage with the highest predicted relevance scores s(q,p)s(q, p), respectively.
    • Lrank2\mathcal{L}_{\text{rank2}} (Random Mining): ppp_p and pnp_n are sampled uniformly at random from the positive and negative passages present within the candidate list P\mathcal{P}.
  4. Knowl 4 — Multiview Identifier Generation Loss and Constrained Decoding

    equation

    In the learning-to-generate stage and as regularizer in the learning-to-rank stage, the autoregressive language model AM\text{AM} parameterized by θ\theta is trained on query-identifier pairs (q,I)(q, I) using the negative log-likelihood of target identifier tokens:

    Lgen=−∑j=1llog⁡pθ(ij∣q;I<j)\mathcal{L}_{\text{gen}} = -\sum_{j=1}^{l} \log p_\theta(i_j \mid q; I_{<j})

    where I=(i1,i2,…,il)I = (i_1, i_2, \dots, i_l) represents a target identifier of length ll, I<j=(i0,i1,…,ij−1)I_{<j} = (i_0, i_1, \dots, i_{j-1}) denotes the sequence of preceding tokens, and i0i_0 is a predefined start token. The target identifier II is prefixed with a view indicator designating one of three identifier views:

    1. "title": the passage title,
    2. "substring": a random text substring from the passage body, or
    3. "pseudo-query": a synthetic question generated for the passage.

    During inference, valid generation is enforced via prefix-constrained decoding using an FM-index over corpus identifiers:

    I=AM(q;b;FM-index)\mathcal{I} = \text{AM}(q; b; \text{FM-index})

    where bb denotes the beam search width, and the FM-index restricts token successors at each decoding step to substrings that exist in the corpus C\mathcal{C}.

  5. Knowl 5 — LTRGR Training and Retrieval Procedure

    algorithm

    The complete training and inference procedure for LTRGR operates as follows:

    Input: Training query set QtrainQ_{train}, passage corpus CC, pretrained generative retriever AMθAM_\theta, margin m=500m = 500, balance weight λ=1000\lambda = 1000, learning rate η=10−5\eta = 10^{-5}, beam size b=15b = 15
    Output: Trained autoregressive ranking model AMθAM_\theta
    Construct FM-index over all valid multiview identifiers (titles, substrings, pseudo-queries) in CC
    for each epoch in 1…31 \dots 3 do
        for each batch of queries q∈Qtrainq \in Q_{train} do
            Generate candidate identifiers I=AMθ(q;b;FM-index)I = AM_\theta(q; b; \text{FM-index})
            Retrieve top-200 candidate passages P={p1,…,p200}P = \{p_1, \dots, p_{200}\} using FM-index
            for each retrieved passage p∈Pp \in P do
                Find subset of covered identifiers Ip⊆II_p \subseteq I (retaining at most 40 identifiers)
                Compute passage relevance score s(q,p)=∑ip∈Ipsips(q, p) = \sum_{i_p \in I_p} s_{i_p}
            end for
            Identify positive passages Ppos⊂PP_{pos} \subset P and negative passages Pneg⊂PP_{neg} \subset P
            Select pp(1)=arg⁡max⁡p∈Pposs(q,p)p_p^{(1)} = \arg\max_{p \in P_{pos}} s(q, p) and pn(1)=arg⁡max⁡p∈Pnegs(q,p)p_n^{(1)} = \arg\max_{p \in P_{neg}} s(q, p)
            Sample pp(2)∼Pposp_p^{(2)} \sim P_{pos} and pn(2)∼Pnegp_n^{(2)} \sim P_{neg} uniformly at random
            Compute Lrank1=max⁡(0,s(q,pn(1))−s(q,pp(1))+m)L_{rank1} = \max(0, s(q, p_n^{(1)}) - s(q, p_p^{(1)}) + m)
            Compute Lrank2=max⁡(0,s(q,pn(2))−s(q,pp(2))+m)L_{rank2} = \max(0, s(q, p_n^{(2)}) - s(q, p_p^{(2)}) + m)
            Compute sequence generation loss Lgen=−∑j=1llog⁡pθ(ij∣q,I<j)L_{gen} = -\sum_{j=1}^l \log p_\theta(i_j \mid q, I_{<j})
            Compute total loss L=Lrank1+Lrank2+λLgenL = L_{rank1} + L_{rank2} + \lambda L_{gen}
            Update parameters θ←θ−η∇θL\theta \leftarrow \theta - \eta \nabla_\theta L using Adam optimizer
        end for
    end for
    return AMθAM_\theta
  6. Knowl 6 — Retrieval Performance on Open-Domain QA Benchmarks

    data/table

    The retrieval performance of LTRGR (built on BART-Large initialized with MINDER) was evaluated on the full 21-million passage corpus of Natural Questions (NQ) and TriviaQA in open-domain QA settings, using Hits@5, Hits@20, and Hits@100:

    Methods Natural Questions TriviaQA
    @5 @20 @100 @5 @20 @100
    BM25 43.6 62.9 78.1 67.7 77.3 83.9
    DPR 68.3 80.1 86.1 72.7 80.2 84.8
    GAR 59.3 73.9 85.0 73.1 80.4 85.7
    DSI-BART 28.3 47.3 65.5 – – –
    SEAL-LM 40.5 60.2 73.1 39.6 57.5 80.1
    SEAL-LM+FM 43.9 65.8 81.1 38.4 56.6 80.1
    SEAL 61.3 76.2 86.3 66.8 77.6 84.6
    MINDER 65.8 78.3 86.7 68.4 78.1 84.8
    LTRGR 68.8 80.3 87.1 70.2 79.1 85.1
    % improve +4.56% +2.55% +0.46% +2.63% +1.28% +0.35%

    LTRGR achieves state-of-the-art performance among all generative retrieval approaches across both benchmarks. On Natural Questions, LTRGR outperforms dense retrieval (DPR) across all metrics (Hits@5, Hits@20, Hits@100), marking the first instance of a generative retrieval system surpassing DPR under full-corpus open-domain QA evaluation.

  7. Knowl 7 — Retrieval Performance on the MS MARCO Web Search Benchmark

    data/table

    Evaluation on the full MS MARCO passage ranking benchmark measures retrieval accuracy in web search, where documents lack structured Wikipedia titles and require semantic text matching. Performance is evaluated using Recall@kk (R@kk) and MRR@10 (M@10):

    Methods Model Size R@5 R@20 R@100 M@10
    BM25 – 28.6 47.5 66.2 18.4
    SEAL BART-Large 19.8 35.3 57.2 12.7
    MINDER BART-Large 29.5 53.5 78.7 18.6
    NCI T5-Base – – – 9.1
    DSI (scaling up) T5-Base – – – 17.3
    DSI (scaling up) T5-Large – – – 19.8
    LTRGR BART-Large 40.2 64.5 85.2 25.5
    % improve – +36.3% +20.6% +8.26% +28.8%

    While prior generative retrieval models (SEAL, NCI) underperformed BM25 on MS MARCO due to web passage noise and lack of metadata, LTRGR with BART-Large achieves 25.5 MRR@10, outperforming the previous best generative retrieval system (DSI scaling up with T5-Large at 19.8 MRR@10) by +5.7 points (+28.8% relative improvement) and improving over baseline MINDER by +6.9 MRR@10 points.

  8. Knowl 8 — Ablation Study of Loss Components in LTRGR

    data/table

    An ablation study conducted on the Natural Questions dataset examines the contribution of each component in the LTRGR multi-task loss L=Lrank1+Lrank2+λLgen\mathcal{L} = \mathcal{L}_{\text{rank1}} + \mathcal{L}_{\text{rank2}} + \lambda \mathcal{L}_{\text{gen}}:

    Methods Hits@5 Hits@20 Hits@100
    w/o learning-to-rank (MINDER baseline) 65.8 78.3 86.7
    w/ rank loss 1 only (single hard pair) 56.1 69.4 78.7
    w/o generation loss (λ=0\lambda = 0) 63.9 76.1 84.4
    w/o rank loss (Lgen\mathcal{L}_{\text{gen}} only) 65.8 78.6 86.5
    w/o rank loss 1 (random pairs only) 68.2 80.8 87.0
    w/o rank loss 2 (hard pairs only) 67.9 79.8 86.7
    LTRGR (Full model) 68.8 80.3 87.1

    Key conclusions include:

    1. Training solely with generation loss ("w/o rank loss") yields essentially no improvement over baseline MINDER (65.8 vs. 65.8 Hits@5), confirming that performance gains stem from the learning-to-rank objective rather than extended training steps.
    2. Removing the generation loss regularizer ("w/o generation loss") degrades Hits@5 from 68.8 to 63.9, below the baseline model, because optimizing ranking loss alone causes the model to collapse into local minima and assign lower scores uniformly.
    3. Combining both hard negative mining (Lrank1\mathcal{L}_{\text{rank1}}) and random negative mining (Lrank2\mathcal{L}_{\text{rank2}}) yields the strongest overall top-rank accuracy (68.8 Hits@5).
  9. Knowl 9 — Generalization of LTRGR to SEAL Substring-Based Generative Retrieval

    empirical result

    The LTRGR learning-to-rank framework is model-agnostic and generalizes across different generative retrieval architectures and identifier designs. When applied to SEAL (which indexes and retrieves passages using purely substring identifiers rather than multiview identifiers) on the Natural Questions benchmark, SEAL-LTR improves performance across all evaluated metrics compared to base SEAL:

    • Hits@5 improves from 61.3 to 63.7 (+2.4 points),
    • Hits@20 improves from 76.2 to 78.1 (+1.9 points),
    • Hits@100 improves from 86.3 to 86.4 (+0.1 points).

    The largest gains occur in top-rank metrics (Hits@5), demonstrating that the ranking objective effectively pushes positive passages up the retrieval list regardless of whether multiview or substring identifiers are used.

  10. Knowl 10 — Empirical Comparison of Margin-Based Ranking Loss and List-Wise InfoNCE in LTRGR

    empirical result

    When comparing different ranking loss formulations within LTRGR on the Natural Questions dataset, the margin-based pairwise loss significantly outperforms the list-wise InfoNCE loss:

    • Margin Loss: Lrank=max⁡(0,s(q,pn)−s(q,pp)+m)\mathcal{L}_{\text{rank}} = \max(0, s(q, p_n) - s(q, p_p) + m) achieves Hits@5 of 68.8, Hits@20 of 80.3, and Hits@100 of 87.1.
    • List-wise InfoNCE Loss: Lrank=−log⁡es(q,pp)es(q,pp)+∑pnes(q,pn)\mathcal{L}_{\text{rank}} = -\log \frac{e^{s(q, p_p)}}{e^{s(q, p_p)} + \sum_{p_n} e^{s(q, p_n)}} with 19 randomly sampled negative passages achieves Hits@5 of 65.4, Hits@20 of 78.5, and Hits@100 of 86.3.

    The underperformance of InfoNCE in generative retrieval is attributed to two factors: (1) the passage relevance scores s(q,p)s(q, p) are unnormalized sums of identifier logits rather than bounded inner products or cosine similarities, making the exponentiated softmax distribution unstable and difficult to optimize; and (2) higher per-step training computational cost which constrained InfoNCE training.

Coverage note — None was omitted; the knowls cover all key methodological components, loss functions, algorithms, benchmark evaluations, ablations, and analytic comparisons.

References

  1. 1.Bevilacqua, M.; Ottaviano, G.; Lewis, P.; Yih, W.-t.; Riedel, S.; and Petroni, F. 2022. Autoregressive search engines: Generating substrings as document identifiers. arXiv preprint arXiv:2204.10628.
  2. 2.Burges, C.; Shaked, T.; Renshaw, E.; Lazier, A.; Deeds, M.; Hamilton, N.; and Hullender, G. 2005. Learning to rank using gradient descent. In Proceedings of the 22nd international conference on Machine learning, 89–96.
  3. 3.Cao, Z.; Qin, T.; Liu, T.-Y.; Tsai, M.-F.; and Li, H. 2007. Learning to rank: from pairwise approach to listwise approach. In Proceedings of the 24th international conference on Machine learning, 129–136.
  4. 4.Chen, D.; Fisch, A.; Weston, J.; and Bordes, A. 2017. Reading Wikipedia to Answer Open-Domain Questions. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 1870–1879.
  5. 5.Cossock, D.; and Zhang, T. 2006. Subset ranking using regression. In Learning Theory: 19th Annual Conference on Learning Theory, COLT 2006, Pittsburgh, PA, USA, June 22-25, 2006. Proceedings 19, 605–619. Springer.
  6. 6.Crammer, K.; and Singer, Y. 2001. Pranking with ranking. Advances in neural information processing systems, 14.
  7. 7.De Cao, N.; Izacard, G.; Riedel, S.; and Petroni, F. 2020. Autoregressive Entity Retrieval. In International Conference on Learning Representations.
  8. 8.Ferragina, P.; and Manzini, G. 2000. Opportunistic data structures with applications. In Proceedings 41st Annual Symposium on Foundations of Computer Science, 390–398.
  9. 9.Freund, Y.; Iyer, R.; Schapire, R. E.; and Singer, Y. 2003. An efficient boosting algorithm for combining preferences. Journal of machine learning research, 4(Nov): 933–969.
  10. 10.Joshi, M.; Choi, E.; Weld, D. S.; and Zettlemoyer, L. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1601–1611.
  11. 11.Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; and Yih, W.-t. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the International Conference on Empirical Methods in Natural Language Processing, 6769–6781. ACL.
  12. 12.Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; et al. 2019. Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics, 7: 452–466.
  13. 13.Lee, K.; Chang, M.-W.; and Toutanova, K. 2019. Latent Retrieval for Weakly Supervised Open Domain Question Answering. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 6086–6096. ACL.
  14. 14.Li, H. 2011. A short introduction to learning to rank. IEICE TRANSACTIONS on Information and Systems, 94(10): 1854–1862.
  15. 15.Li, P.; Wu, Q.; and Burges, C. 2007. Mcrank: Learning to rank using multiple classification and gradient boosting. Advances in neural information processing systems, 20.
  16. 16.Li, Y.; Yang, N.; Wang, L.; Wei, F.; and Li, W. 2023a. Generative retrieval for conversational question answering. Information Processing & Management, 60(5): 103475.
  17. 17.Li, Y.; Yang, N.; Wang, L.; Wei, F.; and Li, W. 2023b. Multiview Identifiers Enhanced Generative Retrieval. arXiv preprint arXiv:2305.16675.
  18. 18.Mao, Y.; He, P.; Liu, X.; Shen, Y.; Gao, J.; Han, J.; and Chen, W. 2021. Generation-Augmented Retrieval for Open-Domain Question Answering. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 4089–4100. ACL.
  19. 19.Nguyen, T.; Rosenberg, M.; Song, X.; Gao, J.; Tiwary, S.; Majumder, R.; and Deng, L. 2016. MS MARCO: A human generated machine reading comprehension dataset. In CoCo@ NIPs.
  20. 20.Nogueira, R.; and Cho, K. 2019. Passage Re-ranking with BERT. arXiv preprint arXiv:1901.04085.
  21. 21.Pradeep, R.; Hui, K.; Gupta, J.; Lelkes, A. D.; Zhuang, H.; Lin, J.; Metzler, D.; and Tran, V. Q. 2023. How Does Generative Retrieval Scale to Millions of Passages? arXiv preprint arXiv:2305.11841.
  22. 22.Ren, R.; Zhao, W. X.; Liu, J.; Wu, H.; Wen, J.-R.; and Wang, H. 2023. TOME: A Two-stage Approach for Model-based Retrieval. arXiv preprint arXiv:2305.11161.
  23. 23.Tay, Y.; Tran, V. Q.; Dehghani, M.; Ni, J.; Bahri, D.; Mehta, H.; Qin, Z.; Hui, K.; Zhao, Z.; Gupta, J.; et al. 2022. Transformer memory as a differentiable search index. arXiv preprint arXiv:2202.06991.
  24. 24.Wang, Y.; Hou, Y.; Wang, H.; Miao, Z.; Wu, S.; Chen, Q.; Xia, Y.; Chi, C.; Zhao, G.; Liu, Z.; et al. 2022. A neural corpus indexer for document retrieval. Advances in Neural Information Processing Systems, 35: 25600–25614.
  25. 25.Xia, F.; Liu, T.-Y.; Wang, J.; Zhang, W.; and Li, H. 2008. Listwise approach to learning to rank: theory and algorithm. In Proceedings of the 25th international conference on Machine learning, 1192–1199.

Citation

MLA
Li, Y., et al. “Learning to Rank in Generative Retrieval”. arXiv, 2023, http://arxiv.org/abs/2306.15222v2.
APA
Li, Y., Yang, N., Wang, L., Wei, F., & Li, W. (2023). Learning to Rank in Generative Retrieval. arXiv. http://arxiv.org/abs/2306.15222v2
Chicago
Li, Y., N. Yang, L. Wang, F. Wei, and W. Li. 2023. “Learning to Rank in Generative Retrieval”. arXiv. http://arxiv.org/abs/2306.15222v2.
Harvard
Li, Y. et al. (2023) “Learning to Rank in Generative Retrieval”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.15222v2.
Vancouver
1. Li Y, Yang N, Wang L, Wei F, Li W (2023) Learning to Rank in Generative Retrieval. arXiv

BibTeX

@article{li2023learning,
  title = {Learning to Rank in Generative Retrieval},
  author = {Li, Yongqi and Yang, Nan and Wang, Liang and Wei, Furu and Li, Wenjie},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.15222v2},
  eprint = {2306.15222}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF