Learning to Rank in Generative Retrieval
Yongqi LiNan YangLiang WangFuru WeiWenjie Li
Proposes LTRGR, a framework that incorporates ranking losses into autoregressive models to bridge the gap between identifier generation and final passage ranking without adding computational overhead during inference.
Modern information retrieval is increasingly exploring generative retrieval, where an autoregressive language model generates text identifier strings—such as titles, substrings, or pseudo-queries—rather than matching indexed dense vectors. However, existing generative retrieval methods suffer from an objective mismatch: they optimize the model purely to generate valid text identifiers and rely on an external heuristic function to convert these generated strings into a final ranked list of documents. This disconnect between string-generation training and document-ranking goals creates a substantial performance bottleneck.
The article demonstrates that integrating classical learning-to-rank techniques directly into generative retrieval resolves this disconnect. The authors introduce LTRGR, a framework that augments existing generative models with an additional training phase driven by a margin-based passage rank loss alongside standard generation loss.
The approach evaluates full corpus benchmarks using two established question-answering datasets (Natural Questions and TriviaQA) and a large-scale web search dataset (MSMARCO). The model builds upon the MINDER architecture using a BART-Large backbone. It first generates identifier candidates via constrained decoding, maps these identifiers to passage rank scores, and directly propagates rank loss gradients through the model logits without altering the inference procedure.
The findings establish new state-of-the-art results across generative retrieval benchmarks. On the MSMARCO web search task, LTRGR improved top-10 Mean Reciprocal Rank by 28.8% (achieving 25.5 compared to the prior best generative benchmark of 18.6) and Recall@5 by 36.3%, outperforming even larger baseline models. On Natural Questions, LTRGR increased Hits@5 by 4.56% and became the first generative system to outperform Dense Passage Retrieval across Hits@5, Hits@20, and Hits@100 on the full corpus. Analysis confirms that these performance gains concentrate primarily at top-ranked positions and generalize across other generative architectures like SEAL.
These results demonstrate that generative retrieval can match or exceed classical dense retrieval when properly aligned with ranking objectives, requiring no extra compute during live query inference. Organizations building search or question-answering systems can adopt this secondary ranking phase to achieve substantial retrieval accuracy without adding operational latency. Future development should explore normalized list-wise loss formulations and enhanced negative sample mining strategies to close remaining gaps on complex benchmarks like TriviaQA.
- Paper: How Does Generative Retrieval Scale to Millions of Passages?, Ronak Pradeep et al. (2023). This paper examines how generative retrieval scales across millions of passages in MS MARCO, establishing the core sequence-to-sequence document retrieval problem that LTRGR directly improves upon.
- Paper: From RankNet to LambdaRank to LambdaMART: An Overview, Christopher J. C. Burges (2010). This work details foundational learning-to-rank algorithms and gradient-based ranking objectives that directly inform LTRGR's passage margin ranking losses.
- Paper: Learning to rank using gradient descent, Chris Burges et al. (2005). This seminal text introduces pairwise ranking loss optimization via gradient descent, providing the mathematical basis for the ranking objectives integrated into the generative retrieval framework.
- Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). This foundational paper establishes standard passage re-ranking methodologies on MS MARCO that classical and generative retrieval models aim to replicate or replace.
- Paper: Learning to rank: from pairwise approach to listwise approach, Zhe Cao et al. (2007). This work outlines listwise learning-to-rank formulations, providing the foundational principles for moving beyond naive generation scoring toward direct ranking optimization.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This paper introduces the standard dense retrieval and generative framework for knowledge-intensive benchmarks like Natural Questions and TriviaQA evaluated in the source.
- Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). This study analyzes metric mismatch and sequence-level training objectives in autoregressive generation, which motivates resolving the objective mismatch in generative retrieval.
- Paper: IR evaluation methods for retrieving highly relevant documents, Kalervo Järvelin et al. (2000). This foundational work formalizes Discounted Cumulative Gain and evaluation metrics designed to prioritize top-ranked documents in information retrieval.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). RankRAG extends the concept of unifying ranking and generation into large language models by instruction-tuning a single model to perform context reranking alongside answer generation.
- Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Siddharth Gollapudi et al. (2026). BlockSearch explores an alternative scaling direction for generative retrieval by testing whether language models can retrieve document identifiers directly within extended million-token contexts.
- Paper: GenRec: An LLM-Backed Recommendation Ranker at Netflix, Ying Li et al. (2026). GenRec applies LLM-backed ranking objectives directly within real-world large-scale recommendation systems, extending generative ranking principles to production recommendation tasks.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). SEARCH-R1 takes the integration of generation and retrieval further by using reinforcement learning to train models to reason and autonomously interact with search engines.
