Built independently by an author, for readers. Read the story and support ChapterPal

keyword

passage ranking

Passage ranking is an information retrieval task that involves scoring and ordering a collection of short text segments, known as passages, according to their relevance to a given user query. Unlike traditional document retrieval, which focuses on returning entire documents or web pages, passage ranking aims to identify specific paragraphs or excerpts that directly address or answer the information need. In modern search and question answering systems, the task is typically addressed using techniques such as sparse lexical matching, dense representation models, neural late interaction mechanisms, and generative retrieval frameworks that evaluate semantic similarity to generate an ordered list of candidate passages.

5 items

Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval

Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval

Luyu Gao, Jamie Callan

OrganizationsCarnegie Mellon University

Why you should read this

Proposes coCondenser, an unsupervised pre-training strategy that combines the Condenser architecture with a corpus-level contrastive loss to structure passage embedding spaces, enabling dense retrievers to match state-of-the-art retrieval accuracy using simple small-batch fine-tuning without complex data engineering.

Recent research demonstrates the effectiveness of using fine-tuned language models (LM) for dense retrieval. However, dense retrievers are hard to train, typically requiring heavily engineered fine-tuning pipelines to realize their full potential. In this paper, we identify and address two underlying problems of dense retrievers: i) fragility to training data noise and ii) requiring large batches to robustly learn the embedding space. We use the recently proposed Condenser pre-training architecture, which learns to condense information into the dense vector through LM pre-training. On top of it, we propose coCondenser, which adds an unsupervised corpus-level contrastive loss to warm up the passage embedding space. Experiments on MS-MARCO, Natural Question, and Trivia QA datasets show that coCondenser removes the need for heavy data engineering such as augmentation, synthesis, or filtering, and the need for large batch training. It shows comparable performance to RocketQA, a state-of-the-art, heavily engineered system, using simple small batch fine-tuning.

Added

2026-09-28

How Does Generative Retrieval Scale to Millions of Passages?

How Does Generative Retrieval Scale to Millions of Passages?

Ronak Pradeep, Kai Hui, Jai Gupta, Ádám D. Lelkes, Honglei Zhuang, Jimmy Lin, Donald Metzler, Vinh Q. Tran

OrganizationsGoogleUniversity of Waterloo

Why you should read this

Presents the first comprehensive empirical evaluation of generative retrieval scaled up to 8.8 million passages and 11 billion parameters, showing that while synthetic queries are vital for indexing, current architectures struggle to match standard dual encoders as corpus size grows.

The emerging paradigm of generative retrieval re-frames the classic information retrieval problem into a sequence-to-sequence modeling task, forgoing external indices and encoding an entire document corpus within a single Transformer. Although many different approaches have been proposed to improve the effectiveness of generative retrieval, they have only been evaluated on document corpora on the order of 100K in size. We conduct the first empirical study of generative retrieval techniques across various corpus scales, ultimately scaling up to the entire MS MARCO passage ranking task with a corpus of 8.8M passages and evaluating model sizes up to 11B parameters. We uncover several findings about scaling generative retrieval to millions of passages; notably, the central importance of using synthetic queries as document representations during indexing, the ineffectiveness of existing proposed architecture modifications when accounting for compute cost, and the limits of naively scaling model parameters with respect to retrieval performance. While we find that generative retrieval is competitive with state-of-the-art dual encoders on small corpora, scaling to millions of passages remains an important and unsolved challenge. We believe these findings will be valuable for the community to clarify the current state of generative retrieval, highlight the unique challenges, and inspire new research directions.

Added

2026-09-26

Learning to Rank in Generative Retrieval

Learning to Rank in Generative Retrieval

Yongqi Li, Nan Yang, Liang Wang, Furu Wei, Wenjie Li

OrganizationsHong Kong Polytechnic UniversityMicrosoft

Why you should read this

Proposes LTRGR, a framework that incorporates ranking losses into autoregressive models to bridge the gap between identifier generation and final passage ranking without adding computational overhead during inference.

Generative retrieval stands out as a promising new paradigm in text retrieval that aims to generate identifier strings of relevant passages as the retrieval target. This generative paradigm taps into powerful generative language models, distinct from traditional sparse or dense retrieval methods. However, only learning to generate is insufficient for generative retrieval. Generative retrieval learns to generate identifiers of relevant passages as an intermediate goal and then converts predicted identifiers into the final passage rank list. The disconnect between the learning objective of autoregressive models and the desired passage ranking target leads to a learning gap. To bridge this gap, we propose a learning-to-rank framework for generative retrieval, dubbed LTRGR. LTRGR enables generative retrieval to learn to rank passages directly, optimizing the autoregressive model toward the final passage ranking target via a rank loss. This framework only requires an additional learning-to-rank training phase to enhance current generative retrieval systems and does not add any burden to the inference stage. We conducted experiments on three public benchmarks, and the results demonstrate that LTRGR achieves state-of-the-art performance among generative retrieval methods. The code and checkpoints are released at https://github.com/liyongqi67/LTRGR.

Added

2026-09-26

ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction

ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction

Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, Matei Zaharia

OrganizationsGeorgia Institute of TechnologyStanford University

Why you should read this

Presents ColBERTv2, a neural retrieval model that combines denoised supervision with residual vector compression to cut indexing storage by 6–10× while outperforming existing dense retrievers across diverse in-domain and zero-shot benchmarks.

Neural information retrieval (IR) has greatly advanced search and other knowledge-intensive language tasks. While many neural IR methods encode queries and documents into single-vector representations, late interaction models produce multi-vector representations at the granularity of each token and decompose relevance modeling into scalable token-level computations. This decomposition has been shown to make late interaction more effective, but it inflates the space footprint of these models by an order of magnitude. In this work, we introduce ColBERTv2, a retriever that couples an aggressive residual compression mechanism with a denoised supervision strategy to simultaneously improve the quality and space footprint of late interaction. We evaluate ColBERTv2 across a wide range of benchmarks, establishing state-of-the-art quality within and outside the training domain while reducing the space footprint of late interaction models by 6–10×.

Added

2026-09-26

Text Embeddings by Weakly-Supervised Contrastive Pre-training

Text Embeddings by Weakly-Supervised Contrastive Pre-training

Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei

OrganizationsMicrosoft

Why you should read this

Introduces E5, a novel family of text embeddings trained with weakly-supervised contrastive pre-training on a curated large-scale dataset, achieving state-of-the-art performance across diverse NLP tasks, including outperforming BM25 in zero-shot retrieval and setting new records on MTEB with significantly fewer parameters.

This paper presents E5, a family of state-of-the-art text embeddings that transfer well to a wide range of tasks. The model is trained in a contrastive manner with weak supervision signals from our curated large-scale text pair dataset (called CCPairs). E5 can be readily used as a general-purpose embedding model for any tasks requiring a single-vector representation of texts such as retrieval, clustering, and classification, achieving strong performance in both zero-shot and fine-tuned settings. We conduct extensive evaluations on 56 datasets from the BEIR and MTEB benchmarks. For zero-shot settings, E5 is the first model that outperforms the strong BM25 baseline on the BEIR retrieval benchmark without using any labeled data. When fine-tuned, E5 obtains the best results on the MTEB benchmark, beating existing embedding models with 40x more parameters.

Added

2026-05-03

Creative Commons License