Built independently by an author, for readers. Read the story and support ChapterPal

topic

passage retrieval

Passage retrieval is an information retrieval process that identifies and extracts specific segments of text, such as paragraphs or sentences, from a collection of documents in response to a user query. Unlike traditional document retrieval, which evaluates and returns entire documents or web pages, passage retrieval operates on granular text fragments to locate the exact portion containing the relevant information. In practice, documents are divided into fixed or semantic units that are indexed and ranked using lexical matching algorithms, dense vector embeddings, or hybrid search techniques. This approach serves as a critical component in automated question answering systems and retrieval-augmented generation pipelines, enabling downstream language models to consume concise, highly targeted context rather than processing large volumes of irrelevant text.

6 items

Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Gangwoo Kim, Sungdong Kim, Byeongguk Jeon, Joonsuk Park, Jaewoo Kang

OrganizationsKorea Advanced Institute of Science and TechnologyKorea UniversityNaver AI LabNAVER CloudUniversity of Richmond

Why you should read this

Proposes Tree of Clarifications, a framework that recursively builds a tree of disambiguated questions guided by retrieved external knowledge and self-verification pruning to generate comprehensive long-form answers to ambiguous open-domain questions.

Questions in open-domain question answering are often ambiguous, allowing multiple interpretations. One approach to handling them is to identify all possible interpretations of the ambiguous question (AQ) and to generate a long-form answer addressing them all, as suggested by Stelmakh et al. (2022). While it provides a comprehensive response without bothering the user for clarification, considering multiple dimensions of ambiguity and gathering corresponding knowledge remains a challenge. To cope with the challenge, we propose a novel framework, Tree of Clarifications (ToC): It recursively constructs a tree of disambiguations for the AQ—via few-shot prompting leveraging external knowledge—and uses it to generate a long-form answer. ToC outperforms existing baselines on ASQA in a few-shot setup across all metrics, while surpassing fully-supervised baselines trained on the whole training set in terms of Disambig-F1 and Disambig-ROUGE. Code is available at github.com/gankim/tree-of-clarifications.

Added

2026-10-03

SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval

SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval

Hossein A. Rahmani, Xi Wang, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas

OrganizationsAlan Turing Institute & AmazonMicrosoftUniversity College LondonUniversity of Sheffield

Why you should read this

Presents SynDL, a large-scale passage retrieval benchmark extending the TREC Deep Learning Track across more than 1,900 queries with synthetic language model labels, providing a cost-effective evaluation resource that yields search system rankings highly correlated with human assessments.

Large-scale test collections play a crucial role in Information Retrieval (IR) research. However, according to the Cranfield paradigm and the research into publicly available datasets, the existing information retrieval research studies are commonly developed on small-scale datasets that rely on human assessors for relevance judgments - a time-intensive and expensive process. Recent studies have shown the strong capability of Large Language Models (LLMs) in producing reliable relevance judgments with human accuracy but at a greatly reduced cost. In this paper, to address the missing large-scale ad-hoc document retrieval dataset, we extend the TREC Deep Learning Track (DL) test collection via additional language model synthetic labels to enable researchers to test and evaluate their search systems at a large scale. Specifically, such a test collection includes more than 1,900 test queries from the previous years of tracks. We compare system evaluation with past human labels from past years and find that our synthetically created large-scale test collection can lead to highly correlated system rankings.

Added

2026-09-30

Improving Passage Retrieval with Zero-Shot Question Generation

Improving Passage Retrieval with Zero-Shot Question Generation

Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, Luke Zettlemoyer

OrganizationsGoogleMcGill UniversityMetaMila – Québec Artificial Intelligence InstituteUniversity of Washington

Why you should read this

Proposes a simple zero-shot passage re-ranking method that scores retrieved texts by their likelihood of generating the input question using off-the-shelf language models, consistently outperforming supervised retrieval pipelines across multiple open-domain question answering benchmarks without any fine-tuning.

We propose a simple and effective re-ranking method for improving passage retrieval in open question answering. The re-ranker re-scores retrieved passages with a zero-shot question generation model, which uses a pre-trained language model to compute the probability of the input question conditioned on a retrieved passage. This approach can be applied on top of any retrieval method (e.g. neural or keyword-based), does not require any domain- or task-specific training (and therefore is expected to generalize better to data distribution shifts), and provides rich cross-attention between query and passage (i.e. it must explain every token in the question). When evaluated on a number of open-domain retrieval datasets, our re-ranker improves strong unsupervised retrieval models by 6%-18% absolute and strong supervised models by up to 12% in terms of top-20 passage retrieval accuracy. We also obtain new state-of-the-art results on full open-domain question answering by simply adding the new re-ranker to existing models with no further changes.

Added

2026-09-28

Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval

Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval

Luyu Gao, Jamie Callan

OrganizationsCarnegie Mellon University

Why you should read this

Proposes coCondenser, an unsupervised pre-training strategy that combines the Condenser architecture with a corpus-level contrastive loss to structure passage embedding spaces, enabling dense retrievers to match state-of-the-art retrieval accuracy using simple small-batch fine-tuning without complex data engineering.

Recent research demonstrates the effectiveness of using fine-tuned language models (LM) for dense retrieval. However, dense retrievers are hard to train, typically requiring heavily engineered fine-tuning pipelines to realize their full potential. In this paper, we identify and address two underlying problems of dense retrievers: i) fragility to training data noise and ii) requiring large batches to robustly learn the embedding space. We use the recently proposed Condenser pre-training architecture, which learns to condense information into the dense vector through LM pre-training. On top of it, we propose coCondenser, which adds an unsupervised corpus-level contrastive loss to warm up the passage embedding space. Experiments on MS-MARCO, Natural Question, and Trivia QA datasets show that coCondenser removes the need for heavy data engineering such as augmentation, synthesis, or filtering, and the need for large batch training. It shows comparable performance to RocketQA, a state-of-the-art, heavily engineered system, using simple small batch fine-tuning.

Added

2026-09-28

Passage Re-ranking with BERT

Passage Re-ranking with BERT

Rodrigo Nogueira, Kyunghyun Cho

OrganizationsCIFARMetaNew York University

Why you should read this

Introduces the indispensable cross-encoder approach for late-stage document selection that explicitly evaluates the deep attention between the user query and each candidate passage.

Recently, neural models pretrained on a language modeling task, such as ELMo (Peters et al., 2017), OpenAI GPT (Radford et al., 2018), and BERT (Devlin et al., 2018), have achieved impressive results on various natural language processing tasks such as question-answering and natural language inference. In this paper, we describe a simple re-implementation of BERT for query-based passage re-ranking. Our system is the state of the art on the TREC-CAR dataset and the top entry in the leaderboard of the MS MARCO passage retrieval task, outperforming the previous state of the art by 27% (relative) in MRR@10. The code to reproduce our results is available at https://github.com/nyu-dl/dl4marco-bert

Added

2026-05-08