Built independently by an author, for readers. Read the story and support ChapterPal

keyword

synthetic query generation

Synthetic query generation is an information retrieval technique in which language models automatically produce artificial search queries corresponding to given documents or passages. By simulating the questions or search phrases a user might submit to find specific content, the process creates pairs of synthetic queries and relevant texts without requiring manual human annotation. This method is primarily used as a data augmentation strategy to train, fine-tune, and evaluate retrieval and reranking models when labeled real-world datasets are scarce, expensive, or unavailable. It also facilitates unsupervised domain adaptation across specialized subject areas and supports the indexing of content in generative retrieval systems by representing documents through the potential search requests they can satisfy.

2 items

UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers

UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers

Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md. Arafat Sultan, Christopher Potts

OrganizationsIBMStanford University

Why you should read this

Proposes a cost-effective unsupervised domain adaptation framework that generates synthetic queries using tiered language models and distills an ensemble of rerankers into a single ColBERTv2 retriever, achieving high zero-shot accuracy without the latency of cross-encoder reranking.

Many information retrieval tasks require large labeled datasets for fine-tuning. However, such datasets are often unavailable, and their utility for real-world applications can diminish quickly due to domain shifts. To address this challenge, we develop and motivate a method for using large language models (LLMs) to generate large numbers of synthetic queries cheaply. The method begins by generating a small number of synthetic queries using an expensive LLM. After that, a much less expensive one is used to create large numbers of synthetic queries, which are used to fine-tune a family of reranker models. These rerankers are then distilled into a single efficient retriever for use in the target domain. We show that this technique boosts zero-shot accuracy in long-tail domains and achieves substantially lower latency than standard reranking methods.

Added

2026-10-03

How Does Generative Retrieval Scale to Millions of Passages?

How Does Generative Retrieval Scale to Millions of Passages?

Ronak Pradeep, Kai Hui, Jai Gupta, Ádám D. Lelkes, Honglei Zhuang, Jimmy Lin, Donald Metzler, Vinh Q. Tran

OrganizationsGoogleUniversity of Waterloo

Why you should read this

Presents the first comprehensive empirical evaluation of generative retrieval scaled up to 8.8 million passages and 11 billion parameters, showing that while synthetic queries are vital for indexing, current architectures struggle to match standard dual encoders as corpus size grows.

The emerging paradigm of generative retrieval re-frames the classic information retrieval problem into a sequence-to-sequence modeling task, forgoing external indices and encoding an entire document corpus within a single Transformer. Although many different approaches have been proposed to improve the effectiveness of generative retrieval, they have only been evaluated on document corpora on the order of 100K in size. We conduct the first empirical study of generative retrieval techniques across various corpus scales, ultimately scaling up to the entire MS MARCO passage ranking task with a corpus of 8.8M passages and evaluating model sizes up to 11B parameters. We uncover several findings about scaling generative retrieval to millions of passages; notably, the central importance of using synthetic queries as document representations during indexing, the ineffectiveness of existing proposed architecture modifications when accounting for compute cost, and the limits of naively scaling model parameters with respect to retrieval performance. While we find that generative retrieval is competitive with state-of-the-art dual encoders on small corpora, scaling to millions of passages remains an important and unsolved challenge. We believe these findings will be valuable for the community to clarify the current state of generative retrieval, highlight the unique challenges, and inspire new research directions.

Added

2026-09-26