UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers
Jon Saad-FalconOmar KhattabKeshav SanthanamRadu FlorianMartin FranzSalim RoukosAvirup SilMd. Arafat SultanChristopher Potts
Proposes a cost-effective unsupervised domain adaptation framework that generates synthetic queries using tiered language models and distills an ensemble of rerankers into a single ColBERTv2 retriever, achieving high zero-shot accuracy without the latency of cross-encoder reranking.
Modern neural information retrieval systems struggle to maintain high accuracy when deployed in new, specialized domains where labeled training data is unavailable. Existing methods attempt to overcome this by using large language models to generate synthetic queries and training high-precision reranker models. However, using these rerankers during live searches causes high computational latency and massive operational costs, making them impractical for real-time, user-facing applications.
The article demonstrates and evaluates an unsupervised domain adaptation framework named UDAPDR (Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers). The main objective is to establish an efficient pipeline that automatically adapts search systems to new domains using only unlabeled text passages, matching the high accuracy of expensive rerankers without suffering from their latency overhead.
The evaluated approach uses a multi-stage process. First, an advanced language model (GPT-3) generates a small seed set of synthetic queries from domain passages. These examples form corpus-adapted prompts that guide a significantly cheaper language model (Flan-T5 XXL) to generate thousands of synthetic queries cheaply. Multiple cross-encoder rerankers (using DeBERTaV3) are trained on distinct subsets of these queries and subsequently distilled into a single, efficient multi-vector search model (ColBERTv2). The authors validated this methodology across standard search benchmarks, including LoTTE, BEIR, Natural Questions, and SQuAD.
The findings show that UDAPDR substantially improves search accuracy across diverse domains without requiring manually labeled target data. On the LoTTE benchmark, UDAPDR improved zero-shot retrieval success by an average of 7.1 points on forum queries and 3.9 points on search queries compared to baseline retrieval. On the BEIR benchmark, it increased accuracy by 5.2 points. Crucially, the distilled system achieved retrieval accuracy that matches or exceeds a pipeline with a live reranker, while operating at a latency of 35 milliseconds per query compared to over 400 to 20,000 milliseconds for reranker-based setups. Additionally, the researchers found that generating 10,000 to 20,000 queries across multiple prompts works better than scaling to millions of synthetic examples, which can saturate or degrade performance.
These results indicate that organizations can achieve state-of-the-art search quality across specialized domains without incurring high inference costs or infrastructure delays. Knowledge distillation from multiple lightweight teacher models effectively eliminates the traditional trade-off between search latency and relevance, allowing low-cost deployment in production environments.
Organizations seeking to adapt neural search to novel domains should adopt multi-teacher synthetic generation pipelines instead of deploying live cross-encoder rerankers. Practitioner teams should use affordable open-source language models to generate synthetic query pools in the range of 10,000 to 100,000 examples and distill them directly into fast retrieval architectures. Future development should explore expanding this adaptation framework to other retrieval backbones and non-English corpora.
The confidence in these findings is high across English benchmarks, but stakeholders should note key limitations. The approach relies on having access to a substantial pool of in-domain target passages, and its effectiveness on extremely small passage collections remains unproven. Furthermore, synthetic data may inherit latent biases from underlying language models, and the findings are currently limited to English datasets.
- Paper: Improving Passage Retrieval with Zero-Shot Question Generation, Devendra Singh Sachan et al. (2022). UPR establishes how synthetic questions and language-model passage reranking can improve retrieval without target labels, the key precursor to UDAPDR’s query-generation and reranking pipeline.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). BEIR defines the zero-shot, multi-domain benchmark framework used to interpret UDAPDR’s cross-domain retrieval evaluation.
- Paper: Unsupervised Dense Information Retrieval with Contrastive Learning, Gautier Izacard et al. (2021). Contriever provides the unsupervised dense-retrieval baseline that clarifies the retrieval setting UDAPDR seeks to improve.
- Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). BERT passage reranking introduces the cross-encoder relevance-scoring setup underlying the teacher rerankers UDAPDR distills.
No sufficiently relevant recommendations were found.
