UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers

Jon Saad-FalconOmar KhattabKeshav SanthanamRadu FlorianMartin FranzSalim RoukosAvirup SilMd. Arafat SultanChristopher Potts

article2023EMNLP68 citations

Proposes a cost-effective unsupervised domain adaptation framework that generates synthetic queries using tiered language models and distills an ensemble of rerankers into a single ColBERTv2 retriever, achieving high zero-shot accuracy without the latency of cross-encoder reranking.

Listen

Modern neural information retrieval systems struggle to maintain high accuracy when deployed in new, specialized domains where labeled training data is unavailable. Existing methods attempt to overcome this by using large language models to generate synthetic queries and training high-precision reranker models. However, using these rerankers during live searches causes high computational latency and massive operational costs, making them impractical for real-time, user-facing applications.

The article demonstrates and evaluates an unsupervised domain adaptation framework named UDAPDR (Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers). The main objective is to establish an efficient pipeline that automatically adapts search systems to new domains using only unlabeled text passages, matching the high accuracy of expensive rerankers without suffering from their latency overhead.

The evaluated approach uses a multi-stage process. First, an advanced language model (GPT-3) generates a small seed set of synthetic queries from domain passages. These examples form corpus-adapted prompts that guide a significantly cheaper language model (Flan-T5 XXL) to generate thousands of synthetic queries cheaply. Multiple cross-encoder rerankers (using DeBERTaV3) are trained on distinct subsets of these queries and subsequently distilled into a single, efficient multi-vector search model (ColBERTv2). The authors validated this methodology across standard search benchmarks, including LoTTE, BEIR, Natural Questions, and SQuAD.

The findings show that UDAPDR substantially improves search accuracy across diverse domains without requiring manually labeled target data. On the LoTTE benchmark, UDAPDR improved zero-shot retrieval success by an average of 7.1 points on forum queries and 3.9 points on search queries compared to baseline retrieval. On the BEIR benchmark, it increased accuracy by 5.2 points. Crucially, the distilled system achieved retrieval accuracy that matches or exceeds a pipeline with a live reranker, while operating at a latency of 35 milliseconds per query compared to over 400 to 20,000 milliseconds for reranker-based setups. Additionally, the researchers found that generating 10,000 to 20,000 queries across multiple prompts works better than scaling to millions of synthetic examples, which can saturate or degrade performance.

These results indicate that organizations can achieve state-of-the-art search quality across specialized domains without incurring high inference costs or infrastructure delays. Knowledge distillation from multiple lightweight teacher models effectively eliminates the traditional trade-off between search latency and relevance, allowing low-cost deployment in production environments.

Organizations seeking to adapt neural search to novel domains should adopt multi-teacher synthetic generation pipelines instead of deploying live cross-encoder rerankers. Practitioner teams should use affordable open-source language models to generate synthetic query pools in the range of 10,000 to 100,000 examples and distill them directly into fast retrieval architectures. Future development should explore expanding this adaptation framework to other retrieval backbones and non-English corpora.

The confidence in these findings is high across English benchmarks, but stakeholders should note key limitations. The approach relies on having access to a substantial pool of in-domain target passages, and its effectiveness on extremely small passage collections remains unproven. Furthermore, synthetic data may inherit latent biases from underlying language models, and the findings are currently limited to English datasets.

Saad-Falcon et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers

Abstract

Many information retrieval tasks require large labeled datasets for fine-tuning. However, such datasets are often unavailable, and their utility for real-world applications can diminish quickly due to domain shifts. To address this challenge, we develop and motivate a method for using large language models (LLMs) to generate large numbers of synthetic queries cheaply. The method begins by generating a small number of synthetic queries using an expensive LLM. After that, a much less expensive one is used to create large numbers of synthetic queries, which are used to fine-tune a family of reranker models. These rerankers are then distilled into a single efficient retriever for use in the target domain. We show that this technique boosts zero-shot accuracy in long-tail domains and achieves substantially lower latency than standard reranking methods.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 2.1 Data Augmentation for Neural IR
  • 2.2 Pretraining Objectives for IR
  • 3 Methodology
  • 4 Experiments
  • 4.1 Models
  • 4.2 Datasets
  • 4.3 Multi-Reranker Domain Adaptation
  • 4.4 Query Latency
  • 4.5 Impact of Pretrained Components
  • 4.6 Different Prompting Strategies
  • 4.7 LoTTE and BEIR Test Results
  • 4.8 Additional Results
  • 5 Discussion & Future Work
  • 6 Conclusion
  • 7 Limitations
  • References
  • Appendix
  • A Reranker Configurations
  • B Fine-tuning Rerankers and Retriever

Knowls

  1. Knowl 1 — UDAPDR adapts a retriever using synthetic-query-trained rerankers

    model/method

    UDAPDR adapts a neural retriever to a target domain using in-domain passages, without requiring target-domain queries or relevance labels. Its six-stage workflow, also depicted in the page-1 diagram, is: (1) sample XX target-domain passages and use GPT-3 text-davinci-002 with five query-generation prompts to produce 5X5X initial synthetic queries; (2) use these examples to construct YY corpus-adapted prompts, each containing passage examples paired with good and bad queries; (3) give each prompt to Flan-T5 XXL to generate ZZ additional synthetic queries per prompt, treating the passage that prompted each query as its positive passage; (4) train one passage reranker for each prompt's query set; (5) use the rerankers to label further synthetic questions and distill their judgments into one ColBERTv2 retriever; and (6) deploy the retriever in the target domain. A synthetic query is retained only if zero-shot ColBERTv2 retrieves its associated passage among its top 20 results. The experiments explored XX values of 5, 10, 50, and 100; YY values of 1, 5, and 10; and ZZ values of 1,000, 10,000, 100,000, and 1,000,000. The primary multi-reranker configuration used five prompts and 20,000 queries per prompt.

  2. Knowl 2 — Reranker training and multi-teacher distillation implementation

    experimental setup

    UDAPDR's passage teachers are DeBERTaV3-Large cross-encoders trained separately on synthetic query–passage examples generated from their corresponding corpus-adapted prompts. Each reranker uses the final hidden state of the [CLS] token followed by a linear classification layer, cross-entropy loss, Adam, 0.1 dropout on Transformer outputs, one training epoch, learning rate 5×10−65\times10^{-6} with linear warmup and decay, and batch size 32. For distillation, the trained rerankers label additional synthetic questions to create training triples; the teachers' judgments are distilled into a single ColBERTv2 retriever, so the rerankers are not needed at inference. The student uses a BERT-Base encoder, learning rate 1×10−51\times10^{-5}, batch size 32, and maximum document length of 300 tokens.

  3. Knowl 3 — Multi-reranker distillation improves dev-set retrieval across domains

    data/table

    On the dev sets, UDAPDR improves Success@5 over zero-shot ColBERTv2 on all reported LoTTE Forum, Natural Questions, and SQuAD tasks. The page-6 main-results table compares the zero-shot retriever, distilled students with one, five, or ten teachers, and zero-shot ColBERTv2 paired with a non-distilled reranker. The three distilled settings use 100,000 synthetic queries total, divided evenly among teachers; the comparison reranker also uses 100,000 queries. Scores are listed in this order: zero-shot ColBERTv2; one teacher trained on 100,000 queries; five teachers trained on 20,000 each; ten teachers trained on 10,000 each; zero-shot ColBERTv2 plus reranker.

    LoTTE Lifestyle: 64.5, 73.0, 74.8, 74.4, 73.5. LoTTE Technology: 44.5, 50.2, 51.3, 51.1, 50.6. LoTTE Writing: 80.0, 84.3, 85.7, 86.2, 85.5. LoTTE Recreation: 70.8, 76.9, 80.4, 79.8, 79.1. LoTTE Science: 61.5, 65.6, 67.9, 68.0, 67.2. LoTTE Pooled: 63.7, 70.0, 72.1, 72.2, 71.1. Natural Questions: 68.9, 72.4, 73.7, 74.0, 73.9. SQuAD: 65.0, 71.8, 73.8, 73.6, 72.6.

    The five- and ten-teacher students match or exceed the reranker-assisted baseline on several tasks while requiring no reranker at inference. All reported scores are dev-set Success@5.

  4. Knowl 4 — UDAPDR improves LoTTE test-set Success@5

    empirical result

    With five rerankers trained on 20,000 distinct synthetic queries each, then distilled into one ColBERTv2 retriever, UDAPDR improves test-set Success@5 over zero-shot ColBERTv2 in every LoTTE Forum and Search category. The page-8 results report the following zero-shot-to-UDAPDR scores: Forum—Lifestyle 76.2 to 84.9, Technology 54.0 to 59.9, Writing 75.8 to 83.2, Recreation 69.8 to 78.6, Science 45.6 to 48.8, and Pooled 62.3 to 70.8; Search—Lifestyle 82.4 to 86.8, Technology 65.9 to 67.7, Writing 80.4 to 84.3, Recreation 73.2 to 77.9, Science 57.5 to 61.0, and Pooled 71.5 to 76.6. The authors report average improvements of 7.1 points for Forum and 3.9 points for Search.

  5. Knowl 5 — UDAPDR improves BEIR zero-shot nDCG@10 on all evaluated datasets

    data/table

    On the BEIR test benchmark, the page-8 results show UDAPDR outperforming zero-shot ColBERTv2 on all 11 listed datasets. Scores below are zero-shot ColBERTv2 nDCG@10 followed by UDAPDR nDCG@10: ArguAna 46.1 and 57.5; Touché 26.3 and 32.4; TREC-COVID 84.7 and 88.0; NFCorpus 33.8 and 34.1; HotpotQA 70.3 and 75.3; DBPedia 44.6 and 47.4; Climate-FEVER 27.1 and 33.7; FEVER 78.0 and 83.2; SciFact 66.0 and 72.2; SCIDOCS 15.4 and 17.8; FiQA 45.8 and 53.5. UDAPDR uses five rerankers, each trained on 20,000 distinct synthetic queries, and uses only the distilled ColBERTv2 retriever at inference. The authors report an average zero-shot accuracy increase of 5.2 points.

  6. Knowl 6 — Distillation retains reranker-level accuracy at ColBERTv2 latency

    empirical result

    For LoTTE Lifestyle, the page-6 latency comparison measures the full single-query retrieval process on one NVIDIA V100 GPU using PyTorch 1.13. Zero-shot ColBERTv2 achieves Success@5 64.5 at 35 ms; a ColBERTv2 student distilled from five rerankers trained on 20,000 queries each achieves 74.8 at the same 35 ms. By comparison, zero-shot ColBERTv2 reranking 20, 100, or 1,000 passages takes 412 ms, 2,060 ms, or 20,600 ms, with Success@5 of 73.3, 73.5, or 73.5, respectively. Thus, the distilled retriever exceeds the measured reranking configurations in accuracy while avoiding their much higher query latency.

  7. Knowl 7 — Corpus-adapted prompts and larger query generators improve adaptation

    empirical result

    LoTTE Pooled dev-set experiments show that query-generation choices affect adaptation quality. In the prompting comparison, Success@5 was 65.8 for the InPars prompt with GPT-3 and one reranker, 67.6 for InPars with Flan-T5 XXL and one reranker, 67.1 for InPars with Flan-T5 XXL and five rerankers, 67.4 for corpus-adapted prompts with GPT-3 plus Flan-T5 XXL and one reranker, and 71.1 for corpus-adapted prompts with those generators and five rerankers; zero-shot ColBERTv2 scored 63.7. The first configuration used 5,000 synthetic queries because of GPT-3 API cost; the others used 100,000 queries total. The 71.1 score is 3.5 points above the one-reranker InPars/Flan-T5 XXL result, though the compared configurations differ in both prompt strategy and reranker count.

    A separate model-configuration comparison trained one non-distilled reranker on 100,000 synthetic queries per setting. Success@5 was 71.1 for GPT-3 plus Flan-T5 XXL with DeBERTaV3-Large; 66.7 for GPT-3 plus Flan-T5 XL with DeBERTaV3-Large; 68.0 for Flan-T5 XXL alone with DeBERTaV3-Large; 65.9 for Flan-T5 XL alone with DeBERTaV3-Large; 67.0 for GPT-3 plus Flan-T5 XXL with DeBERTaV3-Base; and 64.1 for GPT-3 plus Flan-T5 XL with DeBERTaV3-Base. Zero-shot ColBERTv2 scored 63.7. These results indicate that GPT-3 is not necessary to exceed the baseline, while replacing Flan-T5 XXL with the substantially smaller XL model or replacing DeBERTaV3-Large with Base lowers performance.

  8. Knowl 8 — Directly fine-tuning ColBERTv2 on synthetic queries was less reliable than distillation

    empirical result

    The authors also fine-tuned ColBERTv2 directly on synthetic query datasets instead of first training passage rerankers and distilling their judgments. On the LoTTE Forum dev set, direct fine-tuning produced only limited gains—at best 1–3 accuracy points—and sometimes decreased zero-shot retrieval accuracy. They report that distilling rerankers yielded more substantial gains and better adaptation across target domains, motivating the teacher-to-student route in UDAPDR.

  9. Knowl 9 — More distillation triples help, while selecting teachers offers limited added value

    empirical result

    Two sensitivity experiments examine the amount of distillation supervision and the use of teacher selection. With one DeBERTaV3-Large reranker trained on 10,000 synthetic queries, increasing the number of labeled distillation triples from 1,000 to 10,000 to 100,000 raised LoTTE Pooled dev Success@5 from 66.0 to 68.4 to 70.4, compared with 63.7 for zero-shot ColBERTv2; Natural Questions rose from 70.5 to 72.7 to 73.4, compared with 68.9; and SQuAD rose from 67.2 to 70.6 to 71.5, compared with 65.0. The gains also rose with triple count across the individual LoTTE categories reported in the experiment.

    In a separate teacher-count experiment, each reranker used 2,000 synthetic queries. On LoTTE Pooled, distilling the best 1, 5, or 10 rerankers selected from 50 achieved Success@5 of 66.3, 70.3, and 70.7, respectively; distilling all 5 rerankers trained in the unfiltered setting achieved 69.6. The latter uses 10 times fewer synthetic queries than training 50 rerankers and is only 0.7 points below selecting five from 50 on this pooled score. The authors report a 0.6-point average drop for the unfiltered approach and note that selecting teachers requires an annotated in-domain dev set, which the general method does not assume.

  10. Knowl 10 — UDAPDR requires substantial target-domain passages and has resource and coverage constraints

    limitation

    Although UDAPDR does not need target-domain queries or relevance labels, it does require a substantial collection of target-domain passages for synthetic-query generation; its effectiveness on extremely small passage collections remains untested. Generated queries may inherit language-model biases, and unknown pretraining data creates a risk of contamination. The authors specifically note possible overlap between the pretraining data of Flan-T5, GPT-3, and DeBERTa and evaluation material such as NQ, SQuAD, and Wikipedia passages, which could affect measured adaptation gains. The approach also benefits from costly GPU hardware and fast storage, and the experiments cover English-language retrieval only, leaving performance in low-resource languages unestablished.

Coverage note — Prompt-initialization examples and the individual prompt wordings were omitted because they are implementation detail rather than separate findings; other substantial contributed methods, evaluations, and stated limitations are represented.

References

  1. 1.Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling. 2020. Do not have enough data? deep learning to the rescue! In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7383–7390.
  2. 2.Akari Asai, Timo Schick, Patrick Lewis, Xilun Chen, Gautier Izacard, Sebastian Riedel, Hannaneh Hajishirzi, and Wen-tau Yih. 2022. Task-aware retrieval with instructions. arXiv preprint arXiv:2211.09260.
  3. 3.Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022. Inpars: Data augmentation for information retrieval using large language models. arXiv preprint arXiv:2202.05144.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Wei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar. 2020. Pre-training Tasks for Embedding-based Large-scale Retrieval. In International Conference on Learning Representations.
  6. 6.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  7. 7.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020. ELECTRA: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555.
  8. 8.Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B Hall, and Ming-Wei Chang. 2022. Promptagator: Few-shot dense retrieval from 8 examples. arXiv preprint arXiv:2209.11755.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Dheeru Dua, Emma Strubell, Sameer Singh, and Pat Verga. 2022. To adapt or to annotate: Challenges and interventions for domain adaptation in open-domain question answering. arXiv preprint arXiv:2212.10381.
  11. 11.Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. 2021. Splade v2: Sparse lexical and expansion model for information retrieval. arXiv preprint arXiv:2109.10086.
  12. 12.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International Conference on Machine Learning, pages 3929–3938. PMLR.
  13. 13.Christophe Van Gysel, Maarten de Rijke, and Evangelos Kanoulas. 2018. Neural Vector Spaces for Unsupervised Information Retrieval. ACM Trans. Inf. Syst., 36(4).
  14. 14.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. DeBERTaV3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. arXiv preprint arXiv:2111.09543.
  15. 15.Xuanli He, Islam Nassar, Jamie Kiros, Gholamreza Haffari, and Mohammad Norouzi. 2022. Generate, Annotate, and Learn: NLP with Synthetic Text. Transactions of the Association for Computational Linguistics, 10:826–842.
  16. 16.Sebastian Hofstätter, Sophia Althammer, Michael Schröder, Mete Sertkan, and Allan Hanbury. 2021. Improving efficient neural ranking models with cross-architecture knowledge distillation.
  17. 17.Jeremy Howard and Sebastian Ruder. 2018. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328–339, Melbourne, Australia. Association for Computational Linguistics.
  18. 18.Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston. 2019. Poly-encoders: Transformer architectures and pre-training strategies for fast and accurate multi-sentence scoring. arXiv preprint arXiv:1905.01969.
  19. 19.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning.
  20. 20.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299.
  21. 21.Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira. 2023. Inpars-v2: Large language models as efficient dataset generators for information retrieval. arXiv preprint arXiv:2301.01820.
  22. 22.Ehsan Kamalloo, Xinyu Zhang, Odunayo Ogundepo, Nandan Thakur, David Alfonso-Hermelo, Mehdi Rezagholizadeh, and Jimmy Lin. 2023. Evaluating embedding apis for information retrieval.
  23. 23.Omar Khattab, Christopher Potts, and Matei Zaharia. 2021. Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval. In Thirty-Fifth Conference on Neural Information Processing Systems.
  24. 24.Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive nlp. arXiv preprint arXiv:2212.14024.
  25. 25.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, China, July 25-30, 2020, pages 39–48. ACM.
  26. 26.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  27. 27.Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. Data Augmentation using Pre-trained Transformer Models. In Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems, pages 18–26, Suzhou, China. Association for Computational Linguistics.
  28. 28.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural Questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  29. 29.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086–6096, Florence, Italy. Association for Computational Linguistics.
  30. 30.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  31. 31.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  32. 32.Ji Ma, Ivan Korotkov, Yinfei Yang, Keith Hall, and Ryan McDonald. 2020. Zero-shot neural passage retrieval via domain-targeted synthetic question generation. arXiv preprint arXiv:2004.14503.
  33. 33.Rui Meng, Ye Liu, Semih Yavuz, Divyansh Agarwal, Lifu Tu, Ning Yu, Jianguo Zhang, Meghana Bhat, and Yingbo Zhou. 2022. Unsupervised dense retrieval deserves better positive pairs: Scalable augmentation with query extraction and generation. arXiv preprint arXiv:2212.08841.
  34. 34.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. In CoCo@ NIPs.
  35. 35.Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with BERT. arXiv preprint arXiv:1901.04085.
  36. 36.Rodrigo Nogueira, Jimmy Lin, and AI Epistemic. 2019. From doc2query to doctttttquery. Online preprint, 6.
  37. 37.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  38. 38.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2523–2544, Online. Association for Computational Linguistics.
  39. 39.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789, Melbourne, Australia. Association for Computational Linguistics.
  40. 40.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  41. 41.Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022a. PLAID: an efficient engine for late interaction retrieval. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, pages 1747–1756.
  42. 42.Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022b. ColBERTv2: Effective and efficient retrieval via lightweight late interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3715–3734, Seattle, United States. Association for Computational Linguistics.
  43. 43.Keshav Santhanam, Jon Saad-Falcon, Martin Franz, Omar Khattab, Avirup Sil, Radu Florian, Md Arafat Sultan, Salim Roukos, Matei Zaharia, and Christopher Potts. 2022c. Moving beyond downstream task accuracy for information retrieval benchmarking. arXiv preprint arXiv:2212.01340.
  44. 44.Nandan Thakur, Nils Reimers, Johannes Daxenberger, and Iryna Gurevych. 2020. Augmented SBERT: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring Tasks. arXiv preprint arXiv:2010.08240.
  45. 45.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, R. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 5998–6008. Curran Associates, Inc.
  47. 47.Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. TREC-COVID: constructing a pandemic information retrieval test collection. In ACM SIGIR Forum, volume 54, pages 1–12. ACM New York, NY, USA.
  48. 48.Kexin Wang, Nandan Thakur, Nils Reimers, and Iryna Gurevych. 2022. GPL: Generative pseudo labeling for unsupervised domain adaptation of dense retrieval. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2345–2360, Seattle, United States. Association for Computational Linguistics.
  49. 49.Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Michael Kinney, Yunyao Li, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Nancy Xin Ru Wang, Christopher Wilhelm, Boya Xie, Douglas M. Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier. 2020. CORD-19: The COVID-19 open research dataset. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020, Online. Association for Computational Linguistics.
  50. 50.Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021. MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2140–2151, Online. Association for Computational Linguistics.
  51. 51.Ji Xin, Chenyan Xiong, Ashwin Srinivasan, Ankita Sharma, Damien Jose, and Paul Bennett. 2022. Zero-shot dense retrieval with momentum adversarial domain invariant representations. In Findings of the Association for Computational Linguistics: ACL 2022, pages 4008–4020, Dublin, Ireland. Association for Computational Linguistics.
  52. 52.Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey. 2020. Generative data augmentation for commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1008–1025, Online. Association for Computational Linguistics.

Citation

MLA
Saad-Falcon, J., et al. “UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 11265–79, https://doi.org/10.18653/v1/2023.emnlp-main.693.
APA
Saad-Falcon, J., Khattab, O., Santhanam, K., Florian, R., Franz, M., Roukos, S., Sil, A., Sultan, M. A., & Potts, C. (2023). UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 11265–11279. https://doi.org/10.18653/v1/2023.emnlp-main.693
Chicago
Saad-Falcon, J., O. Khattab, K. Santhanam, et al. 2023. “UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 11265–79. https://doi.org/10.18653/v1/2023.emnlp-main.693.
Harvard
Saad-Falcon, J. et al. (2023) “UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 11265–11279. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.693.
Vancouver
1. Saad-Falcon J, Khattab O, Santhanam K, Florian R, Franz M, Roukos S, Sil A, Sultan MA, Potts C (2023) UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 11265–11279

BibTeX

@inproceedings{saad-falcon-etal-2023-udapdr,
    title = "{UDAPDR}: Unsupervised Domain Adaptation via {LLM} Prompting and Distillation of Rerankers",
    author = "Saad-Falcon, Jon  and
      Khattab, Omar  and
      Santhanam, Keshav  and
      Florian, Radu  and
      Franz, Martin  and
      Roukos, Salim  and
      Sil, Avirup  and
      Sultan, Md  and
      Potts, Christopher",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.693/",
    doi = "10.18653/v1/2023.emnlp-main.693",
    pages = "11265--11279"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/