Large Dual Encoders Are Generalizable Retrievers

Jianmo NiChen QuJing LuZhuyun DaiGustavo Hernández ÁbregoJi MaVincent Y. ZhaoYi LuanKeith B. HallMing-Wei Chang

article2022EMNLP677 citations

Demonstrates that scaling dual encoder parameters up to billions while maintaining a fixed-size dot-product bottleneck significantly improves out-of-domain retrieval generalization across diverse benchmarks with remarkable data efficiency.

Listen

Dual encoder models are widely favored in neural information retrieval due to their conceptual simplicity and computational efficiency during search. However, a prevailing belief in the field holds that dual encoders generalize poorly to new domains because their comparison mechanism relies on a single, fixed-size mathematical product. Consequently, researchers have shifted toward computationally intensive multi-vector or late-interaction architectures. The article investigates whether scaling the parameter capacity of standard dual encoders, while keeping the output bottleneck fixed, can overcome these out-of-domain generalization limits.

The authors develop the Generalizable T5-based dense Retriever (GTR) series, testing model architectures ranging from 110 million to 4.8 billion parameters. They employ a multi-stage training pipeline comprising large-scale pre-training on two billion web-mined question-answer pairs, followed by fine-tuning on human-curated search data. The models maintain a fixed embedding dimension of 768. The framework is evaluated on in-domain search benchmarks and tested for zero-shot generalization across 18 distinct retrieval tasks covering nine diverse domains.

The findings show that increasing model capacity significantly and consistently improves zero-shot retrieval performance across domains. The largest model (GTR-XXL) outperforms prior sparse, dense, and late-interaction retrieval baselines, showing a substantial boost on the evaluation benchmark. Scaling also enables dual encoders to surpass traditional keyword-based matching on tasks where dense models historically lagged. Additionally, the analysis demonstrates high data efficiency: fine-tuning on only 10% of the curated search data matches or exceeds the zero-shot performance obtained from using the entire dataset. In contrast, expanding the embedding vector size yields diminishing returns compared to expanding model parameters.

These results demonstrate that the architectural bottleneck of a single dot-product does not inherently limit out-of-domain retrieval. For organizations building search and information retrieval systems, standard dual encoders can deliver state-of-the-art accuracy across diverse topics without requiring complex, multi-vector indexing pipelines. Furthermore, high data efficiency implies that organizations can substantially reduce human labeling costs when adapting dense retrievers to new applications.

Decision-makers should consider scaled dual encoders when high retrieval quality across variable domains is required, but they must balance performance gains against latency. Inference latency increases significantly with scale, rising from 17 milliseconds for the smallest model to 349 milliseconds for the largest. Where real-time speed is essential, teams should deploy intermediate model sizes or explore compression techniques such as model distillation, prompt tuning, and sparsification. Future development should also validate these methods across non-English languages and investigate variations in similarity functions for datasets with uneven document lengths.

arXiv: 2112.07899
Cover for Large Dual Encoders Are Generalizable Retrievers

Abstract

It has been shown that dual encoders trained on one domain often fail to generalize to other domains for retrieval tasks. One widespread belief is that the bottleneck layer of a dual encoder, where the final score is simply a dot-product between a query vector and a passage vector, is too limited compared to models with fine-grained interactions between the query and the passage. In this paper, we challenge this belief by scaling up the size of the dual encoder model while keeping the bottleneck layer as a single dot-product with a fixed size. With multi-stage training, scaling up the model size brings significant improvement on a variety of retrieval tasks, especially for out-of-domain generalization. We further analyze the impact of the bottleneck layer and demonstrate diminishing improvement when scaling up the embedding size. Experimental results show that our dual encoders, Generalizable T5-based dense Retrievers (GTR), outperform previous sparse and dense retrievers on the BEIR dataset (Thakur et al., 2021) significantly. Most surprisingly, our ablation study finds that GTR is very data efficient, as it only needs 10% of MS Marco supervised data to match the out-of-domain performance of using all supervised data.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Dual Encoder and dense retrieval
  • 2.2 BEIR generalization task
  • 3 Generalizable T5 Retriever
  • 3.1 T5 dual encoder
  • 3.2 Multi-stage training
  • 4 Experimental setup
  • 4.1 Training Data
  • 4.2 Configurations
  • 4.3 Models for comparison
  • 5 Evaluation Results
  • 5.1 Results on MS Marco
  • 5.2 Results on BEIR generalization tasks
  • 5.3 Data efficiency for large retrievers
  • 6 Ablation Study and Analysis
  • 6.1 Scaling up in different training stages
  • 6.2 Importance of the fine-tuning dataset
  • 6.3 Different pre-training strategies
  • 6.4 Document length vs. model capacity
  • 6.5 Scaling up with different bottleneck size
  • 7 Related Work
  • 8 Inference latency
  • 9 Conclusion
  • 10 Limitations
  • Acknowledgments
  • References
  • A More results
  • A.1 Comparisons on MS Marco
  • A.2 Recall on BEIR

Knowls

  1. Knowl 1 — Generalizable T5 Retriever Architecture and Formulation

    model/method

    The Generalizable T5 Retriever (GTR) is a dense dual-encoder information retrieval model built from the encoder component of pre-trained T5 language models, omitting the autoregressive decoder. The query encoder and the document (passage) encoder share identical weights.

    Given an input query qq and a candidate passage pp, both text sequences are tokenized using the SentencePiece vocabulary model and encoded by the shared T5 encoder. Text representations are obtained by computing the mean-pooling over the sequence of token output representations from the final layer of the encoder. The resulting query representation eq∈RD\mathbf{e}_q \in \mathbb{R}^D and passage representation ep∈RD\mathbf{e}_p \in \mathbb{R}^D are fixed to an embedding dimension of D=768D = 768 across all model parameter scales.

    The retrieval score s(q,p)s(q, p) is computed as the cosine similarity (or inner product of L2L_2-normalized vectors) between the single dense query and passage vectors:

    s(q,p)=eq⊤ep∥eq∥2∥ep∥2s(q, p) = \frac{\mathbf{e}_q^\top \mathbf{e}_p}{\|\mathbf{e}_q\|_2 \|\mathbf{e}_p\|_2}

    GTR models are instantiated in four parameter configurations based on T5 checkpoints: GTR-Base (110 million parameters), GTR-Large (335 million parameters), GTR-XL (1.24 billion parameters), and GTR-XXL (4.8 billion parameters).

  2. Knowl 2 — Bi-directional In-Batch Sampled Softmax Loss with Hard Negatives

    equation

    GTR models are trained using a bi-directional sampled softmax cross-entropy loss over mini-batches that incorporates both in-batch negative passages and explicitly mined hard negative passages.

    Let B={(qi,pi+)}i=1∣B∣\mathcal{B} = \{(q_i, p_i^+)\}_{i=1}^{|\mathcal{B}|} denote a mini-batch of paired training instances where qiq_i is a query and pi+p_i^+ is its corresponding ground-truth relevant passage. Let pj−p_j^- denote an additional hard negative passage paired with query qjq_j. The query-to-passage contrastive loss for query qiq_i is given by:

    Lq→p(qi,pi+)=−log⁡exp⁡(sim(eqi,epi+)/τ)∑j∈B(exp⁡(sim(eqi,epj+)/τ)+exp⁡(sim(eqi,epj−)/τ))\mathcal{L}_{q \to p}(q_i, p_i^+) = - \log \frac{\exp\left(\text{sim}(\mathbf{e}_{q_i}, \mathbf{e}_{p_i^+}) / \tau\right)}{\sum_{j \in \mathcal{B}} \left( \exp\left(\text{sim}(\mathbf{e}_{q_i}, \mathbf{e}_{p_j^+}) / \tau\right) + \exp\left(\text{sim}(\mathbf{e}_{q_i}, \mathbf{e}_{p_j^-}) / \tau\right) \right)}

    where sim(u,v)=u⊤v∥u∥2∥v∥2\text{sim}(\mathbf{u}, \mathbf{v}) = \frac{\mathbf{u}^\top \mathbf{v}}{\|\mathbf{u}\|_2 \|\mathbf{v}\|_2} is the cosine similarity function and τ>0\tau > 0 is the softmax temperature hyperparameter, fixed to τ=0.01\tau = 0.01.

    The total training loss L\mathcal{L} applies this objective symmetrically in both directions (query-to-passage and passage-to-query):

    L=12∣B∣∑i=1∣B∣(Lq→p(qi,pi+)+Lp→q(pi+,qi))\mathcal{L} = \frac{1}{2 |\mathcal{B}|} \sum_{i=1}^{|\mathcal{B}|} \left( \mathcal{L}_{q \to p}(q_i, p_i^+) + \mathcal{L}_{p \to q}(p_i^+, q_i) \right)

  3. Knowl 3 — Multi-Stage Training Framework for Dense Dual Encoders

    model/method

    GTR employs a two-stage training approach to combine large-scale generic pre-training with domain-specific search fine-tuning:

    1. Generic Pre-Training Stage: The dual encoder is initialized from pre-trained T5 encoder checkpoints and trained on a web-mined corpus comprising 2 billion semi-structured question-answer and input-response pairs sourced from online forums and question-answering websites (e.g., Reddit and StackOverflow). Pre-training runs for 800,000 optimization steps.

    2. Curated Fine-Tuning Stage: The pre-trained encoder is fine-tuned on human-annotated search corpora. By default, fine-tuning uses the MS MARCO dataset (532,000 query-passage pairs) augmented with hard negatives mined via RocketQA. Alternatively, Natural Questions (NQ, 130,000 query-passage pairs) with mined hard negatives is used to evaluate transfer under smaller training distributions. Fine-tuning runs for 20,000 optimization steps.

    All training stages use the Adafactor optimizer with an initial learning rate of 1×10−31 \times 10^{-3} and a linear learning rate decay schedule, a global batch size of 2,048, and a temperature of τ=0.01\tau = 0.01 on Cloud TPU-V3 accelerators. Input sequence lengths are set to 64 tokens for queries and 512 tokens for documents (extended to 768 tokens for documents in TREC-News and Robust04, and 512 tokens for queries in ArguAna).

  4. Knowl 4 — Zero-Shot Out-of-Domain Retrieval Generalization on BEIR

    data/table

    Across 18 diverse retrieval tasks spanning 9 domains in the BEIR benchmark, scaling the dual-encoder backbone model from Base (110M) to XXL (4.8B) parameters produces continuous gains in out-of-domain zero-shot retrieval effectiveness, outperforming both competitive sparse retrievers (BM25, docT5query) and prior dense retrievers (DPR, ANCE, TAS-B, ColBERT).

    Dataset BM25 docT5query DPR TAS-B ColBERT GTR-Base GTR-Large GTR-XXL
    MS MARCO 0.228 0.338 0.177 0.408 0.401 0.420 0.430 0.442
    TREC-COVID 0.656 0.713 0.332 0.481 0.677 0.539 0.557 0.501
    BioASQ 0.465 0.431 0.127 0.383 0.474 0.271 0.320 0.324
    NFCorpus 0.325 0.328 0.189 0.319 0.305 0.308 0.329 0.342
    NQ 0.329 0.399 0.474 0.463 0.524 0.495 0.547 0.568
    HotpotQA 0.603 0.580 0.391 0.584 0.593 0.535 0.579 0.599
    FiQA-2018 0.236 0.291 0.112 0.300 0.317 0.349 0.424 0.467
    Signal-1M 0.330 0.307 0.155 0.289 0.274 0.261 0.265 0.273
    TREC-News 0.398 0.420 0.161 0.377 0.393 0.337 0.343 0.346
    Robust04 0.408 0.437 0.252 0.427 0.391 0.437 0.470 0.506
    ArguAna 0.315 0.349 0.175 0.429 0.233 0.511 0.525 0.540
    Touché-2020 0.367 0.347 0.131 0.162 0.202 0.205 0.219 0.256
    Quora 0.789 0.802 0.248 0.835 0.854 0.881 0.890 0.892
    DBPedia-entity 0.313 0.331 0.263 0.384 0.392 0.347 0.391 0.408
    SCIDOCS 0.158 0.162 0.077 0.149 0.145 0.149 0.158 0.161
    FEVER 0.753 0.714 0.562 0.700 0.771 0.660 0.712 0.740
    Climate-FEVER 0.213 0.201 0.148 0.228 0.184 0.241 0.262 0.267
    SciFact 0.665 0.675 0.318 0.643 0.671 0.600 0.639 0.662
    CQADupStack 0.299 0.325 0.153 0.314 0.350 0.357 0.384 0.399
    Avg (All) 0.413 0.429 0.234 0.414 0.429 0.416 0.444 0.457
    Avg (w/o MS MARCO) 0.423 0.434 0.237 0.415 0.431 0.416 0.445 0.458

    GTR-Large (NDCG@10 = 0.445 w/o MS MARCO) outperforms the prior best dense model (TAS-B, 0.415) and sparse model (docT5query, 0.434). GTR-XXL achieves the highest overall zero-shot retrieval NDCG@10 (0.458), demonstrating that single-dot-product dual encoders generalize effectively when model capacity is scaled.

  5. Knowl 5 — In-Domain Ranking Performance on MS MARCO

    data/table

    When evaluated on the MS MARCO passage ranking development set, scaling GTR dual-encoder model capacity yields monotonic improvements across all in-domain retrieval metrics (NDCG@10, MRR@10, and Recall@1000), outperforming existing sparse and dense retrievers.

    Model NDCG@10 MRR@10 Recall@1000
    ANCE 0.388 0.330 0.959
    TAS-Balanced 0.408 0.340 0.975
    ColBERT 0.401 0.360 0.968
    RocketQA — 0.370 0.979
    GTR-Base (110M) 0.420 0.366 0.983
    GTR-Large (335M) 0.430 0.379 0.991
    GTR-XL (1.24B) 0.439 0.385 0.989
    GTR-XXL (4.8B) 0.442 0.388 0.990

    GTR-XXL achieves an NDCG@10 of 0.442, an MRR@10 of 0.388, and a Recall@1000 of 0.990, surpassing RocketQA without utilizing any auxiliary data augmentation during MS MARCO fine-tuning.

  6. Knowl 6 — Data Efficiency of Large Dual Encoders Under Reduced Supervision

    empirical result

    Large dual encoders exhibit high data efficiency during supervised fine-tuning. When the amount of MS MARCO fine-tuning data is subsampled to only 10% of the training queries (retaining their positive and hard-negative passages):

    1. In-Domain Performance: As expected, NDCG@10 on MS MARCO drops when training data is reduced (e.g., GTR-Large decreases from 0.430 to 0.428; GTR-XL decreases from 0.439 to 0.426; GTR-XXL decreases from 0.442 to 0.379).

    2. Out-of-Domain Zero-Shot Performance: On the BEIR benchmark (average NDCG@10 excluding MS MARCO), fine-tuning with only 10% of MS MARCO matches or slightly improves generalization compared to using 100% of the dataset:

      • GTR-Large: 0.452 (10% data) vs. 0.445 (100% data)
      • GTR-XL: 0.462 (10% data) vs. 0.453 (100% data)
      • GTR-XXL: 0.465 (10% data) vs. 0.458 (100% data)

    This indicates that extensive supervised query data is unnecessary for effective out-of-domain transfer when pre-trained large dual encoders are fine-tuned.

  7. Knowl 7 — Ablation of Training Stages: Pre-Training vs. Fine-Tuning Across Scales

    data/table

    An ablation study comparing models trained with fine-tuning only (GTR-FT, trained only on MS MARCO), pre-training only (GTR-PT, trained only on Community QA), and the full multi-stage pipeline (GTR) reveals how scaling interacts with training stages.

    Training Setting Base (110M) Large (335M) XL (1.24B) XXL (4.8B)
    MS MARCO In-Domain NDCG@10
    GTR-FT (Fine-Tuning Only) 0.400 0.415 0.418 0.422
    GTR-PT (Pre-Training Only) N/A N/A N/A N/A
    GTR (Pre-Training + Fine-Tuning) 0.420 0.430 0.439 0.442
    BEIR Zero-Shot Average NDCG@10 (w/o MS MARCO)
    GTR-FT (Fine-Tuning Only) 0.387 0.412 0.433 0.430
    GTR-PT (Pre-Training Only) 0.295 0.315 0.315 0.332
    GTR (Pre-Training + Fine-Tuning) 0.416 0.445 0.453 0.458

    GTR-FT XL (0.433) outperforms the previous best dual-encoder baseline (TAS-B, 0.415) without any pre-training stage. However, pre-training provides a consistent additive improvement (+0.025 to +0.033 zero-shot NDCG@10) across all model sizes when combined with fine-tuning.

  8. Knowl 8 — Model Capacity Scaling vs. Bottleneck Dimension Scaling

    empirical result

    When evaluating dual encoders on out-of-domain retrieval tasks across different bottleneck embedding dimensions D∈{256,768,2048,4096}D \in \{256, 768, 2048, 4096\} and model sizes (Base, Large, XL, XXL) fine-tuned on MS MARCO without pre-training across six BEIR datasets (BioASQ, DBPedia-entity, NQ, HotpotQA, FEVER, SciFact):

    1. Encoder Capacity Scaling: At every evaluated bottleneck dimensionality, increasing the model capacity from Base (110M) to XXL (4.8B) produces substantial, monotonic increases in average out-of-domain NDCG@10 (from ≈0.39\approx 0.39 to ≈0.48\approx 0.48).

    2. Bottleneck Dimension Scaling: For a fixed model capacity, expanding the bottleneck embedding dimension beyond D=768D = 768 (i.e., to D=2048D = 2048 or D=4096D = 4096) yields negligible performance gains.

    Scaling the parameter capacity of the backbone encoder is significantly more effective for generalizable retrieval than expanding the bottleneck vector dimensionality.

  9. Knowl 9 — Model Scaling Compensates for Domain-Restricted Fine-Tuning Data

    empirical result

    When GTR models are fine-tuned on Natural Questions (NQ, 130K pairs restricted to Wikipedia) rather than MS MARCO (532K pairs spanning diverse web domains):

    • GTR-Base fine-tuned on NQ achieves a zero-shot average BEIR NDCG@10 of 0.360, significantly outperforming the baseline DPR model (NDCG@10 = 0.237), which also uses NQ.
    • Scaling GTR fine-tuned on NQ increases zero-shot NDCG@10 from 0.360 (Base) to 0.379 (Large) and 0.407 (XL).
    • The performance improvement from scaling from Large to XL on NQ (+0.028+0.028) is larger than the gain observed when fine-tuning on MS MARCO (+0.008+0.008), indicating that model scaling has a stronger mitigating effect when training on smaller or domain-constrained supervision.
  10. Knowl 10 — Effect of Dual Encoder Scaling on Retrieved Document Length Distributions

    empirical result

    While dual encoders trained with cosine similarity have historically favored shorter documents, scaling dual-encoder capacity alters the median word length of top-10 retrieved documents across multiple BEIR datasets:

    • On DBPedia-entity, FEVER, HotpotQA, Signal-1M, TREC-News, and Web-Touché-2020, scaling from Base to XXL leads to an increase in the median length of retrieved documents.
    • On Web-Touché-2020, the median retrieved length doubles from 32 words for GTR-Base to 62 words for GTR-XXL, aligning with the benchmark's ground-truth relevant documents and improving NDCG@10 from 0.205 to 0.256.
    • Conversely, on TREC-COVID, GTR-XXL retrieves substantially shorter documents (median 9 words) compared to GTR-Base (median 15 words), which correlates with a performance drop for GTR-XXL on that specific task.
  11. Knowl 11 — Inference Latency Profile of Scaled Dual Encoders

    empirical result

    Measuring single-query encoding latency on Cloud TPU-V3 with a batch size of 1 and an input token length of 128 shows the computational trade-off associated with parameter scaling in GTR:

    • GTR-Base (110M parameters): 17 ms per query (comparable to TAS-B).
    • GTR-Large (335M parameters): 34 ms per query.
    • GTR-XL (1.24B parameters): 96 ms per query.
    • GTR-XXL (4.8B parameters): 349 ms per query (comparable to cross-encoder re-ranking latency).

    While scaling to billions of parameters enables high zero-shot retrieval accuracy, it introduces substantial latency overhead during embedding generation.

  12. Knowl 12 — Limitations of GTR Architecture and Evaluation Scope

    limitation

    The study and design of GTR have several specific limitations:

    1. Language Coverage: The models and evaluations are restricted exclusively to English-language corpora; multilingual and cross-lingual retrieval generalization are not evaluated.
    2. Absence of Knowledge Distillation: GTR relies solely on standard dual-encoder contrastive training and does not incorporate knowledge distillation from cross-encoders or multi-vector models, which is an established technique for boosting dense retrieval.
    3. Inference Latency: The largest model variant (GTR-XXL, 4.8B parameters) incurs an encoding latency of 349 ms per query, posing challenges for real-time serving without downstream model compression (such as sparsification, distillation, or quantization).

Coverage note — None was omitted; all key architectural components, training procedures, benchmark results on BEIR and MS MARCO, scaling ablations, document length analyses, and limitations are represented.

References

  1. 1.Amin Ahmad, Noah Constant, Yinfei Yang, and Daniel Cer. 2019. ReQA: An evaluation for end-to-end answer retrieval models. In Workshop on Machine Reading for Question Answering.
  2. 2.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL.
  3. 3.Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020. Language-agnostic bert sentence embedding.
  4. 4.Thibault Formal, C. Lassance, Benjamin Piwowarski, and Stéphane Clinchant. 2021. Splade v2: Sparse lexical and expansion model for information retrieval. ArXiv, abs/2109.10086.
  5. 5.D. Gillick, A. Presta, and Gaurav Singh Tomar. 2018. End-to-end retrieval in continuous space. ArXiv, abs/1811.08008.
  6. 6.Daniel Gillick, Sayali Kulkarni, Larry Lansing, Alessandro Presta, Jason Baldridge, Eugene Ie, and Diego Garcia-Olano. 2019. Learning dense representations for entity retrieval. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 528–537, Hong Kong, China. Association for Computational Linguistics.
  7. 7.Matthew Henderson, Rami Al-Rfou, B. Strope, Yun-Hsuan Sung, László Lukács, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and R. Kurzweil. 2017. Efficient natural language response suggestion for smart reply. ArXiv, abs/1705.00652.
  8. 8.Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021. Efficiently teaching an effective dense retriever with balanced topic aware sampling. arXiv preprint arXiv:2104.06967.
  9. 9.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Towards unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118.
  10. 10.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2021. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 7(3):535–547.
  11. 11.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  12. 12.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, pages 39–48.
  13. 13.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  14. 14.Jing Lu, Gustavo Hernández Ábrego, Ji Ma, Jianmo Ni, and Yinfei Yang. 2021. Multi-stage training with improved negative contrast for neural passage retrieval. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6091–6103.
  15. 15.Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2020. Sparse, dense, and attentional representations for text retrieval. arXiv preprint arXiv:2005.00181.
  16. 16.Ji Ma, Ivan Korotkov, Yinfei Yang, Keith Hall, and Ryan McDonald. 2020. Zero-shot neural retrieval via domain-targeted synthetic query generation. arXiv e-prints, pages arXiv–2004.
  17. 17.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas A. Tezak, Jong Wook Kim, Chris Hallacy, Johannes Heidecke, Pranav Shyam, Boris Power, Tyna Eloundou Nekoul, Girish Sastry, Gretchen Krueger, David P. Schnurr, Felipe Petroski Such, Kenny Sai-Kin Hsu, Madeleine Thompson, Tabarak Khan, Toki Sherbakov, Joanne Jang, Peter Welinder, and Lilian Weng. 2022. Text and code embeddings by contrastive pre-training. ArXiv, abs/2201.10005.
  18. 18.Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human generated machine reading comprehension dataset.
  19. 19.Jianmo Ni, Gustavo Hernández Ábrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang. 2021. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. arXiv preprint arXiv:2108.08877.
  20. 20.Chen Qu, Liu Yang, Cen Chen, Minghui Qiu, W. Bruce Croft, and Mohit Iyyer. 2020. Open-retrieval conversational question answering. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval.
  21. 21.Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. Rocketqa: An optimized training approach to dense passage retrieval for open-domain question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5835–5847.
  22. 22.Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, W. Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21/140.
  23. 23.Stephen Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: Bm25 and beyond. Found. Trends Inf. Retr., 3(4):333–389.
  24. 24.Devendra Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant, Wei Ping, William L. Hamilton, and Bryan Catanzaro. 2021. End-to-end training of neural retrievers for open-domain question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6648–6662, Online. Association for Computational Linguistics.
  25. 25.Keshav Santhanam, O. Khattab, Jon Saad-Falcon, Christopher Potts, and Matei A. Zaharia. 2021. Colbertv2: Effective and efficient retrieval via lightweight late interaction. ArXiv, abs/2112.01488.
  26. 26.Noam Shazeer and Mitchell Stern. 2018. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR.
  27. 27.Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  28. 28.Dan Vanderkam, Rob Schonberger, Henry Rowley, and Sanjiv Kumar. 2013. Nearest neighbor search in google correlate. Technical report, Google.
  29. 29.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  30. 30.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate nearest neighbor negative contrastive learning for dense text retrieval. arXiv preprint arXiv:2007.00808.
  31. 31.Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2021. Byt5: Towards a token-free future with pre-trained byte-to-byte models. arXiv preprint arXiv:2105.13626.
  32. 32.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934.
  33. 33.Yinfei Yang, Gustavo Hernández Abrego, Steve Yuan, Mandy Guo, Qinlan Shen, Daniel Cer, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2019. Improving multilingual sentence embedding using bidirectional dual encoder with additive margin softmax. arXiv preprint arXiv:1902.08564.
  34. 34.Yinfei Yang, Daniel Matthew Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, G. Ábrego, Steve Yuan, C. Tar, Yun-Hsuan Sung, B. Strope, and R. Kurzweil. 2020. Multilingual universal sentence encoder for semantic retrieval. In ACL.
  35. 35.Wen-tau Yih, Kristina Toutanova, John C. Platt, and Christopher Meek. 2011. Learning discriminative projections for text similarity measures. In Proceedings of the Fifteenth Conference on Computational Natural Language Learning, pages 247–256, Portland, Oregon, USA. Association for Computational Linguistics.

Citation

MLA
Ni, J., et al. “Large Dual Encoders Are Generalizable Retrievers”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 9844–55, https://doi.org/10.18653/v1/2022.emnlp-main.669.
APA
Ni, J., Qu, C., Lu, J., Dai, Z., Abrego, G. H., Ma, J., Zhao, V., Luan, Y., Hall, K., Chang, M.-W., & Yang, Y. (2022). Large Dual Encoders Are Generalizable Retrievers. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 9844–9855. https://doi.org/10.18653/v1/2022.emnlp-main.669
Chicago
Ni, J., C. Qu, J. Lu, et al. 2022. “Large Dual Encoders Are Generalizable Retrievers”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 9844–55. https://doi.org/10.18653/v1/2022.emnlp-main.669.
Harvard
Ni, J. et al. (2022) “Large Dual Encoders Are Generalizable Retrievers”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9844–9855. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.669.
Vancouver
1. Ni J, Qu C, Lu J, et al (2022) Large Dual Encoders Are Generalizable Retrievers. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 9844–9855

BibTeX

@inproceedings{ni-etal-2022-large,
    title = "Large Dual Encoders Are Generalizable Retrievers",
    author = "Ni, Jianmo  and
      Qu, Chen  and
      Lu, Jing  and
      Dai, Zhuyun  and
      Hernandez Abrego, Gustavo  and
      Ma, Ji  and
      Zhao, Vincent  and
      Luan, Yi  and
      Hall, Keith  and
      Chang, Ming-Wei  and
      Yang, Yinfei",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.669/",
    doi = "10.18653/v1/2022.emnlp-main.669",
    pages = "9844--9855"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/