Large Dual Encoders Are Generalizable Retrievers
Jianmo NiChen QuJing LuZhuyun DaiGustavo Hernández ÁbregoJi MaVincent Y. ZhaoYi LuanKeith B. HallMing-Wei Chang
Demonstrates that scaling dual encoder parameters up to billions while maintaining a fixed-size dot-product bottleneck significantly improves out-of-domain retrieval generalization across diverse benchmarks with remarkable data efficiency.
Dual encoder models are widely favored in neural information retrieval due to their conceptual simplicity and computational efficiency during search. However, a prevailing belief in the field holds that dual encoders generalize poorly to new domains because their comparison mechanism relies on a single, fixed-size mathematical product. Consequently, researchers have shifted toward computationally intensive multi-vector or late-interaction architectures. The article investigates whether scaling the parameter capacity of standard dual encoders, while keeping the output bottleneck fixed, can overcome these out-of-domain generalization limits.
The authors develop the Generalizable T5-based dense Retriever (GTR) series, testing model architectures ranging from 110 million to 4.8 billion parameters. They employ a multi-stage training pipeline comprising large-scale pre-training on two billion web-mined question-answer pairs, followed by fine-tuning on human-curated search data. The models maintain a fixed embedding dimension of 768. The framework is evaluated on in-domain search benchmarks and tested for zero-shot generalization across 18 distinct retrieval tasks covering nine diverse domains.
The findings show that increasing model capacity significantly and consistently improves zero-shot retrieval performance across domains. The largest model (GTR-XXL) outperforms prior sparse, dense, and late-interaction retrieval baselines, showing a substantial boost on the evaluation benchmark. Scaling also enables dual encoders to surpass traditional keyword-based matching on tasks where dense models historically lagged. Additionally, the analysis demonstrates high data efficiency: fine-tuning on only 10% of the curated search data matches or exceeds the zero-shot performance obtained from using the entire dataset. In contrast, expanding the embedding vector size yields diminishing returns compared to expanding model parameters.
These results demonstrate that the architectural bottleneck of a single dot-product does not inherently limit out-of-domain retrieval. For organizations building search and information retrieval systems, standard dual encoders can deliver state-of-the-art accuracy across diverse topics without requiring complex, multi-vector indexing pipelines. Furthermore, high data efficiency implies that organizations can substantially reduce human labeling costs when adapting dense retrievers to new applications.
Decision-makers should consider scaled dual encoders when high retrieval quality across variable domains is required, but they must balance performance gains against latency. Inference latency increases significantly with scale, rising from 17 milliseconds for the smallest model to 349 milliseconds for the largest. Where real-time speed is essential, teams should deploy intermediate model sizes or explore compression techniques such as model distillation, prompt tuning, and sparsification. Future development should also validate these methods across non-English languages and investigate variations in similarity functions for datasets with uneven document lengths.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). This benchmark establishes the heterogeneous zero-shot evaluation framework (BEIR) used by the source to measure out-of-domain retrieval generalization.
- Paper: Dense Passage Retrieval for Open-Domain Question Answering, Vladimir Karpukhin et al. (2020). This work introduces standard dual-encoder architectures for dense passage retrieval, establishing the baseline model paradigm that the source scales up.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). This paper presents dynamic hard negative mining for dense retrieval contrastive learning, establishing core training principles that the source adapts.
- Paper: Unsupervised Dense Information Retrieval with Contrastive Learning, Gautier Izacard et al. (2021). This study demonstrates unsupervised contrastive pretraining for dense retrievers across zero-shot benchmarks, which directly informs multi-stage dual encoder training.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). This foundational paper introduces masked-language-model pre-training specifically coupled with dense retrieval components for knowledge-intensive NLP.
- Paper: SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking, Thibault Formal et al. (2021). This work details learned sparse retrieval models that serve as a primary comparison point for dense dual encoders on MS MARCO and BEIR.
- Paper: ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction, Keshav Santhanam et al. (2022). This paper advances retrieval generalization beyond single-vector dual encoders by compressing token-level late interactions while maintaining superior out-of-domain transfer.
- Paper: Precise Zero-Shot Dense Retrieval without Relevance Labels, Luyu Gao et al. (2023). This work extends zero-shot dense retrieval by using generative models to synthesize hypothetical passages, bypassing the need for labeled supervised training data.
- Paper: M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation, Jianlv Chen et al. (2024). This research builds on dense embedding scaling by unifying dense, sparse, and multi-vector retrieval across multilingual and long-document contexts via self-knowledge distillation.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). This paper scales text embedding and reranking foundation models using massive synthetic pre-training pipelines to further push general-purpose retrieval performance.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). This work extends dense retrieval encoders by introducing in-context learning capabilities directly into the embedding model.
- Paper: How Does Generative Retrieval Scale to Millions of Passages?, Ronak Pradeep et al. (2023). This study evaluates scaling properties in the alternative paradigm of generative retrieval across millions of passages where dual encoders traditionally dominate.
