keyword
dual encoders
A dual encoder is a neural network architecture designed for information retrieval and semantic matching tasks that processes two separate inputs, such as a search query and a target document, using two distinct or parameter-sharing encoding networks. Each network independently transforms its input into a low-dimensional dense vector representation within a shared embedding space. The semantic similarity or relevance between the query and candidate items is then computed through simple mathematical operations, most commonly a dot product or cosine similarity. Because document embeddings can be precomputed and indexed offline, dual encoders enable highly scalable and computationally efficient retrieval over massive datasets using nearest neighbor search, contrasting with joint architectures that require simultaneous cross-attention over input pairs at query time.
2 items

Large Dual Encoders Are Generalizable Retrievers
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, Yinfei Yang
Why you should read this
Demonstrates that scaling dual encoder parameters up to billions while maintaining a fixed-size dot-product bottleneck significantly improves out-of-domain retrieval generalization across diverse benchmarks with remarkable data efficiency.
It has been shown that dual encoders trained on one domain often fail to generalize to other domains for retrieval tasks. One widespread belief is that the bottleneck layer of a dual encoder, where the final score is simply a dot-product between a query vector and a passage vector, is too limited compared to models with fine-grained interactions between the query and the passage. In this paper, we challenge this belief by scaling up the size of the dual encoder model while keeping the bottleneck layer as a single dot-product with a fixed size. With multi-stage training, scaling up the model size brings significant improvement on a variety of retrieval tasks, especially for out-of-domain generalization. We further analyze the impact of the bottleneck layer and demonstrate diminishing improvement when scaling up the embedding size. Experimental results show that our dual encoders, Generalizable T5-based dense Retrievers (GTR), outperform previous sparse and dense retrievers on the BEIR dataset (Thakur et al., 2021) significantly. Most surprisingly, our ablation study finds that GTR is very data efficient, as it only needs 10% of MS Marco supervised data to match the out-of-domain performance of using all supervised data.
Added
2026-09-28

Transformer Memory as a Differentiable Search Index
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Prakash Gupta, Tal Schuster, William W. Cohen, Donald Metzler
Why you should read this
Proposes the Differentiable Search Index, demonstrating that a single Transformer can internalize an entire text corpus within its parameters to map queries directly to document identifiers, outperforming standard dual-encoder retrieval systems.
In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search Index (DSI), a new paradigm that learns a text-to-text model that maps string queries directly to relevant docids; in other words, a DSI model answers queries directly using only its parameters, dramatically simplifying the whole retrieval process. We study variations in how documents and their identifiers are represented, variations in training procedures, and the interplay between models and corpus sizes. Experiments demonstrate that given appropriate design choices, DSI significantly outperforms strong baselines such as dual encoder models. Moreover, DSI demonstrates strong generalization capabilities, outperforming a BM25 baseline in a zero-shot setup.
Added
2026-09-26
