Transformer Memory as a Differentiable Search Index
Yi TayVinh TranMostafa DehghaniJianmo NiDara BahriHarsh MehtaZhen QinKai HuiZhe ZhaoJai Prakash Gupta
Proposes the Differentiable Search Index, demonstrating that a single Transformer can internalize an entire text corpus within its parameters to map queries directly to document identifiers, outperforming standard dual-encoder retrieval systems.
Modern information retrieval systems typically rely on multi-stage, retrieve-then-rank pipelines that combine external vector indexes or search algorithms with separate ranking models. While effective, managing distinct external indices and search modules introduces architectural complexity and limits full end-to-end learning. In response to these operational and technical bottlenecks, the article evaluates whether information retrieval can be fully parameterized within a single sequence-to-sequence neural model, termed the Differentiable Search Index (DSI).
The article's main objective is to demonstrate that a single pre-trained Transformer language model can memorize an entire corpus within its model weights and directly map user queries to document identifiers without needing an external index. To evaluate this approach, the authors tested various document representations, identifier structures, and training setups on the Natural Questions benchmark, scaling across corpus sizes ranging from 10,000 to 320,000 document pairs and model sizes up to 11 billion parameters.
The findings show that DSI consistently outperforms state-of-the-art dual encoder and traditional BM25 baselines. On the largest corpus evaluated (320,000 pairs), an 11-billion-parameter DSI model using semantically structured identifiers achieved a 40.4% top-1 retrieval accuracy, outperforming the dual encoder baseline by approximately 66% in relative terms. In zero-shot retrieval scenarios where the model saw no supervised query-document training pairs, DSI with atomic identifiers achieved a top-1 score of 25.1%, substantially exceeding standard keyword and unsupervised contrastive baselines. Furthermore, the analysis established that assigning semantically structured identifiers through hierarchical clustering works best for scaling, direct indexing of the first 32 tokens provides the optimal input representation, and co-training indexing and retrieval simultaneously in a multi-task setup is essential for stable optimization.
These results demonstrate that collapsing external indexes and retrieval algorithms directly into a single neural model can significantly improve search performance while simplifying system architectures. Importantly, unlike dual encoders whose accuracy quickly plateaued as model size grew, DSI displayed strong scaling characteristics, yielding marked gains as model capacity expanded. However, training these models requires substantial computational resources, with larger models requiring upwards of a full day of training across hundreds of specialized hardware accelerators.
Organizations evaluating this paradigm should treat it as an effective proof of concept for specialized, moderate-sized document collections before attempting enterprise-wide deployment. Practitioners are advised to adopt multi-task co-training and semantically structured identifiers if deploying DSI architectures. Before broad commercial deployment, further research is required to evaluate scaling behavior on multi-million-document corpora and develop mechanisms for dynamically adding, updating, and removing documents without requiring full model retraining.
- Paper: Dense Passage Retrieval for Open-Domain Question Answering, Vladimir Karpukhin et al. (2020). Introduces the dual-encoder dense passage retrieval paradigm that the Differentiable Search Index directly challenges and outperforms as an end-to-end parametric alternative.
- Paper: The Probabilistic Relevance Framework: BM25 and Beyond, Stephen Robertson et al. (2009). Details the classic BM25 lexical ranking framework that serves as the foundational baseline against which neural generative indexing methods are measured.
- Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). Pioneers the use of Transformer architectures for passage scoring, establishing the multi-stage retrieve-and-rank baseline that DSI consolidates into a single model.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). Explores augmenting neural language models with corpus retrieval, framing the dichotomy between parametric memorization and external index-based retrieval.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). Demonstrates sequence-to-sequence conditioning for question answering over retrieved texts, providing foundational insights for generative retrieval setups.
- Paper: BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, Nandan Thakur et al. (2021). Establishes standard zero-shot and out-of-distribution evaluation benchmarks across diverse information retrieval architectures.
- Paper: ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT, Omar Khattab et al. (2020). Presents late-interaction multi-vector passage search, representing the state of the art in decoupled neural retrieval systems.
- Paper: Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval, Lee Xiong et al. (2021). Introduces approximate nearest neighbor negative sampling for training dense neural retrievers, contextualizing dual-encoder optimization limits.
- Paper: How Does Generative Retrieval Scale to Millions of Passages?, Ronak Pradeep et al. (2023). Extends the generative retrieval paradigm established by DSI to multi-million passage corpora to evaluate scaling behaviors and architectural bottlenecks.
- Paper: Learning to Rank in Generative Retrieval, Yongqi Li et al. (2024). Improves generative retrieval by integrating explicit learning-to-rank objectives into the sequence-to-sequence document identifier generation process.
- Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale, Siddharth Gollapudi et al. (2026). Investigates in-context retrieval where language models generate document identifiers directly from long contexts rather than memorized parameters.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). Examines the fundamental reliability and limitations of parametric memory in language models versus non-parametric retrieval augmentation.
- Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). Provides formal theoretical bounds and geometric limits of embedding-based retrieval that motivate alternative paradigms like generative retrieval.
