Transformer Memory as a Differentiable Search Index

Yi TayVinh TranMostafa DehghaniJianmo NiDara BahriHarsh MehtaZhen QinKai HuiZhe ZhaoJai Prakash Gupta

article2022NeurIPS400 citations

Proposes the Differentiable Search Index, demonstrating that a single Transformer can internalize an entire text corpus within its parameters to map queries directly to document identifiers, outperforming standard dual-encoder retrieval systems.

Listen

Modern information retrieval systems typically rely on multi-stage, retrieve-then-rank pipelines that combine external vector indexes or search algorithms with separate ranking models. While effective, managing distinct external indices and search modules introduces architectural complexity and limits full end-to-end learning. In response to these operational and technical bottlenecks, the article evaluates whether information retrieval can be fully parameterized within a single sequence-to-sequence neural model, termed the Differentiable Search Index (DSI).

The article's main objective is to demonstrate that a single pre-trained Transformer language model can memorize an entire corpus within its model weights and directly map user queries to document identifiers without needing an external index. To evaluate this approach, the authors tested various document representations, identifier structures, and training setups on the Natural Questions benchmark, scaling across corpus sizes ranging from 10,000 to 320,000 document pairs and model sizes up to 11 billion parameters.

The findings show that DSI consistently outperforms state-of-the-art dual encoder and traditional BM25 baselines. On the largest corpus evaluated (320,000 pairs), an 11-billion-parameter DSI model using semantically structured identifiers achieved a 40.4% top-1 retrieval accuracy, outperforming the dual encoder baseline by approximately 66% in relative terms. In zero-shot retrieval scenarios where the model saw no supervised query-document training pairs, DSI with atomic identifiers achieved a top-1 score of 25.1%, substantially exceeding standard keyword and unsupervised contrastive baselines. Furthermore, the analysis established that assigning semantically structured identifiers through hierarchical clustering works best for scaling, direct indexing of the first 32 tokens provides the optimal input representation, and co-training indexing and retrieval simultaneously in a multi-task setup is essential for stable optimization.

These results demonstrate that collapsing external indexes and retrieval algorithms directly into a single neural model can significantly improve search performance while simplifying system architectures. Importantly, unlike dual encoders whose accuracy quickly plateaued as model size grew, DSI displayed strong scaling characteristics, yielding marked gains as model capacity expanded. However, training these models requires substantial computational resources, with larger models requiring upwards of a full day of training across hundreds of specialized hardware accelerators.

Organizations evaluating this paradigm should treat it as an effective proof of concept for specialized, moderate-sized document collections before attempting enterprise-wide deployment. Practitioners are advised to adopt multi-task co-training and semantically structured identifiers if deploying DSI architectures. Before broad commercial deployment, further research is required to evaluate scaling behavior on multi-million-document corpora and develop mechanisms for dynamically adding, updating, and removing documents without requiring full model retraining.

arXiv: 2202.06991
Cover for Transformer Memory as a Differentiable Search Index

Abstract

In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search Index (DSI), a new paradigm that learns a text-to-text model that maps string queries directly to relevant docids; in other words, a DSI model answers queries directly using only its parameters, dramatically simplifying the whole retrieval process. We study variations in how documents and their identifiers are represented, variations in training procedures, and the interplay between models and corpus sizes. Experiments demonstrate that given appropriate design choices, DSI significantly outperforms strong baselines such as dual encoder models. Moreover, DSI demonstrates strong generalization capabilities, outperforming a BM25 baseline in a zero-shot setup.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Differentiable Search Index
  • 3.1 Indexing Strategies
  • 3.1.1 Indexing Method
  • 3.1.2 Document Representation Strategies
  • 3.2 Representing Docids for Retrieval
  • 3.3 Training and Optimization
  • 4 Experiments
  • 4.1 Baselines
  • 4.2 Experimental Results
  • 5 Conclusion
  • 6 Acknowledgements
  • References
  • 7 Appendix
  • 7.1 Dataset Statistics
  • 7.2 Extended Results
  • 7.2.1 Indexing/Memorization Performance
  • 7.2.2 Discussion of DSI Training Dynamics

Knowls

  1. Knowl 1 — Differentiable Search Index Paradigm

    model/method

    The Differentiable Search Index (DSI) is an information retrieval framework that parameterizes the entire indexing and retrieval pipeline within a single sequence-to-sequence Transformer model, storing all corpus knowledge within the model parameters rather than in external data structures such as inverted indexes or dense vector indexes with Maximum Inner Product Search (MIPS).

    DSI operates in two modes:

    1. Indexing: The model learns to associate the text tokens of each document djd_j with a document identifier (docid) jj via sequence-to-sequence learning.
    2. Retrieval: Given a natural language query qq, the model directly generates a ranked list of relevant docids jj autoregressively using beam search or by ranking output softmax logits.

    This framework enables fully differentiable end-to-end training and eliminates the separation between index creation and model inference.

  2. Knowl 2 — Semantically Structured Docid Generation via Hierarchical Clustering

    algorithm

    Semantically structured docids assign identifiers to documents such that semantically similar documents share common identifier prefixes, organizing the search space as a decimal trie.

    The procedure takes document embeddings X1:N={X1,…,XN}X_{1:N} = \{X_1, \dots, X_N\} with Xi∈RdX_i \in \mathbb{R}^d (obtained from an encoder model such as an 8-layer BERT) and a maximum cluster threshold cc (e.g., c=100c = 100). It recursively partitions documents into 10 clusters using kk-means clustering until each leaf cluster contains cc or fewer documents. For leaf clusters, documents are assigned arbitrary integer suffixes from 00 to ∣C∣−1|C| - 1.

    Input: Document embeddings X1:NX_{1:N}, where Xi∈RdX_i \in \mathbb{R}^d, cluster capacity threshold cc
    Output: Corresponding docid strings J1:NJ_{1:N}
    function GenerateSemanticIDs(X1:NX_{1:N})
        C1:10←Cluster(X1:N,k=10)C_{1:10} \leftarrow \text{Cluster}(X_{1:N}, k = 10)
        J←empty listJ \leftarrow \text{empty list}
        for i=0i = 0 to 99 do
            Jcurrent←[i]×∣Ci+1∣J_{\text{current}} \leftarrow [i] \times |C_{i+1}|
            if ∣Ci+1∣>c|C_{i+1}| > c then
                Jrest←GenerateSemanticIDs(Ci+1)J_{\text{rest}} \leftarrow \text{GenerateSemanticIDs}(C_{i+1})
            else
                Jrest←[0,1,…,∣Ci+1∣−1]J_{\text{rest}} \leftarrow [0, 1, \dots, |C_{i+1}| - 1]
            end if
            Jcluster←elementwiseStrConcat(Jcurrent,Jrest)J_{\text{cluster}} \leftarrow \text{elementwiseStrConcat}(J_{\text{current}}, J_{\text{rest}})
            J←J.appendElements(Jcluster)J \leftarrow J.\text{appendElements}(J_{\text{cluster}})
        end for
        J←reorderToOriginal(J,X1:N,C1:10)J \leftarrow \text{reorderToOriginal}(J, X_{1:N}, C_{1:10})
        return JJ
    end function
  3. Knowl 3 — Unstructured Atomic Docid Output Layer Formulation

    equation

    When docids are represented as unstructured atomic identifiers, each document in the corpus is treated as a unique token in an extended vocabulary. The output probability distribution OO over all vocabulary tokens and docids is computed by extending the final linear projection layer of the Transformer decoder stack:

    O=Softmax([Wtokens;Wdocs]Thlast)O = \text{Softmax}\left([W_{\text{tokens}}; W_{\text{docs}}]^T h_{\text{last}}\right)

    where:

    • [;][;] denotes the row-wise concatenation operator.
    • hlast∈Rdmodelh_{\text{last}} \in \mathbb{R}^{d_{\text{model}}} is the final hidden state from the last layer of the Transformer decoder stack.
    • Wtokens∈Rdmodel×∣Ntokens∣W_{\text{tokens}} \in \mathbb{R}^{d_{\text{model}} \times |N_{\text{tokens}}|} is the weight matrix for standard text vocabulary tokens.
    • Wdocs∈Rdmodel×∣Ndocuments∣W_{\text{docs}} \in \mathbb{R}^{d_{\text{model}} \times |N_{\text{documents}}|} is the learned weight matrix corresponding to the corpus docids, where ∣Ndocuments∣|N_{\text{documents}}| is the number of documents in the index.

    Top-kk retrieval under this representation is obtained by sorting the logits across all document entries in OO given the query encoding.

  4. Knowl 4 — Indexing Tasks and Document Representations in DSI

    model/method

    Differentiable Search Index models evaluate multiple training tasks and document representations to memorize document-to-identifier mappings:

    Indexing Task Formulations:

    • Inputs2Target: Doc tokens are inputs and the docid is the target (doc_tokens →\to docid). This aligns the indexing task length balance with retrieval.
    • Targets2Inputs: Docid is the input and document tokens are the target (docid →\to doc_tokens), modeling conditional generation of document content.
    • Bidirectional: Jointly trains Inputs2Target and Targets2Inputs using task prefixes.
    • Span Corruption: Concatenates docid tokens to document tokens and applies masked span denoising across both.

    Document Representation Strategies:

    • Direct Indexing: Uses the first LL contiguous tokens of a document in their original sequential order.
    • Set Indexing: Deduplicates document words into a unique word set and removes stopwords before passing tokens to the model.
    • Inverted Index: Randomly samples contiguous kk-token chunks from anywhere in the document to pair with the docid.
  5. Knowl 5 — Multi-Task Optimization and Indexing-to-Retrieval Ratio for DSI

    model/method

    DSI models are optimized using sequence-to-sequence cross-entropy loss with teacher forcing. Training indexing (document →\to docid) and retrieval (query →\to docid) jointly via multi-task co-training with task prefixes outperforms sequential training (indexing memorization followed by retrieval fine-tuning).

    Because the retrieval task is functionally dependent on the indexing task (docids are ungrounded tokens without indexing), the sampling ratio rr of indexing examples to retrieval examples is critical during co-training. An extreme ratio (too high or too low) degrades performance, and a ratio of r=32r = 32 indexing examples per retrieval example achieves the strongest empirical performance on retrieval tasks.

  6. Knowl 6 — Supervised Document Retrieval Performance of DSI on Natural Questions

    data/table

    Across Natural Questions sub-corpora of varying sizes (NQ10K, NQ100K, and full NQ320K), DSI models consistently outperform BM25 and T5-based Dual Encoders (DE). Semantic String Docids achieve the highest performance among docid representations in supervised retrieval settings.

    Model Size Method NQ10K NQ100K NQ320K
    Hits@1 Hits@10 Hits@1 Hits@10 Hits@1 Hits@10
    BM25 – – 12.4 33.5 20.9 46.4 11.6 34.4
    T5 Base (220M) Dual Encoder 16.2 48.6 18.7 55.2 20.5 58.3
    T5 Large (800M) Dual Encoder 18.8 55.7 22.3 60.5 22.4 63.3
    T5 XL (3B) Dual Encoder 20.8 59.6 23.3 63.2 23.9 65.8
    T5 XXL (11B) Dual Encoder 22.1 61.6 24.1 64.5 24.3 67.3
    DSI Base (250M) Atomic Docid 13.0 38.4 23.8 58.6 20.7 40.9
    DSI Large (800M) Atomic Docid 31.3 59.4 17.1 52.3 11.6 37.6
    DSI XL (3B) Atomic Docid 40.1 76.9 19.0 55.3 28.1 61.9
    DSI XXL (11B) Atomic Docid 39.4 77.0 25.3 67.9 24.0 55.1
    DSI Base (250M) Naive String Docid 28.1 48.0 18.7 44.6 6.7 21.0
    DSI Large (800M) Naive String Docid 34.7 60.5 21.2 50.7 13.3 33.6
    DSI XL (3B) Naive String Docid 44.7 66.4 24.0 55.1 16.7 58.1
    DSI XXL (11B) Naive String Docid 46.7 77.9 27.5 62.4 23.8 55.9
    DSI Base (250M) Semantic String Docid 33.9 57.3 19.0 44.9 27.4 56.6
    DSI Large (800M) Semantic String Docid 37.5 65.1 20.4 50.2 35.6 62.6
    DSI XL (3B) Semantic String Docid 41.9 67.1 22.4 52.2 39.1 66.8
    DSI XXL (11B) Semantic String Docid 48.5 72.1 26.9 59.5 40.4 70.3

    On NQ320K, DSI XXL with Semantic String Docids achieves 40.4% Hits@1 compared to 24.3% for T5 XXL Dual Encoder, representing a 66% relative gain.

  7. Knowl 7 — Zero-Shot Document Retrieval with DSI Models

    data/table

    In a zero-shot retrieval setting where the model is trained purely on the indexing task (doc_tokens →\to docid) without any labeled query-document pairs, DSI models outperform unsupervised dense baselines (SentenceT5 and raw T5) and the lexical BM25 baseline.

    Model Size Method NQ10K NQ100K NQ320K
    Hits@1 Hits@10 Hits@1 Hits@10 Hits@1 Hits@10
    BM25 – – 12.4 33.5 20.9 46.4 11.6 34.4
    T5 XXL Dual Encoder 0.3 1.3 1.9 8.0 1.1 5.9
    SentenceT5 Large Dual Encoder 17.6 50.7 17.4 50.8 16.9 51.0
    DSI XXL Atomic Docid 25.7 60.1 23.0 57.3 25.1 56.6
    DSI XXL Naive String Docid 43.4 67.4 17.4 41.5 9.2 22.6
    DSI XXL Semantic String Docid 43.9 68.8 11.4 26.6 13.9 31.1

    Unlike in supervised retrieval, Unstructured Atomic Docids achieve the best zero-shot generalization on larger corpora (NQ100K and NQ320K), scoring 25.1% Hits@1 on NQ320K versus 11.6% for BM25 and 16.9% for SentenceT5 Large.

  8. Knowl 8 — Parameter Scaling Dynamics in DSI vs. Dual Encoders

    empirical result

    As model capacity increases from Base (220M/250M parameters), Large (800M), XL (3B), to XXL (11B), DSI shows continuous, steep performance improvements on retrieval (Hits@1), whereas Dual Encoders experience diminishing returns and plateau across parameter scales.

    On the NQ320K corpus, scaling DSI with Semantic String Docids from Base to XXL raises Hits@1 from 27.4% to 40.4% (+13.0 points), whereas scaling T5 Dual Encoders from Base to XXL only improves Hits@1 from 20.5% to 24.3% (+3.8 points).

  9. Knowl 9 — Ablation of Indexing Strategies and Document Lengths in DSI

    empirical result

    Empirical evaluations on indexing task configurations and document representations reveal the following properties:

    1. Indexing Task: Inputs2Target (doc_tokens →\to docid) performs best (13.5% Hits@1 on NQ100K with Naive Docids), closely matched by Bidirectional training (13.2% Hits@1). Targets2Inputs (docid →\to doc_tokens) and Span Corruption with docids fail entirely, yielding 0% retrieval accuracy. Training without indexing examples also yields 0% Hits@1.
    2. Document Representation: Direct indexing (preserving sequential word order) outperforms set indexing and inverted indexing. Applying stopword filtering or set deduplication yields no performance improvement.
    3. Document Length: For direct indexing on NQ320K, shorter document prefixes yield superior retrieval accuracy: L=32L = 32 achieves ≈20%\approx 20\% Hits@1, whereas performance dips significantly at L=64L = 64 and further drops at L=128L = 128 (<15%< 15\% Hits@1), indicating that parameter-based memorization degrades when exposed to longer document token sequences.
  10. Knowl 10 — Scalability and Dynamic Updating Limitations of DSI

    limitation

    The Differentiable Search Index architecture has several inherent limitations:

    1. Corpus Scale: Evaluations are restricted to moderate-sized corpora (up to 320k documents). Scaling to web-scale collections is bounded by Transformer parameter memory capacity and output vocabulary/decoding constraints.
    2. Dynamic Corpus Updating: In DSI, document additions, deletions, or modifications require continuous model updates or retraining, presenting challenges regarding catastrophic forgetting and index consistency compared to traditional inverted indexes.
    3. Optimization Instability: Unstructured atomic docids with large softmax vocabularies exhibit training instability, intermittent non-convergence, and high variance across corpus splits.

Coverage note — Hardware and environment details (TPUv3/TPUv4 allocations, JAX/T5X setup), standard hyperparameter tuning bounds, and dataset partition sizes from the appendix were omitted as standard non-contributed implementation details.

References

  1. 1.Samira Abnar, Mostafa Dehghani, Behnam Neyshabur, and Hanie Sedghi. Exploring the limits of large scale pre-training. arXiv preprint arXiv:2110.02095, 2021.
  2. 2.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. arXiv preprint arXiv:2112.04426, 2021.
  3. 3.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  4. 4.Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. Autoregressive entity retrieval. arXiv preprint arXiv:2010.00904, 2020.
  5. 5.Mostafa Dehghani, Hamed Zamani, Aliaksei Severyn, Jaap Kamps, and W Bruce Croft. Neural ranking models with weak supervision. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 65–74, 2017.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  7. 7.Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al. Glam: Efficient scaling of language models with mixture-of-experts. arXiv preprint arXiv:2112.06905, 2021.
  8. 8.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:2101.03961, 2021.
  9. 9.Tianyu Gao, Xingcheng Yao, and Danqi Chen. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.552. URL https://aclanthology.org/2021.emnlp-main.552.
  10. 10.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. arXiv preprint arXiv:2012.14913, 2020.
  11. 11.Daniel Gillick, Alessandro Presta, and Gaurav Singh Tomar. End-to-end retrieval in continuous space. arXiv preprint arXiv:1811.08008, 2018.
  12. 12.Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. Accelerating large-scale inference with anisotropic vector quantization. In International Conference on Machine Learning, 2020. URL https://arxiv.org/abs/1908.10396.
  13. 13.Kelvin Guu, Kenton Lee, Zora Tung, and Panupong Pasupat. REALM: Retrieval-Augmented Language Model Pre-Training. In Proceedings of ICML 2020, 2020.
  14. 14.Martin Josifoski, Nicola De Cao, Maxime Peyrard, and Robert West. Genie: Generative information extraction. arXiv preprint arXiv:2112.08340, 2021.
  15. 15.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  16. 16.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906, 2020.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin Kenton Lee, Kristina Toutanova, Llion Jones Matthew Kelcey, Ming-Wei Chang, Andrew M Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural Questions: a Benchmark for Question Answering Research. In Transactions of the ACL, 2019.
  18. 18.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.
  19. 19.Donald Metzler, Yi Tay, Dara Bahri, and Marc Najork. Rethinking search: making domain experts out of dilettantes. In ACM SIGIR Forum, volume 55, pages 1–27. ACM New York, NY, USA, 2021.
  20. 20.Jianmo Ni, Gustavo Hernández Ábrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. arXiv preprint arXiv:2108.08877, 2021.
  21. 21.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019.
  22. 22.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  23. 23.Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlovic, Geir Kjetil Sandve, et al. Hopfield networks is all you need. arXiv preprint arXiv:2008.02217, 2020.
  24. 24.Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020.
  25. 25.Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt. Test-time training with self-supervision for generalization under distribution shifts. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 9229–9248. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.press/v119/sun20b.html.
  26. 26.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. arXiv preprint arXiv:1409.3215, 2014.
  27. 27.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. Scale efficiently: Insights from pre-training and fine-tuning transformers. arXiv preprint arXiv:2109.10686, 2021.
  28. 28.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  29. 29.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.

Citation

MLA
Tay, Y., et al. “Transformer Memory as a Differentiable Search Index”. arXiv, 2022, http://arxiv.org/abs/2202.06991v3.
APA
Tay, Y., Tran, V. Q., Dehghani, M., Ni, J., Bahri, D., Mehta, H., Qin, Z., Hui, K., Zhao, Z., Gupta, J., Schuster, T., Cohen, W. W., & Metzler, D. (2022). Transformer Memory as a Differentiable Search Index. arXiv. http://arxiv.org/abs/2202.06991v3
Chicago
Tay, Y., V. Q. Tran, M. Dehghani, et al. 2022. “Transformer Memory as a Differentiable Search Index”. arXiv. http://arxiv.org/abs/2202.06991v3.
Harvard
Tay, Y. et al. (2022) “Transformer Memory as a Differentiable Search Index”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2202.06991v3.
Vancouver
1. Tay Y, Tran VQ, Dehghani M, et al (2022) Transformer Memory as a Differentiable Search Index. arXiv

BibTeX

@article{tay2022transformer,
  title = {Transformer Memory as a Differentiable Search Index},
  author = {Tay, Yi and Tran, Vinh Q. and Dehghani, Mostafa and Ni, Jianmo and Bahri, Dara and Mehta, Harsh and Qin, Zhen and Hui, Kai and Zhao, Zhe and Gupta, Jai and Schuster, Tal and Cohen, William W. and Metzler, Donald},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2202.06991v3},
  eprint = {2202.06991}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission