Retrieval with Learned Similarities

Bailu DingJiaqi Zhai

article2025WWW6 citations

Develops Mixture-of-Logits and an approximate top-k search algorithm to enable efficient retrieval with complex learned similarities, cutting latency by up to 66x while maintaining over 99% recall across recommendation and question answering tasks.

Listen

Large-scale internet applications, including recommendation engines, search systems, and natural language question answering, rely on fast initial retrieval to identify relevant candidates from catalogs containing millions or billions of items. Most standard systems, such as vector databases, employ simple dot-product similarity search to match queries with items. While efficient, dot products impose an expressive bottleneck that degrades retrieval accuracy. Recent advanced methods adopt learned similarity functions or generative indexing to boost relevance, but these models are computationally expensive, complex to index, and difficult to deploy under strict latency limits.

The article evaluates Mixture-of-Logits as a universal similarity framework and demonstrates how to achieve highly accurate, low-latency top-item retrieval across diverse applications. To improve practical training, the authors introduce a mutual information-based load balancing loss for conditional computations in Mixture-of-Logits. They also propose and benchmark exact two-pass and approximate two-stage search algorithms designed to operate efficiently on modern graphics processing units.

The research evaluates these methods across recommendation benchmarks containing up to 674,000 items and the Natural Questions benchmark of roughly 110,000 documents. The authors test sequential recommendation architectures, such as SASRec and HSTU, alongside language models finetuned from T5. They measure retrieval accuracy using Hit Rate and Mean Reciprocal Rank, and evaluate query latency on graphics processors across exact brute-force and approximate retrieval techniques.

The evaluation yielded several key findings. First, the article establishes mathematically that Mixture-of-Logits is a universal approximator capable of representing any complex similarity function. Second, adding Mixture-of-Logits to recommendation systems improved top-1 Hit Rate by an average of 29.1%, top-10 Hit Rate by 16.3%, and Mean Reciprocal Rank by 18.1% compared to standard dot-product baselines. Third, on the Natural Questions dataset, the approach achieved a top-100 Hit Rate of 97.0%, outperforming state-of-the-art generative and dense retrieval models. Fourth, the proposed approximate search methods achieved up to a 66-fold reduction in latency compared to exact brute-force search while preserving over 99% of the retrieval accuracy.

These results demonstrate that organizations do not need to choose between the high accuracy of complex learned similarity models and the fast execution of dot-product vector search. Mixture-of-Logits leverages the parallel compute power of modern accelerators to provide high throughput at latency budgets comparable to standard vector search. Furthermore, because the retrieval algorithms interface naturally with existing vector search operations, organizations can upgrade retrieval quality without completely redesigning their underlying database infrastructure.

Engineering and infrastructure teams should consider adopting the proposed Retrieval with Learned Similarities framework for accelerator-based search pipelines. Teams can select approximate search heuristics based on their specific workload needs: the average dot-product heuristic offers the lowest latency and is agnostic to the number of embedding components, while combined candidate generation provides the highest accuracy when recall requirements are exceptionally stringent. Before enterprise-scale rollout, organizations should conduct pilot testing on internal billion-scale datasets to validate memory overheads, indexing costs, and custom hardware kernel optimizations.

The findings are supported by consistent results across both structured recommendation datasets and unstructured language tasks. However, the study focuses on datasets of up to hundreds of thousands of items; exact performance at extreme multi-billion-item scales will depend on network bandwidth, distributed infrastructure, and hardware configurations.

Cover for Retrieval with Learned Similarities

Abstract

Retrieval plays a fundamental role in recommendation systems, search, and natural language processing (NLP) by efficiently finding relevant items from a large corpus given a query. Dot products have been widely used as the similarity function in such tasks, enabled by Maximum Inner Product Search (MIPS) algorithms for efficient retrieval. However, state-of-the-art retrieval algorithms have migrated to learned similarities. These advanced approaches encompass multiple query embeddings, complex neural networks, direct item ID decoding via beam search, and hybrid solutions. Unfortunately, we lack efficient solutions for retrieval in these state-of-the-art setups. Our work addresses this gap by investigating efficient retrieval techniques with expressive learned similarity functions. We establish Mixture-of-Logits (MoL) as a universal approximator of similarity functions, demonstrate that MoL's expressiveness can be realized empirically to achieve superior performance on diverse retrieval scenarios, and propose techniques to retrieve the approximate top-k results using MoL with tight error bounds. Through extensive experimentation, we show that MoL, enhanced by our proposed mutual information-based load balancing loss, sets new state-of-the-art results across heterogeneous scenarios, including sequential retrieval models in recommendation systems and finetuning language models for question answering; and our approximate top-kk algorithms outperform baselines by up to 66x in latency while achieving >.99 recall rate compared to exact algorithms.

Table of Contents

  • 1 Introduction
  • 2 Mixture of Logits
  • 2.1 Expressiveness of Mixture of Logits
  • 2.2 Applying MoL to Heterogeneous Use Cases
  • 3 Retrieval Algorithms
  • 3.1 Exact algorithm
  • 3.2 Approximate algorithms
  • 4 Evaluation
  • 4.1 Workloads
  • 4.2 Quality of MoL-based Learned Similarity
  • 4.3 Top KK retrieval performance
  • 5 Related work
  • 6 Conclusion
  • References
  • A Experiment Setups
  • A.1 Reproducibility
  • A.2 Parameterization of low-rank (“component-level”) embeddings
  • A.2.1 Recommendation Systems
  • A.2.2 Question Answering (QA)
  • A.3 Parameterization of πp​(q,x)\pi_{p}(q,x) matrices
  • A.4 Hyperparameter settings
  • A.4.1 Recommendation Systems
  • A.4.2 Question Answering (QA)
  • B Examples for exact and approximate top KK retrieval algorithms

Knowls

  1. Knowl 1 — Universal Expressiveness of Mixture-of-Logits for Similarity Functions

    theoretical result

    Any similarity matrix A∈Rn×mA \in \mathbb{R}^{n \times m} (with n≤mn \le m) representing query-item relevance scores can be approximated to arbitrary precision by a Mixture-of-Logits (MoL) decomposition over low-rank matrices.

    Formally, for any ϵ>0\epsilon > 0, there exist PP component matrices B1,B2,…,BP∈Rn×mB_1, B_2, \dots, B_P \in \mathbb{R}^{n \times m} each having Rank(Bp)≤d\mathrm{Rank}(B_p) \le d (d≤nd \le n), and PP weight matrices π1,π2,…,πP∈[0,1]n×m\pi_1, \pi_2, \dots, \pi_P \in [0, 1]^{n \times m} satisfying ∑p=1Pπp(i,j)=1\sum_{p=1}^P \pi_p(i, j) = 1 for all (i,j)(i, j), such that:

    ∣A−∑p=1Pπp∘Bp∣<ϵ\left| A - \sum_{p=1}^P \pi_p \circ B_p \right| < \epsilon

    where ∘\circ denotes the Hadamard (element-wise) product. When Rank(A)=n\mathrm{Rank}(A) = n, an exact approximation with error ≤ϵ\le \epsilon is achievable using only P=2P = 2 low-rank matrices of rank dd. This proves that Mixture-of-Logits is a universal approximator for similarity functions, overcoming the low-rank bottleneck inherent in single dot-product retrieval models without requiring increased embedding dimensions.

  2. Knowl 2 — Mixture-of-Logits Formulation for High-Throughput Learned Retrieval

    model/method

    Mixture-of-Logits (MoL) computes the learned similarity score ϕ(q,x)\phi(q, x) between a query qq and an candidate item xx. The query and item are represented by PP pairs of low-rank component-level embeddings fp(q),gp(x)∈RdPf_p(q), g_p(x) \in \mathbb{R}^{d_P} for 1≤p≤P1 \le p \le P, where dPd_P is the component embedding dimension. MoL combines their dot products using adaptive, query- and item-dependent gating weights πp(q,x)∈[0,1]\pi_p(q, x) \in [0, 1]:

    ϕ(q,x)=∑p=1Pπp(q,x)⟨fp(q),gp(x)⟩\phi(q, x) = \sum_{p=1}^P \pi_p(q, x) \langle f_p(q), g_p(x) \rangle

    To enable high-throughput implementations on hardware accelerators like GPUs, the PP pairs are decomposed as an outer product of PqP_q query embeddings and PxP_x item embeddings (P=Pq×PxP = P_q \times P_x) with L2L_2-normalized vectors:

    ϕ(q,x)=∑pq=1Pq∑px=1Pxπpq,px(q,x)⟨fpq(q)∥fpq(q)∥2,gpx(x)∥gpx(x)∥2⟩\phi(q, x) = \sum_{p_q=1}^{P_q} \sum_{p_x=1}^{P_x} \pi_{p_q, p_x}(q, x) \left\langle \frac{f_{p_q}(q)}{\|f_{p_q}(q)\|_2}, \frac{g_{p_x}(x)}{\|g_{p_x}(x)\|_2} \right\rangle

    The gating weights πp(q,x)\pi_p(q, x) are parameterized by a two-layer Multi-Layer Perceptron (MLP) with SiLU activations, taking the PP inner products ⟨fp(q),gp(x)⟩\langle f_p(q), g_p(x) \rangle (and optionally user/item feature vectors) as input.

  3. Knowl 3 — Mutual Information-Based Load Balancing Loss for Mixture-of-Logits

    model/method

    In Mixture-of-Logits (MoL), the gating distribution πp(q,x)\pi_p(q, x) directs conditional computation across PP low-rank embedding pairs. To ensure even utilization of all PP components globally across training batches while keeping routing decisions sparse and confident for individual query-item pairs, a mutual information-based regularization loss LMI\mathcal{L}_{MI} is employed:

    LMI=−H(p)+H(p∣(q,x))\mathcal{L}_{MI} = -H(p) + H(p \mid (q, x))

    where H(p)H(p) is the marginal entropy over the component indices p∈{1,…,P}p \in \{1, \dots, P\} across training instances, and H(p∣(q,x))H(p \mid (q, x)) is the conditional entropy of component weights for an individual pair (q,x)(q, x).

    The overall model training objective combines sampled softmax loss with LMI\mathcal{L}_{MI}:

    L=−log⁡exp⁡(ϕ(q,x))exp⁡(ϕ(q,x))+∑x′∈Xnegexp⁡(ϕ(q,x′))+αLMI\mathcal{L} = -\log \frac{\exp(\phi(q, x))}{\exp(\phi(q, x)) + \sum_{x' \in \mathcal{X}_{\text{neg}}} \exp(\phi(q, x'))} + \alpha \mathcal{L}_{MI}

    where ϕ(q,x)\phi(q, x) is the MoL score, Xneg\mathcal{X}_{\text{neg}} is a set of sampled negative items, and α>0\alpha > 0 is a loss weight hyperparameter (set to α=0.001\alpha = 0.001 across all experimental configurations).

  4. Knowl 4 — Exact Two-Pass Top-K Retrieval Algorithm for Mixture-of-Logits

    algorithm

    To retrieve the exact top-KK items under the Mixture-of-Logits similarity function ϕ(q,x)=∑p=1Pπp(q,x)⟨fp(q),gp(x)⟩\phi(q, x) = \sum_{p=1}^P \pi_p(q, x) \langle f_p(q), g_p(x) \rangle without exhaustive brute-force evaluation over all items, a two-pass threshold-pruning algorithm is used. Because ϕ(q,x)\phi(q, x) is a convex combination of component dot products, max⁡1≤p≤P⟨fp(q),gp(x)⟩≥ϕ(q,x)\max_{1 \le p \le P} \langle f_p(q), g_p(x) \rangle \ge \phi(q, x) holds for every item.

    In Pass 1, top-KK dot-product searches are executed across each of the PP component embedding spaces to form a candidate set GG. The minimum MoL score among candidates in GG is computed as Smin⁡=min⁡x∈Gϕ(q,x)S_{\min} = \min_{x \in G} \phi(q, x). In Pass 2, range dot-product searches retrieve all items x∈Xx \in X satisfying ⟨fp(q),gp(x)⟩≥Smin⁡\langle f_p(q), g_p(x) \rangle \ge S_{\min} across any component pp into a set G′G'. Finally, exact top-KK scoring with MoL is evaluated over G′G'.

    Input: query qq, item set XX, embedding encoders fp(⋅)f_p(\cdot) and gp(⋅)g_p(\cdot) for component p∈{1,…,P}p \in \{1, \dots, P\}, target rank KK
    Output: Exact top-KK items by MoL similarity score ϕ(q,x)\phi(q, x)
    G←∅G \leftarrow \emptyset
    for p∈{1,…,P}p \in \{1, \dots, P\} do
        Xp←{gp(x)∣x∈X}X_p \leftarrow \{g_p(x) \mid x \in X\}
        G←G∪TopKDotProduct(fp(q),Xp,K)G \leftarrow G \cup \text{TopKDotProduct}(f_p(q), X_p, K)
    Smin⁡←∞S_{\min} \leftarrow \infty
    for x∈Gx \in G do
        s←ϕ(q,x)s \leftarrow \phi(q, x)
        if s<Smin⁡s < S_{\min} then
            Smin⁡←sS_{\min} \leftarrow s
    G′←∅G' \leftarrow \emptyset
    for p∈{1,…,P}p \in \{1, \dots, P\} do
        G′←G′∪RangeDotProduct(fp(q),Smin⁡,Xp)G' \leftarrow G' \cup \text{RangeDotProduct}(f_p(q), S_{\min}, X_p)
    return BruteForceTopKMol(q,G′,K)\text{BruteForceTopKMol}(q, G', K)
  5. Knowl 5 — Approximate Top-K Retrieval Algorithms for Mixture-of-Logits

    algorithm

    Approximate top-KK retrieval under Mixture-of-Logits (MoL) uses a two-stage paradigm: fast candidate selection via dot-product heuristics followed by full MoL re-ranking over the selected candidate set.

    Three candidate selection heuristics are defined:

    1. TopKPerEmbedding: Retrieves top-NN items by dot product ⟨fp(q),gp(x)⟩\langle f_p(q), g_p(x) \rangle independently for each of the PP embedding pairs and returns their deduplicated union (N×P≥KN \times P \ge K).
    2. TopKAvg: Computes the mean dot product 1P∑p=1P⟨fp(q),gp(x)⟩=1P⟨∑pq=1Pqfpq(q),∑px=1Pxgpx(x)⟩\frac{1}{P} \sum_{p=1}^P \langle f_p(q), g_p(x) \rangle = \frac{1}{P} \langle \sum_{p_q=1}^{P_q} f_{p_q}(q), \sum_{p_x=1}^{P_x} g_{p_x}(x) \rangle. Pre-materializing average item vectors 1Px∑px=1Pxgpx(x)\frac{1}{P_x} \sum_{p_x=1}^{P_x} g_{p_x}(x) reduces retrieval to a single vector search per query regardless of PP.
    3. CombinedTopK: Unions candidates from TopKPerEmbedding with candidate count N1N_1 and TopKAvg with candidate count N2N_2.
    Input: query qq, item set XX, target rank KK, candidate selection parameters
    Output: Approximate top-KK items by MoL similarity score ϕ(q,x)\phi(q, x)
    function TopKPerEmbedding(q,X,Nq, X, N):
        G←∅G \leftarrow \emptyset
        for p∈{1,…,P}p \in \{1, \dots, P\} do
            Xp←{gp(x)∣x∈X}X_p \leftarrow \{g_p(x) \mid x \in X\}
            G←G∪TopKDotProduct(fp(q),Xp,N)G \leftarrow G \cup \text{TopKDotProduct}(f_p(q), X_p, N)
        return GG
    function TopKAvg(q,X,Nq, X, N):
        q′←∑p=1Pfp(q)q' \leftarrow \sum_{p=1}^P f_p(q)
        X′←{1P∑p=1Pgp(x)∣x∈X}X' \leftarrow \{\frac{1}{P} \sum_{p=1}^P g_p(x) \mid x \in X\}
        return TopKDotProduct(q′,X′,N)\text{TopKDotProduct}(q', X', N)
    function CombTopK(q,X,N1,N2q, X, N_1, N_2):
        G1←TopKPerEmbedding(q,X,N1)G_1 \leftarrow \text{TopKPerEmbedding}(q, X, N_1)
        G2←TopKAvg(q,X,N2)G_2 \leftarrow \text{TopKAvg}(q, X, N_2)
        return G1∪G2G_1 \cup G_2
    function ApproxTopK(q,X,K,Gq, X, K, G):
        return BruteForceTopKMol(q,G,K)\text{BruteForceTopKMol}(q, G, K)
  6. Knowl 6 — Tight Error Bound for Per-Embedding Top-K Retrieval under Mixture-of-Logits

    theoretical result

    Let qq be a query, XX be the candidate item set, and PP be the number of component embedding pairs. For each embedding pair p∈{1,…,P}p \in \{1, \dots, P\}, let XK,pX_{K, p} denote the top-KK items retrieved by component dot product ⟨fp(q),gp(x)⟩\langle f_p(q), g_p(x) \rangle.

    Define SS as the maximum component dot product among all (K+1)(K+1)-th items across the PP embedding sets:

    S=max⁡{⟨fp(q),gp(x)⟩∣x∈XK+1,p∖XK,p,1≤p≤P}S = \max \{ \langle f_p(q), g_p(x) \rangle \mid x \in X_{K+1, p} \setminus X_{K, p}, 1 \le p \le P \}

    Let SKS_K be the KK-th largest MoL score among items in the candidate union set ⋃p=1PXK,p\bigcup_{p=1}^P X_{K, p}, and let S′S' be the actual KK-th largest MoL score across the entire set XX. The retrieval score gap SΔ=S′−SKS_\Delta = S' - S_K satisfies:

    SΔ≤S−SKS_\Delta \le S - S_K

    This bound is tight; equality SΔ=S−SKS_\Delta = S - S_K occurs when the gating weights assign πp(q,x)=1\pi_p(q, x) = 1 for the specific item xx and component pp that achieve SS. In addition, any item x∈⋃p=1PXK,px \in \bigcup_{p=1}^P X_{K, p} with MoL score ϕ(q,x)≥S\phi(q, x) \ge S is guaranteed to belong to the exact top-KK item set.

  7. Knowl 7 — Tight Error Bound for Combined Top-K Retrieval under Mixture-of-Logits

    theoretical result

    Let qq be a query, XX be the item set, and PP be the number of component embedding pairs. Let XK,pX_{K, p} denote the top-KK items retrieved by component dot product for embedding pair pp, and let XK′X'_K be the top-KK items retrieved by the average dot product 1P∑p=1P⟨fp(q),gp(x)⟩\frac{1}{P} \sum_{p=1}^P \langle f_p(q), g_p(x) \rangle.

    Let SKS_K be the KK-th largest MoL score evaluated over the combined candidate set (⋃p=1PXK,p)∪XK′\left( \bigcup_{p=1}^P X_{K, p} \right) \cup X'_K. Define S′S' as the maximum dot product over any component for items outside the combined set:

    S′=max⁡{⟨fp(q),gp(x)⟩∣x∈X∖((⋃p=1PXK,p)∪XK′),1≤p≤P}S' = \max \left\{ \langle f_p(q), g_p(x) \rangle \mid x \in X \setminus \left( (\bigcup_{p=1}^P X_{K, p}) \cup X'_K \right), 1 \le p \le P \right\}

    Then the retrieval gap SΔS_\Delta between the true KK-th largest MoL score and SKS_K satisfies:

    SΔ≤S′−SKS_\Delta \le S' - S_K

    This upper bound is tight and ensures that no item omitted from the combined candidate pool can have an MoL score exceeding S′S'.

  8. Knowl 8 — Special Aggregation Tokens and Parameterized Pooling for Language Model Retrieval

    model/method

    To adapt Mixture-of-Logits (MoL) to homogeneous text sequences (such as pre-trained language models like T5 or BERT for question answering) where heterogeneous feature tables are unavailable, two architectural mechanisms are used:

    1. Special Aggregation Tokens: The tokenizer vocabulary is expanded with PQP_Q query aggregation tokens {Q1,…,QPQ}\{Q_1, \dots, Q_{P_Q}\} and PXP_X item aggregation tokens {X1,…,XPX}\{X_1, \dots, X_{P_X}\}. These tokens are prepended to the input token sequences ([Q1,…,QPQ,SP1,…,SPN][Q_1, \dots, Q_{P_Q}, SP_1, \dots, SP_N]) prior to self-attention in bidirectional models (or appended in unidirectional models). They function as multiple specialized [CLS] tokens co-trained to aggregate distinct semantic facets of the sequence.

    2. Parameterized Pooling: A learned pooling layer is placed directly after the language model encoder. For each component position p∈{1,…,PQ}p \in \{1, \dots, P_Q\} (or PXP_X), the pooling layer takes the DD-dimensional representation of the first output token as an input conditioning vector to parameterize an example-level probability distribution over all token positions {0,…,max_seq_len−1}\{0, \dots, \text{max\_seq\_len} - 1\}. The resulting weighted sum of token representations produces the pp-th component embedding fp(q)∈RdPf_p(q) \in \mathbb{R}^{d_P} or gp(x)∈RdPg_p(x) \in \mathbb{R}^{d_P}.

  9. Knowl 9 — Empirical Performance of Mixture-of-Logits in Sequential Recommendation

    data/table

    Mixture-of-Logits (MoL) with the mutual information load balancing loss LMI\mathcal{L}_{MI} was evaluated against standard dot-product (cosine similarity) baselines across three sequential recommendation benchmark datasets: MovieLens-1M (ML-1M), MovieLens-20M (ML-20M), and Amazon Reviews Books. Two sequential backbones were used: SASRec and HSTU. Across all datasets and backbones, MoL consistently improved retrieval accuracy, averaging gains of +29.1% in HR@1, +16.3% in HR@10, and +18.1% in MRR over baseline dual encoders. Ablating LMI\mathcal{L}_{MI} led to consistent performance degradation across all datasets.

    Method HR@K MRR
    K=1 K=10 K=50 K=200
    ML-1M dataset
    SASRec .0610 .2818 .5470 .7540 .1352
    SASRec + MoL .0697 .3036 .5617 .7667 .1441
    HSTU .0750 .3332 .5956 .7824 .1579
    HSTU + MoL .0884 .3465 .6022 .7935 .1712
    HSTU + MoL abl. LMI\mathcal{L}_{MI} .0847 .3417 .6011 .7942 .1662
    ML-20M dataset
    SASRec .0653 .2883 .5484 .7658 .1375
    SASRec + MoL .0778 .3102 .5682 .7779 .1535
    HSTU .0962 .3557 .6146 .8080 .1800
    HSTU + MoL .1010 .3698 .6260 .8132 .1881
    HSTU + MoL abl. LMI\mathcal{L}_{MI} .0994 .3670 .6241 .8128 .1866
    Books dataset
    SASRec .0058 .0306 .0754 .1431 .0153
    SASRec + MoL .0095 .0429 .0915 .1635 .0212
    HSTU .0101 .0469 .1066 .1876 .0233
    HSTU + MoL .0156 .0631 .1308 .2173 .0324
    HSTU + MoL abl. LMI\mathcal{L}_{MI} .0153 .0625 .1286 .2172 .0321
  10. Knowl 10 — Question Answering Retrieval Performance on Natural Questions (NQ320K)

    data/table

    On the Natural Questions benchmark (NQ320K, 320k query-item pairs), MoL (finetuned from T5-base with PQ=4,PX=4,dP=768P_Q=4, P_X=4, d_P=768) was evaluated against sparse, dense, and generative retrieval methods. MoL with LMI\mathcal{L}_{MI} achieved state-of-the-art results across all metrics (.685 HR@1, .919 HR@10, .970 HR@100, .773 MRR), outperforming all generative retrieval baselines (including GenRet, NCI, DSI, and SEAL) and dense dual-encoder models (DPR, GTR-Base, Sentence-T5). Ablating LMI\mathcal{L}_{MI} caused drops across all evaluation metrics.

    Method HR@K MRR
    K=1 K=10 K=100
    Sparse retrieval
    BM25 .297 .603 .821 .402
    DocT5Query .380 .693 .861 .489
    Dense retrieval
    DPR .502 .777 .909 .599
    Sentence-T5 .536 .830 .938 .641
    GTR-Base .560 .844 .937 .662
    Generative retrieval
    GENRE .552 .673 .754 .599
    DSI .552 .674 .780 .596
    SEAL .570 .800 .914 .655
    DSI+QG .631 .807 .880 .695
    NCI .659 .852 .924 .731
    GenRet .681 .888 .952 .759
    Learned similarities
    MoL .685 .919 .970 .773
    MoL abl. LMI\mathcal{L}_{MI} .673 .919 .968 .767
  11. Knowl 11 — Efficiency and Recall Trade-offs of Approximate MoL Top-K Retrieval on GPU

    data/table

    Top-KK retrieval performance was evaluated on a single NVIDIA RTX 6000 Ada GPU with a batch size of 32 queries across ML-20M, Books, and NQ320K datasets. Hit rate (HR) was normalized relative to exact brute-force MoL scoring.

    TopKAvg achieved >0.99>0.99 relative hit rate across all KK while significantly reducing latency. On NQ320K, TopKAvg100 reached 1.00 relative HR across all metrics with a latency of 0.57 ms0.57\text{ ms}, representing a 66×66\times speedup over BruteForce (37.74 ms37.74\text{ ms}). On Books (674k674\text{k} items), TopKAvg500 maintained 1.17 ms1.17\text{ ms} latency compared to 128.36 ms128.36\text{ ms} for BruteForce (109×109\times speedup), demonstrating sub-linear latency scaling with dataset size due to precomputed average embeddings.

    Method HR@1 HR@5 HR@10 HR@50 HR@100 Latency / ms
    ML-20M
    BruteForce 1.00 1.00 1.00 1.00 1.00 2.730.01
    TopKPerEmbd5 0.69 0.67 0.63 0.46 0.40 1.610.05
    TopKPerEmbd10 0.95 0.88 0.82 0.65 0.57 1.650.06
    TopKPerEmbd50 1.00 1.00 0.99 0.95 0.92 1.640.05
    TopKPerEmbd100 1.00 1.00 1.00 0.99 0.98 2.310.02
    TopKAvg200 1.00 1.00 1.00 0.99 0.97 1.190.04
    TopKAvg500 1.00 1.00 1.00 1.00 1.00 1.220.05
    CombTopK5_200 1.00 1.00 1.00 1.00 1.00 1.860.06
    Books
    BruteForce 1.00 1.00 1.00 1.00 1.00 128.360.30
    TopKPerEmbd5 1.00 0.92 0.90 0.61 0.48 20.470.07
    TopKPerEmbd50 1.00 1.00 1.00 0.98 0.93 21.600.08
    TopKPerEmbd100 1.00 1.00 1.00 0.99 0.97 23.320.08
    TopKPerEmbd200 1.00 1.00 1.00 1.00 1.00 26.550.12
    TopKAvg200 0.97 0.92 0.91 0.76 0.67 1.130.04
    TopKAvg500 0.97 0.98 0.97 0.86 0.81 1.170.04
    TopKAvg1000 0.99 0.98 1.00 0.92 0.88 1.120.05
    TopKAvg2000 1.00 0.99 1.00 0.95 0.92 1.200.02
    TopKAvg4000 1.00 1.00 1.00 0.96 0.95 2.050.01
    TopKAvg8000 1.00 1.00 1.00 0.97 0.97 3.790.01
    CombTopK5_200 1.00 1.00 1.00 0.96 0.95 20.750.07
    CombTopK50_500 1.00 1.00 1.00 0.99 0.96 22.120.07
    CombTopK100_1000 1.00 1.00 1.00 0.99 0.98 24.020.13
    CombTopK200_2000 1.00 1.00 1.00 1.00 1.00 28.010.11
    NQ320K
    BruteForce 1.00 1.00 1.00 1.00 1.00 37.740.47
    TopKPerEmbd5 1.00 1.00 1.00 0.96 1.00 4.710.08
    TopKPerEmbd10 1.00 1.00 1.00 0.98 1.00 4.830.08
    TopKPerEmbd50 1.00 1.00 1.00 1.00 1.00 6.310.09
    TopKAvg100 1.00 1.00 1.00 1.00 1.00 0.570.05
    CombTopK5_100 1.00 1.00 1.00 1.00 1.00 5.280.08

Coverage note — None was omitted; all major theoretical bounds, algorithms, embedding parameterization methods, and empirical benchmark evaluations were converted into standalone knowls.

References

  1. 1.[n. d.]. ANN Benchmarks. https://ann-benchmarks.com/. Accessed: 2024-08-06.
  2. 2.Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup. 2016. Conditional Computation in Neural Networks for faster models. arXiv:1511.06297 [cs.LG] https://arxiv.org/abs/1511.06297
  3. 3.Jon Louis Bentley. 1975. Multidimensional binary search trees used for associative searching. Commun. ACM 18, 9 (sep 1975), 509–517. https://doi.org/10.1145/361002.361007
  4. 4.Michele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih, Sebastian Riedel, and Fabio Petroni. 2022. Autoregressive Search Engines: Generating Substrings as Document Identifiers. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 31668–31683. https://proceedings.neurips.cc/paper_files/paper/2022/file/cd88d62a2063fdaf7ce6f9068fb15dcd-Paper-Conference.pdf
  5. 5.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. 2022. Improving Language Models by Retrieving from Trillions of Tokens. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA (Proceedings of Machine Learning Research, Vol. 162), Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato (Eds.). PMLR, 2206–2240. https://proceedings.mlr.press/v162/borgeaud22a.html
  6. 6.Fedor Borisyuk, Qingquan Song, Mingzhou Zhou, Ganesh Parameswaran, Madhu Arun, Siva Popuri, Tugrul Bingol, Zhuotao Pei, Kuang-Hsuan Lee, Lu Zheng, Qizhan Shao, Ali Naqvi, Sen Zhou, and Aman Gupta. 2024. LiNR: Model Based Neural Retrieval on GPUs at LinkedIn. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (Boise, ID, USA) (CIKM ’24). Association for Computing Machinery, New York, NY, USA, 4366–4373. https://doi.org/10.1145/3627673.3680091
  7. 7.Haw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, and Andrew McCallum. 2023. Multi-CLS BERT: An Efficient Alternative to Traditional Ensembling. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 821–854. https://doi.org/10.18653/v1/2023.acl-long.48
  8. 8.Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik G. Learned-Miller, and Chuang Gan. 2023. Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task Learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 11828–11837.
  9. 9.Felix Chern, Blake Hechtman, Andy Davis, Ruiqi Guo, David Majnemer, and Sanjiv Kumar. 2022. TPU-KNN: K Nearest Neighbor Search at Peak FLOP/s. In Advances in Neural Information Processing Systems.
  10. 10.Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. In Proceedings of the 10th ACM Conference on Recommender Systems (RecSys ’16). 191–198.
  11. 11.Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2021. Autoregressive Entity Retrieval. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. https://openreview.net/forum?id=5k8F6UU39V
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, 4171–4186. https://doi.org/10.18653/v1/n19-1423
  13. 13.Chantat Eksombatchai, Pranav Jindal, Jerry Zitao Liu, Yuchen Liu, Rahul Sharma, Charles Sugnet, Mark Ulrich, and Jure Leskovec. 2018. Pixie: A System for Recommending 3+ Billion Items to 200+ Million Users in Real-Time. In Proceedings of the 2018 World Wide Web Conference (WWW ’18). 1775–1784.
  14. 14.Stefan Elfwing, Eiji Uchibe, and Kenji Doya. 2017. Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning. CoRR abs/1702.03118 (2017). arXiv:1702.03118 http://arxiv.org/abs/1702.03118
  15. 15.Weihao Gao, Xiangjun Fan, Chong Wang, Jiankai Sun, Kai Jia, Wenzi Xiao, Ruofan Ding, Xingyan Bin, Hui Yang, and Xiaobing Liu. 2021. Learning An End-to-End Structure for Retrieval in Large-Scale Recommendations. In Proceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM ’21). 524–533.
  16. 16.Daniel Gillick, Alessandro Presta, and Gaurav Singh Tomar. 2018. End-to-End Retrieval in Continuous Space. arXiv:1811.08008 [cs.IR]
  17. 17.Aristides Gionis, Piotr Indyk, and Rajeev Motwani. 1999. Similarity Search in High Dimensions via Hashing. In Proceedings of the 25th International Conference on Very Large Data Bases (VLDB ’99). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 518–529.
  18. 18.Ruiqi Guo, Sanjiv Kumar, Krzysztof Choromanski, and David Simcha. 2016. Quantization based Fast Inner Product Search. In Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, AISTATS 2016, Vol. 51. 482–490.
  19. 19.Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020. Accelerating large-scale inference with anisotropic vector quantization. In Proceedings of the 37th International Conference on Machine Learning (ICML’20). JMLR.org, Article 364, 10 pages.
  20. 20.F. Maxwell Harper and Joseph A. Konstan. 2015. The MovieLens Datasets: History and Context. ACM Trans. Interact. Intell. Syst. 5, 4, Article 19 (dec 2015), 19 pages. https://doi.org/10.1145/2827872
  21. 21.Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural Collaborative Filtering. In Proceedings of the 26th International Conference on World Wide Web (Perth, Australia) (WWW ’17). 173–182.
  22. 22.Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016. Session-based Recommendations with Recurrent Neural Networks. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http://arxiv.org/abs/1511.06939
  23. 23.Kalina Jasinska, Krzysztof Dembczynski, Robert Busa-Fekete, Karlson Pfannschmidt, Timo Klerx, and Eyke Hullermeier. 2016. Extreme F-measure Maximization using Sparse Probability Estimates. In Proceedings of The 33rd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 48), Maria Florina Balcan and Kilian Q. Weinberger (Eds.). PMLR, New York, New York, USA, 1435–1444. https://proceedings.mlr.press/v48/jasinska16.html
  24. 24.Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi. 2019. DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2019/file/09853c7fb1d3f8ee67a61b6bf4a7f8e6-Paper.pdf
  25. 25.Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Trans. Pattern Anal. Mach. Intell. 33, 1 (jan 2011), 117–128. https://doi.org/10.1109/TPAMI.2010.57
  26. 26.J. Johnson, M. Douze, and H. Jegou. 2021. Billion-Scale Similarity Search with GPUs. IEEE Transactions on Big Data 7, 03 (Jul 2021), 535–547.
  27. 27.Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In 2018 International Conference on Data Mining (ICDM). 197–206.
  28. 28.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 6769–6781. https://doi.org/10.18653/v1/2020.emnlp-main.550
  29. 29.Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (Virtual Event, China) (SIGIR ’20). Association for Computing Machinery, New York, NY, USA, 39–48. https://doi.org/10.1145/3397271.3401075
  30. 30.Yehuda Koren, Robert Bell, and Chris Volinsky. 2009. Matrix Factorization Techniques for Recommender Systems. Computer 42, 8 (2009), 30–37. https://doi.org/10.1109/MC.2009.263
  31. 31.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Eduardo Blanco and Wei Lu (Eds.). Association for Computational Linguistics, Brussels, Belgium, 66–71. https://doi.org/10.18653/v1/D18-2012
  32. 32.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics 7 (2019), 452–466. https://doi.org/10.1162/tacl_a_00276
  33. 33.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (, Vancouver, BC, Canada,) (NIPS ’20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages.
  34. 34.Chen Li, E. Chang, H. Garcia-Molina, and G. Wiederhold. 2002. Clustering for approximate similarity search in high-dimensional spaces. IEEE Transactions on Knowledge and Data Engineering 14, 4 (2002), 792–808.
  35. 35.Chao Li, Zhiyuan Liu, Mengmeng Wu, Yuchi Xu, Huan Zhao, Pipei Huang, Guoliang Kang, Qiwei Chen, Wei Li, and Dik Lun Lee. 2019. Multi-Interest Network with Dynamic Routing for Recommendation at Tmall. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM ’19). 2615–2623.
  36. 36.Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In International Conference on Learning Representations. https://openreview.net/forum?id=Bkg6RiCqY7
  37. 37.Yu A. Malkov and D. A. Yashunin. 2020. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Trans. Pattern Anal. Mach. Intell. 42, 4 (apr 2020), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473
  38. 38.Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel. 2015. Image-Based Recommendations on Styles and Substitutes. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval (Santiago, Chile) (SIGIR ’15). Association for Computing Machinery, New York, NY, USA, 43–52. https://doi.org/10.1145/2766462.2767755
  39. 39.Stanislav Morozov and Artem Babenko. 2018. Non-metric Similarity Graphs for Maximum Inner Product Search. In Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2018/file/229754d7799160502a143a72f6789927-Paper.pdf
  40. 40.Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022. Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models. In Findings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 1864–1874. https://doi.org/10.18653/v1/2022.findings-acl.146
  41. 41.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large Dual Encoders Are Generalizable Retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 9844–9855. https://doi.org/10.18653/v1/2022.emnlp-main.669
  42. 42.Rodrigo Nogueira and Jimmy Lin. 2019. From doc2query to doctttttquery. https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf
  43. 43.Hiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang, Tamas Feher, and Yong Wang. 2024. CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs.
  44. 44.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 [cs.LG] https://arxiv.org/abs/1910.10683
  45. 45.Parikshit Ram and Alexander G. Gray. 2012. Maximum Inner-Product Search Using Cone Trees. In Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’12). 931–939.
  46. 46.Steffen Rendle, Walid Krichene, Li Zhang, and John Anderson. 2020. Neural Collaborative Filtering vs. Matrix Factorization Revisited. In Fourteenth ACM Conference on Recommender Systems (RecSys’20). 240–248.
  47. 47.Stephen Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Found. Trends Inf. Retr. 3, 4 (April 2009), 333–389. https://doi.org/10.1561/1500000019
  48. 48.Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022. ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (Eds.). Association for Computational Linguistics, Seattle, United States, 3715–3734. https://doi.org/10.18653/v1/2022.naacl-main.272
  49. 49.Noam Shazeer, *Azalia Mirhoseini, *Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations. https://openreview.net/forum?id=B1ckMDqlg
  50. 50.Yikang Shen, Zheyu Zhang, Tianyou Cao, Shawn Tan, Zhenfang Chen, and Chuang Gan. 2023. ModuleFormer: Modularity Emerges from Mixture-of-Experts. arXiv:2306.04640 [cs.CL] https://arxiv.org/abs/2306.04640
  51. 51.Anshumali Shrivastava and Ping Li. 2014. Asymmetric LSH (ALSH) for Sublinear Time Maximum Inner Product Search (MIPS). In Advances in Neural Information Processing Systems, Vol. 27.
  52. 52.Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via Self-Attentive Neural Networks. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (Beijing, China) (CIKM ’19). Association for Computing Machinery, New York, NY, USA, 1161–1170. https://doi.org/10.1145/3357384.3357925
  53. 53.Weiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang, Haichao Zhu, Pengjie Ren, Zhumin Chen, Dawei Yin, Maarten Rijke, and Zhaochun Ren. 2023. Learning to Tokenize for Generative Retrieval. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 46345–46361. https://proceedings.neurips.cc/paper_files/paper/2023/file/91228b942a4528cdae031c1b68b127e8-Paper-Conference.pdf
  54. 54.Shulong Tan, Zhixin Zhou, Zhaozhuo Xu, and Ping Li. 2020. Fast Item Ranking under Neural Network based Measures. In Proceedings of the 13th International Conference on Web Search and Data Mining (Houston, TX, USA) (WSDM ’20). Association for Computing Machinery, New York, NY, USA, 591–599. https://doi.org/10.1145/3336191.3371830
  55. 55.Yi Tay, Vinh Q. Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, Tal Schuster, William W. Cohen, and Donald Metzler. 2022. Transformer Memory as a Differentiable Search Index. In Advances in Neural Information Processing Systems, Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.). https://openreview.net/forum?id=Vu-B0clPfq
  56. 56.Yiwei Wang, Bryan Hooi, Yozen Liu, Tong Zhao, Zhichun Guo, and Neil Shah. 2022. Flashlight: Scalable Link Prediction With Effective Decoders. In Proceedings of the First Learning on Graphs Conference (Proceedings of Machine Learning Research, Vol. 198), Bastian Rieck and Razvan Pascanu (Eds.). PMLR, 14:1–14:17. https://proceedings.mlr.press/v198/wang22a.html
  57. 57.Yujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao, Shibin Wu, Qi Chen, Yuqing Xia, Chengmin Chi, Guoshuai Zhao, Zheng Liu, Xing Xie, Hao Sun, Weiwei Deng, Qi Zhang, and Mao Yang. 2022. A Neural Corpus Indexer for Document Retrieval. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 25600–25614. https://proceedings.neurips.cc/paper_files/paper/2022/file/a46156bd3579c3b268108ea6aca71d13-Paper-Conference.pdf
  58. 58.Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018. Breaking the Softmax Bottleneck: A High-Rank RNN Language Model. In International Conference on Learning Representations (ICLR’18).
  59. 59.Jiaqi Zhai, Zhaojie Gong, Yueming Wang, Xiao Sun, Zheng Yan, Fu Li, and Xing Liu. 2023. Revisiting Neural Retrieval on Accelerators. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Long Beach, CA, USA) (KDD ’23). Association for Computing Machinery, New York, NY, USA, 5520–5531. https://doi.org/10.1145/3580305.3599897
  60. 60.Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 58484–58509. https://proceedings.mlr.press/v235/zhai24a.html
  61. 61.Jiaqi Zhai, Yin Lou, and Johannes Gehrke. 2011. ATLAS: A Probabilistic Algorithm for High Dimensional Similarity Search. In Proceedings of the 2011 ACM SIGMOD International Conference on Management of Data (SIGMOD ’11). 997–1008.
  62. 62.Han Zhu, Xiang Li, Pengye Zhang, Guozheng Li, Jie He, Han Li, and Kun Gai. 2018. Learning Tree-Based Deep Model for Recommender Systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (London, United Kingdom) (KDD ’18). 1079–1088.
  63. 63.Shengyao Zhuang, Houxing Ren, Linjun Shou, Jian Pei, Ming Gong, Guido Zuccon, and Daxin Jiang. 2023. Bridging the Gap Between Indexing and Retrieval for Differentiable Search Index with Query Generation. arXiv:2206.10128 [cs.IR] https://arxiv.org/abs/2206.10128
  64. 64.Jingwei Zhuo, Ziru Xu, Wei Dai, Han Zhu, Han Li, Jian Xu, and Kun Gai. 2020. Learning optimal tree models under beam search. In Proceedings of the 37th International Conference on Machine Learning (ICML’20). JMLR.org, Article 1080, 10 pages.

Citation

MLA
Ding, B., and J. Zhai. “Retrieval with Learned Similarities”. Proceedings of the ACM on Web Conference 2025, 2025, pp. 1626–37, https://doi.org/10.1145/3696410.3714822.
APA
Ding, B., & Zhai, J. (2025). Retrieval with Learned Similarities. Proceedings of the ACM on Web Conference 2025, 1626–1637. https://doi.org/10.1145/3696410.3714822
Chicago
Ding, B., and J. Zhai. 2025. “Retrieval with Learned Similarities”. Proceedings of the ACM on Web Conference 2025, 1626–37. https://doi.org/10.1145/3696410.3714822.
Harvard
Ding, B. and Zhai, J. (2025) “Retrieval with Learned Similarities”, Proceedings of the ACM on Web Conference 2025. ACM, pp. 1626–1637. Available at: https://doi.org/10.1145/3696410.3714822.
Vancouver
1. Ding B, Zhai J (2025) Retrieval with Learned Similarities. In: Proceedings of the ACM on Web Conference 2025. ACM, pp 1626–1637

BibTeX

@inproceedings{Ding_2025, series={WWW ’25}, title={Retrieval with Learned Similarities}, url={http://dx.doi.org/10.1145/3696410.3714822}, DOI={10.1145/3696410.3714822}, booktitle={Proceedings of the ACM on Web Conference 2025}, publisher={ACM}, author={Ding, Bailu and Zhai, Jiaqi}, year={2025}, month=Apr, pages={1626–1637}, collection={WWW ’25} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/